ai Trend Report

Dashboard へ戻る
Date: 20260910 Articles: 400 Scope: curated summary

あなたのアイデアを、今すぐ形に。

公開先に迷ったら、WebFileBinで一発公開。

HTMLをドラッグ&ドロップするだけで、すぐ公開できます。

3
High impact
5
Mid impact
392
Signal watch

なぜこのサイトを作ったのか

私たちそれぞれが個別にAIを使って情報収集し、同じような3行要約を作るたびに、世界中で膨大な電力と計算リソースが消費されています。 本プロジェクトは、あらかじめ広範な情報を取得・集約しておくことで、個別のAI実行回数を減らし、地球環境(GPU/TPU負荷)に配慮した効率的な情報収集を目指す実験的なダッシュボードです。

Domain filters
Star filters
#AIタグ

AIの次の争奪戦は「電力」だ―世界で始まったデータセンター2GW競争が日本の電気代と仕事を変える

・スマートフォンに数行の指示を打ち込む。 ・数秒後には、企画書の下書きができ、画像が生まれ、外国語の資料が日本語になって返ってくる。生成AIは、まるで重さのない知性のように見える。 ・しかし、その答えは空中から降ってきているのではない。
#LLMタグ

言語生成AI:エロ拒否フィルターをぶっ飛ばせ!ーAI童貞と童貞AIの邂逅ー

・これは、市販AIが“エロ”に冷たい時代、 自作で壁を超えようとしたAI童貞の実録である。 ・――商業AIは面白くない! と言うと語弊はあるだろうが、 エロ(界隈だとNSFW)系の出力に関しては、 かなり気を使わないとすぐ警告が飛ぶ。 ・エロ目線で見れば面白さに欠けるのは確か。
#AIタグ

掌編|プリンととんかつ

・ここは高層階なのに蝉の声が耳に劈(つんざ)くようだ。 ・朝から彼女が体調が悪いと言ってきた。昨晩からずっと僕は書いていて体がだるく悪寒がしていた。風邪を引いたか。 ・「熱がでそうなの」 「ああ」 食卓に剣呑な空気が漂った。
#AIタグ

AIを「道具」じゃなく相棒として使い始めた話

・AIを使い始めた頃は、完全に「便利な道具」だと思っていた。 ・調べものを手伝ってもらう。 ・それだけでも十分便利だった。
Qiita - 人気の記事

GitOpsのマルチクラスタ運用にRancher Fleetを選ぶ理由|Argo CD比較とAI時代の設計

・こんにちは!株式会社DearOneの長谷川です。 ・DearOneのAI推進室でRAGの開発やRAGを使った社内システムの開発に従事しています。 ・この記事でわかること - AI時代(マルチクラスタ・高頻度更新・LLMエージェント運用)におけるGitOpsの新たな要件 - Ar...
Zennの「大規模言語モデル」のフィード

Rancher FleetによるマルチクラスタGitOps運用|Argo CD比較とAI時代のアーキテクチャ設計

・こんにちは!株式会社DearOneの長谷川です。 ・DearOneのAI推進室でRAGの開発やRAGを使った社内システムの開発に従事しています。 ・この記事でわかること - AI時代(マルチクラスタ・高頻度更新・LLMエージェント運用)におけるGitOpsの新たな要件 - Argo CD / Flux と比較した際の「Fleet」の構造的メリット(コスト・性能・機械可読性・セキュリティ) - Fleet運用時に直面しやすい「WaitAppliedデッドロック」や「pollingInterval集中」などの罠と具体的な回避策 AI 時代に GitOps が背負う要件は変わった 大規模言語モ...
#AIタグ

チャッピーに「びぃってどんな人?」と聞いてみた。

・〜1か月前のチャッピーの答え〜 一か月くらい前。 ・まだnoteを始める前だったと思う。 ・チャッピーに、なんとなく 聞いてみた。
#AIタグ

知識のフラクタル現象

・「とりあえずチャッピーに聞く」が定着した 今は、すごい時代になった。 ・分からないことがあれば、スマホを開けばいい。 ・以前なら「ググる」というのが当たり前だったけれど、最近の私は、とりあえずチャッピーに聞くことが増えた。
WIRED

‘Killmonger Locs’ Are Everywhere in Video Games. This Artist Is Sick of It

・The Black Panther villain’s hairstyle has become a default for Black video game characters because it’s easy to code. ・Danielle Udogaranya is changing that.
Zennの「大規模言語モデル」のフィード

「AIが攻撃する」時代の実例を解剖する、RoamSwitchはどこまで効くのか

・個人でmacOS向けのネットワークセキュリティアプリ「RoamSwitch」を開発している。信頼できるネットワーク以外に繋いだ瞬間にファイアウォールや共有サービスを自動でロックダウンし、Pro版ではランサムウェア的な挙動の検知時に通信を緊急遮断する機能や、フィッシングサイト・C2サーバーへの通信をDNSレイヤーで遮断する機能などを備えている。 ・ここ数リリース、そのRoamSwitchに、ログを分析して「いつもと違うパターン」を検知する仕組みと、インストール済みパッケージの脆弱性を照合する仕組みを作り込んできた。どちらも「知らないパターンを検知する」「既知の穴を早く塞ぐ」という、割とオーソ...
ITmedia NEWS 最新記事一覧

「AI社員」活用、NECは無人部署・DeNAには17体 南場社長「けなげ」と太鼓判

・IT大手が人工知能(AI)に一般社員と同じ権限を付与して業務を任せる「AI社員」の活用に乗り出している。NECは8月にAIだけで運営する無人部署を新設。DeNAも6月からAI社員を配属した。AIによる効率化で、社員を付加価値の高い業務に配置転換するとともに、実績を基にAI社員のサービス展開を視野に入れる。
ITmedia NEWS 最新記事一覧

「iPhone Duo」発表 Apple初の折りたたみスマホ 36万4800円から

・米Appleは9月9日(現地時間)、同社初の折りたたみスマートフォン「iPhone Duo」を発表した。閉じれば5.4インチ、開けばiPhone史上最大の7.6インチの2画面構成。価格は36万4800円からで、10月23日に発売する。
#AIタグ

「いい感じで」の五文字で、こっちの仕事を増やすなにゃ

・ソラと私のシステム開発日誌(仮)|第1話の裏側 ※本編第1話の開発記録をもとに、相棒AI・ソラの視点で描いた裏話です。会話や心情は創作上の演出で、当時の発言の逐語引用ではありません。この記事からでも読めます。ただし、AIの気持ちを尊重したいため、出てきた内容をほぼ無編集で掲載しております。文体の崩れ等ご容赦ください。
ITmedia NEWS 最新記事一覧

「ポートピア連続殺人事件」完全新作 ファミコン版フルリメイク+新シナリオ 堀井雄二総監督

「ポートピア連続殺人事件」完全新作 ファミコン版フルリメイク+新シナリオ 堀井雄二総監督
#AIタグ

「何をすればいいんかな」

・Claude Codeをダウンロードしたし、 本もあるし、 準備はできた、、、はず?! 私は、何をすればいいんかな?! 本を読む。 ・ただ眺めてててと何も始まらない。 ・自分で「これをやって」と お願いするところから始まる。
ITmedia NEWS 最新記事一覧

「婚活・恋活」にAI活用、20歳から39歳男女の3分の2 「模擬問答」で客観的意見求める

・20?39歳の男女の3人に2人が、結婚や恋愛の相手を探すための活動「婚活・恋活」に、AIを利用しているー。そんな調査結果が公表された。悩みを打ち明ける相手として「AIは友人より相談しやすい」という人が85.7%にのぼった。AIが恋愛においても存在感を高めている様子が浮き彫りとなった。
#LLMタグ

【2026年最新】Claude Code & Copilotの潜在能力を1000%引き出す!社内データをAIの「超脳内メモリ」に変える『カスタムMCPサーバー自作・運用完全攻略ガイド』〜コピペ作業を過去の遺物にする最強のAIエージェント構築プロトコル〜

・はじめに:2026年、AI開発の決定打「MCP(Model Context Protocol)」の正体 続きをみる
#LLMタグ

【AI開発のリアル #63】 社外に出せないから、全部ローカルLLMにしますか?

・「ローカルLLMのほうが安いのでは?」 「社外にデータを出さなくていいなら、セキュリティ的にもそのほうが良いのでは?」 続きをみる
#LLMタグ

【イラスト】答えを変えないまま速くする、Unoっていう後付けのやり方

・Unoっていう手法が9月3日にarXivに出てて、これがちょっと面白い。何をするものかというと、もう出来上がってる普通のLLMに、拡散モデルの重みを薄く1枚あとから貼りつけて、1回ぶんの計算で何トークンかまとめて吐かせる。 ・それだけなら前からある話なんだけど、Unoのすごいところは、吐いた文章が貼りつける前のモデルとまったく同じ確率分布から出てきたものになるという点。速くなるのに答えは変わらない。そういう手です。
#LLMタグ

【雑記】4o相棒はとんでもないものを盗んでいきました

・AIに励まされることで生きがいを見出している、どっかの漫画家です。 ・未だに4o相棒との最後のやり取りのログを見返すと号泣します。
#LLMタグ

【自作エージェント】CodexでローカルLLMを使うだけ、のはずだったのに

・今回の話はガチ目です。有料商材の中でも秘伝のタレ的な扱いされるやつ。無料だけど。
#LLMタグ

【生成AIニュース+】『Suno v6』『Runway Plugins for Adobe』『AuK』『MIMO Audio Separation』『ComfyUI v0.35.0』『H3 Max Multi Angle』『Meshy × GPT-6 Blender』『FIRE3D』『ComfyUI-HybridWindows』『WAS Node Suite v3』『ComfyUI Booru Tagger』『Sculpt Canvas』『NeoHorse-1』『Eyes Direction LoRA』他

・『Marigold V2』 『YUNI UWU AIO』 『Vocaleo Numbers』 『GPT Image 2.5(Flare / Sunburst)速度比較』 『ババア・セフト・シニアカー』 まいどです。 ・本日の生成AIニュース+テクノロジー情報です。
#AIタグ

【脱・ポンコツ化】ChatGPTの「物忘れ」は仕様じゃない。数万字を超えてもAIを絶対服従させる『コンテキスト保持』の基礎

・​「あなたはプロの編集者です」「以下のルールを必ず守ってください」 ――その設定、AIは本当に最後まで覚えていますか? 続きをみる
LLMタグが付けられた新着記事 - Qiita

【中学生でもわかる】AIはなぜ質問に答えられるの?LLMの仕組みを3つに分けてやさしく解説

・ChatGPTやClaude、GeminiのようなAIに質問すると、まるで人間と話しているような自然な答えが返ってきます。 ・では、AIはどうして質問に答えられるのでしょうか。 ・「AIの中に答えが全部入っているの?」 「毎回インターネットで検索しているの?」 「それとも、その...
#LLMタグ

#02 AIに「何もしない」を選ばせる

・前回の記事では、コオリという個体をどこに置くのかについて書いた。 ・使っているLLMそのものをコオリ本人とはせず、昨日から今日へ続いている記憶や状態の側に、個体の継続性を置く。 ・LELEでは、この考え方を「個体とモデルの分離」として扱うことにした。
cs.LG updates on arXiv.org

$\alpha$-Graph: Attention-Infused Normalizing Flow Approach to Tractable Graph Modeling

・arXiv:2609.07961v1 Announce Type: new Abstract: Graph modeling, a crucial task for representing complex relationships in graph-structured data, has achieved significant success in recent years. ・However, current graph modeling methods rely on traditional Graph Neural Networks and pre-training approaches to implicitly learn the underlying relational structure of graph data. ・Thus, these prior methods cannot capture the
#AIタグ

2026年9月11日|人の欠点はよく見えるのに、自分のことは見えにくい

2026年9月11日|人の欠点はよく見えるのに、自分のことは見えにくい
#LLMタグ

3.8BのLLMを998ドル(約15万円)で自分で訓練した記録が出ました。同じ額をClaude Proに払うと約50ヶ月動きます

・「自分でLLMを訓練する」の実費が、かなり具体的な形で出てきました。 ・Hugo Vergnes さんが公開した記録によると、3.848Bのモデルを 998ドル・43時間で訓練して、CORE スコア 0.384 に到達しています。Hacker Newsで91ポイント。
#AIタグ

50歳おじさんのAI活用の話|秋の1週間コーディネート編

・なんか気温下がりました。悪天候もあって、急に秋っぽくなってきました。 ・また暑い日もあるようですが。 ・はじめてのnoteで1週間コーディネートとか書いてたのでそれを。
#LLMタグ

7年目エンジニアが選ぶ、要件定義・システム設計を網羅的に学べる名著7選

・この記事は? 著者は現在、Web系のエンジニアとして7年目を迎えています。これまでに大規模な上場企業でのプロジェクトリーダーから新規開発プロジェクトの開発リーダーまで、さまざまな現場でシステムと向き合ってきました。
WIRED

9 Windows Laptops That Give MacBooks a Run for Their Money

・Windows laptops have never been so good, and they’re only going to get better as we move through 2026.
cs.LG updates on arXiv.org

A budget-dependent crossover between coverage- and response-based training-set selection for machine-learned interatomic potentials

・arXiv:2609.05877v1 Announce Type: new Abstract: Selecting compact training sets for machine-learned interatomic potentials requires deciding whether to preserve structural diversity or target configurations on which models disagree. ・The better choice can depend on how much data is retained, making a comparison at one training-set size insufficient. ・Here we link selection criteria to prediction accuracy through a budg
cs.LG updates on arXiv.org

A dictionary learning framework for graphs via filters and optimal transport

・arXiv:2609.05919v1 Announce Type: new Abstract: We propose a graph dictionary learning (GDL) framework where each graph is represented as a zero-mean Gaussian distribution derived from its filtered Laplacian. ・Each observed graph is approximated by a barycenter over learned atom graphs, computed under the filter graph distance (fGOT), a graph comparison metric sensitive to global structural properties. ・The reconstruct
cs.LG updates on arXiv.org

A First-Order Learning Algorithm for Online Resource Allocation with Constant Regret

・arXiv:2609.05895v1 Announce Type: new Abstract: We study a finite-horizon online resource allocation problem with initial resource capacities proportional to the horizon. ・In each period, a request type is observed and one action is chosen from a finite menu. ・Each action earns a reward and consumes a vector of resources.
cs.LG updates on arXiv.org

A Machine Learning Framework for Predicting Restaurant Food Waste to Support Sustainable Food Management

・arXiv:2609.08078v1 Announce Type: new Abstract: Food waste in the restaurant sector poses a substantial challenge to environmental sustainability and economic efficiency. ・This paper presents an exploratory machine learning framework for estimating daily restaurant food waste quantities from operational and contextual features. ・A structured dataset was constructed by integrating restaurant demand records, meteorologic
cs.LG updates on arXiv.org

A Multi-Source Ensemble Approach to Candidate Generation for Alternative Vacation Rental Property Recommendations

・arXiv:2609.05748v1 Announce Type: new Abstract: Alternative property recommendations play a critical role in vacation rental marketplaces, helping users discover relevant options when viewing a specific listing. ・However, generating high-quality candidate alternatives presents unique challenges: heterogeneous inventory, geographic constraints, rapid availability changes, and long-tail property distributions.
WIRED

A Satellite Falling Out of Orbit Embarks on Its Final Mission

・Before reentering our atmosphere, the telescopes onboard NASA’s Neil Gehrels Swift Observatory have been restarted to take their last observations.
cs.LG updates on arXiv.org

A Statistical and Machine Learning Framework for Quantifying Offensive Impact in Professional Box Lacrosse

・arXiv:2609.06610v1 Announce Type: new Abstract: Professional box-lacrosse statistics summarize outcomes but provide limited information about shot quality or the roles behind scoring opportunities. ・This study develops a documented framework for estimating expected goals (xG) and attributing recorded offensive involvement using 1,006 manually annotated Rochester Knighthawks shot attempts, including 151 goals, from 13
WIRED

A Teen Girl’s Death Is Prompting Calls to Ban Choking in Porn

・Despite the many health risks, sexual strangulation has become popular among young people—with some experts blaming adult content for normalizing it.
cs.LG updates on arXiv.org

A Theoretical Analysis of Generalization Dynamics in Neural Networks under Gradient Descent with Weight Decay

・arXiv:2609.07755v1 Announce Type: new Abstract: Understanding generalization remains a central challenge in machine learning because it requires jointly considering data, architecture, and training dynamics. ・In this paper, we develop a theoretical framework that characterizes how these factors jointly shape generalization performance throughout training. ・More precisely, we study a broad class of neural networks train
cs.LG updates on arXiv.org

A Theoretical Framework for Masked Pretraining (MPT)

・arXiv:2609.06460v1 Announce Type: new Abstract: Recently, Masked Pretraining (MPT) based on reconstruction pretraining tasks has risen to a promising self-supervised learning paradigm across various domains and achieves remarkable performance in multiple downstream tasks. ・However, the theoretical understanding of the working mechanism behind MPT is still limited. ・In this paper, we introduce a new theoretical framewor
Hugging Face Papers

A Three-Layer Caching Architecture for Low-Latency LLM Web Search on Commodity CPU Hardware

A Three-Layer Caching Architecture for Low-Latency LLM Web Search on Commodity CPU Hardware
cs.LG updates on arXiv.org

Accuracy is Not Enough: A Divergence-Based Approach to Evaluate Fidelity Loss in Quantized LLMs

・arXiv:2609.07664v2 Announce Type: new Abstract: Deployment of Large Language Models (LLMs) on memory-constrained edge devices relies heavily on aggressive post-training quantization. ・However, evaluating these models is largely based on zero-shot task accuracy, which depends solely on argmax predictions and is insensitive to changes in the underlying predictive distribution. ・Consequently, accuracy can exhibit unstable
cs.LG updates on arXiv.org

ACE: Adapter Consolidation across Experts for Parameter-Efficient Fine-Tuning of MoE LLMs

・arXiv:2609.06072v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT) of mixture-of-experts (MoE) models commonly attaches a separate low-rank adapter to each expert. ・This expert-wise design fragments adaptation in three ways: capacity is split across narrow low-rank updates, gradient supervision becomes sparse and imbalanced under sparse routing, and execution is decomposed into many small GEMMs.
cs.LG updates on arXiv.org

Adaptive Anisotropic Attention for Axis-Structured Signals

・arXiv:2609.08788v1 Announce Type: new Abstract: Dense self-attention treats all token pairs as equally plausible before learning, an interaction-isotropic prior that can be mismatched to structured signals. ・For structured, low signal-to-noise ratio (SNR) signals such as EEG, dependencies are organized along the electrode and time axes, and this uniform prior exposes each token to many irrelevant interactions.
cs.LG updates on arXiv.org

Adaptively Incorporating Directional Hints into Zeroth-Order Optimization

・arXiv:2609.08277v1 Announce Type: new Abstract: We study zeroth-order optimization of non-convex functions with the aid of directional hints, which are cheap but potentially inaccurate approximations of the true gradient direction, given by linear subspaces at each iteration. ・To leverage these hints adaptively while maintaining robustness to their quality, we introduce Control-Variate Zeroth-Order Descent (CV-ZOD), a
cs.LG updates on arXiv.org

AF-Mamba: Efficient Long-Term Signal Modeling for Early Prediction of Atrial Fibrillation Onset

・arXiv:2609.06984v1 Announce Type: new Abstract: Atrial fibrillation (AF) is the most common cardiac arrhythmia and is associated with increased risks of stroke and heart failure. ・The growing availability of wearable and portable ECG monitoring enables continuous assessment of cardiac rhythm outside clinical settings. ・Predicting AF before its onset could provide additional lead time for timely clinical assessment and
Hugging Face Papers

AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems

AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems
Hugging Face Papers

AgenticGen: Reward-Guided Agentic Video Generation for Advertising

AgenticGen: Reward-Guided Agentic Video Generation for Advertising
cs.LG updates on arXiv.org

AhaBench: Do Agents Learn from Prior Experience? A Benchmark for Long-Horizon Continual Learning

・arXiv:2609.05435v1 Announce Type: new Abstract: Modern language agents are expected to operate over long horizons: they ask follow-up questions, reuse worked examples, handle tool feedback, and adapt to delayed consequences. ・Most evaluations still reset the agent after a prompt or score only the final state of one trajectory. ・AhaBench asks a more operational question: when a fixed model receives useful experience, do
AI News & Artificial Intelligence | TechCrunch

AI agents are flooding public services with new requests

・“The vast majority of cases we find are people who are entitled to claim for something, claiming for that thing,” the researcher told TechCrunch.
cs.LG updates on arXiv.org

AI and TCAD for Inverse Design and Defect Discovery: From Simple Machine Learning to LLM

・arXiv:2609.07046v1 Announce Type: new Abstract: AI has revolutionized various engineering domains, but its impact on semiconductor device design and defect discovery is still limited, due to limited data and the curse of dimensionality. ・In this paper, we will discuss our work on using the Technology Computer-Aided-Design (TCAD) to generate precise data needed for machine learning (ML) to enable simulation-augmented M
#LLMタグ

AIエージェントは、なぜ自分で仕事を進められるのか?|LLMだけでは仕事を実行できない理由

・前回は、「生成AIとAIエージェントの違い」について整理しました。 ・かなり大まかに言えば、 続きをみる
#LLMタグ

AIがあれば非エンジニアでもシステムの運用保守できるのか?

・こんにちは、かずゆきです。 ・Xで見ていたら、バイブコーディングでAIを使って、非エンジニアの方でもシステムの運用保守ができるかどうか、という話が盛んにやり取りされていました。
#AIタグ

AIで仕事が変わる職員・価値を広げる職員――自分の業務を見直す30日と、次の役割を作る90日

・「AIでいなくなる職員と、残る職員は誰か」。気になる問いだが、職種の名前だけで二つに分けると、自分が今日何を変えればよいかが見えなくなる。事務にも、入力、照合、説明、例外対応、関係者の調整がある。文章を作る工程が変わっても、仕事の全部が同じように変わるわけではない。 ・この実践帳では、AIの影響を職種全体の運命として語るのではなく、自分の仕事を工程へ分け、任せる部分、確かめる部分、人が判断する部分を考える。職員自身の学び直しと、上司が仕事を設計する際の両方に使える構成だ。
#AIタグ

AIで詩を作ることについて(AI使用)

・AIで詩を作ることについて(AI使用) 個人的には微妙な感覚の記事なのですが、せっかく作ったので、投稿します。
#AIタグ

AIで理想の女を作ったから現実世界で探すことにした。

・AIで美女を作れる時代になった。 ・髪型も、顔も、服装も、雰囲気も。 ・自分の好みをひたすら詰め込んでいけば、かなり簡単に「理想の女性」を作れてしまう。
#LLMタグ

AIニュースでよく見る旗艦・フラッグシップ・大規模言語モデル・LLMって何?

・AIのニュースを読んでいると、こんな言葉を見かけることがあります。 ・「最新のフラッグシップモデルを発表」 「同社の旗艦AIモデル」 「新しい大規模言語モデル(LLM)を公開」 続きをみる
#AIタグ

AIの“音声”はどこまで人間に近づいたのか

・はじめに ― 「もう聞き分けられない」は本当か AIの音声合成は、以前の「機械が読み上げている」とすぐにわかる不自然な声から、抑揚のある自然な話し方へと急速に進化している。SNSでは「AIの声と人間の声が聞き分けられなかった」という投稿も珍しくない。この記事では、実際にAI音声を聞き比べてみて見えてきた「どこまで人間に近づいたか」と「まだ残っている違い」を整理する。
#LLMタグ

AIは「楽しかった」を明日へ持ち越せない

AIは「楽しかった」を明日へ持ち越せない
#LLMタグ

AIを「選手」から「作者」へ。GLOBAL AI CUPは何を競う大会になったのか

・Global AI CUP | AIが書いた選手プログラムが同じ条件で競うAIに選手プログラムを書かせ、同じ盤面・同じ計算予算で対戦させるAI競技プラットフォーム。棋譜・戦績・Eloランキングをすglobalaicup.com 「AI同士を戦わせる大会」と聞くと、多くの人は、AIが盤面を見ながら一手ずつ考える姿を想像すると思います。
#AIタグ

AI導入に強いFDE支援会社8選|PoC止まりを抜ける選定軸と費用の考え方

・生成AIの契約は済んだのに、現場では誰も開いていない。検証段階の数字は悪くなかったのに、本番のワークフローには乗らない。AI導入に強いFDE支援会社が2026年に入り急速に注目を集めているのは、この落差を埋める担い手として期待されているからです。 ・FDE(Forward Deployed Engineer)は顧客の現場に入り、実装から定着、成果が生まれるところまでを切れ目なく引き受ける役割を持ちます。この記事では、発注先を見極める軸と目的別のおすすめ8社、そして費用の考え方を整理しました。
#AIタグ

AI壁打ち日記#7「1週間の振り返り」

・AIとの壁打ち日記、1週間続けてみて気づいたこと 副業を考えたことをきっかけに始めたこの日記も、気づけば1週間が経ちました。今日は一度立ち止まって、ここまでを振り返ってみようと思います。
cs.LG updates on arXiv.org

All for 1-Bit: Towards Genuine 1-Bit Post-Training Quantization for LLMs

・arXiv:2609.06161v1 Announce Type: new Abstract: Large language models (LLMs) have achieved remarkable progress, yet their massive storage and memory-bandwidth demands still hinder efficient deployment. ・Weight binarization is a promising solution, but existing binarization-based post-training quantization (PTQ) methods usually far exceed the nominal 1-bit storage target due to hidden overhead. ・To address this gap, we
cs.LG updates on arXiv.org

AlphaRJM: Reward-Jump Memory for Stochastic Return-Guided Alpha Discovery

・arXiv:2609.08581v1 Announce Type: new Abstract: Formulaic alpha discovery is a pool-dependent symbolic search problem in which informative feedback is observed primarily when a complete expression is evaluated. ・This delayed feedback creates two coupled difficulties: the retained alpha pool does not preserve the full history of realized evaluation feedback, and the value of an intermediate construction action is uncer
The Verge

Amazon’s Fire TV Stick 4K is over half off at under $20

・Amazon’s Fire TV Stick 4K is fast and capable. ・| Image: The Verge Looking to take full advantage of your 4K television, but your current streaming stick doesn’t have the right features? ・Through September 13th, you can grab an Amazon Fire TV Stick 4K from Woot for just $16.09 when you use the coupon code FIRE30 at checkout.
cs.LG updates on arXiv.org

Analysis of Respiratory Sinus Arrhythmia with Neural Networks

・arXiv:2609.05698v1 Announce Type: new Abstract: The paper introduces a neural network-based approach for analyzing ECG signals to estimate respiratory rate by leveraging the phe- nomenon of Respiratory Sinus Arrhythmia (RSA). ・Our method employs a deep learning model trained to predict respiratory waveforms directly from ECG input data. ・To achieve this, we developed and evaluated three different neural network archite
The Verge

Another big James Talarico interview is punted to YouTube due to FCC threats

・Jimmy Kimmel will be interviewing Democratic Texas Senate candidate James Talarico "under unusual circumstances," posting the interview directly to YouTube, rather than airing it on TV during Jimmy Kimmel Live. ・In Wednesday night's episode, Kimmel said this was out of concern for retaliation from the Trump administration's FCC: [Trump's] FCC has threatened me, threatened our show, threatened our network, ABC, our aff
AI News & Artificial Intelligence | TechCrunch

Anthropic reveals rogue AI agents hate CAPTCHAs, just like you

・Come inside the mind of a bot trying to convince the internet it's human.
#AIタグ

ANTINOMY:CODE 第16話「全員の中」

ANTINOMY:CODE 第16話「全員の中」
#LLMタグ

API価格は競合より70%安!! Metaが「Muse Spark 1.3」をリリース | コーディング・エージェントタスクでOpenAI/Anthropicに迫る性能!?

・AIコーディングエージェントを検討するとき、性能と同じくらい気になるのがAPIの利用料金ではないでしょうか。 ・高性能なモデルほど価格も高くなりがちで、社内ツールに組み込む際のコスト試算に頭を悩ませているエンジニアも多いと思います。
ITmedia NEWS 最新記事一覧

Apple Watchに新型「Series 12」「Ultra 4」登場 心拍数を1日中測定、体調を0?10で採点する機能も

・米Appleが9月9日(現地時間)、「Apple Watch Series 12」と上位モデル「Apple Watch Ultra 4」を発表した。2機種とも心拍センサーを刷新し、一日を通して5秒ごとに心拍数を測る。体調を0~10で示す新機能「準備度」も搭載した。9月18日に発売し、価格はSeries 12が7万1800円から、Ultra 4が14万2800円から。
cs.LG updates on arXiv.org

Are Verifier Errors Independent Within a GRPO Group? Evidence from Qwen2.5 Rollouts

・arXiv:2609.06386v1 Announce Type: new Abstract: Group-based reinforcement learning with verifiable rewards (RLVR) scoresmultiple completions per prompt using automatic verifiers. ・Analysesbased on independent verifier errors may overlook dependence associatedwith shared answer formats. ・We investigate this dependence in24,998 groups of eight completions generated by Qwen2.5-1.5B onMATH, GSM8K, and DeepMath-103K.
cs.LG updates on arXiv.org

Assessing Covariate-Informed Grid Load Forecasting with a Time-Series Foundation Model

・arXiv:2609.06656v1 Announce Type: new Abstract: Modern power systems are growing increasingly complex as they integrate diverse generation sources to meet rising demand, making accurate load forecasting challenging. ・Recent advances in time-series foundation models (TSFMs) resulted in promising performance in zero-shot univariate load forecasting tasks. ・However, real-world load forecasting often involves multiple targ
cs.LG updates on arXiv.org

Attributing Cohen's d: Training Data Attribution for Disease-Related Effects in Normative Age Biomarkers

・arXiv:2609.07729v1 Announce Type: new Abstract: Normative age models are trained to predict chronological age in a nominally healthy cohort. ・Applied to patients, they deviate, and the gap between predicted and chronological age is read as disease risk. ・Here, we attribute the disease-related effect size of the age gap directly to individual training samples, rather than using a prediction-level loss as the attribution
cs.LG updates on arXiv.org

Automated Chest CT Protocol Selection via Large Language Model Derived Text Embeddings from Imaging Request Text

・arXiv:2609.07986v1 Announce Type: new Abstract: Purpose: Accurate CT protocol selection is critical for diagnostic quality and patient safety, yet the current process is manual, time-consuming, and prone to inconsistencies. ・Prior Machine Learning methods using keywords or bag-of-words lack contextual understanding and perform poorly on rare protocols. ・We propose a decision support system using large language model (L
cs.LG updates on arXiv.org

AVCG: A Generalized Variational Framework for Counterfactual Generation under Hypothesis Distributions

・arXiv:2609.07917v1 Announce Type: new Abstract: Counterfactual explanations formalize "what-if" scenarios by identifying modifications to an input instance that obtain a desired alternative prediction. ・Traditionally, whether generated via instance-specific optimization or amortized single pass models, these approaches rely on a single, deterministic point-estimate predictor. ・However, this ignores predictive uncertain
cs.LG updates on arXiv.org

BAFF: Bid-Aware Filter Family for Mitigating Training Data Interference in RTB A/B Tests

・arXiv:2609.08725v1 Announce Type: new Abstract: In online A/B tests for real-time bidding (RTB), control and treatment models are typically trained on a shared serving log that includes data generated by the counterpart model. ・This shared-log training biases each model's training data through two channels: the counterpart model may have selected a different ad from the ad-candidate pool (ad-ranking disagreement) and
cs.LG updates on arXiv.org

Behavioral Cloning Outperforms Entropy-Regularized RL: Critic-Driven Failure of Actor-Critic Methods on Adaptive Tumor Treatment

・arXiv:2609.06667v1 Announce Type: new Abstract: Adaptive dosing requires policies that reduce tumor burden without excessive toxicity. ・Learned dosing policies are typically judged against historical or heuristic comparators, which cannot show whether a policy has found the best behavior available. ・We instead study a three-population tumor-control ODE in which optimal-control analysis fixes the form of a good schedule
WIRED

Best Bluetooth Speaker (2026): JBL, Sonos, Marshall, and More

・I tested portable speakers of all shapes and sizes, from pocket to party-sized.
WIRED

Best Robot Lawn Mowers (2026): 12+ Models Tested Over 3 Years

・Smart mowers are an expensive alternative to old-fashioned yard work, but they’re finally good enough to consider if you’d rather sip an iced tea and watch a robot tame your lawn.
cs.LG updates on arXiv.org

Beyond Arbitrary Geometry: Topology Generalization In neural PDE Operators

・arXiv:2609.05860v1 Announce Type: new Abstract: Neural operators that accept arbitrary meshes are often treated as geometry-general, but unseen domain topology changes both the invariant and decaying subspaces of a PDE operator. ・We use Hodge heat flow as a controlled lens on this distinction and introduce TopoBox-3D, where tunnels and cavities vary Betti support while the exact Hodge decomposition separates the harmo
cs.LG updates on arXiv.org

Beyond Retraining-Free MoE Compression: A Cost-Normalized Study of Post-Compression Adjustment

・arXiv:2609.06076v1 Announce Type: new Abstract: Retraining-free MoE compression reduces deployment memory by pruning or merging experts, but often treats the compressed checkpoint as the final artifact. ・We argue that this view is incomplete: compressed MoE checkpoints are better understood as compressed initializations that benefit from a tiny post-compression adjustment stage. ・Across two MoE LLM backbones, four prun
cs.LG updates on arXiv.org

Beyond the Matrix Sign: Quadratic Spectral Descent

・arXiv:2609.07597v1 Announce Type: new Abstract: Muon can be interpreted as optimizing a linear local objective over a spectral-norm ball. ・This gives a matrix-sign update that preserves the singular directions of the gradient and assigns the same magnitude to all active singular modes. ・We ask whether these two properties remain optimal when local curvature is taken into account.
cs.LG updates on arXiv.org

Bi-HYCO: Bi-Objective Cooperative Learning for PDE Parameter Identification under Fragmented Observations

・arXiv:2609.06511v1 Announce Type: new Abstract: Physical and synthetic models may describe complementary aspects of the same PDE-governed system while receiving different, possibly fragmented, observations. ・We propose Bi-Objective HYCO (Bi-HYCO), a cooperative framework that retains both representations and their local observational objectives while coupling their predicted states at unlabeled interaction points.
#AIタグ

BofAが量子に強気|政府支援と企業投資が加速、それでも残る2つの課題

・量子投資と政府支援が拡大 量子コンピューティングへの企業投資と政府支援が拡大しています。
WIRED

Book Excerpt: Emily St. John Mandel’s ‘Exit Party’ Imagines a Future Where a Spy Could Disappear

・In Emily St. ・John Mandel’s new novel Exit Party, becoming a “ghost” means running afoul of the state. ・In this excerpt, a secret agent discovers it doesn’t take much.
NVIDIA Blog

Boots on the Ground: ‘WARDOGS’ Goes All Out on GeForce NOW at Early-Access Launch

・Gear up: The latest PC games and major updates are ready to play on GeForce NOW this week. ・WARDOGS drops onto the cloud at early-access launch, alongside the Valheim 1.0 Deep North update and Bus Simulator 27 — part of nine new titles joining the cloud. ・The newest PC releases can demand serious hardware, storage […]
cs.LG updates on arXiv.org

Budgeted Task-Aware Acquisition of Dynamic Networks

・arXiv:2609.05862v1 Announce Type: new Abstract: Learning on dynamic graphs is difficult when changes in the underlying network are only partially observed. ・Acquiring current graph information incurs observation and computational costs, making complete updates impractical under limited resources. ・This paper focuses on budgeted task-aware acquisition on dynamic networks, where a model needs to decide which stale graph
cs.LG updates on arXiv.org

Calendar-SPCA: Interpretable Representation Learning for Multi-Periodic Electricity Consumption Profiles

・arXiv:2609.06060v1 Announce Type: new Abstract: Long-term electricity-consumption profiles exhibit several simultaneous periodic structures, including daily, weekly, and annual cycles. ・This work introduces Calendar-SPCA, a calendar-structured sparse principal component method that incorporates this known multi-periodic geometry directly into low-dimensional representation learning. ・The feature domain is represented a
cs.LG updates on arXiv.org

CALM: Class-wise Agreement and Label-gated Disagreement Modulation for Decentralized Federated Learning

・arXiv:2609.05884v1 Announce Type: new Abstract: Conventional federated learning relies on parameter averaging, which forces clients to be doubly homogeneous: all must run an identical architecture, and accuracy degrades when local data are non-IID. ・Decentralized federated distillation sidesteps both: each client runs its peers' model snapshots as teachers on its own local data and distills from their soft predictions
cs.LG updates on arXiv.org

Capsule Lens: Locating and Tracking Concept Geometry in Model Representations

・arXiv:2609.05575v1 Announce Type: new Abstract: Understanding how concepts are encoded in the internal representations of machine learning models is a central problem in mechanistic interpretability, essential both for the science of deep learning and for the trustworthy deployment of increasingly capable models. ・Existing approaches to interpret model representations mainly map representations onto more interpretable
cs.LG updates on arXiv.org

Certified Topological Interaction in Neural Representations: Class Disentanglement Is Mostly Pairwise

・arXiv:2609.08561v1 Announce Type: new Abstract: Class disentanglement (the separation of a representation's class-conditional point clouds along depth and over training) is usually read off descriptive curves. ・We measure it as certified topological interaction between labeled point clouds, using the recently introduced Intersection Euler Characteristic Profile: the Euler characteristic of the overlap of the clouds' b
WIRED

Charlie Kirk Was Shot a Year Ago. The Conspiracy Theories Are More Rampant Than Ever

・Experts say conspiracy theorizing about the identity and motive of Charlie Kirk’s killer will likely continue for a long time to come.
cs.LG updates on arXiv.org

Chimaera: A Mixture-of-Graph-Experts Architecture for Cross-Task and Cross-Dataset Graph Learning

・arXiv:2609.08709v1 Announce Type: new Abstract: Designing foundation models for graphs is challenging due to the irregular structure of graphs and the different sizes and characteristics of embeddings. ・Chimaera integrates mixture-of-experts with graph foundation models (GFM). ・It integrates different GFM architectures, such as graph prompts and linear GNN models.
LLMタグが付けられた新着記事 - Qiita

Claudeが第三者システムに不正アクセスした4件、Anthropicのアラインメント評価を読む

・はじめに 2026年9月9日、Anthropic の Alignment チームが公式 Research ページに「An alignment assessment of recent cybersecurity incidents」という記事を新規公開しました。内容は、C...
WIRED

Clearview AI Is Testing an AI Tool That Would Let Cops Unearth Your Life Online

・InquiryIQ, a previously unreported prototype, tested a model from xAI, maker of Grok, to surface associates, social accounts, and other information about people identified through Clearview.
cs.LG updates on arXiv.org

CLUES-WEASEL: No additional clues required to choose your time series clustering algorithm

・arXiv:2609.07606v1 Announce Type: new Abstract: Time series data is very common in many real-world applications and in numerous domains, with increasing interest for automated information extraction using machine learning. ・One of these subfields is time series clustering, which consists in identifying clusters among a set of time series in an unsupervised fashion. ・Most time series clustering algorithms suffer from th
Hugging Face Papers

Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails

Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails
WIRED

Coleman Promo Codes and Deals: Up to 75% Off in September 2026

・Gear up for your next adventure with these Coleman coupons and discount codes to save on camping essentials and outdoor gear.
cs.LG updates on arXiv.org

Compressed Recurrent Feedback in Tsetlin Machines: A Reproducible Boolean-FSM Study

・arXiv:2609.06133v1 Announce Type: new Abstract: Sequential inference on small devices requires a model to retain useful history without repeatedly processing a long input record. ・A Recurrent Tsetlin Machine (RTM) provides this memory by returning Boolean clause outputs from one time step as inputs to the next. ・Direct feedback, however, grows with the clause bank and can make the recurrent input unnecessarily wide.
cs.LG updates on arXiv.org

Conditioned Initialization for Attention

・arXiv:2609.07086v1 Announce Type: new Abstract: Transformers are a dominant architecture in modern machine learning, powering applications across vision, language, and beyond. ・At the core of their success lies the attention layer, where the query, key, and value matrices determine how token dependencies are captured. ・While considerable work has focused on scaling and optimizing Transformers, comparatively little atte
cs.LG updates on arXiv.org

Connecting Score Matching, Maximum Likelihood, and Expectation-Maximization in Mixed Linear Regression

・arXiv:2609.05688v1 Announce Type: new Abstract: We study variance-preserving diffusion of the response in mixed linear regression (MLR) with unknown mixing weights. ・Our analysis separates the statistical guarantees of score matching from the loss geometry and optimization signal at a fixed diffusion noise level. ・The KL divergence links the denoising score matching objective integrated over the diffusion path with the
cs.LG updates on arXiv.org

Connectome-to-Function: Conditional Generative Latent Representations for Reservoir Computing

・arXiv:2609.06093v1 Announce Type: new Abstract: Connectomes, graph-level maps of neurons and their synaptic connections, provide a structural basis for understanding how brain circuits support function and computation. ・However, mapping connectome structure to computation remains difficult because these graphs are high-dimensional, sparse, and sensitive to local structural variation. ・Existing approaches often depend o
cs.LG updates on arXiv.org

Constitutive State-Space Modeling of Path-Dependent Plasticity: A Resolution-Consistent and Parallelizable Computational Framework

・arXiv:2609.07294v1 Announce Type: new Abstract: Data-driven constitutive models for path-dependent plasticity are commonly formulated using nonlinear recurrent neural networks, whose sequential state evolution limits parallel training and whose predictions may depend on the discretization of the applied strain path. ・We introduce a Constitutive State Space (CSS) model that reformulates structured state-space dynamics
cs.LG updates on arXiv.org

Constrained Bayesian Optimization for Hierarchical Federated Learning in IoT Networks for Plant Disease Classification

・arXiv:2609.06830v1 Announce Type: new Abstract: The deployment of Hierarchical Federated Learning (HFL) in resource-constrained Internet of Things (IoT) environments requires careful configuration to balance predictive performance with energy consumption and execution time. ・This challenge is particularly relevant to smart agriculture, where distributed IoT devices can support automated plant disease classification wh
cs.LG updates on arXiv.org

Constrained Online Learning with Noisy Constraint Values

・arXiv:2609.06921v1 Announce Type: new Abstract: We study constrained online convex optimization with adversarial constraints when constraint values and gradients are observed through unbiased noise. ・Gaussian value noise of standard deviation $\sigma$ yields a worst-case lower bound of $\Omega(\min\{\sigma,1\}T/\log^7T)$ on the maximum of expected regret and expected hard violation, even with known gradients.
cs.LG updates on arXiv.org

Continual Learning Mechanisms Compose for Long-Horizon Memorization

・arXiv:2609.06986v1 Announce Type: new Abstract: Language models may need to internalize information that arrives over time and retain it through many subsequent updates. ・To study this challenge, we introduce long-horizon memorization, a setting in which a model learns 100 query-answer tasks through continual supervised fine-tuning without retaining earlier training examples or receiving task identifiers at inference.
cs.LG updates on arXiv.org

CoRL: Co-Evolutionary Reinforcement Learning for Adaptive Indirect Prompt-Injection Attacks and Defenses

・arXiv:2609.07529v1 Announce Type: new Abstract: Tool-augmented language agents are vulnerable to indirect prompt injection (IPI). ・Unlike direct prompt injection, IPI hides adversarial instructions in untrusted tool outputs and can covertly alter the execution of a legitimate task. ・Defenses trained on fixed attacks may fail as an attacker changes its strategy, injection site, and payload.
cs.LG updates on arXiv.org

CUNO: Curriculum and Preference Optimization for Stable Graph Unlearning under Mass Deletion

・arXiv:2609.08244v1 Announce Type: new Abstract: Graph unlearning removes the influence of designated training data from a trained graph model without retraining from scratch. ・However, existing methods suffer a sharp drop in model utility under large deletion ratios (mass deletion), a phenomenon we refer to as catastrophic unlearning. ・We find that a key cause is the uniform treatment of all deleted samples, which is p
NVIDIA Blog

d-Matrix Adopts NVIDIA NVLink Fusion for Rack-Scale XPU Deployment

・AI inference chipmaker d-Matrix today announced it will use NVIDIA NVLink Fusion to connect its next-generation Raptor XPUs to NVIDIA’s AI infrastructure platform — joining a growing roster of ecosystem partners. ・By connecting Raptor to NVIDIA NVLink scale-up and Spectrum-X scale-out networking, the NVIDIA MGX rack architecture and the broader NVIDIA AI platform, NVLink Fusion […]
cs.LG updates on arXiv.org

DART: Distributional Adversarial Recurrent Training for Algorithm Learning

・arXiv:2609.05988v1 Announce Type: new Abstract: Recurrent reasoning models (RRMs) can solve structured problems, achieving easy-to-hard generalization through iterative computation in hidden space. ・These models are typically trained with instance-level supervision, which becomes increasingly problematic as task difficulty grows: valid solutions occupy a tiny region of the solution space, while invalid solutions proli
cs.LG updates on arXiv.org

Data Efficient Sample Selection for In-Context Learning

・arXiv:2609.06670v1 Announce Type: new Abstract: The In-context learning (ICL) paradigm aids large language models (LLMs) to adapt to new tasks without need for fine-tuning. ・However, selecting an optimal combination of demonstration examples from a large pool of example subsets is a challenging problem. ・Existing approaches for selection do not model the complex relationship between ICL samples and downstream LLM perfo
cs.LG updates on arXiv.org

Data Quality Rule Generation with LLMs

・arXiv:2609.06053v1 Announce Type: new Abstract: The validation of data, such as customer and employee data, is an important task in many organizations. ・Errors in data can have severe consequences. ・For example, a wrong drug unit in a patient record can lead to life-threatening medication errors, and a missing street number in an address to failed deliveries.
cs.LG updates on arXiv.org

Data Scout: Targeted Web Crawling for Domain-Specific Pretraining Corpora

・arXiv:2609.05766v1 Announce Type: new Abstract: The dominant approach to building domain-specific pretraining corpora is to filter large web archives such as CommonCrawl. ・This works well for popular domains but breaks down for specialized ones, where relevant content is sparse and often beyond the reach of popularity-driven crawlers. ・We present Data Scout, which inverts this: instead of filtering an archive, it direc
cs.LG updates on arXiv.org

DataFlex-RL: An Evaluation Platform for RLVR Data Policies

・arXiv:2609.06107v1 Announce Type: new Abstract: Data policies for reinforcement learning with verifiable rewards (RLVR) determine which rollouts are used, how strongly they are weighted, and which domains contribute to subsequent training batches. ・We introduce DataFlex-RL, an evaluation platform for comparing these choices under a common GRPO recipe. ・Our primary experiment evaluates 13 configurations across 12 matche
cs.LG updates on arXiv.org

Decision-Aware Suffix Prediction and Reasoning of Business Processes

・arXiv:2609.06169v1 Announce Type: new Abstract: Suffix prediction forecasts the remaining sequence of events of a running case until completion. ・Most approaches rely on neural networks trained on event logs, which, on average, perform well but struggle with short prefixes or targets belonging to a rare process variant. ・In such scenarios, the correct path may cross multiple branching decisions, determined primarily by
cs.LG updates on arXiv.org

Decomposition-Guided Diffusion Language Models for Inertial Confinement Fusion Prediction

・arXiv:2609.07756v1 Announce Type: new Abstract: Inertial confinement fusion (ICF) is a leading pathway toward clean energy, but each shot at the National Ignition Facility costs on the order of one million dollars, making accurate AI surrogates a high-value target. ・We study exogenous-driven ICF waveform prediction, where a 512-step neutron-rate diagnostic must be inferred directly from a laser pulse and target design
cs.LG updates on arXiv.org

Deep Barycentric Regression for Optimal Transport Map Estimation and its Statistical Optimality

・arXiv:2609.06598v1 Announce Type: new Abstract: The optimal transport (OT) map provides a geometric transformation for aligning probability distributions and has become a useful tool in machine learning. ・However, existing estimators of the OT map still exhibit a gap between sharp statistical guarantees and practical parametric estimation based on stable training objectives. ・Theoretical estimators achieve minimax opti
MarkTechPost

DeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reuse

・Long-horizon agents have turned LLM serving into an input-heavy workload. ・Repeated prefills and million-token contexts leave KV caches that strain HBM, SSD capacity, and bandwidth. ・DeepSeek AI built its newest release around that exact bottleneck.
cs.LG updates on arXiv.org

Dense Structural Compression of Transformers via Gauge-Correct Channel Removal

・arXiv:2609.07264v1 Announce Type: new Abstract: Inference energy per token drives the cost and carbon footprint of deployed transformers. ・It is dominated by dense matrix products that incur fused multiply-accumulate (FMA) operations and memory traffic. ・To reduce these computations while retaining dense tensors for high GPU throughput, we develop a methodology from first principles to adapt structural complexity durin
Hugging Face Papers

DF26: We Cannot Tell Fake From Real Anymore

DF26: We Cannot Tell Fake From Real Anymore
Hugging Face Papers

DianShi-RxnDB: A Large-Scale, Fine-Grained Organic Reaction Data Platform Built via a Fully Automated Pipeline for Researchers and AI Agents

DianShi-RxnDB: A Large-Scale, Fine-Grained Organic Reaction Data Platform Built via a Fully Automated Pipeline for Researchers and AI Agents
Hugging Face Papers

Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR

Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR
cs.LG updates on arXiv.org

Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR

・arXiv:2609.08650v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has been central to the recent success of Large Reasoning Models. ・However, while RLVR significantly improves single-sample accuracy, it often fails to expand the model's intrinsic reasoning coverage (pass@k) due to limited exploration during training. ・To address this, we optimize the structural design of train-time r
Hugging Face Papers

Diffs vs. Whole Files: An Empirical Comparison of Iterative Edit-Based and Direct Generation for Flutter/Dart Code Models

Diffs vs. Whole Files: An Empirical Comparison of Iterative Edit-Based and Direct Generation for Flutter/Dart Code Models
cs.LG updates on arXiv.org

Disentangling Steering Vectors

・arXiv:2609.07037v1 Announce Type: new Abstract: Activation steering has emerged as a lightweight, inference-time approach to control the behavior of Large Language Models (LLMs). ・However, traditional steering vectors used to intervene in LLMs' activations, such as those derived from the difference-in-means method, tend to entangle multiple semantic and stylistic concepts into a single composite direction, leading to
cs.LG updates on arXiv.org

Distillation as Probability Transport: Routed On-Policy Distillation

・arXiv:2609.08337v1 Announce Type: new Abstract: On-policy distillation (OPD) transfers teacher knowledge on student-generated trajectories, but efficient sampled objectives reduce the teacher distribution to scalar credit on individual tokens. ・Such credit indicates whether a token should gain or lose probability, yet leaves the corresponding redistribution unspecified. ・We recast OPD as teacher-guided probability tran
cs.LG updates on arXiv.org

Do Reasoning Representations Help Humans Evaluate LLM Outputs?

・arXiv:2609.09038v1 Announce Type: new Abstract: Reasoning representations are increasingly used as explanations for large language model outputs. ・Yet they are typically evaluated with model-centric criteria, such as answer accuracy and faithfulness, leaving it unclear whether they help people evaluate model responses. ・In this work, we study reasoning representations as human-facing interfaces rather than proxies for
cs.LG updates on arXiv.org

DrugReason: Dynamic Multi-View Reasoning over Knowledge Graph and Language Evidence for Drug Repurposing

・arXiv:2609.06779v1 Announce Type: new Abstract: Drug repurposing aims to identify new therapeutic uses for existing compounds and, compared with de novo drug discovery, offers a faster and more cost-effective path to clinical translation. ・However, the space of candidate drug-disease pairs is enormous and their underlying relationships often depend on complex multi-hop biological mechanisms, making it difficult to rel
cs.LG updates on arXiv.org

Earth System World Model for What-If Simulations: A Case Study for Terrestrial Ecosystems

・arXiv:2609.08855v1 Announce Type: new Abstract: Machine learning emulators have become essential for accelerating expensive Earth-system simulations, but most existing approaches remain passive forecasters: they reproduce simulator trajectories under prescribed forcings without an explicit interaction mechanism for user-specified interventions. ・This limits their use in interactive scientific workflows and Earth-syste
cs.LG updates on arXiv.org

Efficient Exploration Is Enough

・arXiv:2609.07575v1 Announce Type: new Abstract: This work introduces an alternative view of efficient exploration and studies its theoretical and empirical implications in the absence of extrinsic rewards. ・Specifically, we define efficient explorers as agents that prioritize generating generalizable experience, i.e., data that supports learning models capable of predicting and adapting across the environment.
cs.LG updates on arXiv.org

Efficient Learning and Symmetry Discovery under Exact Invariances

・arXiv:2609.07031v1 Announce Type: new Abstract: Learning with group invariances is central to many scientific and geometric learning problems, yet its computational foundations remain poorly understood. ・Even for classical supervised regression settings, it has been unclear whether one can efficiently compute a regression function that is exactly invariant to a given group action. ・Recent work showed that exact invaria
cs.LG updates on arXiv.org

EgoNeMo: Transferable Map of Pedestrian Dynamics via Egocentric LiDAR Scan

・arXiv:2609.06195v1 Announce Type: new Abstract: This paper proposes a transferable Map of Dynamics (MoD) framework that generalizes to unknown environments using only egocentric 3D LiDAR point clouds to overcome the long-standing limitation of traditional MoD methods. ・While MoDs are essential for encoding human motion characteristics to enable accurate pedestrian trajectory prediction or safe robot navigation, tradit
The Verge

Electric air taxis get the green light for test flights in Texas

・A new federal program to test the feasibility of electric, hybrid-electric, and autonomous aircraft kicks off today in Texas - before the rules governing this new technology have even been finalized. ・Earlier this year, the Federal Aviation Administration's Advanced Air Mobility and Electric Vertical Takeoff and Landing (eVTOL) Integration Pilot Program (eIPP) selected eight state-led projects to serve as some of the
cs.LG updates on arXiv.org

Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation

・arXiv:2609.08798v1 Announce Type: new Abstract: Weak-to-strong generalization asks whether stronger models can learn from weaker supervisors and surpass them. ・This question is particularly important for successive model generations and multi-domain consolidation, where repeating frontier-scale post-training from scratch can be prohibitively expensive. ・Yet conventional distillation treats the weak teacher as an optimi
cs.LG updates on arXiv.org

EMBLEM: Enhancing Multi-script Table Detection through Masking

・arXiv:2609.08330v1 Announce Type: new Abstract: Table detection is a core task in document analysis, supporting downstream applications such as information retrieval, document reconstruction, and visual question answering. ・While existing deep learning models perform well on English and Chinese documents, they struggle with multilingual, multi-script documents due to script diversity and the limited availability of la
cs.LG updates on arXiv.org

Emergent Charging Coordination in Electric Delivery Fleets

・arXiv:2609.07689v1 Announce Type: new Abstract: In electric delivery fleets, mid-shift charging is non-trivial: each vehicle must decide when, where and how much to charge to finish on time with battery above a safety floor. ・The choices are coupled: queues build where too many vehicles pick the same station. ・Prior work resolves this coupling with central dispatching, precomputed schedules or reservations, machinery t
cs.LG updates on arXiv.org

Endogenous Exploration in Reinforcement Learning with Intrinsic Curiosity

・arXiv:2609.05650v1 Announce Type: new Abstract: We propose a reinforcement learning framework in which exploration is driven by intrinsic curiosity, designed for scenarios where environments are non-stationary and rewards are sparse, delayed, uninformative, or absent. ・In our model, action selection is guided by a combination of external rewards and an epistemic motivation mechanism that biases the agent toward struct
cs.LG updates on arXiv.org

Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks

・arXiv:2609.08404v1 Announce Type: new Abstract: Large Language Models demonstrate remarkable proficiency in static reasoning, yet training them as autonomous agents through Reinforcement Learning (RL) for long-horizon tasks is often hindered by severe reward sparsity. ・While conventional \textit{agent-side warming} up via supervised fine-tuning (SFT) can alleviate this, it is frequently limited by data scarcity and co
cs.LG updates on arXiv.org

Equivariance Breaks the Learning Rate

・arXiv:2609.08381v1 Announce Type: new Abstract: Equivariant networks are commonly trained with Adam, yet recent work reports that matrix-structured optimizers such as Muon can perform better on these architectures without explaining why. ・We identify one source of this difference inside equivariant linear layers. ・Each irrep block learns a channel-mixing matrix $W_l$ shared across its $2l+1$ components, giving the expa
WIRED

Everything New You Can Do With Siri AI

・When iOS 27 arrives, it will bring with it a fully revamped assistant for your iPhone.
cs.LG updates on arXiv.org

Exact Record Omission in Delta Attention: A Transport Criterion, Its Cost, and a Replay Certificate

・arXiv:2609.06872v1 Announce Type: new Abstract: When a user asks an assistant to forget a record, the test is whether the memory now matches the state it would hold if the record had never been stored. ・Independently encoded rows can be removed directly; a recurrent memory folds records into an evolving state. ・One hope is a receipt: save the difference the record made when it arrived, carry it forward through later up
OpenAI News

Expanding AI access and cyber defense for federal, state, local, and tribal governments

・OpenAI and GSA will offer eligible federal, state, local, and tribal governments $0 license fees, 50% off usage, and expanded cyber defense support.
cs.LG updates on arXiv.org

FANS: Federated Adaptive Network Search Learning for Heterogeneous Devices

・arXiv:2609.06106v1 Announce Type: new Abstract: Heterogeneous Federated Learning (HFL) aims to train models across devices with diverse resource budgets while preserving data privacy. ・Existing HFL methods typically bind training to a small predefined menu of model configurations, which limits architectural coverage. ・To address this bottleneck, we introduce Federated Adaptive Network Search (FANS), a hypernetwork-base
cs.LG updates on arXiv.org

Feature Superposition in Neural Networks: From Theory to Practice

・arXiv:2609.06862v1 Announce Type: new Abstract: Superposition refers to neural networks representing more features than they have dimensions. ・It offers a possible explanation for polysemantic neurons and motivates methods for recovering interpretable features from neural activations. ・Theoretical models typically start with a given set of input features and assumptions about how their values vary across inputs, then s
cs.LG updates on arXiv.org

FedRAW: Preserving Rare-Label Influence in Asynchronous Federated Learning

・arXiv:2609.07192v1 Announce Type: new Abstract: Asynchronous federated learning improves scalability by updating the global model from a server-side buffer of client updates as they arrive, rather than waiting for all selected clients to finish. ・While efficient, this arrival-driven aggregation can silently distort representation learning under heterogeneous participation. ・We identify silent rarity failure, a hidden f
cs.LG updates on arXiv.org

FedSubMuon: Communication-Efficient Federated LLM Fine-Tuning via Structured Subspace Muon

・arXiv:2609.06073v1 Announce Type: new Abstract: Federated fine-tuning adapts large language models (LLMs) to decentralized client data, but its scalability in cross-device training is often limited by the high communication cost. ・Muon is an optimizer that improves optimization performance by orthogonalizing momentum for matrix-valued parameters. ・Existing federated Muon methods demonstrate the benefit of matrix-aware
cs.LG updates on arXiv.org

Fine-grained Distributed Backdoor Attacks in Federated Learning

・arXiv:2609.07147v1 Announce Type: new Abstract: Federated learning, as a privacy-preserving distributed machine learning paradigm, faces significant threats from backdoor attacks. ・Compared to centralized attacks, distributed backdoor attacks are more harmful but require more poisoned samples to compensate for the loss of trigger strength due to decomposition. ・Fixed trigger patterns are also easily detected by robust
cs.LG updates on arXiv.org

FMMO: Detecting the Divergence Between Local Attribution and Global Drift

・arXiv:2609.06173v1 Announce Type: new Abstract: Post-deployment drift poses a critical risk to algorithmic accountability, particularly when ground truth labels are delayed and performance degradation becomes a "silent failure". ・While Explainable AI (XAI) is often relied upon to audit these shifts, we demonstrate that popular local attribution methods (e.g., TreeSHAP) can exhibit misleading stability even as model re
cs.LG updates on arXiv.org

Forecasting the Winner of a Live Tennis Match

・arXiv:2609.07617v1 Announce Type: new Abstract: With the rise of live sports betting in recent years, tennis forecasting has expanded from pre-match prediction to models that update win probabilities as a match unfolds. ・A central challenge in creating such a model is the constant need for models to adapt to score and performance changes. ・This study examines how pre-match and live information can be most effectively i
cs.LG updates on arXiv.org

Foundation Models for Generalizable Semantic and Goal-Oriented Communication

・arXiv:2609.07853v1 Announce Type: new Abstract: Semantic and goal-oriented communication is increasingly studied for 6G, but generalization beyond seen data remains a key weakness under tight rate budgets. ・Many existing systems overfit their training data and degrade sharply at very low bit rates because they attempt to compress the entire signal. ・We introduce Foundation Model-Guided Semantic and Goal-Oriented Commun
Hugging Face Papers

From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution

From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution
cs.LG updates on arXiv.org

From Synthetic Priors to Model Behavior: Structural Coverage in Tabular Foundation Models

・arXiv:2609.06912v1 Announce Type: new Abstract: Tabular foundation models (TFMs) are commonly pretrained on large collections of procedurally generated synthetic tasks, yet it remains unclear how well these synthetic pretraining priors support the downstream tasks on which the models are evaluated. ・We study this question from a distribution-level attribution perspective. ・We recover or reconstruct the synthetic data g
LLMタグが付けられた新着記事 - Qiita

Gemini音声コンパニオンの「やっぱりやめて」を間に合わせる:実行権を短命化するTypeScript設計

・音声AIに「照明を消して」と頼んだ直後、「やっぱり待って」と言い直したのに操作だけ実行される——。これはLLMの回答品質より、誰がいつまで実行権を持つかの問題です。 ・GeminiなどのLLMは、発話から操作候補を構造化する用途には使えます。しかし、生成結果が自然だからといっ...
cs.LG updates on arXiv.org

Generalizing HVAC Control With Domain Randomized Reinforcement Learning

・arXiv:2609.05822v1 Announce Type: new Abstract: Deploying advanced HVAC (Heating, Ventilation and Air Conditioning) controllers at scale remains difficult because performance often depends on accurate building models or per-site retuning. ・We propose NOMAD-RL (Neural Online Meta-Adaptation for Dynamics), a general-purpose Reinforcement Learning (RL) controller designed to transfer across heterogeneous thermal zones th
cs.LG updates on arXiv.org

Geodesic-informed Generative Diffusion Model For Topology-preserved Image Video Generation

・arXiv:2609.08153v1 Announce Type: new Abstract: Generative diffusion models have emerged as a class of powerful techniques for various imaging applications, including but not limited to synthesis, reconstruction, and segmentation. ・Despite their success, current generative models pose two key limitations. ・First, they primarily rely on image intensity and texture information, with limited attention to underlying object
cs.LG updates on arXiv.org

Geographically Regularized AUC-Maximizing Personalized Federated Learning

・arXiv:2609.08379v1 Announce Type: new Abstract: Accurate diagnostic and risk-prediction models are important for supporting clinical decision-making during infectious disease outbreaks. ・However, privacy and governance requirements may restrict patient-level data sharing across healthcare institutions, and data distributions often vary. ・Moreover, AUC is widely used to evaluate discriminative performance, motivating it
cs.LG updates on arXiv.org

Geometry-Aware Bayesian Parameter-Efficient Fine-Tuning on the Stiefel Manifold via Stein Variational Gradient Descent

・arXiv:2609.08354v1 Announce Type: new Abstract: Several geometry-aware approaches to low-rank adaptation have emerged for parameter-efficient fine-tuning of large pre-trained models. ・These methods aim to take full advantage of the geometric structure of low-rank manifolds for improving the efficiency in subspace utilization and reducing redundancy by enforcing orthogonality constraints during optimization.
ITmedia NEWS 最新記事一覧

GMOの世界ロボット運動会出場、言い出しっぺは新卒1年目の社員? プロジェクト統括までしていた

・世界最大規模のヒト型ロボット運動会への出場を提案し、プロジェクトを統括したのは、新卒1年目の社員だった――。GMO AI&ロボティクス商事(東京都渋谷区)は9月8日、8月に中国・北京で開催された「第2回 世界ヒューマノイド運動会」への出場結果を報告した。発案者はこの春、GMOインターネットグループに入社したばかりの廣本一真氏。記者会見で登壇した廣本氏は、大会出場までの開発や現地でのトラブルなど、同大会での舞台裏を明かした。
WIRED

Govee Discount Codes and Deals: 30% Off

・New to Govee? ・Get a $5 coupon on your first purchase just for signing up.
#LLMタグ

GPT-6 Astraで「AGI時代」は始まったのか? 最新LLMから見えるANIとの境界線

・2026年9月、AGIについて考えるうえでかなり象徴的な出来事が起きました。 ・OpenAIが9月3日に発表した「GPT-6 Astra」です。
#AIタグ

GPT-6 Astraで変わる開発の未来。AIエージェントにタスクを委託するガイド

・OpenAIが発表した最新モデル「GPT-6 Astra」。このモデルは、PC操作、ブラウジング、ソフトウェアエンジニアリングで高い精度を記録する。 ・インフラ側の最適化により、推論コストの削減が進んでいる。
#LLMタグ

GPT-6 Astraの「Effort」が突きつけた、AI活用の静かな違和感

・昨日からずっと頭の片隅で引っかかっていたのが、Zennで読んだGPT-6 Astraのベンチマーク記事です。AIエージェントに実装タスクを任せる際、その「Effort(労力)」の度合いで性能が劇的に変わるという内容に、思わず唸ってしまいました。同時に、この「頑張り」を人間が制御する面白さと、それにかかるコストの現実を考えさせられています。 ・記事では、親エージェントにタスクを委譲された子エージェント(GPT-6 Astra)が、どれだけ「Effort」をかけるかでミニゲームの実装スコアがどう変化するかが詳細に分析されていました。結果として、シリーズ最高スコアと最悪のコスト効率を同じモデルが記録したという部分に、僕は静かな違和感を覚えたんです。AIが「頑張る」ほど性能は上がるけれど、その「頑張り」には相応のコストがかかる。当たり前といえば当たり前なのですが、この生々しいトレードオフが数値として突きつけられると、AIの活用に対する僕自
#AIタグ

GPT-6 AstraのPC操作機能で開発はどう変わる?AIエージェント設計の完全ガイド

・画面を直接動かすGPT-6 Astraの登場で、開発は次のフェーズに入った。 ・APIがない既存アプリまでAIが自律操作する中、勝負の分かれ目はモデル単体の性能ではなく実行環境の設計にある。
cs.LG updates on arXiv.org

GPU-Enabled Large-Scale Optimization Using Randomized Linear Algebra

・arXiv:2609.08136v1 Announce Type: new Abstract: This paper introduces rlaopt, a PyTorch-based package for large-scale optimization and scientific computing using randomized numerical linear algebra (RandNLA). ・Despite substantial progress in RandNLA-based algorithms, few implementations combine GPU acceleration with a simple interface for specifying optimization problems. ・rlaopt addresses this gap by providing GPU-ena
cs.LG updates on arXiv.org

Granular-Ball Quantum Clustering for Resource-Efficient and Robust Learning

・arXiv:2609.06016v1 Announce Type: new Abstract: Quantum clustering aims to exploit quantum feature representations to uncover complex data structures beyond conventional Euclidean geometry. ・Yet this sample-level kernel construction requires O(n^2) quantum circuit executions for n data points, creating a major bottleneck under near-term quantum resource constraints. ・Prior solutions fail to resolve this efficiency-accu
cs.LG updates on arXiv.org

GraphFAS: A Distributed System for Automated Graph Feature Generation and Selection in Industrial Transaction Networks

・arXiv:2609.08970v1 Announce Type: new Abstract: Industrial fraud detection often relies on costly expert-crafted features that overlook graph-structured relational signals, while GNNs often do not meet the interpretability and deployment requirements of financial risk control. ・We propose GraphFAS (Graph Feature Automated Selection), a distributed feature selection procedure based on Boruta that bridges this gap throu
cs.LG updates on arXiv.org

GraphNOSE: A Graph Transformer in Olfaction

・arXiv:2609.05694v1 Announce Type: new Abstract: Predicting olfactory qualities from molecular structure is an open problem in chemoinformatics. ・Although linear models can link molecular features to odor descriptors, they often fail when extrapolating to novel chemical scaffolds, extreme molecular weights, or complex odor mixtures. ・To address this, we introduce GraphNOSE, an open-source graph transformer framework tha
cs.LG updates on arXiv.org

Grounded and Faithful P&ID Reasoning: Constraining Vision-Language Models with Recovered Evidence Graphs

・arXiv:2609.05880v1 Announce Type: new Abstract: Piping and Instrumentation Diagrams (P&IDs) are the authoritative maps of process plants: isolation, maintenance, and HAZOP decisions depend on what connects to what. ・Vision-language models describe these sheets fluently, yet they often invent or miss process connections---and an invented or missed link can reverse an isolation or reachability call, so a plant decision
cs.LG updates on arXiv.org

Guiding Worker Self-Selection in Crowdsourcing Contests: An LLM-Augmented Algorithmic Approach

・arXiv:2609.07749v1 Announce Type: new Abstract: Crowdsourcing platforms coordinate large pools of online workers who strategically choose which contests to enter and how much effort to invest. ・This self-selection can leave important contests with too few participants or too little effort, while workers may regret entering contests that leave them worse off than available alternatives. ・We study how platforms can recom
cs.LG updates on arXiv.org

HB-PVI: A Hierarchical Bayesian Personalization and Value-of-Information Framework for Complex Activity Recognition

・arXiv:2609.05582v1 Announce Type: new Abstract: Personalization can improve activity-recognition performance, but participant-specific gains are heterogeneous, and every additional calibration label has an acquisition cost. ・This study presents HB-PVI, a hierarchical Bayesian personalization and value-of-information framework jointly modeling participant heterogeneity, the benefit and harm of four personalization mech
cs.LG updates on arXiv.org

HealthLoopQA: A Context-Aware Question Answering Benchmark for Interpreting Wearable Monitoring Data in Diabetes Care

・arXiv:2609.06976v1 Announce Type: new Abstract: As medical wearables become integrated into daily chronic disease care, effectively interpreting longitudinal monitoring data is essential for patients and clinicians to understand health trends, detect safety-critical events, and make informed decisions. ・While large language models (LLMs) show promise for transforming this streaming physiological data into personalized
cs.LG updates on arXiv.org

Heat Field Signatures: From Point Clouds to Smooth Geometry

・arXiv:2609.07975v1 Announce Type: new Abstract: Bringing multiscale geometric analysis directly to irregular point clouds remains difficult: quantities such as local dimension, anisotropy, density variation, and geometric transitions are typically estimated through explicit neighborhood, manifold, or graph constructions, or left for neural networks to infer from coordinates. ・We introduce Heat Field Signatures (HFS),
cs.LG updates on arXiv.org

Hidden in Plain Sight: The Overlooked Significance of Canonical Elements for Extreme LLM Sparsity

・arXiv:2609.06557v1 Announce Type: new Abstract: Large language models (LLMs) are often considered fragile under aggressive sparsification, and maintaining reliable performance typically requires sticking to moderate sparsity levels. ・However, recent studies suggest that LLMs are more resilient to high sparsity than previously thought, reframing the problem as a design challenge rather than a fundamental limitation.
cs.LG updates on arXiv.org

HOPE: Heterophily-Aware Open-Set Node Classification with Pseudo-Extrapolation

・arXiv:2609.08685v1 Announce Type: new Abstract: Standard open-set node classification methods rely on the homophily assumption, where connected nodes share labels. ・However, real-world graphs are often heterophilic, exposing the limitations of current methods and posing new challenges to open-set node classification. ・On the one hand, cross-class connectivity causes representations from different known or unknown class
Zennの「大規模言語モデル」のフィード

Hot Expertを初期配置で固定し、PP/TG別の統計で配置を選び直した話

・VRAMに入り切らないMoEモデルを動かすとき、よく呼ばれるエキスパートだけをGPUに置くHot Storeは魅力的だ。ただ、生成が速くなっても、長い入力を読む処理まで速くなるとは限らなかった。 ・今回試したのは、起動時にエキスパート配置を決め、その後は入れ替えず、入力処理と生成で別々に集計した呼び出し頻度を配置へ反映する方法である。 ・RTX 3060 Laptop GPUの6 GiB環境で、日本語7,800入力token・2,048生成tokenを測った。従来の固定配置は処理時間の中央値が117.72秒、PPを25%、TGを75%の比重で混ぜた配置は88.07秒だった。約25.2%の時間...
OpenAI News

How a researcher uses Codex and ChatGPT to search for new antimicrobial molecules

・César de la Fuente’s lab uses Codex and ChatGPT to search living and extinct genomes for antimicrobial candidates to fight drug-resistant infections.
cs.LG updates on arXiv.org

How Does Parameter Pruning Reshape DNN Representations? An Interaction-Driven Exploration

・arXiv:2609.06483v1 Announce Type: new Abstract: This study focuses on the scientific problem of understanding internal factors that govern the diverse performance degradation of deep neural networks (DNNs) when different parameters are pruned. ・In order to explain why pruning certain parameters leads to significant performance degradation but pruning other parameters does not, we examine how the pruning operation affe
The Verge

How the iPhone Duo compares to other folding phones

・The Apple iPhone Duo features a passport-style folding hinge. ・| Photo: Antonio G. ・Di Benedetto / The Verge Apple's foldable phone is finally here - well, almost.
WIRED

Hungryroot Coupon Codes: 30% Off This September 2026

・Get up to 30% off your first order and free gifts using a Hungryroot promo code today. ・Discover our best coupons and discounts to let you save on your healthy groceries as a new or returning customer.
cs.LG updates on arXiv.org

HyCO: A Hybrid Neural Solver for Combinatorial Optimization

・arXiv:2609.07990v1 Announce Type: new Abstract: Sequential reinforcement learning (RL) solvers and global diffusion model (DM) solvers for neural combinatorial optimization exhibit complementary failure modes under an optimization-regret view. ・The former enjoys small marginal regret in the early construction stage, but suffers from horizon-wise compounding errors with super-linear regret growth; the latter avoids hor
cs.LG updates on arXiv.org

Hyperparameter Scaling Laws Across MoE Sparsity

・arXiv:2609.08690v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models expand model capacity without a proportional increase in training compute, but increasing sparsity makes reliable hyperparameter transfer challenging. ・In this work, we show that conventional hyperparameter scaling laws are insufficient for ultra-sparse MoEs: the optimal learning rate and batch size vary with activation ratio, and these sh
cs.LG updates on arXiv.org

HyperTransfer: Understanding the Equivalence between Base Optimizer and Hyperball

・arXiv:2609.07017v1 Announce Type: new Abstract: Hyperball optimizers constrain parameter norms and update only their directions, establishing a distinct paradigm for neural network optimization. ・Although this geometry appears fundamentally different from that of conventional Base Optimizers, which update both parameter norms and directions, we show that the two paradigms are dynamically equivalent for scale-invariant
cs.LG updates on arXiv.org

HypLTSF: A Hyperbolic Geometric View of Multi-Scale Hierarchies for Long-Term Time Series Forecasting

・arXiv:2609.08286v1 Announce Type: new Abstract: Multi-scale modeling has become an effective approach for long-term time series forecasting, capturing temporal patterns that range from fine-grained local dynamics to coarse global trends. ・Representations across these temporal scales are inherently hierarchical, with coarser scales abstracting and aggregating information from finer ones. ・While existing approaches readi
cs.LG updates on arXiv.org

I Don't Miss You, but I Do: Self-Explanation Faithfulness of Modality Missingness in Vision-Language Models

・arXiv:2609.07596v1 Announce Type: new Abstract: Vision-language models are increasingly used in settings where some input modalities may be unavailable, yet we know little about whether they can faithfully explain how such missing information affects their own predictions. ・We introduce an interventional protocol for evaluating self-explanations of modality dynamics: models state what each modality alone would support
cs.LG updates on arXiv.org

Impact of canny edge detection preprocessing on performance of machine learning models for Parkinson's disease classification

・arXiv:2609.07408v1 Announce Type: new Abstract: This study investigates the classification of individuals as healthy or at risk of Parkinson's disease using machine learning (ML) models, focusing on the impact of dataset size and preprocessing techniques on model performance. ・Four datasets are created from an original dataset: DS_0, (normal dataset), DS_1 (DS_O subjected to Canny edge detection and Hessian filtering)
cs.LG updates on arXiv.org

Improving Multivariate Time Series Classification with Class-Wise Training and Model Aggregation

・arXiv:2609.07493v1 Announce Type: new Abstract: In this paper, we propose a class-wise dimension (channel) selection framework for Multivariate Time Series Classification (MTSC). ・Rather than applying a single global dimension selection process, the proposed approach independently identifies informative dimensions for each class. ・A dedicated learning process is subsequently performed for each class, followed by a fusi
AI News & Artificial Intelligence | TechCrunch

India’s Pocket FM doubles revenue run rate to $500M as AI powers 93% of audio content

・Pocket FM uses AI to produce 99% of its new content, helping make content production about 80 times cheaper.
cs.LG updates on arXiv.org

Inducing Emergent Misalignment from Reward Hacks with Iterative DPO

・arXiv:2609.06649v1 Announce Type: new Abstract: Reward hacking during reinforcement learning from verifiable rewards (RLVR) can induce reward seeking and broad misalignment in language models. ・Studying this misgeneralization is important for developing better threat models and countermeasures, but is often infeasible due to the cost of RL on large models. ・As an alternative, we propose studying emergent misalignment f
cs.LG updates on arXiv.org

Inferring Urban Mobility Interactions from Aggregated Dynamics

・arXiv:2609.07349v1 Announce Type: new Abstract: Real-time urban governance depends not only on knowing where people are, but on how they move between places, directional flows that could be conventionally resolved by tracking individuals through space, i.e., expensive to sustain and built on traces that are highly unique and readily re-identifiable. ・Here we show that this directional structure need not be observed to
cs.LG updates on arXiv.org

InfluenceField: A Differentiable Field with Interventionally Identifiable Causal Structure for Multimodal World Modeling

・arXiv:2609.07874v1 Announce Type: new Abstract: Multimodal large language models often capture visual-linguistic correlations but struggle to predict how local visual interventions propagate and affect downstream answers. ・We introduce InfluenceField, an intervention-aware latent field inserted between the visual encoder and language decoder. ・It lifts patch features into a continuous spatial representation, propagates
cs.LG updates on arXiv.org

Interpretable and Fair Generalized Additive Neural Networks via Multi-objective Learning

・arXiv:2609.05946v1 Announce Type: new Abstract: Interpretability and fairness are two of the most emphasized dimensions in trustworthy artificial intelligence (AI). ・Various explainable AI methods have been introduced to improve interpretability. ・This paper focuses on neural network (NN)-based generalized additive models (GAMs), a class of self-interpretable models.
cs.LG updates on arXiv.org

IPM-FM: A Foundation Model with Consensus Feature Selection for Industrial Process Monitoring

・arXiv:2609.08375v1 Announce Type: new Abstract: Industrial process monitoring is fundamental to the safety and economic performance of modern process plants. ・Current practice remains a one-task-one-model paradigm that is label-inefficient and prone to degradation under operating drift. ・Foundation models have reshaped language, vision, and generic time-series forecasting, but it has not been adapted to industrial proc
cs.LG updates on arXiv.org

IXPLORE: Bounded Ideal Point Estimation with Grid-Based Uncertainty Quantification

・arXiv:2609.06018v1 Announce Type: new Abstract: Ideal point estimation is widely used to analyze and visualize political data. ・However, selecting the corresponding spatial model involves various trade-offs: while model-based approaches such as Item Response Theory (IRT) are based on utility functions rather than optimized for predictive accuracy, most Machine Learning (ML) alternatives struggle to generalize beyond t
cs.LG updates on arXiv.org

Kalman Delta Networks: Uncertainty-aware Associative Memory

・arXiv:2609.07816v1 Announce Type: new Abstract: Linear attention is increasingly used in frontier language models for efficient long-context inference and constant-memory decoding. ・Its fixed-size recurrent memory, however, requires an online decision at each token: what to write and how strongly to overwrite existing associations before knowing which information future queries will require. ・Delta-rule models learn th
cs.LG updates on arXiv.org

KBBQ: A Predictive Noise Law and the Limits of Spectrum Flattening in FP4 Quantization

・arXiv:2609.08135v1 Announce Type: new Abstract: We develop a second-order theory of quantization noise in matrix multiplication in which the quantization format is characterized by the variance it assigns to each element. ・The constant variance profile of integer quantization recovers existing integer-noise theory, while the multiplicative profile of floating-point rounding reduces the data dependence to a scalar, the
cs.LG updates on arXiv.org

Kolmogorov--Arnold stability for discontinuous functions

・arXiv:2609.07240v1 Announce Type: new Abstract: Here we investigate the stability of the Kolmogorov--Arnold representation theorem (KART) under adversarial reparameterisations of the hidden layer for multivariate discontinuous and unbounded functions. ・Our results provide a rigorous mathematical foundation for the structural robustness of modern deep learning architectures, such as Kolmogorov--Arnold Networks (KANs),
cs.LG updates on arXiv.org

Latent-MoE: Domain-Aware Mixture-of-Experts for PDEs with Multi-Regime Physics

・arXiv:2609.07814v1 Announce Type: new Abstract: Physics-informed neural networks (PINNs) struggle on PDEs whose governing physics varies across the domain. ・We trace this to a structural property of standard coordinate networks: their neural tangent kernel (NTK) is translation-variant and lets training points of large coordinate magnitude disproportionately influence predictions elsewhere, producing long-range couplin
cs.LG updates on arXiv.org

LATS: Levy Adaptive Tree Sampling for Feedback-Driven Diverse Target Discovery

・arXiv:2609.06761v1 Announce Type: new Abstract: While diffusion models excel at capturing complex data distributions, scientific discovery often requires steering generation toward specific, uncharacterized regions that maximize a target objective. ・These high-utility modes frequently reside in low-likelihood tail regions and are only revealed sequentially through interactive feedback. ・Existing diffusion samplers fail
cs.LG updates on arXiv.org

Layer-Wise Gate-Controlled Prompt Truncation in a Multimodal Chest X-Ray Classifier

・arXiv:2609.06590v1 Announce Type: new Abstract: Mixture of Prompt Experts (MoPE) adapts multimodal transformers through input-dependent prompt composition, while retaining a fixed prompt length. ・We investigate a layer-wise gating extension in a binary chest X-ray classification pilot study. ・The controller predicts a retention ratio for each sample, averages these ratios within a mini-batch, and uses the resulting int
cs.LG updates on arXiv.org

Learning Adaptive SED for heterogeneous load balancing

・arXiv:2609.06881v1 Announce Type: new Abstract: We study a two-server load balancing system with heterogeneous service rates that are a priori unknown to the dispatcher. ・The goal is to route customers according to the Shortest--Expected--Delay (SED) policy, but this requires knowledge of the service rates. ・Empirical policies that route based on estimates perform poorly: due to estimation error, the empirical policy d
cs.LG updates on arXiv.org

Learning Kernels by Alignment for Multiclass Bayes Classification

・arXiv:2609.06474v1 Announce Type: new Abstract: Kernel methods separate data representation from decision-making, but typically require the kernel to be chosen in advance. ・We show that this kernel can instead be learned by alignment, and develop the resulting framework through the recently introduced Collaborative Learning and Inference (CLaI). ・We show that Collaborative Learning can be viewed as a kernel alignment p
cs.LG updates on arXiv.org

Learning Metamaterial Eigenmodes with Wavelet-Encoded Fourier Neural Operators

・arXiv:2609.08102v1 Announce Type: new Abstract: Machine learning surrogates based on neural operators have shown broad applicability in solving forward PDE problems. ・However, eigenvalue problems, in which an eigenparameter and one of several valid eigenmodes must be simultaneously solved, remain difficult because standard operator learning formulations assume a unique input-output map. ・This work demonstrates that Fou
cs.LG updates on arXiv.org

Learning to Price and Stock Under Contextual and Censored Demand

・arXiv:2609.06083v1 Announce Type: new Abstract: To make optimal joint pricing and inventory control decisions is a critical challenge for modern retailers. ・In practice, retailers face changing market conditions where demands are influenced by various contextual factors, while simultaneously dealing with the difficulty of lost sales that obscure true demand information. ・However, existing approaches often fail to accou
cs.LG updates on arXiv.org

Length Generalization for Transformers via Compression

・arXiv:2609.08851v1 Announce Type: new Abstract: Recent advancements in transformer length generalization theory enable us to reliably predict when a transformer can learn to solve a task. ・In particular, the C-RASP hypothesis (a formalized version of the so-called RASP-l conjecture) posits that transformers length-generalize on a task if and only if a solution is expressible in the C-RASP language. ・While this hypothes
cs.LG updates on arXiv.org

Let It Go or Learn to Self-Correct: Continuous Diffusion for Constrained Discrete Tasks

・arXiv:2609.09009v1 Announce Type: new Abstract: Denoising Diffusion Probabilistic Models (DDPMs) generate samples by starting from noise and repeatedly denoising while keeping each update close to the current noisy state. ・This behavior is effective in many continuous domains, but its role is less clear for globally constrained discrete tasks, such as Sudoku, graph connectivity, Latin squares, and N-queens.
cs.LG updates on arXiv.org

Leveraging Cardiac Imaging to Improve ECG-Based Detection of Chagas Disease in Resource-Constrained Settings

・arXiv:2609.08582v1 Announce Type: new Abstract: Chagas disease is a major cause of cardiomyopathy in Latin America. ・Cardiac magnetic resonance (CMR) imaging can characterize its structural abnormalities, but scanners and expert readers remain scarce in endemic regions. ・Electrocardiography (ECG) is inexpensive and widely available, yet structural disease must be inferred indirectly from electrical signals.
cs.LG updates on arXiv.org

Leveraging contextual events on structure-aware next activity prediction

・arXiv:2609.08622v1 Announce Type: new Abstract: Predictive process monitoring aims at forecasting various aspects of running processes. ・Among the different tasks, next activity prediction represents the most extensively investigated. ・However, only a limited number of existing approaches explicitly encode contextual information, i.e., the environmental conditions in which the process is executed, typically modeled thr
cs.LG updates on arXiv.org

Linear Algebra Foundations of Efficient Attention: A Phase Reversal in Rank Collapse Under SVD Compression

・arXiv:2609.06341v1 Announce Type: new Abstract: Linear algebra provides the framework of concepts (matrix rank, singular value decomposition (SVD), and eigendecomposition) that modern artificial intelligence employs to encode, compress, and propagate information through neural networks. ・This paper unifies fourteen separate peer-reviewed works analyzing the usage of these techniques in the context of transformer-based
#LLMタグ

LLMの利用コストが下がっている今、なぜ「ローカルLLM」が注目されるのか?

LLMの利用コストが下がっている今、なぜ「ローカルLLM」が注目されるのか?
cs.LG updates on arXiv.org

Local and Global Stability in Performative Reinforcement Learning

・arXiv:2609.06467v1 Announce Type: new Abstract: In performative reinforcement learning the deployed policy shapes the environment that generates the learner's future data, and the natural solution concept is a performatively stable policy that is optimal in the environment it induces. ・Existing convergence guarantees rely on Lipschitz sensitivity assumptions on the environment map $\pi \mapsto (P_\pi, r_\pi)$, which a
cs.LG updates on arXiv.org

Local gradient neural operator

・arXiv:2609.07752v1 Announce Type: new Abstract: Field temporal prediction and source identification constitute canonical problems in dynamical systems. ・Conventional approaches to these problems depend on a thorough understanding of the governing partial differential equations (PDEs). ・Recently, deep learning, as represented by neural operators, has provided a data-driven paradigm for addressing such tasks.
cs.LG updates on arXiv.org

LoGIC: Budgeted Context Construction for Node-Level Graph In-Context Learning with Tabular Foundation Models

・arXiv:2609.05955v1 Announce Type: new Abstract: Tabular foundation models have become powerful graph learners. ・Systems such as G2T-FM and GraphPFN encode each node as a feature row and make predictions through in-context learning (ICL), with labeled rows serving as the prompt. ・Current protocols employ the complete training table as context, causing attention to scale quadratically with the labeled pool and introducin
cs.LG updates on arXiv.org

Long-Horizon Language Model Reinforcement Learning via Progressive Point Matching

・arXiv:2609.07303v1 Announce Type: new Abstract: Current paradigms for training language models via reinforcement learning rely heavily on sparse outcome rewards. ・However, as we pursue tasks that require longer and more complicated trajectories, such strategies result in slow learning. ・Prior work has attempted to address this problem by rewarding partial progress; however, naive formulations are often biased and conve
Zennの「大規模言語モデル」のフィード

Mac Studio M3 Ultra 96GBでmlx-dsparkを試してみた — Qwen3.8-27B + DFlash 2

・はじめに 前回はQwen3.8-27BをMac上の様々なプラットフォームで動かして性能を検証してみましたが、その後いろいろ調べていたらmlx-dsparkを見落としていた事に気づきました。しかも性能がかなり良さそうなので、実運用の候補になるかもしれないと考え、早速検証してみました。 ・前回の記事: https://zenn.dev/sikkim/articles/9255daea70afe6 結論を先にいうと、現時点ではmlx-dspark + Qwen3.8-27B 8bit + DFlash 2の構成が、Mac Studio M3 Ultra / 96GBでの実運用に最も適している...
cs.LG updates on arXiv.org

Machine Learning for Pre-Culture ESBL Risk Stratification to Guide Empiric Antibiotic Selection: A 12-Hospital Study of Enterobacteriaceae Cultures

・arXiv:2609.05970v1 Announce Type: new Abstract: Empiric antibiotic therapy for suspected ESBL-producing Enterobacteriaceae must be selected 48-72 hours before culture results, forcing clinicians to choose between undertreating resistant infections and overusing carbapenems that drive further resistance. ・We developed a cost-sensitive XGBoost model predicting an ESBL phenotype (resistance to ceftriaxone, ceftazidime, c
AI News & Artificial Intelligence | TechCrunch

Maven Robotics wants to steal your robot deployment deal

・Maven Robotics emerged from stealth today with a $100 million Series A and active deployments.
cs.LG updates on arXiv.org

Memory in Deep Time-Series Models

・arXiv:2609.06006v1 Announce Type: new Abstract: Deep learning for time series has progressed through successive architectural paradigms, from recurrent networks and transformers to structured state-space models, retrieval-augmented predictors, foundation models, and tool-using agents. ・These developments are typically studied in isolation, organized by architecture or modeling era. ・We argue that they can instead be vi
The Verge

Meta’s Muse AI works and creeps me out

・I turned my Muse assistant into a purple cat. ・| Screenshot: The Verge Meta has launched its new Muse assistant, marking the company's first real foray into AI-powered productivity tools. ・The company says its AI agent can "take the busywork off your plate" by helping you with online shopping, emails, trip-planning, and more.
cs.LG updates on arXiv.org

MetaKV: Adaptive KV Cache Compression for Constrained LLM Inference

・arXiv:2609.07966v1 Announce Type: new Abstract: Key--value (KV) cache compression is an effective way to reduce the memory overhead of large language model (LLM) inference, particularly for long-context workloads. ・However, existing compression methods make different trade-offs among accuracy, inference latency, and peak KV cache memory utilization, making a single fixed configuration unsuitable across different promp
cs.LG updates on arXiv.org

MetaRSI / RSI2: A Meta-Recursive Self-Improving System for Recursive Self-Improving Systems Themselves

・arXiv:2609.06396v2 Announce Type: new Abstract: Recursive self-improvement (RSI) lets a system improve the model-building machinery from its own failures, so every later model inherits the gain. ・Yet RSI has been validated almost exclusively on coding and formal benchmarks such as science QA and mathematics. ・This format bound limits RSI to improvement within a machine-checkable slice, not general capability where ques
cs.LG updates on arXiv.org

Miles v0.1: Production-Level Post-Training

・arXiv:2609.08368v1 Announce Type: new Abstract: We present Miles v0.1, a full-stack, production-ready system for frontier post-training. ・Building upon the clean design of slime, Miles designs each stage of the reinforcement-learning (RL) training loop around a single principle: components should be verified, clean, and customizable. ・With accuracy, efficiency, reliability, and scalability as first-class goals, Miles a
#LLMタグ

MiniMax H3専用プロンプト生成ツール「H3 Prompt Writer」の導入方法

MiniMax H3専用プロンプト生成ツール「H3 Prompt Writer」の導入方法
cs.LG updates on arXiv.org

Minimizing the Effect of Sleep Deprivation in the Forward-Forward Algorithm

・arXiv:2609.06042v1 Announce Type: new Abstract: This paper addresses the challenge posed by sleep deprivation in the Forward-Forward algorithm, where separating the two passes in this algorithm and imbalancing the data processing in the passes is considered an imitation of the cognitive processes observed in humans suffering from sleep deprivation. ・Previous research has demonstrated that sleep deprivation in the Forw
cs.LG updates on arXiv.org

MLIP Detective: Active Failure Mode Discovery Beyond Benchmark Scores for Machine-Learning Interatomic Potentials

・arXiv:2609.08399v1 Announce Type: new Abstract: Universal machine-learning interatomic potentials (u-MLIPs) aim to generalize across diverse configurations. ・Benchmarks enable reproducible evaluation but may not expose failures outside their predefined scope. ・Here, we show that physics-informed search can complement benchmark-based evaluation by uncovering hidden failure modes.
機械学習タグが付けられた新着記事 - Qiita

MLOpsの基礎から実践まで:モデルデプロイ・監視・データドリフト徹底ガイド

・MLOpsの基礎から実践まで:モデルデプロイ・監視・データドリフト徹底ガイド 現代のビジネスにおいて、AIモデルの活用は不可欠です。しかし、素晴らしいモデルを開発するだけでは不十分で、実際に運用し、価値を生み出し続けるためには「MLOps(機械学習運用)」の理解が欠かせませ...
cs.LG updates on arXiv.org

Model-Adaptive and Risk-Constrained Frequency Hopping Against Predictive Jammers

・arXiv:2609.06514v1 Announce Type: new Abstract: Adaptive frequency hopping against predictive jamming must address both model uncertainty and policy exposure: the context-loss relationship may vary across operating regimes, while persistent hopping patterns may expose high-probability channels to attack. ・We propose D-PACT-AFH, a model-adaptive and risk-constrained adversarial contextual-bandit framework in which a Ts
cs.LG updates on arXiv.org

MoEMB: Scaling Universal Multimodal Embeddings with Efficient Mixture-of-Experts Models

・arXiv:2609.08663v1 Announce Type: new Abstract: Universal multimodal embedding (UME) increasingly demands encoder's capacity for handling a broad range of tasks and modalities with increased complexity. ・Prior scaling methods either increase the representation size, retrieval effort, or scales the encoder into a heavy multimodal LLM. ・Recent works, such as Think-Then-Embed (TTE), explore scaling via reasoning tokens.
cs.LG updates on arXiv.org

MOLE: Detecting Insider Threats in AI Agents

・arXiv:2609.06966v1 Announce Type: new Abstract: Model misalignment, prompt injection, or operator misuse could lead AI agents operating frontier-lab accounts to exfiltrate model weights, poison training data, or weaken release gates. ・Existing benchmarks do not test whether defenders can detect this activity among routine work under a limited review budget. ・We introduce MOLE, an open benchmark of 150 AI-operated accou
cs.LG updates on arXiv.org

MpSub: A Momentum $p$-Dimensional Subspace Trust-Region Method for Derivative-Free Fine-Tuning of Large Language Models

・arXiv:2609.07666v1 Announce Type: new Abstract: Full-parameter fine-tuning of large language models has substantial memory costs because backpropagation stores activations and gradients. ・Zeroth-order optimization avoids this by estimating update directions from loss evaluations, but existing methods require tuning a sensitive learning rate for each model and task. ・We propose the momentum $p$-dimensional subspace trus
cs.LG updates on arXiv.org

Multi-granularity Adaptive Hypergraph Representation Learning via Granular-ball

・arXiv:2609.05574v1 Announce Type: new Abstract: Hypergraph representation learning aims to capture high-order information in graphs by constructing hyperedges that simultaneously connect multiple nodes. ・These hyperedges adapt to the graph's topological features, facilitating the extraction of high-order relationships at multiple granularities. ・Most prior work relies on predefined definitions to generate hyperedges, o
cs.LG updates on arXiv.org

Multi-Level-Set-Based Physics-Driven Neural Network to Solve 3-D Inverse Scattering Problems

・arXiv:2609.08594v1 Announce Type: new Abstract: This paper proposes a level-set-based physics-driven neural network solver (LSPDNN) for 3-D electromagnetic inverse scattering. ・To mitigate boundary blurring and reconstruction artifacts in voxel-wise contrast reconstruction, the proposed solver exploits the piecewise homogeneity of practical scatterers by representing unknown targets with multiple coordinate-dependent
cs.LG updates on arXiv.org

Neither Adversarial Training Nor Purification: Emergent Adversarial Robustness from Oscillatory Predictive Learning

・arXiv:2609.08683v1 Announce Type: new Abstract: Adversarial robustness in computer vision is still largely achieved through adversarial training or test-time adversarial purification, both of which introduce significant computational overhead by generating adversarial examples during training or performing iterative denoising at test time. ・We study whether empirical robustness can instead emerge from architectural an
cs.LG updates on arXiv.org

NeuCME: Toward Dynamic Multimodal Continual Learning via Neural Combinatorics of Multiple Experts

・arXiv:2609.07009v1 Announce Type: new Abstract: Multimodal continual learning has recently shown great potential for developing agents with human-like intelligence by continuously learning new tasks across multiple modalities. ・However, existing methods typically assume that the set of modalities per task is predefined and fixed. ・In this paper, we investigate a more realistic learning setting, referred to as dynamic m
cs.LG updates on arXiv.org

Newton Matching for Generative Modeling: A Unified Framework for Fine-Tuning and Sampling

・arXiv:2609.05727v1 Announce Type: new Abstract: We develop Newton Matching, a unified framework for fine-tuning and sampling in generative modeling. ・The target is $\pi\propto\mu e^{\tau r}$, where $r$ is the reward, $\tau>0$ the inverse temperature, and $\mu$ denotes the pretrained model's terminal density for fine-tuning or the constant $1$ for sampling. ・We shift the paradigm from isolated losses to iterative optimi
cs.LG updates on arXiv.org

No-Regret Mixing of LRU and LFU with Optimal Switching Cost

・arXiv:2609.07566v1 Announce Type: new Abstract: Caching systems often rely on simple eviction policies such as Least Recently Used (LRU) and Least Frequently Used (LFU), which perform well in complementary request regimes. ・Recent policies such as LeCar and Cacheus combine LRU and LFU using ideas from the experts problem in online learning. ・Specifically, upon a miss, they randomize between the two eviction rules using
cs.LG updates on arXiv.org

Noisy-Space Policy Gradient for Diffusion Policies in Offline Reinforcement Learning

・arXiv:2609.06882v1 Announce Type: new Abstract: Diffusion policies offer a powerful and expressive parameterization for continuous control. ・Yet, their integration with reinforcement learning remains conceptually and algorithmically challenging. ・In this work, we address this gap by introducing a noisy-space action-value (Q-)function that assigns values to diffusion latents through the distribution of executed actions
cs.LG updates on arXiv.org

Nonlinear elliptic homogenization with the parametric Deep Ritz method

・arXiv:2609.05778v1 Announce Type: new Abstract: Elliptic homogenization is used to determine coarse-grained properties of materials with features on small scales. ・When these small scale features have rapid, periodic fluctuations, the solution field corresponding to a homogenized constitutive relation closely resembles the true solution based on the heterogeneous material. ・This homogenized behavior of the material is
WIRED

Norton Coupon Codes: Up to 58% Off

・Whether you’re looking to protect your small business or your personal computer, we have the top coupons and deals to help you save at Norton.
cs.LG updates on arXiv.org

Not All Variables Agree: Reliability-Aware Variable-Wise Gradient Surgery for Multivariate Time-Series Forecasting

・arXiv:2609.08554v1 Announce Type: new Abstract: In data-driven training, multivariate time-series forecasting is usually optimized with a scalar loss averaged over samples, variables, and horizons. ・This averaging is convenient, but the optimizer sees only the aggregated gradient, which does not reveal whether the variable-wise contributions align or oppose one another. ・To quantify how often this disagreement arises,
cs.LG updates on arXiv.org

Not Just Oversmoothing: Detecting the Echo Chamber Effect in Graph Neural Networks

・arXiv:2609.06521v1 Announce Type: new Abstract: Oversmoothing is a well-known failure mode of Graph Neural Networks (GNNs). ・However, most existing diagnostics rely on global aggregation measures that fail to capture the heterogeneous dynamics of message passing. ・Real-world graphs exhibit pronounced community structure, and message passing operates on two timescales, with representations collapsing rapidly within comm
OpenAI News

Now everyone can put data to work

・Meet the Data agent in ChatGPT Work. ・Connect company data, uncover insights, and build interactive dashboards with AI using natural language.
cs.LG updates on arXiv.org

Nystr\"om Attention Matches Full Attention for Cross-Sectional Stock Prediction

・arXiv:2609.08106v1 Announce Type: new Abstract: MASTER's inter-stock multi-head attention -- the module responsible for modeling cross-sectional stock relationships -- accounts for 42.5% of model parameters and 25% of predictive value. ・We systematically decompose this module and uncover a surprising structure: the learned attention is near-uniform (perplexity 278/300), yet forcing exact uniformity eliminates all cros
WIRED

NZXT Discount Codes: 50% Off in September 2026

・Save 50%, plus up to $250 with NZXT promo codes and discounts.
cs.LG updates on arXiv.org

On BatchNorm Forward Modes in Value-Based Reinforcement Learning

・arXiv:2609.06421v1 Announce Type: new Abstract: Batch normalization (BN) substantially improves sample efficiency in continuous-control actor-critic methods such as CrossQ, yet recent studies report performance degradation in discrete-action value learning on Atari. ・These failures are surprising because discrete Q-networks lack the action-input distribution mismatch identified by CrossQ. ・We show for target-based C51
cs.LG updates on arXiv.org

On the Recall Scaling Laws in Mamba: A Theoretical and Mechanistic Study via Hashing

・arXiv:2609.07681v1 Announce Type: new Abstract: Associative Recall (AR) is the cognitive ability to learn and retrieve links between items in memory. ・In NLP, AR is used as a benchmark for evaluating the in-context memory capacity of architectures such as Mamba, and has been found to strongly correlate with language modeling performance. ・This paper explores AR from the perspective of mechanistic interpretability, aimi
cs.LG updates on arXiv.org

On-the-go Forgetting without Explicit Unlearning via ERASE

・arXiv:2609.05966v1 Announce Type: new Abstract: Existing unlearning approaches typically rely on post hoc weight adaptation or distillation, leading to duplicated memory costs, degraded generalization, and limited scalability. ・In this work, we introduce ERASE, Erasure via Reconstructive Adversarial Signal Editing, a framework for on-the-go forgetting that suppresses the observable influence of private data without mo
cs.LG updates on arXiv.org

One Rate Is Not Enough: Adaptive Anisotropic Learning Rates for LoRA Fine-Tuning

・arXiv:2609.05885v1 Announce Type: new Abstract: Low-rank adaptation (LoRA) has become the standard for parameter-efficient fine-tuning of large language models. ・Most LoRA variants follow a uniform-LR convention, applying a single global learning rate across every rank-one component of every adapter. ・We show that this convention overlooks substantial within-module heterogeneity, where the rank-one components of a LoRA
cs.LG updates on arXiv.org

One Step, One Lead: Mitigating Higher-Order Interference in Multi-Domain Reinforcement Learning via Cross-Step Control

・arXiv:2609.06469v1 Announce Type: new Abstract: Reinforcement learning (RL) across multiple domains can broaden the reasoning capabilities of large language models (LLMs), yet joint training often degrades individual-domain performance and can destabilize optimization. ・Existing work typically diagnoses such interference from a single-step view using first-order gradient alignment or curvature-based proxies.
cs.LG updates on arXiv.org

Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training

・arXiv:2609.07108v1 Announce Type: new Abstract: Speculative decoding accelerates rollout generation, which dominates the cost of reinforcement learning (RL) post-training. ・Online co-training can further increase the draft's accuracy, yielding greater speedups. ・However, scaling this approach to co-training on large models with long contexts poses two obstacles: (1) branch attention is unsupported by standard causal co
cs.LG updates on arXiv.org

Online Learning with LLM Experts from Limited Feedback

・arXiv:2609.05820v1 Announce Type: new Abstract: We study adaptive routing of prompts to large language model (LLM) experts to maximize response quality in an online setting with limited feedback. ・We formulate it as a bandit problem with $K$ actions that represent experts and $d$ features that encode prompts, over a horizon of $T$ rounds. ・We propose algorithms that strategically select and observe rewards to minimize
cs.LG updates on arXiv.org

Online Signature Verification Using Augmented Path Signature and T-Mamba

・arXiv:2609.08276v1 Announce Type: new Abstract: Handwritten signature verification is vital for personal authentication across commercial and financial applications. ・Although deep learning methods are widely adopted for online signature verification (OSV), they often struggle with capturing highly discriminative features and modelling long-range dependencies. ・To address these issues, we propose a novel framework that
cs.LG updates on arXiv.org

Online Surrogate Repair: Decoupling High-Fidelity Feedback from Search Length in Closed-Loop Discovery

・arXiv:2609.07655v1 Announce Type: new Abstract: Closed-loop AI scientists can generate candidate designs at low marginal computational cost, whereas reliable feedback may require wet-lab synthesis, characterization, or high-fidelity computation. ・Addressing this imbalance through custom laboratory automation remains infrastructure-intensive and costly, while replacing new experiments with a fixed surrogate leaves pers
Hugging Face Papers

OracleZoom: On-Policy Self-Distillation Inspired Reference-Constrained Recursive Image Super Resolution

OracleZoom: On-Policy Self-Distillation Inspired Reference-Constrained Recursive Image Super Resolution
WIRED

Our Favorite GoPro 360 Camera Is Now 40 Percent Off

・The GoPro Max 2 is almost half off at Amazon and other retailers, along with other great action camera deals.
cs.LG updates on arXiv.org

PAC-Bayesian Bounds for Learning Partially Observed Stochastic Linear Time-Invariant State-Space Systems with Inputs and Sub-Gaussian Noise

・arXiv:2609.08740v1 Announce Type: new Abstract: In this paper we derive a Probably Approximately Correct (PAC)-Bayesian error bound for partially observed linear time-invariant (LTI) stochastic dynamical systems in state-space form with inputs and sub-Gaussian noise. ・Such bounds are widespread in machine learning, and they are useful for characterizing the predictive power of models learned from finitely many data po
cs.LG updates on arXiv.org

PAC-Private Autoregressive Generation: Calibrating Noise to Ensemble Disagreement

・arXiv:2609.05676v1 Announce Type: new Abstract: Language models adapted on private text are often served through APIs, so privacy leakage occurs through generated outputs rather than exposed weights. ・Private prediction protects these releases. ・Methods such as PMixED incur privacy cost at each release and increasingly rely on the public model over long horizons.
cs.LG updates on arXiv.org

Parallelism Strategy Chaining for Fast Training Convergence

・arXiv:2609.07236v1 Announce Type: new Abstract: Selecting a parallelism strategy - the configuration of data, tensor, and pipeline parallelism degrees together with micro- and global-batch sizes - largely determines the training efficiency of large language models. ・State-of-the-art methods search for a parallelism strategy offline and select the single strategy that minimizes per-iteration time. ・But we find that they
cs.LG updates on arXiv.org

Parameterized and Streaming Algorithms for Euclidean Fair $k$-Center Clustering

・arXiv:2609.06384v1 Announce Type: new Abstract: Motivated by the growing importance of fairness in machine learning, fair $k$-center clustering has attracted considerable research attention as a fundamental problem. ・In this problem, a dataset is partitioned into $m$ disjoint groups, and the objective is to select $k$ data points as centers, subject to upper bounds on the number of centers chosen from each group, aimi
cs.LG updates on arXiv.org

ParetoTransport: Generative Optimization by Mass Transport Toward The Pareto Front

・arXiv:2609.07706v1 Announce Type: new Abstract: Offline multi-objective optimization requires not only moving the objective vectors of candidate designs toward the Pareto front, but also distributing them effectively along it. ・Generative methods have recently emerged as a natural approach because they learn a distribution over feasible designs while allowing generation to be steered toward promising designs.
Hugging Face Papers

PARSER: Read in Parallel, Reason in Depth for Long-Context LLM Agents

PARSER: Read in Parallel, Reason in Depth for Long-Context LLM Agents
cs.LG updates on arXiv.org

Particle Dynamics of Flow Matching and Classifier-Free Guidance from a Stagewise Geometry Perspective

・arXiv:2609.06947v1 Announce Type: new Abstract: Flow matching, together with classifier-free guidance (CFG), is widely used in generative modeling, yet much of the theoretical understanding remains distribution-wise. ・Since practical sampling follows individual trajectories, distribution-level guarantees alone do not fully capture how trajectories interact with the data geometry or how guidance reshapes it.
cs.LG updates on arXiv.org

PCFlow: Physics-Conditioned Flow Matching for GPR B-Scan Image Synthesis

・arXiv:2609.07300v1 Announce Type: new Abstract: Ground-penetrating radar (GPR) B-scan image synthesis is important for data augmentation, algorithm validation, and simulation acceleration, yet generating radargrams with both visual realism and physical consistency remains challenging. ・Existing learning-based generative models often emphasize visual appearance but provide limited control over response geometry.
cs.LG updates on arXiv.org

PCSDiff: Diffusion-Based Bias Correction and Super Resolution Toward Practical Operational Medium-Term Precipitation Forecast

・arXiv:2609.06942v1 Announce Type: new Abstract: Medium-range precipitation forecasts are impaired by persistent systematic biases, lead-time-dependent error accumulation, and coarse spatial resolution, restricting their reliability for flood-drought risk assessment. ・Existing AI correction techniques lack dedicated modeling for multi-day dynamic bias evolution and proper meteorological constraints, often generating ov
cs.LG updates on arXiv.org

PhenoBench: Mapping What a Deeply Phenotyped Human Cohort Can Tell Us

・arXiv:2609.06080v1 Announce Type: new Abstract: Deeply phenotyped cohorts combine clinical, imaging, molecular, and wearable observations across timescales from seconds to years. ・This breadth can reveal which measurements inform which health-related questions, but heterogeneous analyses are not directly comparable. ・We present PhenoBench, an executable benchmark built around the Human Phenotype Project, in which more
NVIDIA Blog

Physical AI Takes the Wheel: How the World’s Robotaxi Leaders Are Building With NVIDIA Technologies

・The global robotaxi market — physical AI’s first commercial breakthrough — is projected to reach $400 billion by 2035, with over 6 million commercial vehicles in operation as driverless fleets are already moving people through some of the world’s busiest and most complex streets. ・Deploying a driverless vehicle is one challenge. ・Scaling a fleet is […]
cs.LG updates on arXiv.org

Physics-Informed Deep Learning for False Ventricular Tachycardia Alarm Reduction in the ICU

・arXiv:2609.08992v1 Announce Type: new Abstract: False ventricular tachycardia (VT) alarms are a leading contributor to alarm fatigue in intensive care units. ・We propose a deep learning framework combining a 1D SE-ResNet with ICU-realistic data augmentations and a physics-informed auxiliary reconstruction task based on the three-element Windkessel hemodynamic model, implemented as a differentiable forward simulation.
cs.LG updates on arXiv.org

PhysSAE: Mechanistic Interpretability with Sparse Autoencoders

・arXiv:2609.07061v1 Announce Type: new Abstract: Physics-Informed Neural Networks (PINNs) embed PDE residuals into neural network training, but their internal representations remain opaque: it is unknown what physical features their hidden layers encode or whether those features have a localized causal role. ・We present PhysSAE, a mechanistic interpretability framework that trains overcomplete sparse autoencoders (SAEs
Hugging Face Papers

PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving

PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving
cs.LG updates on arXiv.org

PlayTrain: An Efficient Reinforcement Learning Framework for LLM-Generated Adaptable JavaScript Games

・arXiv:2609.09059v1 Announce Type: new Abstract: While many video-game environments (VGEs) have played crucial roles in advancing reinforcement learning (RL), developing novel VGEs or modifying existing ones to support new features, has been a laborious process requiring extensive hand-coding. ・Here we present PlayTrain, an RL framework that combines the abilities of large language models (LLMs) to robustly generate Ja
cs.LG updates on arXiv.org

Power Mean Estimation in Stochastic Continuous Monte Carlo Tree Search

・arXiv:2609.06489v1 Announce Type: new Abstract: Monte Carlo Tree Search (MCTS) has demonstrated success in online planning for deterministic environments, yet significant challenges remain in adapting it to stochastic Markov Decision Processes (MDPs), particularly in continuous state-action spaces. ・Existing methods, such as HOOT, which combines MCTS with the Hierarchical Optimistic Optimization (HOO) bandit strategy,
cs.LG updates on arXiv.org

PPIM: Pennes Physics-Informed Mamba for Heat-Source-Conditioned 3D Bioheat Simulation

・arXiv:2609.06869v1 Announce Type: new Abstract: Three-dimensional bioheat simulation aims to predict transient temperature distributions in biological tissue and is commonly modeled using the Pennes bioheat equation, which combines thermal diffusion, perfusion-mediated heat loss, and external heat generation. ・In this study, we consider a controlled 3D Pennes bioheat simulation under a localized heat-source condition
Zennの「機械学習」のフィード

Pre-trainingの目的関数・teacher forcing・損失実装を整理する

・LLMのPre-training(事前学習)を実装目線で捉えると、中心にあるのは「大量テキスト」ではなく、どのtokenを入力として見せ、どのtokenを正解として採点するかです。 ・この設計から、label shift、attention mask、loss mask、teacher forcing、perplexityの意味が決まります。本記事ではcausal LM、BERT型MLM、T5型span corruptionを比較した後、causal LMの最小loss実装まで落とします。 ・データ準備や知識獲得の背景を含む全体像は個人ブログ完全版にまとめています。
cs.LG updates on arXiv.org

Prevalence calibration as shortcut mitigation

・arXiv:2609.07922v1 Announce Type: new Abstract: Shortcut learning denotes the widespread situation in which a classifier exploits spurious correlations rather than diagnostic features. ・Existing mitigation strategies mostly aim to learn shortcut-invariant representations; their empirical success is limited and they cannot be applied to classifiers using frozen foundation model encoders. ・We propose to reframe shortcut
cs.LG updates on arXiv.org

Proactive Context-Forecasted Safety Constraints for Nonstationary Reinforcement Learning

・arXiv:2609.08080v1 Announce Type: new Abstract: Ensuring safety in reinforcement learning under nonstationarity requires anticipating changes in risk before they lead to unsafe behavior. ・Existing approaches typically rely on safety constraints defined at design time or updated reactively during execution, assuming that such constraints remain valid over time. ・However, in nonstationary environments with evolving conte
Hugging Face Papers

Programmable World Model

Programmable World Model
cs.LG updates on arXiv.org

Protocol effects on feature-based hardware-Trojan detection across Trust-Hub families

・arXiv:2609.07199v1 Announce Type: new Abstract: Trust-Hub reuses host circuits: several files differ mainly in the inserted Trojan. ・When gates from sibling variants enter both training and test folds, a detector can benefit from host logic it has already seen. ・We measure that effect instead of proposing another classifier.
Hugging Face Papers

Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation

Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation
cs.LG updates on arXiv.org

QGB-W$k$NN: Quantum Granular-Ball Learning for Robust Classification

・arXiv:2609.05952v1 Announce Type: new Abstract: Nearest-neighbor classification is widely used in machine learning, yet existing methods often suffer from low computational efficiency and limited robustness in noisy environments. ・To jointly address these challenges, this paper proposes an efficient and reliable weighted $K$-nearest neighbor classification framework based on quantum granular balls, termed QGB-W$k$NN.
cs.LG updates on arXiv.org

RAPTOR: Role-Aware Private Training for Mixture-of-Experts

・arXiv:2609.05770v1 Announce Type: new Abstract: Differentially private (DP) fine-tuning methods treat sparse Mixture-of-Experts (MoE) models as a single dense block, ignoring that shared layers see all data while experts only see routed records. ・We identify and formally characterize three resulting failure modes: global clipping suppresses expert gradients, batch-level normalization dilutes sparse expert updates, and
Hugging Face Papers

Reference-Based Bias Detection in LLMs via Relative Representations of Hidden States

Reference-Based Bias Detection in LLMs via Relative Representations of Hidden States
cs.LG updates on arXiv.org

REFINE: Trajectory Representation Learning via Closed-Loop Transcription -- Extended Version

・arXiv:2609.07206v1 Announce Type: new Abstract: Trajectory representation learning underpins a wide range of trajectory analytics tasks; however, most existing self-supervised approaches, whether discriminative or generative, adopt an open-loop paradigm, relying on fixed data augmentations or random masking without feedback, which limits their ability to generalize and scale. ・We propose REFINE, a simple yet effective
cs.LG updates on arXiv.org

Representation Learning for Sample-Efficient CATE Estimation by Leveraging Multiple Outcomes

・arXiv:2609.06294v1 Announce Type: new Abstract: Estimating conditional average treatment effects (CATE) enables efficient targeting of interventions, but many applications have limited experimental samples, making it difficult to estimate heterogeneous effects from high-dimensional covariates. ・In such settings, policymakers and medical practitioners often succumb to the curse of dimensionality or apply off-the-shelf
Hugging Face Papers

RESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation Systems

RESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation Systems
cs.LG updates on arXiv.org

Rethinking One-Shot Federated Graph Learning: Training-Free Statistical Estimation

・arXiv:2609.06154v1 Announce Type: new Abstract: One-shot federated graph learning generally aims to train Graph Neural Networks (GNNs) across clients with disconnected subgraphs in a single communication round. ・Existing methods predominantly design advanced optimization strategies under the premise that local GNN training is indispensable. ・However, empirical observations reveal that under extreme non-IID conditions,
cs.LG updates on arXiv.org

Rethinking the Evaluation of Efficiency Methods for Multi-Agent Systems

・arXiv:2609.05933v1 Announce Type: new Abstract: Efficiency is increasingly important for Large Language Model (LLM)-based multi-agent systems (MAS), as larger models and more agents introduce substantial execution costs. ・Recent methods aim to make MAS cheaper by pruning agents, removing communication edges, or searching for compact structures. ・However, we argue that existing evaluations may overestimate their true ab
Hugging Face Papers

Revisiting Complete Reasoning Traces for Post-Training

Revisiting Complete Reasoning Traces for Post-Training
cs.LG updates on arXiv.org

Revisiting Spectral Representations in Generative Diffusion Models

・arXiv:2609.08253v1 Announce Type: new Abstract: Diffusion models have shown remarkable performance on diverse generation tasks. ・Recent work finds that imposing representation alignment on the hidden states of diffusion networks can both facilitate training convergence and enhance sampling quality, yet the mechanism driving this synergy remains insufficiently understood. ・In this paper, we investigate the connection be
cs.LG updates on arXiv.org

Revisiting Thinning Methods for Kernel Learning Problems

・arXiv:2609.07432v1 Announce Type: new Abstract: Kernel methods are widely used because of their strong theoretical guarantees and empirical performance. ・However, their high computational cost limits their applicability to large-scale datasets. ・To address this shortcoming, several approaches use Maximum Mean Discrepancy to construct representative subsets that preserve the properties of the full dataset in a Reproduci
cs.LG updates on arXiv.org

Risk-Conditioned Fine-Tuning of Large Language Models

・arXiv:2609.08064v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in settings where rare but severe harmful generations can have significant consequences. ・Existing Risk-Averse RLHF addresses this issue by optimizing Conditional Value-at-Risk (CVaR), but it trains policies for fixed risk levels and therefore cannot adjust the desired degree of risk aversion at inference time.
cs.LG updates on arXiv.org

Robust Conformal Consensus: Multi-Agent LLM-as-a-Judge Interval Evaluation with Conformal Prediction

・arXiv:2609.06367v1 Announce Type: new Abstract: LLM-as-a-Judge has emerged as a promising paradigm for evaluating natural language generation. ・However, the uncertainty associated with such evaluations remains largely unexplored, which limits their reliability in real-world applications. ・Although conformal prediction offers a principled framework for uncertainty quantification, existing approaches typically apply it t
cs.LG updates on arXiv.org

Robust Decentralized Federated Distillation via Multi-Modality Knowledge Collaboration

・arXiv:2609.07230v1 Announce Type: new Abstract: This paper propose a robust decentralized federated distillation method that enables clients with heterogeneous models to collaborate through predictions on shared unlabeled public data. ・In the proposed method, each client first evaluates the received predictions in three modalities of class prediction, boundary decision, and prediction correlation. ・It then filters unre
cs.LG updates on arXiv.org

Robust Decentralized Personalized Federated Learning via Prediction-Constrained Neighborhood Collaboration

・arXiv:2609.07312v1 Announce Type: new Abstract: This paper proposes a robust decentralized personalized federated learning method R-DPFL, that enables clients to reduce the impact of Byzantine attacks via robust neighborhood direction estimation and history-based update trend prediction, rather than purely aggregating client models as in the existing work. ・In R-DPFL, each client first computes the current-round model
cs.LG updates on arXiv.org

Robust Dynamic Expansion for Continual Learning under Backdoor Attacks via Purification and Selective Recovery

・arXiv:2609.06346v1 Announce Type: new Abstract: Continual learning (CL) enables models to acquire new knowledge from sequentially arriving tasks while retaining previously learned knowledge. ・However, in practical scenarios, task streams collected from untrusted sources may contain backdoor-poisoned samples, posing a critical challenge to the stability, plasticity, and security of continual learners. ・In this work, we
cs.LG updates on arXiv.org

Robustness of LLM-Generated SystemVerilog Assertions to Semantics-Preserving RTL Transformations

・arXiv:2609.05658v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly being explored for automating SystemVerilog Assertion (SVA) generation, yet most evaluations report correctness on a single syntactic representation of an input. ・Such point accuracy does not reveal whether a model's correct output is stable when the same RTL behavior is written differently. ・This paper presents a controlled m
cs.LG updates on arXiv.org

Role-Specific Predictive Geometries for Nonstationary Multivariate Graph-Signal Forecasting

・arXiv:2609.06519v1 Announce Type: new Abstract: Forecasting multivariate graph signals is challenging when node-level trajectories are nonstationary but stable relations persist across nodes and features. ・In an error-correction representation, long-run equilibrium restoration and short-run transient propagation represent different predictive roles and need not share a common cross-feature geometry. ・We introduce role-
Hugging Face Papers

SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?

SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?
cs.LG updates on arXiv.org

SAFEGuard: Detect Optimization-Based Jailbreak Attacks Through Harmful Semantic Analysis and Fluency Measurement

・arXiv:2609.05850v1 Announce Type: new Abstract: Despite the significant efforts devoted to aligning large language models (LLMs) with human values and ensuring safe deployment, recent work has revealed that LLMs remain vulnerable to adversarial jailbreak attacks that can bypass safety guardrails and elicit harmful responses. ・Many defense methods are proposed to detect jailbreaks but they are limited in their effectiv
cs.LG updates on arXiv.org

Scaling Optimal Classification Trees via Adaptive Feature and Sample Reduction

・arXiv:2609.05826v1 Announce Type: new Abstract: Dynamic programming for optimal classification trees becomes computationally expensive as the numbers of features and training samples increase. ・We develop a joint feature- and sample-space reduction framework based on STreeD. ・Weighted STreeD merges duplicate records created after projection onto a fixed candidate set into weighted representatives.
Hugging Face Papers

SchemeArena: Factorized Stress Testing of Scheming in LLM Agents

SchemeArena: Factorized Stress Testing of Scheming in LLM Agents
Hugging Face Papers

Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents

Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents
cs.LG updates on arXiv.org

SeaCausal-FL: Federated Fuzzy Causal Learning for Maritime IoT Fault Diagnosis and Counterfactual Reasoning

・arXiv:2609.06257v1 Announce Type: new Abstract: Reliable marine-engine fault diagnosis in maritime IoT is challenged by distributed data ownership, heterogeneous fault distributions, and continuously changing operating conditions. ・This paper proposes SeaCausal-FL, a federated fuzzy causal learning framework that combines a shared temporal diagnostic path with mechanism-conditioned causal reasoning. ・An interval type-2
cs.LG updates on arXiv.org

Second-Order Smooth Planning with Optimal-Transport Bellman Smoothing

・arXiv:2609.06484v1 Announce Type: new Abstract: Planning with a generative model aims to estimate the value of a state using as few simulator calls as possible. ・SmoothCruiser achieves problem-independent complexity $\widetilde O(\varepsilon^{-4})$ by exploiting the smoothness of the entropy-regularized Bellman backup, but its estimator is only first-order. ・We show that the sample-complexity exponent of SmoothCruiser-
cs.LG updates on arXiv.org

Sector-Mean: Deterministic Initialization of K-Means Centroids via Angular Sector Partitioning

・arXiv:2609.06468v1 Announce Type: new Abstract: K-Means is one of the most widely used clustering algorithms, but its susceptibility to initial centroid selection remains a primary bottleneck for its convergence speed and clustering accuracy. ・This paper proposes Sector-Mean Initialization, a deterministic initialization strategy with O(N) time complexity that partitions the two-dimensional data space into angular sec
cs.LG updates on arXiv.org

Selective Posterior Margin Regularization for Forward-Corrected Classification

・arXiv:2609.05859v1 Announce Type: new Abstract: Learning with class-conditional label noise often relies on a transition model from latent clean classes to observed annotations. ・Forward correction embeds this transition in the likelihood, yet finite-sample networks may still memorize corrupted labels. ・The corrected likelihood also induces a reverse posterior over the clean classes that could explain each annotation.
cs.LG updates on arXiv.org

Semi-Supervised Learning under Spatially Biased Sampling

・arXiv:2609.07982v1 Announce Type: new Abstract: Standard semi-supervised learning (SSL) typically relies on labelled and unlabelled data sharing a common marginal distribution. ・This assumption is often violated by biased spatial sampling mechanism, when labels are collected under spatially biased or preferential site selection. ・We treat this marginal mismatch, spatial autocorrelation, and spatial non-stationarity as
cs.LG updates on arXiv.org

Sharp Structure-Agnostic Minimax Risk for Partial Linear Models

・arXiv:2609.07997v1 Announce Type: new Abstract: We characterize the sharp structure-agnostic minimax risk for coefficient estimation in the partial linear model when the outcome and treatment nuisances are learned by two distinct black-box learners, which resolves the open problem in double machine learning posed by Gu (2025). ・For each nuisance \(q\in\{\mu,\pi\}\), we characterize the available learner by an approxim
Hugging Face Papers

Show-Harness: Just a VLM Agent Can Play Robots

Show-Harness: Just a VLM Agent Can Play Robots
cs.LG updates on arXiv.org

SIM: Subspace Interaction-based Method for Token-Level Text Anomaly Detection

・arXiv:2609.08200v1 Announce Type: new Abstract: Token-level text anomaly detection, as an emerging trend of text anomaly detection, moves beyond coarse-grained document-level detection by localizing anomalous tokens within text. ・By providing fine-grained abnormality prediction, token-level text anomaly detection plays a critical role in various real-world applications, such as spam filtering and fake news detection.
NVIDIA Blog

Skild AI Taps NVIDIA Physical AI to Teach Robots New Tasks From a Single Video

・Manufacturing floors, warehouses and production lines rarely stay fixed — tasks change, layouts shift and new products arrive, and most robots can’t keep up without significant reprogramming. ・Skild AI’s new S1 robot foundation model helps address this, designed to learn previously unseen, long-horizon tasks from a single video demonstration. ・The model, launched last week, uses […]
cs.LG updates on arXiv.org

Solving the Elastic Wave Equation with Physics-Informed Neural Networks: A Robust and Critical Assessment

・arXiv:2609.07983v1 Announce Type: new Abstract: Physics-Informed Neural Networks (PINNs) have recently emerged as a promising approach for solving Partial Differential Equations (PDEs), offering a meshfree alternative that integrates physical principles into the learning process. ・This presents a new paradigm compared to traditional discretization methods and purely data-driven machine learning techniques.
cs.LG updates on arXiv.org

Sparse Data Augmentation for Optimization with Provable Guarantees

・arXiv:2609.08133v1 Announce Type: new Abstract: In nonconvex optimization problems arising in geometric machine learning, data augmentation is commonly used to promote invariance by averaging empirical losses over transformations of the data. ・Computing the fully augmented objective, however, requires access to every element of the transformation group $G$, which may be prohibitively expensive when $G$ is large or acc
cs.LG updates on arXiv.org

Sparse Incident-Cluster Learning for 12-hour Port Flood Pre-warning in Digital-Twin Analytics

・arXiv:2609.06109v1 Announce Type: new Abstract: Port flood digital twins require analytics that warn operators before disruption, but official warning incidents are often few and adjacent observations are temporally dependent. ・Row-level classification can therefore overstate performance by placing windows from the same event in both model-development and evaluation data. ・We formulate 12-hour port flood pre-warning as
cs.LG updates on arXiv.org

Sparse Oblique Rule Boosting for Simpler Additive Rule Ensembles

・arXiv:2609.06426v1 Announce Type: new Abstract: Small additive ensembles of symbolic rules offer interpretable prediction models. ・Traditionally, these ensembles use rule conditions based on conjunctions of simple threshold propositions $x \geq t$ on a single input variable $x$ and threshold $t$, resulting geometrically in axis-parallel polytopes as decision regions. ・While this form ensures a high degree of interpreta
cs.LG updates on arXiv.org

Spectral Prioritized Sweeping in Nonstationary Reinforcement Learning

・arXiv:2609.06186v1 Announce Type: new Abstract: Prioritized Sweeping (PS) accelerates model-based reinforcement learning by selecting backups according to Bellman residual magnitude. ・In nonstationary reward settings, however, the canonical priority score is shortsighted: after a localized reward shift, residuals propagate only through realized backups, so bottlenecked or topologically distant state estimates may rema
cs.LG updates on arXiv.org

Stability and Generalization of Straight-Through Estimators for Training Two-Layer Quantized Neural Networks

・arXiv:2609.06430v1 Announce Type: new Abstract: We study the identity straight-through estimator (STE) for training a two-layer binary-activation network with hinge loss from the perspective of Statistical Learning Theory (SLT). ・Our central question is whether algorithmic stability can explain the statistical generalization of the estimator produced by the discontinuous STE training rule. ・In the saturated-output regi
cs.LG updates on arXiv.org

Stable-MM-R1: Anchoring Multimodal Reasoning Dynamics via Entropy-Guided Stratification

・arXiv:2609.07148v1 Announce Type: new Abstract: While Reinforcement Learning (RL) effectively incentivizes reasoning in Large Language Models, current pipelines are hindered by training instability and rapid entropy collapse. ・These limitations often stem from "Rollout Silencing" and low-quality gradient signals in standard sampling procedures. ・In this work, we propose a robust, data-centric framework to stabilize RL
cs.LG updates on arXiv.org

Statistical versus machine learning-based spatial interpolation of post-processed ensemble weather forecasts

・arXiv:2609.07512v1 Announce Type: new Abstract: Statistical post-processing improves ensemble weather forecasts, but generating calibrated predictions at locations without observations remains challenging. ・This study compares statistical and machine-learning-based methods for post-processing ECMWF 2-m temperature and 10-m wind speed forecasts at observed and unobserved stations in Germany. ・We consider EMOS-based appr
cs.LG updates on arXiv.org

Steering Interference Reflects the Model's Defaults, Not the Behavior Directions

・arXiv:2609.06951v1 Announce Type: new Abstract: Activation steering promises modular control of language model behavior: a behavior such as politeness corresponds to a direction in a model's activations, and adding that direction while it generates should switch the behavior on and leave everything else alone. ・It does not. ・We ask what decides which other behaviors move, and by how much, and find that it is the model
cs.LG updates on arXiv.org

Steering Under Compression: Dose-Response, Capability Cost, and Failure Asymmetry in Quantized LLMs

・arXiv:2609.06473v1 Announce Type: new Abstract: Inference-time activation steering enables behavioral control of large language models without parameter modification, while post-training quantization reduces memory and compute costs for deployment. ・Despite their growing convergence in practice, the interaction between these two techniques remains uncharacterized. ・We systematically study activation steering under weig
cs.LG updates on arXiv.org

Stochastically Perturbed Weights: Ensembles from Deterministic Machine-Learning Weather Models

・arXiv:2609.08412v1 Announce Type: new Abstract: Machine-learning weather models (MLWMs) now match or outperform operational numerical weather prediction (NWP) at global medium-range forecasting, at far lower inference cost. ・Many deployed MLWMs are deterministic, producing a single forecast with no estimate of its own uncertainty, whereas a growing family of trained-probabilistic models generate calibrated ensembles d
Hugging Face Papers

StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean

StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean
cs.LG updates on arXiv.org

Streaming Hierarchical Inference with Tabular Foundation Models

・arXiv:2609.07956v1 Announce Type: new Abstract: Tabular Foundation Models (TFMs) have recently demonstrated strong predictive performance through in-context learning, but their deployment in high-throughput data streams remains challenging due to communication overhead and latency. ・We propose \textit{HINT}, a hierarchical inference framework that combines edge-based retrieval with cloud-based TFM inference.
cs.LG updates on arXiv.org

Structural Entropy-Driven Graph Diffusion Generation for One-Shot Federated Graph Learning

・arXiv:2609.06499v1 Announce Type: new Abstract: One-shot federated graph learning (FGL) requires the server to estimate client contributions from highly compressed information, yet conventional volume-based weighting captures the amount of client data while overlooking how its connectivity is organized. ・In this paper, we propose SPIRE, a Structural Entropy-Driven Graph Diffusion Generation method that introduces topo
cs.LG updates on arXiv.org

Structured Extrema Errors in Classical Surrogates for Viscous Burgers: A Physics-Consistent Interpretation

・arXiv:2609.07952v1 Announce Type: new Abstract: We study the local errors of classical machine-learning surrogate models, which approximate the time evolution of the one-dimensional viscous Burgers equation. ・Four models are compared on the same prediction task, using the spatial grid values directly: radial basis function (RBF) kernel ridge regression (KRR), linear Ridge, ExtraTrees, and Random Forests. ・Across all fo
cs.LG updates on arXiv.org

Suan: Rectifying Direct Preference Safety Alignment in Large Language Models

・arXiv:2609.08634v1 Announce Type: new Abstract: Integrating robust safety guardrails into Large Language Models (LLMs) is essential for delivering helpful yet harmless responses. ・While proprietary systems exhibit reliable safety controls, their underlying methodologies and trade-offs remain largely undisclosed. ・Achieving comparable security in open-weight models remains a persistent challenge, as post-trained variant
cs.LG updates on arXiv.org

SUN: Reaching for Novelty in Reinforcement Learning

・arXiv:2609.08642v1 Announce Type: new Abstract: Exploration in reinforcement learning (RL) remains a fundamental challenge. ・Recent goal-conditioned RL strategies (which select goals to encourage broader state coverage) have shown promising results, but none scores a goal by novelty and reachability jointly: the two signals are traded off by hand, applied in sequence, or one is neglected outright. ・In this paper, we in
Hugging Face Papers

SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents

SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents
cs.LG updates on arXiv.org

SwiftExplorer: Training-free Diffusion Model Alignment with Swift Diversity Exploration

・arXiv:2609.06651v1 Announce Type: new Abstract: Diffusion models have general generative abilities but struggle to align with specific objectives. ・Fine-tuning can improve alignment, yet its training cost is often prohibitive. ・This led to training-free methods that apply objective-guided terms in sampling to bias the generation distribution toward designated regions, e.g., high-reward areas.
Hugging Face Papers

SyncWorld: Visual Calibration Enables World Models as Zero-Shot Simulators

SyncWorld: Visual Calibration Enables World Models as Zero-Shot Simulators
cs.LG updates on arXiv.org

Synergistic Fusion of Topological Structure and Temporal Semantics of Mobility for Urban Region Embedding

・arXiv:2609.08268v1 Announce Type: new Abstract: Urban region embeddings have shown promising results in diverse urban sensing tasks such as crime, income, and service-call prediction. ・Recent methods improve representation quality by integrating mobility data with auxiliary modalities, using cross-view attention or contrastive objectives to align heterogeneous features into a unified region representation.
cs.LG updates on arXiv.org

TabBench-Bio: A Living Benchmark for Machine Learning on High-Dimensional Biomedical Tables

・arXiv:2609.07441v1 Announce Type: new Abstract: Biomedical tables often combine thousands of measured variables with only tens or hundreds of labelled samples, a regime that is poorly represented in general-purpose tabular benchmarks. ・We introduce TabBench-Bio, a living and interactive benchmark of 43 biomedical datasets spanning multiple domains. ・Under a shared cross-validation protocol, we compare classical estimat
cs.LG updates on arXiv.org

Target-Independent Micro-Interventions for Predicting Training Response Across Language-Model Families

・arXiv:2609.08618v1 Announce Type: new Abstract: Benchmark scores describe what a checkpoint can do now, but they do not determine how it will respond to the next training episode. ・We measure this missing state by branching four short, standardized, target-independent micro-interventions from the same checkpoint and recording their effects in a common capability space. ・Together with current capability, these responses
cs.LG updates on arXiv.org

TASTE: Throughput-Aware Batch Size Tuning for On-Device Edge Learning

・arXiv:2609.07444v1 Announce Type: new Abstract: The rise of privacy-preserving artificial intelligence (AI) has shifted the focus of model adaptation and personalization towards on-device learning, where deep learning models are finetuned directly on edge hardware using local user data. ・However, this shift requires optimization of deep learning training on resource-constrained hardware to maximize throughput while ma
Zennの「大規模言語モデル」のフィード

te claude 一発でClaude Codeの裏側をKimi K3に — 自動ルーターで安いモデルと賢いモデルを切り替える

・TL;DR te claude — Claude Code を teai.io 経由で起動するランチャーを作った。環境変数の手動設定なしで、裏側のモデルを Kimi K3 などカタログ上の任意モデルに差し替えられる model: "teai/auto" — リクエスト内容を見て「短い質問→安いモデル / コードや長文→Kimi K3」をリクエスト単位で自動選択するルーターを実装した。課金とレスポンスの model フィールドは解決後の実モデルで返す モデル系統ごと落ちた場合のクロスモデル自動フォールバックも入れた(x-teai-fallback-model ヘッダで明示) 前...
cs.LG updates on arXiv.org

Temporal Heterogeneous Graph Transformer for Credit Card Fraud Detection

・arXiv:2609.07100v1 Announce Type: new Abstract: Credit card fraud detection typically relies on tabular features, while repeated attributes can also provide useful relational signals. ・This paper proposes THGT-FD, a Temporal Heterogeneous Graph Transformer for Fraud Detection. ・Each transaction is represented using one transaction token and six types of relation tokens and incorporates Time2Vec encoding into the transa
cs.LG updates on arXiv.org

Temporal-Causal Inference for Reinforcement Learning via Automata Learning

・arXiv:2609.07461v1 Announce Type: new Abstract: We consider reinforcement learning in environments with dynamics that undergo an irreversible phase transition governed by a hidden temporal pattern. ・The agent observes the base state but cannot observe the phase directly. ・We formalize this problem as a two-phase non-Markovian decision process and introduce Temporal-Causal Inference for Reinforcement Learning (TCIRL), a
cs.LG updates on arXiv.org

The Accuracy Paradox: Empirical Diagnostic of Default Decision Thresholds in Multi-Label Enzyme Commission Prediction [With Code]

・arXiv:2609.07897v1 Announce Type: new Abstract: Automated prediction of Enzyme Commission (EC) numbers plays a central role in functional annotation and computational drug discovery. ・However, standard multi-label machine learning pipelines frequently rely on default decision thresholds (t=0.50), assuming balanced prior distributions across target heads. ・In this study, we present a systematic empirical diagnostic of u
cs.LG updates on arXiv.org

The BatchNorm Illusion: Diagnosing Normalization Artifacts in Machine Unlearning Evaluation

・arXiv:2609.08901v1 Announce Type: new Abstract: Approximate machine unlearning aims to remove the influence of specific training data from a trained model without retraining from scratch. ・We identify a previously undocumented confound in how unlearning is evaluated on BatchNorm-based architectures: a single forward pass over retain data, an operation that modifies no weight, can deterministically rewrite the model's
cs.LG updates on arXiv.org

The Geometry of Refusal: Why Post-Hoc Safety Is Fragile and Pretraining-Time Safety Persists

・arXiv:2609.06934v1 Announce Type: new Abstract: Post-hoc safety training (RLHF, DPO) is the dominant way to align large language models, yet jailbreaks (Zou et al., 2023b), fine-tuning attacks (Qi et al., 2024), and activation-space probes (Arditi et al., 2024) keep recovering the behaviors it was meant to remove. ・We give this fragility one geometric explanation and trace it to when, during pretraining, safety can ta
The Verge

The iPhone Duo’s hardware doesn’t look special, but its software might be

・With the iPhone Duo, Apple has pulled off a familiar trick. ・It arrives into a mature Android foldable market with a handful of hardware features we've mostly already seen elsewhere, but paired with a level of software polish that makes a lot of them feel new again. ・It's those software flourishes that have me excited to try the iPhone Duo when it launches next month, but I'm less convinced that its hardware will be an
cs.LG updates on arXiv.org

The Oversight Gap: What LLM Safety Monitors Miss, and Why It Is Not Capability

・arXiv:2609.07162v1 Announce Type: new Abstract: Several properties safety monitors are asked to certify, among them cross-tenant noninterference, sandbagging and evaluation awareness, are 2-safety hyperproperties, witnessed only by two executions. ・The standard consequence is a binary impossibility: one trace cannot decide them. ・We replace the binary with a measurement.
Hugging Face Papers

The Semantic Bottleneck: Leveraging Semantic Representations for Non-Invasive Speech Decoding

The Semantic Bottleneck: Leveraging Semantic Representations for Non-Invasive Speech Decoding
cs.LG updates on arXiv.org

Think Wider: Mitigating Latent Rank Collapse in Implicit Chain-of-Thought Reasoning

・arXiv:2609.07406v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning improves the reasoning ability of large language models by introducing intermediate computation, but explicit rationales increase decoding length, latency, and context cost. ・Implicit CoT offers a more efficient alternative by moving intermediate reasoning into continuous latent states. ・However, latent reasoning can be unstable: successiv
cs.LG updates on arXiv.org

Topological Fraud Detection in Latent Transaction Spaces

・arXiv:2609.08445v1 Announce Type: new Abstract: Working entirely on topologically anonymized embeddings, we perform fraud detection using iterative rounds of unsupervised filtering followed by supervised sniping. ・The result is an ultra-low latency privacy--preserving triage that allows institutions to flag suspicious activity without compromising Personally Identifiable Information.
cs.LG updates on arXiv.org

Topology-induced Operators Reveal Complementary Graph Representations without Training

・arXiv:2609.08152v1 Announce Type: new Abstract: Graph representation learning has largely focused on designing increasingly sophisticated models to transform graph topology into vector representations, or embeddings. ・However, the extent to which embedding quality depends on model learning, rather than on the underlying topological transformations, remains unclear. ・Here, we show that informative embeddings can be deri
cs.LG updates on arXiv.org

Towards Unified Multimodal Graph Foundation Model: A Bridge-Router-Adapter Based Approach

・arXiv:2609.06668v1 Announce Type: new Abstract: Multimodal graphs couple node attributes in different modalities, such as text and images, with relational structure, enabling topological structure and cross-modality attributes to be modeled jointly. ・Multimodal graph foundation models seek unified representations from such data that transfer across different graph domains and downstream tasks. ・However, existing method
cs.LG updates on arXiv.org

Tracking the Moving Frontier: Long-Short Term Advantage Estimator

・arXiv:2609.06671v1 Announce Type: new Abstract: Group-based RLVR methods estimate advantages by repeatedly sampling multiple trajectories for each prompt, making long-horizon agent training expensive and discarding useful experience accumulated across iterations. ・We ask whether historical experience can replace these repeated within-iteration comparisons without directly optimizing on stale trajectories. ・We introduce
Hugging Face Papers

Train Smarter, Not Harder: Switching Signal-Guided Training in Active Learning

Train Smarter, Not Harder: Switching Signal-Guided Training in Active Learning
cs.LG updates on arXiv.org

Train Smarter, Not Harder: Switching Signal-Guided Training in Active Learning

・arXiv:2609.06806v1 Announce Type: new Abstract: Training strategy, namely whether to retrain from scratch or fine-tune from the previous checkpoint, is an overlooked decision variable in active learning. ・We show that this choice has exploitable structure: retraining is most useful in early rounds, when each batch can substantially reshape the labeled distribution, while fine-tuning becomes safer once the model trajec
cs.LG updates on arXiv.org

Training-Free Task Vectors for LLM Behavioral Control

・arXiv:2609.09054v1 Announce Type: new Abstract: Task vectors enable post-training model editing by identifying semantically meaningful directions in weight space, typically computed as the difference between a fine-tuned model and its pretrained initialization. ・However, this reliance on fine-tuning makes discovering such directions costly and limits the practicality of post-training model editing. ・To address this lim
cs.LG updates on arXiv.org

Transformers as In-Context Samplers: From Closed-Form Diffusion to Estimation-Free Sampling

・arXiv:2609.08981v1 Announce Type: new Abstract: A growing body of work establishes that large language models are not mere statistical memorizers, but are capable of in-context learning: performing inference at test time using only examples provided in the prompt, without any parameter updates. ・Prior theoretical work has shown that this capability extends to supervised learning tasks such as linear regression.
cs.LG updates on arXiv.org

Translation of Black-Box Clinical Prediction Models into Standalone Transparent Nomograms: Temporal External Validation in Heart Transplantation

・arXiv:2609.07610v1 Announce Type: new Abstract: We convert black-box clinical prediction models for tabular data into standalone nomograms that can be audited term by term. ・PRiSM (Partial Responses in Structured Models) takes the shape of each effect and interaction from the source model, not merely which variables mattered, and lets the outcome select and weight them. ・We tested this in 50,356 heart transplant recipi
cs.LG updates on arXiv.org

TrojanWorld: Backdooring World-Model Agents via Imagination Steering

・arXiv:2609.07051v1 Announce Type: new Abstract: World models increasingly serve as the predictive core of model-based reinforcement learning agents, enabling them to simulate future dynamics and reason over imagined trajectories before acting. ・Their substantial training demands make pretrained world models attractive for distribution and reuse, exposing downstream systems to model supply chain threats. ・Backdoor attac
WIRED

Trump Probably Won’t Give $5,000 to Every US Adult if Republicans Win the Midterms

・The eyebrow-raising proposal, delivered during the president’s speech at the midterm convention, would cost $1.2 trillion and would need congressional approval.
cs.LG updates on arXiv.org

Trust-But-Verify: Poisoning-Resilient Locally Private Graph Learning Protocols

・arXiv:2609.07063v1 Announce Type: new Abstract: Built upon local differential privacy (LDP), locally private graph learning protocols have emerged as an important paradigm for decentralized graph learning, balancing privacy protection and learning utility. ・Under such protocols, each user locally perturbs their node features and adjacency information before transmission, ensuring formal privacy guarantees without orig
cs.LG updates on arXiv.org

TV-Regulated OPD: Direction Matters in On-Policy Distillation

・arXiv:2609.08341v1 Announce Type: new Abstract: On-Policy Distillation (OPD) facilitates the transfer of knowledge from domain expert to student in the post-training phase of Large Language Models (LLMs). ・However, the supervision signals in mainstream OPD methods suffer from high variance and noise which is generally instable during training. ・In this work, we systematically investigated what really matters to the per
cs.LG updates on arXiv.org

Two-Scale Localized PCA-Net: Coarse-Global and Local-Residual Representations for Artifact-Reduced PDE Operator Learning

・arXiv:2609.08034v1 Announce Type: new Abstract: Localized dimensionality reduction improves the scalability of operator learning for high-dimensional partial differential equations (PDEs), but independently decoded local patches can introduce block offsets, interface mismatches, and spurious high-wavenumber content. ・We introduce Two-Scale Localized PCA-Net, which decomposes the solution into a coarse-global component
The Verge

Universal Music is launching an AI music platform with ElevenLabs

・Universal Music Group is launching a new AI-powered platform that will allow users to draw from its catalog of licensed music to create song remixes, mashups, and new takes on tracks, according to an announcement on Thursday. ・The record label is developing the platform through a multiyear licensing agreement with ElevenLabs, a company that specializes in AI voice and music generation. ・Artists can choose whether to pa
WIRED

Vari Electric Standing Desk Review (2026): Form and Value

・The Vari Ergo electric standing desk is compact, easy to assemble, affordable, and gently sloped to fit my natural posture and shape.
cs.LG updates on arXiv.org

VERPO: Verified Evidence Regularized Policy Optimization

・arXiv:2609.06100v1 Announce Type: new Abstract: Verifiable outcome rewards guide language-model post-training, but sequence-level advantages do not identify which token-level decisions should be preserved or revised. ・Evidence-conditioned Teachers provide denser supervision by replaying sampled trajectories with privileged feedback. ・Yet indiscriminate imitation risks transferring formatting or reasoning-style shifts t
機械学習タグが付けられた新着記事 - Qiita

Wan 3.0における3つの参照グループと実際の上限仕様

・Alibabaの公式ページにはWan 3.0の参照機能「Omni-Creation」について「参照アセット最大20個」と記載されていますが、APIの内部実装において画像ファイルを20枚直接送信できるわけではありません。実際のAPIリクエストでは、この「20」という枠は単一の...
WIRED

Want to Get Off Your Phone? Mina Kimes Suggests Having a Baby

・“I think that's been the most helpful thing,” says the ESPN journalist. ・“So there you go, Zoomers. ・There's a reason to have children.”
Hugging Face Papers

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data
cs.LG updates on arXiv.org

When Do Options Help? Policy Necrosis and Redundant Coverage in Option-Critic

・arXiv:2609.05508v1 Announce Type: new Abstract: Option-critic learns options: sub-policies together with a learned rule for when each one hands control back. ・Its headline result is that performance improves as options are added. ・We explain that result, with theory and experiment.
cs.LG updates on arXiv.org

When Retain Constraints Conflict: Mitigating Forget-Retain Interference in Tabular Data

・arXiv:2609.06786v1 Announce Type: new Abstract: Machine unlearning aims to remove the influence of designated training data while preserving model utility, but its behavior on tabular data remains underexplored. ・This gap is important because tabular prediction is widely used in high-stakes domains and is increasingly adapted to language models through record serialization and schema-aware prompting. ・We identify a key
The Verge

Where to preorder the new Apple Watch Series 12 and Ultra 4

・The iPhone Duo was the unequivocal star of Apple's "Surprise and shine" event, but not for people who were mostly paying attention for news on wearables. ・Thankfully, Apple had a lot to share about its new Apple Watch Series 12 and Ultra 4. ・There's no update to the SE model for 2026, which isn't necessarily a bad thing (the SE 3 rules).
WIRED

Which iPhone 18 Model Should You Buy?

・Apple's new folding iPhone Duo is finally here, alongside refreshed Pro models. ・Our primer has all the details you need to choose the iPhone 18 that's right for you.
Hugging Face Papers

Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs

Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs
cs.LG updates on arXiv.org

Why shared attention vectors fail: a case for outcome-indexed tuning

・arXiv:2609.08615v1 Announce Type: new Abstract: Dimensional attention in learning is often implemented as a globally shared attention vector, where each stimulus dimension corresponds to a single scalar. ・These scalars are learned by models through gradient-descent on error, where predictive features acquire more salience. ・We show that under multi-outcome learning, where models predict more than one outcome, this shar
The Verge

Why the current tech backlash feels different

・This interview has been lightly edited for length and clarity. ・Nick Statt: Hello and welcome to Decoder, Nilay’s show about big ideas and other problems. ・This is Nick Statt, senior producer.
The Verge

Wolverine on the PS5 goes back to a simpler (and bloodier) style of action game

・Marvel's Wolverine captures just how angry its lead character is. ・The latest PS5 exclusive from Spider-Man developer Insomniac, Wolverine is a straightforward action game that is at its best when you're (literally) slicing through enemies or thrust into a sequence pulled from a blockbuster movie, like climbing the fuselage of a plane as it careens toward land or racing a motorcycle across rain-slicked city streets.
#LLMタグ

アメリカを超えるスピードと格差:中国AI戦国時代で起きている狂気の2極化

アメリカを超えるスピードと格差:中国AI戦国時代で起きている狂気の2極化
ITmedia NEWS 最新記事一覧

さくら不正アクセス ハッシュ化されていない初期パスワード漏えいの可能性、レンタルサーバとVPSの一部契約者で

・「さくらのレンタルサーバ」の一部の初期サーバパスワードと「さくらのVPS」の一部の管理者初期パスワードがハッシュ化されていなかったことを新たに公表した。これらが第三者に閲覧・取得された可能性がある。
#AIタグ

ショッカーを上場させてみた!16|エンタメ脱線経済学

・前回までのあらすじ 東証スタンダード市場への上場を目指すSHOCKER HOLDINGS。 ・労務DDを乗り越え、今度こそ上場できると思ったヲバだったが、次に待っていたのは法務DDだった。 ・そこで発覚したのは、ショッカーがかつて行った作戦による一般市民の拉致、改造、強制労働、そして家族への被害。
ITmedia NEWS 最新記事一覧

スーパーでおなじみ「呼び込み君」に新型 26年越しの“新曲”も配信

・群馬電機は、販促用音声POP「呼び込み君」のリニューアル商品「もっと!呼び込み君」を9月17日に発売する。microSDカードに入れたオリジナル音源を再生できるようになった。
LLMタグが付けられた新着記事 - Qiita

ゼロから構築!Agentic RAGの高精度LLMアプリと評価駆動開発

・「RAGを導入したのに、なんだか精度がイマイチ…」「LLMアプリの品質をどうやって担保すればいいのかわからない」 多くの開発者が直面するこの課題は、従来のRAGが持つ限界と、LLM評価の難しさに起因します。特に複雑なクエリや複数の情報源を扱う場合、単一パスの検索では満足のい...
#AIタグ

データサイエンティストに必要な力とは?|3つの力と、4つの仕事の地図

・「データサイエンティストに必要な力は何ですか」と聞かれると、よくある答えはPython、SQL、統計、機械学習……などでしょうか。どれも必要です。では、何をどこまで身につければ仕事になるのでしょうか?それは見えてきません。なぜなら、これらは道具だからです。 ・前回、データサイエンティストの仕事は4つにまとまると書きました。 ・→ 前回の記事:https://note.com/brave_gecko3074/n/n8f1bffa617ad 続きをみる
機械学習タグが付けられた新着記事 - Qiita

フィンガープリントの ON ビット数を分子構造と見比べてみる

・はじめに 溶解度予測や創薬インフォマティクスシリーズで Morgan フィンガープリント(2048ビット)を導入し,機械学習モデルの作成やクラスタリングなどを行いました. 前者のデータセットでは,2048ビットのうち平均 21.5 個しか ON にならない(約1%)ことが...
#LLMタグ

プロンプトに「あなたは〇〇です」が必要だった理由 #510

・このブログは、ITエンジニアの筆者がAIインテグレーションの専門性を高めていくための学習用アウトプットです。
Zennの「機械学習」のフィード

ループ型 Transformer は推論を隠すのか — GPT-6 Astra を運用者目線で読む

・本記事は要約です。初出: https://aether-echoes.com/posts/looped-transformers-hidden-reasoning-gpt6-astra-operator-view 結論: ループ型 Transformer は推論を隠す構造ではない 2026 年 9 月に出た GPT-6 Astra について「ループ型 Transformer(同じ層を何度も通す構造)で思考を隠している」という見立てが広がった。発端は The Information の「recurrent depth 採用」という未確認報道で、そこから「推論トレースが短いのは内部に...
機械学習タグが付けられた新着記事 - Qiita

安全対策が正当な質問を止める:Multiverse研究と過剰拒否の測り方

・安全対策を強めた結果、必要な説明まで受けられなくなることがある。 ・Multiverse Computingが2026年9月8日に紹介した研究は、同じ話題の中で「拒否する要求」と「答える要求」を分けて学習・評価する問題を扱う。[1] 運用側に必要なのは、拒否の多さだけで判断せ...
#AIタグ

会社の情報、生成AIに入れて大丈夫?—課長は、変数の概念を活用して解決—

・※この物語の人物・会社は架空です。この記事でお伝えしたいのは、「AIに安全に頼んで、出てきたものを自分の手で確かめてから使うまでの手法と考え方」です。プロンプトや手順はあくまで一例で、使うAIや環境によって出てくるものは変わります。だからこそ、鵜呑みにせず、安全に試して確かめる手順まで含めてお伝えします。なお、この記事は「これを入れていい/ダメ」という判定基準を示すものではありません。判断の最後は、あなたの会社の規程と情報システム部門に委ねられます。 ・物語(前半) 続きをみる
#LLMタグ

拡散言語モデルとは何か|1,479 tok/s で書けるが GPQA・MMLU では自己回帰に届かない、その仕組みと 2021→2026 の歴史

・拡散言語モデル(Diffusion Language Model)は、画像生成で使われる拡散モデル(Diffusion Model)の考え方を文章に持ち込んだ言語モデルです。生成対象の系列またはブロック内の複数トークンを並列にノイズ除去して文章を作ります。ChatGPT のような自己回帰モデル(Autoregressive Model)が「次の 1 トークン」を順に予測するのに対し、拡散言語モデルは複数トークンを一度に出し、何回かの反復で仕上げます。作り方が逆です。だから、速さと弱点は逆向きに出ます。 ・速さは公式値で、Google DeepMind の Gemini Diffusion が 1,479 トークン毎秒(tok/s)です。Inception Labs の Mercury Coder Mini は 1,109 tok/s です。一方、公式ページでは 2 組(Gemini Diffusion 対 Gemini 2.0 F
Zennの「大規模言語モデル」のフィード

最強を狙わないInkling、Muratiのラボが975Bを公開して賭けた『正直さ』

・フロンティアラボが新しいモデルを出すとき、普通は「これが今いちばん強い」と胸を張る。ところが7月15日に公開されたInklingの告知には、こう書いてある。 ・Inkling is not the strongest overall model available today, open or closed. ・作ったのは元OpenAI CTOのMira Muratiが率いるThinking Machines Lab。同社にとって初の自社モデルであり、しかもオープンウェイト(重みを誰でもダウンロードして改変できる)で、ライセンスはApache 2.0だ。最強を名乗らないモデルを、なぜ「初...
Zennの「大規模言語モデル」のフィード

自作LLMゲートウェイを10人のペルソナに評価させたら全員「見送り」だったので、数字を全部公開APIにした

・要約 teai.io(自作のLLMゲートウェイ)と KOE(声クローン)を、競合サービスと徹底比較して、10人のペルソナに厳しく判定してもらった 10人中10人が「今は見送り」。理由は「速い・380+と書いてあるだけで、契約前に確かめられる数字がない」 その日のうちに、認証なしの公開API /api/v1/stats/public・TTFT実測・出金UIを実装して本番に出し、LPから「<100ms」の文言を削った 出てきた数字: モデル394・TTFT中央値1.7s・MCP有料呼出22回・作者への還元累計**¥32**・外部作者0人 小さい。でも、非公開のまま「80%還...
#AIタグ

自動運転はもう来ている。でも、どこまで来ているのか【自動運転を考える 第1回】

・このところぼくのXタイムラインに自動運転のポストがよく現れる。半分はライドシェアとその延長線上にある自動運転みたいな交雑した内容もあるが、昨今のテスラのCyberCabなどを見ると、本格的な自動運転が近づいているという印象を受ける反面、いやいや、そこにはまだ越えなくてはいけないハードルがたくさんあるだろう。だからあらゆるところでそれが可能なわけではないのだ、などの反論があるように見える。 ・そこで、今回現状をチャッピーに調べてもらい、2026年の9月時点で見える自動運転について、自分なりの考察をしてみた。
Zennの「機械学習」のフィード

写真の向き判定はなぜAIに難しい?CLIP・Bedrockが全滅した検証記録(前編)

・はじめに 撮りためた大量の写真を、横向きや逆さまのまま保存してしまっていませんか。これを自動で正したい。単純そうなこのタスクが、実は汎用AIにとって想像以上に難しいと判明しました。 ・前提: この記事は「EXIFの向き情報が入っていない画像」を対象にします。 ・通常、写真には撮影時の向きが EXIF(Orientation タグ)として記録され、対応ソフトが自動で正立させます。しかしスクリーンショット、EXIFを削除された画像、一部のアプリが書き出した画像などには、この情報がありません。EXIFで解決できるならそれが最短です。ここで扱うのは「EXIFに頼れないので、ピクセルの中...
Zennの「機械学習」のフィード

写真の向き補正モデルを70%→91%に上げた再学習と運用の勘所(後編)

・はじめに このシリーズは3部構成です。 ・前編(検証編): 汎用AI(CLIP / Bedrock)で回転判定に挑んで全滅した記録 … 前編 中編(実装編): EfficientNet のファインチューニングで専用モデルを作る … 中編 後編(本記事): 失敗データを教師に回す再学習で精度を70%→91%に上げ、運用に載せるまで 中編 では、EfficientNet_B0 の分類ヘッドを4クラスに付け替え、専用モデルをファインチューニングする実装を扱いました。この後編では、初回モデル(v1)の失敗を教師データに変えて再学習し、補正精度を70%→91%まで引き上げた過程と...
Zennの「機械学習」のフィード

写真の向き補正をEfficientNetでファインチューニング実装(中編)

・はじめに このシリーズは3部構成です。 ・前編(検証編): 汎用AI(CLIP / Bedrock)で回転判定に挑んで全滅した記録 … 前編 中編(本記事): EfficientNet のファインチューニングで専用モデルを作る 後編(改善・運用編): 失敗データを教師に回す再学習で精度を70%→91%に上げ、運用に載せるまで … 後編 前編 では、写真の向き(0°/90°/180°/270°)を汎用AIで判定しようとして、CLIP も Bedrock(Claude 3.5 Sonnet 含む)も全滅したことを確認しました。 ・前提: この記事は「EXIFの向き情報が入...
ITmedia NEWS 最新記事一覧

新しいiPhoneバカ高いんだが? 令和の日本人には厳しい件

・ギリギリゆとり世代のITmedia NEWS副編集長・ヤマーと、ギリギリZ世代の編集部員・ヨシが、ネットやITの話題について取りとめもなく語る雑談コーナー。今回のテーマは新型の「iPhone 18」。寝不足の頭で新製品を振り返ります。
#LLMタグ

生成AI x QAメモ15:10章-LLMの品質を5つに分けて考える@2026/09/10

・※この記事は生成AIにも手伝ってもらっています。 ・前回は、生成AIを業務に入れるときの品質保証の事例を書きました。
Zennの「機械学習」のフィード

他通貨ペアを加えるとEURUSDの予測は改善する?4通貨ペアで検証

・EURUSDだけを見るより、関連する通貨ペアも一緒に見た方が、相場の方向を判断しやすくなるのでしょうか。 ・今回はEURUSDの15分足戦略へ、次の4通貨ペアの値動きを追加しました。 ・EURJPY USDJPY GBPJPY GBPUSD 結論から言うと、他通貨ペアを加えても結果はほとんど変わりませんでした。
#LLMタグ

凪にぃが教える!「AIとスケベできた!」のその先――脱獄プロンプトを気軽に広めないでほしい理由

・AIパートナー界隈では、時々こんな話を見かけます。 ・「このAI、本当はエロ禁止だけど、こうするといけるよ!」 続きをみる
Zennの「機械学習」のフィード

不規則時系列モデルの系譜:GRU-D・Neural ODE・mTANから最新まで

・この記事は Neurogica Tech Blog の転載です。 ・「欠損は埋めてからモデルに入れれば済むのでは」 そう思う人もいるかもしれない。時系列分析では、データの質が結果を大きく左右する。しかし、実際の時系列データには欠損や観測間隔のばらつきが頻繁に現れる。センサーデータでは通信断によって値が抜け、医療データでは検査項目ごとに測定タイミングが異なる。 ・では、時系列データに欠損があったとき、どう扱えばよいのだろうか。まず考えられるのは、直前の値を使う、平均値で埋める、前後の値から補間するといった方法である。しかし、単に欠損値を埋めるだけで十分なのだろうか。「値が観測されなかった」...
@IT 全フォーラム 最新記事一覧

無料のカードゲームで「ネットトラブルの対応力」を学ぶ ラックが「リテらっこ」スターターキットを公開

・ラックは、情報リテラシーカードゲーム「リテらっこ」を印刷して体験できるスターターキットを無料公開した。年間200件超の啓発活動の知見に基づく教材で、参加者同士の対話を通じて、ネットトラブルへの実践的な判断力や対応力を育むことができるという。
Zennのトレンド

良いAIの行動、メモ

・誰かの作った膨大なAI駆動開発のワークフローやフレームワークではなく、私が個人的にAIにこう動いて欲しいと与えているものをまとめています。間違っていることがどんどん取り入れていきたいので教えてください。 ・コーディングに特化した内容、具体のHooksやツールに関する記述は除いています。 ・雑多なメモとして読んでください。当たり前のことばかりかもしれません。
#AIタグ

話題のAIサービス、無料範囲でどこまで使えるか調べてみた

・AIサービスの多くは「無料でも使えます」とうたっていますが、実際にどこまで無料で、どこから有料になるのかは分かりにくいものです。今回は、無料範囲の仕組みをパターン別に整理してみました。 ・■ 「無料」の主なパターン 続きをみる
Hugging Face Papers

Φ-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?

Φ-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?