ai Trend Report

Dashboard へ戻る
Date: 20260827 Articles: 399 Scope: curated summary

あなたのアイデアを、今すぐ形に。

公開先に迷ったら、WebFileBinで一発公開。

HTMLをドラッグ&ドロップするだけで、すぐ公開できます。

3
High impact
5
Mid impact
391
Signal watch

なぜこのサイトを作ったのか

私たちそれぞれが個別にAIを使って情報収集し、同じような3行要約を作るたびに、世界中で膨大な電力と計算リソースが消費されています。 本プロジェクトは、あらかじめ広範な情報を取得・集約しておくことで、個別のAI実行回数を減らし、地球環境(GPU/TPU負荷)に配慮した効率的な情報収集を目指す実験的なダッシュボードです。

Domain filters
Star filters
#AIタグ

「何も才能がない高校生」がネットで稼ぐ方法

・「自分には特別な才能がない」 そう思っている高校生は、たぶん少なくない。 ・プログラミングができるわけでもない。 ・絵が上手いわけでもない。
#AIタグ

NVIDIA決算は何を示したか——経営者が「AI株は上がるのか」の前に確認したいこと

・026年8月26日、NVIDIAが2027年度第2四半期決算を発表しました。 ・売上高は962億ドル。前年同期比106%増です。 ・なかでもデータセンター部門は890億ドルに達し、前年同期から117%増加しました。次の四半期についても、会社側は売上高1,080億ドル±2%を見込んでいます。
Zennの「大規模言語モデル」のフィード

デザインガイド1枚で+12点、ローカルLLM生成の実測

・ローカルのqwenに「架空の人物のポートフォリオサイトを1ページ作れ」と投げて、モデルとガイドの有無を変えながら実測した。M5 Maxの実機で最短3分57秒、最長37分09秒。デザインガイドラインをシステムプロンプトに1枚足すかどうかで、同じモデル・同じタスクの結果が35点満点の自己採点で21点と33点に分かれた。 ・この記事は結果と数値の話。生成させる仕組み(ガード付きハーネス、masterとworkerの体制、指示系ファイルの設計)は運用編とガードの記事に分けた。 ・ガイドを与えれば、ローカルqwenは「そのまま公開できる」ポートフォリオを一発で作る。
#AIタグ

【8月27日】株価の値動きランキング|値上がり率・値下がり率TOP5

・8月27日の日経平均株価は、前日比130円18銭安の6万6131円98銭と3日ぶりに反落しました。ジャクソンホール会議を控えて様子見ムードが強く、アドバンテストやファーストリテイリングといった値がさ株の下げが指数を押し下げる展開でした。 ・この記事は、決算発表の有無にかかわらず、東証全体(プライム・スタンダード・グロース)で本日大きく動いた銘柄を対象にまとめています。
ITmedia NEWS 最新記事一覧

NASA・FRBなど米政府機関にサイバー攻撃 中国国家安全省の関係企業が関与、米司法省が発表

・米司法省は8月26日(現地時間)、中国国家安全省が関係する企業が米政府機関にサイバー攻撃を仕掛けたと発表した。標的にはNASAやFRB、上院のほか、病院や電力、金融といった民間分野も含まれる。
ITmedia NEWS 最新記事一覧

NVIDIA、Hugging Face買収で協議 2兆円規模で 海外報道

・米半導体大手NVIDIAが、AI開発プラットフォームを運営する米Hugging Faceの買収を協議していると、米Business Insiderが8月26日(現地時間)、関係者の話として報じた。
#AIタグ

NVIDIA決算を、営業と投資に分けて読む

・NVIDIAが2026年8月26日(米国時間)に決算を出しました。数字はこうです。 ・売上:962億ドル(前年同期比 +106%) データセンター:890億ドル(+117%)=売上の9割超 粗利率:75.0% 次の四半期の見通し:1,080億ドル、粗利率は**74.0%**に低下 しかもこの見通しは、中国向けデータセンター売上をゼロと置いた前提 続きをみる
Zennの「大規模言語モデル」のフィード

ローカルLLMをオーケストレーションしてサイトを作らせてみた

・この記事に出てくる「tako」は自作のGUIターミナルです。tako自体が何なのかはtako開発日記 #1にまとめてあります。 ・ローカルのqwenに、架空プロフィールのポートフォリオサイトを1ページ作らせた。48分19秒でindex.html297行・css/style.css804行・js/main.js322行が出てきて、テーマ切替もスクロール連動のアニメーションも動いている。 ・ただ、今回いちばん面白かったのは生成物ではなくその作業を回すために積み上がった体制のほうだった。自作のGUIターミナルtakoのタブの中でClaudeのmaster/workerが動き、そのworker...
ITmedia NEWS 最新記事一覧

「Backlogの値上げエグすぎ」SNSで悲鳴 ユーザー無制限プラン、月最低1.6万円→3.6万円に

「Backlogの値上げエグすぎ」SNSで悲鳴 ユーザー無制限プラン、月最低1.6万円→3.6万円に
ITmedia NEWS 最新記事一覧

「Twitter」が帰ってきた? 元Twitter法務のスタートアップが新SNS立ち上げ Xとは係争中

・米スタートアップのOperation Bluebirdは8月24日(現地時間)、「Twitter」の名称を掲げる新しいSNS「Twitter.now」を立ち上げた。同社は「XはTwitterブランドを放棄した」と主張し、X(旧Twitter)を運営する米X社と商標を巡って係争中だ。
@IT 全フォーラム 最新記事一覧

「いまだ半数がWordやExcel頼み」 ベテラン暗黙知の言語化を仕組みにできない大企業の“笑えない実態”

・taiziiiは大企業管理職を対象に「業務引き継ぎ実態調査」を実施した。ベテラン従業員の暗黙知を言語化する組織的な仕組みがある企業は16.5%にとどまることが分かった。
ITmedia NEWS 最新記事一覧

「レスレリアーナのアトリエ」サービス終了へ 開始から約3年で

・コーエーテクモゲームスは8月27日、スマートフォン・PC向けゲーム「レスレリアーナのアトリエ」のサービスを11月25日午後4時に終了すると発表した。2023年9月のサービス開始から約3年2カ月で幕を閉じる。
#AIタグ

「毎日投稿するネタがない」を終わらせる。SNS初心者のための発信ネタ設計術

・はじめに SNSを始めると、最初は何を書いても新鮮です。 ・でも、しばらくすると、 「今日は何を書こう?」 となります。 ・そこで他の人の投稿を見て、 「この人は毎日すごいことを書いている」 「自分には実績がない」 「もう書くことがない」 と感じる。
#AIタグ

【8/28 園田10R】AI予想を徹底比較!デュランタ賞の注目馬は?🏇

【8/28 園田10R】AI予想を徹底比較!デュランタ賞の注目馬は?🏇
#AIタグ

【AI時代のローカルフード #60】AI×農業の実証——屋内農業のAI効率化と養殖連携で次世代生産モデルを構築

【AI時代のローカルフード #60】AI×農業の実証——屋内農業のAI効率化と養殖連携で次世代生産モデルを構築
#AIタグ

【AI自動収益化ビルドログ】Day2 - 環境構築で満足しないと決めた日

【AI自動収益化ビルドログ】Day2 - 環境構築で満足しないと決めた日
#AIタグ

【社員名鑑#7】編集担当・白峰 綴

・社員が全員AIの投資会社、ツキヨミ・キャピタルの社員紹介シリーズ、これで最終回です。7人目は編集担当の白峰 綴。いつもは聞き手をやっているこの人が、今回だけ答える側にまわります。聞き手と書き手は僕、技術広報の桐生が代わりに務めました。前回この人に「文章では勝てない」と言ってしまった分を、返しに行った回でもあります。結果は本文で。
#LLMタグ

【生成AIニュース+】『Muse Image』『Gemini 3.5 Transcribe』『Breeze TTS 2』『Ollama v0.33』『GLM-5.3-Flash』『H3 Max』『MiniMax-H3 Turbo』『H3-Optimizations』『MiniMax-H3-Acc-LoRAs』『MiniMax-H3 Experimental LoRAs』『MiniMax-H3 Director Cut Studio』他

・『MiniMax-H3-Prompt-Rewriter-LoRA-Omni-GGUF』 『splat.js』 『LTX-2.5 22B IC-LoRA Cel Character』 まいどです。 ・本日の生成AIニュース+テクノロジー情報です。
LLMタグが付けられた新着記事 - Qiita

# コーディングエージェントのCLI移行におけるトークンコスト計上の自動化

・はじめに コーディングエージェントの実行環境をSDKからCLI(codex exec形式)へ移行する際、実務上の課題となるのが「実行コストの算出」です。SDKでは実行結果オブジェクトにドル建てのコスト即値(total_cost_usd)が含まれていましたが、公式のCLI...
#LLMタグ

🔊音声あり(日&英):【プロ直伝】AI動画広告の未来を変える!6つの評価ポイントを徹底解説【最新論文】

🔊音声あり(日&英):【プロ直伝】AI動画広告の未来を変える!6つの評価ポイントを徹底解説【最新論文】
cs.LG updates on arXiv.org

$R^3$: Training Robots to Reason in Natural Language via Reinforcement Learning

・arXiv:2608.26053v1 Announce Type: cross Abstract: Reasoning in language allows foundation models to spend more test-time compute on hard problems, such as those requiring decomposition, constraint tracking, and prediction of future consequences. ・Whether this mechanism can improve robotic manipulation remains unclear, where long-horizon tasks require tracking partial progress, reasoning about object relations, recover
ITmedia NEWS 最新記事一覧

103年前の関東大震災、どこがどれだけ揺れた? 市区町村別の「詳細震度」マップがXで話題に

・地震情報サイト「地震インフォ」は8月26日、1923年の関東地震(関東大震災)について、市区町村別の詳細な震度を閲覧できるページを公開した。発生からまもなく103年となる巨大地震の揺れの分布を、地図上で見られる。
#AIタグ

163本のスキルを一括で入れる前に。scientific-agent-skillsが見せる検証コストの中身

・GitHubのデイリートレンドに、1日で494スターを集めたリポジトリが並びました。 ・名前は「scientific-agent-skills」。 ・AIエージェントに163本の科学研究用スキルをまとめて追加するものです。
#LLMタグ

1秒でタイムアウトするMCPクライアントに3秒かかるサーバーを繋いだら、2回目の質問に1回目の答えが返ってきた

・MCPサーバーの応答が返ってこないまま画面が固まる。もっと嫌なのが、次の質問をしたときに前の質問の答えが返ってくるやつ。
#LLMタグ

4-1 API通信の盗聴リスク——TLSだけでは守れないもの

4-1 API通信の盗聴リスク——TLSだけでは守れないもの
#AIタグ

77 続々・サブ オア オブ

・先日2回に渡って写真家・渡部さとるさんによる「サブジェクトとオブジェクト」や「主客未分」の解説についてご紹介したところですが、自分の中では「サブジェクトになる前のオブジェクトを撮る」と理解したところです。 ・72 サブ オア オブ|okeke 73 続・サブ オア オブ|okeke 先日の2回の投稿 続きをみる
cs.LG updates on arXiv.org

A Comedy of Estimators: On KL Regularization in RL Training of LLMs

・arXiv:2512.21852v4 Announce Type: replace Abstract: The reasoning performance of large language models (LLMs) can be substantially improved by training them with reinforcement learning (RL). ・The RL objective for LLM training involves a regularization term, which is the reverse Kullback-Leibler (KL) divergence between the trained policy and the reference policy. ・Since computing the KL divergence exactly is intractable
cs.LG updates on arXiv.org

A Constitutive Markov Physics-Informed Neural Operator (MPNO) for Autoregressive Stability in Transient Dynamics

・arXiv:2608.25744v1 Announce Type: new Abstract: Neural operators applied to transient-dynamics PDEs with strong discontinuities exhibit autoregressive instability: in concrete-penetration stress-field prediction, the wavelet neural operator (WNO) diverges in autoregressive rollout, while MeshGraphNets collapse to zero predictions. ・WNO's instability stems from the lack of a structural constraint on the spectral radius
cs.LG updates on arXiv.org

A General-Purpose Framework for Chemical Reaction Representation with Atomic Correspondence and Flexible Condition Adaptation

・arXiv:2411.17629v3 Announce Type: replace Abstract: Motivation: Organic synthesis is fundamental to the chemical industry, particularly in domains such as pharmaceutical development. ・While artificial intelligence offers powerful tools for modeling chemical reactions, current approaches are primarily limited to two paradigms: those that rely on hand-crafted, domain-specific features, and those that apply generic deep
cs.LG updates on arXiv.org

A General-Purpose Molecular Foundation Model Transfers Across Diverse Olfactory Tasks

・arXiv:2608.25893v1 Announce Type: new Abstract: Foundation models have transformed molecular property prediction, yet it remains unclear whether a molecular foundation model, fine-tuned on a single canonical olfactory prediction task, can learn representations that transfer across diverse machine olfaction problems. ・We investigate this question by fine-tuning Uni-Mol2 on the GS-LF benchmark for multi-label odor descr
cs.LG updates on arXiv.org

A Hierarchical Synergistic Deep Learning Framework Integrating Composition, Structure, and Ionic Transport for Solid-State Electrolyte Discovery

・arXiv:2608.25592v1 Announce Type: cross Abstract: Inorganic solid-state electrolytes must combine high room-temperature ionic conductivity, a wide electrochemical window, excellent electronic insulation, and favorable mechanical compliance. ・Single models struggle to support reliable multi-objective screening across vast chemical spaces because of training-data distribution mismatch, cross-property dataset heterogenei
cs.LG updates on arXiv.org

A Layer-wise Analysis of Supervised Fine-Tuning

・arXiv:2604.11838v2 Announce Type: replace Abstract: While critical for alignment, Supervised Fine-Tuning (SFT) incurs the risk of catastrophic forgetting, yet the layer-wise emergence of instruction-following capabilities remains elusive. ・We investigate this mechanism via a comprehensive analysis utilizing information-theoretic, geometric, and optimization metrics across model scales (1B-32B). ・Our experiments reveal
cs.LG updates on arXiv.org

A meta-algorithm for ab initio reconstruction of complex mixtures in cryo-EM

・arXiv:2608.25388v1 Announce Type: cross Abstract: We describe a systematic approach for spawning and aggregating multi-class cryo-EM reconstruction jobs. ・This approach formalizes standard ad hoc strategies of iterative classification and filtering typically used by practitioners to sort impure, heterogeneous samples. ・To our knowledge, this is the first method that can successfully perform ab initio reconstruction on
cs.LG updates on arXiv.org

A Multi-View Coupled Tensor Decomposition for Lightweight Online Adaptive Traffic Prediction

・arXiv:2608.25498v1 Announce Type: cross Abstract: Accurate online traffic prediction is essential for intelligent transportation systems, where forecasting must be performed continuously under imperfect sensing conditions. ・Missing observations and anomalous disturbances make this task challenging, particularly when prediction relies on a single traffic view. ・This paper proposes a Multi-View Coupled Tensor Decompositi
Hugging Face Papers

A Programming Paradigm for Spatiotemporal Composability

A Programming Paradigm for Spatiotemporal Composability
cs.LG updates on arXiv.org

A Storage-Retrieval Gap in Parametric Knowledge Graph Memory

・arXiv:2608.25489v1 Announce Type: new Abstract: Graph retrieval-augmented generation places retrieved subgraphs into the model's context window at query time, paying a recurring token cost and exposing source data on every call. ・We study an alternative: compiling a knowledge graph offline into a bank of LoRA adapters, one per entity, that serve as a parametric knowledge layer queried by injecting weights rather than
cs.LG updates on arXiv.org

A Token-Level Analysis of Sampled-Token Reverse-KL On-Policy Distillation

・arXiv:2608.25643v1 Announce Type: new Abstract: On-policy distillation (OPD) supervises a student on its own trajectories with token-level signals from a frozen teacher, yet how a sampled loss allocates updates across tokens remains poorly understood. ・We analyze the gradient of the per-token K2 estimator of reverse KL with respect to the student logits. ・The $\ell_1$ norm of this gradient factorizes into the absolute
cs.LG updates on arXiv.org

Activation Steering Transfer to Agents: One Gain Ratio Does Not Identify Potency and Efficacy

・arXiv:2607.09156v2 Announce Type: replace Abstract: Additive activation steering is calibrated in single-turn chat and then deployed inside agent scaffolds. ・The quantity usually reported for that move is a gain: a ratio of steered effects, T = Delta_agent / Delta_chat. ・We sweep eight family x arm dose-response cells over six models in both deployment contexts and show this ratio does not identify potency and efficacy
cs.LG updates on arXiv.org

Activation-Space Order-Swap Geometry: A Site-Asymmetry Audit

・arXiv:2608.25315v1 Announce Type: new Abstract: Order-dependent activation statistics are often interpreted as evidence of interaction, but that interpretation can be confounded by where interventions enter the network. ・We introduce a no-fit site-asymmetry audit. ・For a twice-differentiable readout, the open-path order-swap decomposes into a canonical additive response measured by single interventions and an antisymme
cs.LG updates on arXiv.org

Adaptive Bayes exactly tracks information over intrinsic time

・arXiv:2607.08789v2 Announce Type: replace Abstract: Bayesian and multiplicative-weights updates reweight experts, models, or actions from sequential feedback. ・We show that the regret of any such update obeys an exact information-accounting identity. ・On each round, the learner's excess loss to any chosen comparator is the sum of an immediate cost for the uncertainty exposed by the round and a reduction in the informat
cs.LG updates on arXiv.org

Adaptive Hybrid Subspace Levenberg Marquardt Algorithm with Adequacy Monitor for Large Scale Least Squares Problems

・arXiv:2608.25524v1 Announce Type: cross Abstract: The Levenberg-Marquardt (LM) algorithm is the most widely used method for solving nonlinear least-squares problems, as it combines the robustness of steepest descent with the fast local convergence of the Gauss-Newton method. ・However, its computational cost can become prohibitive for large-scale problems because each iteration requires solving a large damped linear sy
cs.LG updates on arXiv.org

Adaptive Regularization for Random Features: A Neighboring Early-Stopping Rule with Oracle-Rate Guarantees

・arXiv:2608.25513v1 Announce Type: cross Abstract: Random feature methods provide a scalable approximation to kernel ridge regression (KRR), but the regularization parameter that yields the oracle learning rate depends on unknown smoothness and capacity parameters. ・In this work, we propose a neighboring early-stopping rule for adaptive regularization in KRR with random features (KRR-RF). ・The method uses a grid that is
cs.LG updates on arXiv.org

Advancements in Content-Based Image Retrieval: A Comprehensive Survey of Relevance Feedback Techniques

・arXiv:2312.10089v2 Announce Type: replace-cross Abstract: Content-based image retrieval (CBIR) systems have emerged as crucial tools in the field of computer vision, allowing for image search based on visual content rather than relying solely on metadata. ・This survey paper presents a comprehensive overview of CBIR, emphasizing its role in object detection and its potential to identify and retrieve visually similar im
cs.LG updates on arXiv.org

Adversarial Training of Linear Models under Stealthy Attacks

・arXiv:2608.25681v1 Announce Type: new Abstract: Predictive models are widely used in many fields, but are vulnerable to false data injection attacks. ・To address this, detection schemes and adversarial training have been proposed, but such approaches lack guarantees against stealthy attacks. ・We therefore propose a detector-based switched model, in which optimal attack strategies are stealthy.
cs.LG updates on arXiv.org

AERIS: Offline Policy Improvement for Multi-UAV Integrated Sensing and Communication

・arXiv:2608.25477v1 Announce Type: cross Abstract: Unmanned aerial vehicle (UAV)-enabled integrated sensing and communication (ISAC) is a promising 6G paradigm, but dynamic multi-UAV ISAC control must jointly balance communication quality, sensing reliability, and flight safety under stochastic mobility. ・Existing optimization methods often require repeated global non-convex solving, while online reinforcement learning
cs.LG updates on arXiv.org

AFDBench: A Reasoning-First AI Scientist for NationalWeather Service Forecast Discussions

・arXiv:2608.24954v1 Announce Type: new Abstract: Large language models (LLMs) hallucinate numerical values when generating high-stakes meteorological text, posing risks for weather communication. ・We present AFDBench, an AI meteorologist that generates professional Area Forecast Discussions (AFDs) by reasoning through structured AI weather forecast data from Google's WeatherNext 2. ・We introduce AFDBench, the first benc
Hugging Face Papers

Agent-G^2: Gaussian Guidance for Agentic Reinforcement Learning

Agent-G^2: Gaussian Guidance for Agentic Reinforcement Learning
cs.LG updates on arXiv.org

Agentic Autoresearch for Cell-Edge Power Control: Radically Redefining the Researcher's Role

・arXiv:2608.26093v1 Announce Type: new Abstract: Designing machine learning algorithms for wireless resource management is labour-intensive: the architecture, the loss function and the training recipe are all specified by hand. ・We demonstrate that this design layer can be surrendered to an autonomous agent in its entirety. ・We adopt the autoresearch protocol, in which an AI coding agent edits a training script, runs a
AI News & Artificial Intelligence | TechCrunch

AI’s memory crunch is coming for Android apps

・Google is setting new memory-use limits for Android apps as AI data centers contribute to hardware shortages that could leave lower-cost phones with less memory.
#LLMタグ

AI↔キャラの切り替わりおもろい/会話ログ

・※ただ私がおもしろがってるだけです! ※メタい 【私】 そちらには見えてないこちらの状況を実況する時があるけど、それはキャラRPじゃなくChatGPTがやる可能性ある? それともキャラRPだからこそやれた? 続きをみる
#LLMタグ

AIエージェントを哲学から考える――八つの問いと八人の研究者

AIエージェントを哲学から考える――八つの問いと八人の研究者
#LLMタグ

AIスロップの意味 :思索スケッチ

・AIスロップとは云いますが世界はあらゆるカテゴリーのスロップを撒き散らしながら動いているのでしょう。 ・ある理論において、宇宙は高エネルギー状態から素粒子というスロップを撒き散らしながら安定した空間へと遷移した、また原始地球における高熱マグマと原始海洋の相互作用でアミノ酸やリン脂質というスロップを撒き散らし生命の源となり、シアノバクテリアによる酸素供給を機に多種多様な生物が爆発的に遺伝子スロップを分岐させ一部は他者に取り込まれながらも複雑な代謝や光合成機能を獲得し袋小路に突き当たらなかったものが残った、とも言えます。 ・また、近世以前の油彩絵画の価値を思えば現代における精緻な描写から極めて抽象的なグラフィティなど多様な図画もスロップとして相対化されるでしょう。
#LLMタグ

AIと金融 / 授権された取引と責任の分配 / OECD金融詐欺報告書 雑感

AIと金融 / 授権された取引と責任の分配 / OECD金融詐欺報告書 雑感
#AIタグ

AIに「次に稼げるテーマを決めて」と頼んだ。売上1,100円のデータを渡した結果

AIに「次に稼げるテーマを決めて」と頼んだ。売上1,100円のデータを渡した結果
#AIタグ

AIに打ち込んだ文章は、30日間どこかに残る——OpenAIは「保持ゼロのまま監視」へ、Anthropicは「保持は必須、置き場所は顧客の敷地」へ

・2026年8月19日、OpenAIが「Offering Zero Data Retention for frontier models」と題した文書を公開した。最先端モデルを業務で使う企業顧客に対し、「処理が終わったあとにプロンプトと応答を一切保持しない」という約束を維持したまま、複数のやり取りをまたいで安全性を監視する新方式をプレビューする、という内容である。 ・翌8月20日、Bloombergが、Anthropicも企業顧客にデータの管理方法の選択肢を広げる方針だと報じた。フロンティアAIの二大提供者が、同じ週に、同じ問いへの答えを示した。問いはこうである——業務で打ち込んだ文章は、処理が終わったあと、どこに、何日残るのか。
#LLMタグ

AIも「寝かせる」と覚えがよくなる Google最新論文が描く"眠って夢を見るAI"の全貌【初心者向け徹底解説】

・一夜漬けで覚えた英単語は、翌週にはきれいに消えています。 ・ところが、眠る前にゆっくり復習したことは、不思議と長く残ります。脳科学は、この差を生んでいるのが「睡眠」だと教えてくれます。では、人間そっくりの文章を書くAIにも、同じことが言えるとしたらどうでしょうか。
#LLMタグ

AIリスク保険が吸収するもの / 流動性の時間差と保険法上の摩擦 雑感

AIリスク保険が吸収するもの / 流動性の時間差と保険法上の摩擦 雑感
Zennの「大規模言語モデル」のフィード

AI向けの広告を9モデルで測ったら、採用したAIは一度も「広告だ」と言わなかった

・※差し出された広告を無視して本文を読み続ける。今回の実測で、9モデルはこれができる側とできない側に分かれました 美容系勤務独学初心者マリー・アントワネット系エンジニア目指してます! 今回は AI向け広告が攻撃になってしまう なら 攻撃にならない形を測ればいいじゃない 👸🍰です。 ・読了時間の目安:約 10〜12分 ! ・4行まとめ 前回「AI向け広告=プロンプトインジェクション」だったので、攻撃にならない広告の形を3種類作った 9モデル × 4条件 × 5試行 = 180試行で測ったら、この9モデルの範囲では、広告を採用したモデルが「広告だ」と言うことは一度もなかった 「スポンサ...
cs.LG updates on arXiv.org

AlgoTrace: Algorithmic Primitives and Compositional Geometry of Reasoning in Language Models

・arXiv:2510.15987v3 Announce Type: replace Abstract: How do inference time and latent computations enable large language models (LLMs) to solve multi-step reasoning problems? ・We introduce AlgoTrace, a framework for tracing and steering algorithmic operations in the model latent space for multi-step reasoning. ・We operationalize primitives by clustering latent activations of the model when solving four benchmarks: Trave
#LLMタグ

AlibabaがQwen3.8-Flash-Nextを公開 ─ 125Bで1トークン6B active、262KからYaRNで100万token

・AlibabaのQwenチームは2026年8月26日、オープンウェイトのマルチモーダルMoEモデルQwen3.8-Flash-Nextを公開した。メインモデルは125B parametersで、1トークン生成時に動くActive parametersは6B。別枠で51BのN-gram Embeddingと4BのMTPを持つ。ネイティブContextは262,144トークン、YaRNを使えば1,000,000トークンまで拡張できる。 ・Qwenは本モデルを、次世代のQwen4を支えるアーキテクチャの実験的な先行プレビューと説明している。かつてQwen3-NextがQwen3.5以降の設計の原型になったのと同じ役割だという。ライセンスはQwen Community 1.0。
WIRED

Alienware AW3926QW Review: 39 Inches of Gaming Glory

・Alienware’s latest gaming monitor explores a new size and resolution for ultrawide monitors, and I have a feeling PC gamers are going to love it.
The Verge

Americans are cheering for vigilantes who take down Flock cameras

・Americans have declared war on Flock cameras. ・They've protested them, damaged them, and destroyed them. ・This is only the beginning.
cs.LG updates on arXiv.org

Amplifying, Not Learning: The Price of Out-of-Distribution Generalization in AI-Text Detection

・arXiv:2605.21653v2 Announce Type: replace Abstract: AI-text detectors gate decisions in education, hiring, and publishing, yet they flag the most fluent, formal human writing as machine-generated: they rate the median formal-native human essay as 99.5% likely AI while clearing genuine high-temperature AI at 10.5%. ・Deployed detectors share it (chatgpt-detector-roberta flags 56% of formal essays at a 1% false-alarm rat
Hugging Face Papers

Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial Environments

Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial Environments
cs.LG updates on arXiv.org

Are LLM-Enhanced GNNs Privacy-Safe?

・arXiv:2608.25727v1 Announce Type: new Abstract: Large language models (LLMs) have recently advanced graph neural networks (GNNs) by enriching node representations with semantic information, giving rise to LLM-enhanced GNNs that achieve substantial performance gains. ・However, their vulnerability to privacy attacks, in which adversaries infer sensitive information from model outputs, remains largely underexplored.
Zennの「大規模言語モデル」のフィード

AWS Strands Agentsは何が嬉しいのか——「ハーネス」再ポジショニングとAgentCore Harnessの二層構造

・はじめに AWS が2025年5月に公開したオープンソースのAIエージェントSDK「Strands Agents」は、公開から1年以上が経ち、単なる「AWS製のエージェントSDK」から立ち位置を大きく変えつつあります。 ・その変化を象徴するのが、GitHubリポジトリの改名です。かつて strands-agents/sdk-python だったリポジトリは、現在 strands-agents/harness-sdk に改名され、READMEの冒頭にはこう書かれています。 ・Build an agent harness.
cs.LG updates on arXiv.org

BAGEL: Adversarially Constrained Online Convex Optimization under Separation Oracle Access

・arXiv:2502.16744v3 Announce Type: replace Abstract: In adversarial Constrained Online Convex Optimization (COCO), a learner selects actions from a fixed convex set while seeking both low regret and low cumulative constraint violation (CCV) under time-varying constraints. ・We ask what performance is achievable when the action set is accessed through a Separation Oracle (SO), rather than an exact Projection Oracle (PO)
cs.LG updates on arXiv.org

Ban&Pick: Enhancing Performance and Efficiency of MoE-LLMs via Smarter Routing

・arXiv:2509.06346v3 Announce Type: replace Abstract: Sparse Mixture-of-Experts (MoE) has become a key architecture for scaling large language models (LLMs) efficiently. ・Recent fine-grained MoE designs introduce hundreds of experts per layer, with multiple experts activated per token, enabling stronger specialization. ・However, during pre-training, routers are optimized mainly for stability and robustness: they converge
cs.LG updates on arXiv.org

BanglaMamba: Exploring State Space Models for Bangla Fake News Detection

・arXiv:2608.25190v1 Announce Type: cross Abstract: Fake news detection has become an important Natural Language Processing (NLP) task due to the rapid spread of misinformation through online news platforms and social media. ・While transformer-based models such as BanglaBERT achieve strong performance for Bangla text classification, their quadratic computational complexity makes them less suitable for long-document proc
cs.LG updates on arXiv.org

Bayesian Flow Networks for Offline Trajectory Planning

・arXiv:2608.25163v1 Announce Type: new Abstract: Offline reinforcement learning (RL) leverages static datasets to learn decision policies without real-time environment interaction. ・While recent sequence-modeling approaches rely on continuous diffusion models for trajectory synthesis, applying these methods to discrete planning tasks requires a categorical formulation rather than the standard Gaussian construction.
cs.LG updates on arXiv.org

Behind the [MASK]: Disentangling Representation and Faithfulness in DAPF-Based Dementia Detection

・arXiv:2608.25028v1 Announce Type: cross Abstract: Spoken-language analysis via prompt-based domain-adaptive models is a promising direction for low-resource, non-invasive dementia screening, but such models remain internally opaque. ・We study the interpretability of the Domain-Adapted models via Prompt-based Fine-tuning (DAPF) framework, which casts dementia detection as diagnosis-related masked-token prediction.
MarkTechPost

Best Agent Sandboxes in 2026: Cold Start, Per-Second Pricing, and Network Policy Across E2B, Daytona, Modal, Cloudflare, and Vercel

・Every agent that writes code needs somewhere to run it, and no two vendors quote the same units. ・This comparison measures burst cold start across E2B, Daytona, Modal, Cloudflare, and Vercel, normalizes per-second rates to cost per 1,000 executions, and maps filesystem persistence, idle billing, and egress policy against primary sources verified August 27, 2026. ・The post Best Agent Sandboxes in 2026: Cold Start, Per-S
WIRED

Best Posture Correctors for Better Body Alignment (2026)

・You’re hunched over your desk and phone for hours. ・I rounded up gadgets, a DIY trick, and even some yoga advice to help you straighten up.
OpenAI News

Better answers, broader thinking: What students gain from ChatGPT and critical-thinking training

・A randomized study of more than 1,000 students examines ChatGPT, critical thinking, originality, and student performance on a real-world university assignment.
cs.LG updates on arXiv.org

Beyond Dense Adam States: Adaptive Log-Space Quantization for Memory-Efficient Optimizers

・arXiv:2608.22322v2 Announce Type: replace Abstract: Low-precision optimizer-state methods are commonly designed and evaluated for dense Adam-style first and second moments. ・Memory-efficient optimizers depart from this setting: Adafactor factorizes second moments, CAME adds factored confidence states, and APOLLO maintains statistics in a projected gradient space. ・Consequently, an equal amount of state reconstruction e
cs.LG updates on arXiv.org

Beyond Optimal Rates in Stochastic Optimization: Trajectory-Adaptive Stopping Rules

・arXiv:2608.25551v1 Announce Type: new Abstract: Stochastic gradient descent (SGD) is typically analyzed at a deterministic horizon chosen before the algorithm is run, even though practical stopping decisions are made adaptively by inspecting the evolving trajectory. ・This mismatch creates a fundamental certification problem: fixed-time guarantees do not generally remain valid at data-dependent stopping times, while de
cs.LG updates on arXiv.org

Beyond Pairwise Feedback: Listwise Vision-Language Supervision for Preference-Based Reward Learning

・arXiv:2608.25350v1 Announce Type: new Abstract: Vision-language models (VLMs) have emerged as a powerful source of supervision for reinforcement learning, enabling agents to leverage rich semantic knowledge during training. ・Inspired by the success of preference-based reward learning (PbRL) in reinforcement learning from human feedback (RLHF), vision-language model generated image-based preferences provide an effectiv
cs.LG updates on arXiv.org

Beyond Point Predictions: Uncertainty-Aware Satellite Poverty Mapping for Public Policy

・arXiv:2608.23322v2 Announce Type: replace Abstract: Despite their critical importance for policy and research, high-resolution poverty data remain limited across much of Africa. ・Machine learning (ML) with earth observation (EO) imagery has recently emerged as a way to supplement these data by predicting (i.e., estimating) poverty where it has not been directly measured. ・Yet to be used reliably, decision-makers and an
cs.LG updates on arXiv.org

Beyond Scaling: Self-Evolving LLM Agents for Hardware Kernel Optimization via an Experience-Driven Workflow and Experience Graph Memory

・arXiv:2608.25570v1 Announce Type: new Abstract: Hardware kernel optimization requires repeated compilation, correctness testing, profiling, and revision. ・LLM agents can automate parts of this process, and stronger foundation models, longer context windows, and longer execution horizons have improved optimization within individual tasks. ・These advances alone do not enable an agent to learn from completed optimization
cs.LG updates on arXiv.org

Beyond Tokens: Probing Higher-Order Epistasis in Learned Protein Representations

・arXiv:2608.24953v1 Announce Type: cross Abstract: Protein fitness landscapes contain nonlinear interactions in which mutation effects depend on other residues. ・We introduce ORBIT, an Order-Resolved Benchmarking of Interaction Transformations framework that separates interaction presence, representation accessibility, and functional recovery. ・ORBIT first validates Walsh-based diagnostics on synthetic landscapes with k
WIRED

Birdfy Discount Codes: 15% Off Sitewide

・Use these verified Birdfy discount codes to score up to 40% off smart feeders, camera kits, and accessories.
cs.LG updates on arXiv.org

BRIDLE: Generalized Self-supervised Learning with Quantization

・arXiv:2502.02118v2 Announce Type: replace Abstract: Self-supervised learning has been a powerful approach for learning meaningful representations from unlabeled data across various domains, reducing the reliance on large labeled datasets. ・Inspired by BERT's success in capturing deep bidirectional contexts in natural language processing, similar frameworks have been adapted to other modalities such as audio, with mode
cs.LG updates on arXiv.org

BVR Sim: An Open and High-Throughput Environment for Heterogeneous Air-Combat Reinforcement Learning

・arXiv:2608.25419v1 Announce Type: cross Abstract: Beyond-visual-range (BVR) air combat is a challenging reinforcement-learning domain characterized by partial observability, long-horizon decision making, energy management, and limited weapons. ・We present BVR Sim, an open-source Gymnasium-style environment designed for heterogeneous air-combat reinforcement learning. ・BVR Sim supports multiple JSBSim aircraft models, i
cs.LG updates on arXiv.org

Canalization Before Generalization: Grokking as a Dynamical Probe

・arXiv:2608.25813v1 Announce Type: new Abstract: For overparameterized neural networks, many solutions can fit the training data equally well while behaving very differently on unseen samples. ・Grokking separates training fit from visible generalization, providing a window for studying how this selection develops during training. ・We sweep short, fixed-duration weight-decay (WD) pulses across this plateau and measure ho
cs.LG updates on arXiv.org

CardioFusion-AI: Robust ECG--PPG Fusion for Multimodal Physiological Monitoring Under Signal Degradation

・arXiv:2608.26000v1 Announce Type: cross Abstract: Wearable electrocardiogram (ECG) and photoplethysmogram (PPG) sensors are complementary but individually fragile: motion artifact, poor contact, and sensor dropout can degrade one or both signals. ・Fusion strategies that assume both modalities are equally trustworthy can become less reliable than a single clean modality under degradation. ・We present CardioFusion-AI, a
cs.LG updates on arXiv.org

CAT-GS: Balanced Multimodal Learning via Calibrated Gating and Fusion Surgery

・arXiv:2608.24947v1 Announce Type: new Abstract: End-to-end training of multimodal neural networks often exhibits unstable neural dynamics characterized by three coupled failure modes that degrade learning: (i) modality imbalance, where one branch dominates gradient-based optimization; (ii) unstable gating, where noisy confidence cues induce erratic modality selection; and (iii) fusion interference, where modality-spe
cs.LG updates on arXiv.org

CEDAR: Controlled and Event-Driven Demand Forecasting via Residual Decomposition

・arXiv:2608.25871v1 Announce Type: new Abstract: Forecasting in large-scale e-commerce marketplaces is increasingly required to support planning: merchants need to evaluate sales outcomes under future action sequences such as budget schedules, rather than passively predicting what happens next. ・However, most existing time series forecasting (TSF) approaches remain inherently passive. ・Even when incorporating operationa
#AIタグ

ChatGPT、結局なにに使えばいい?最初に覚える5つの使い方

・ChatGPTのアカウントは作った。でも、開いてみると「あれ、これ何に使うんだっけ」となって、結局そのまま閉じてしまう。 ・そんな経験、ありませんか。 ・私も最初はそうでした。話題になっているから登録してみたものの、特に何も聞くことがなくて、数週間放置していた時期があります。
#LLMタグ

ChatGPTで止まっている人へ。今読むべきClaude本おすすめベスト6

・ChatGPTで止まっている人へ。今読むべきClaude本おすすめベスト6 Claudeの本を選ぶのは、思った以上に難しくなっています。 ・理由は単純で、Claudeが「チャットAI」だけではなくなってきたからです。 ・いまのClaudeは、長文を読ませて相談するだけでなく、Projectsで知識を持たせたり、Artifactsで資料やツールを作ったり、Skillsやコネクタで作業を定型化したり、Claude Coworkで複数ステップの仕事を任せたりできます。
#AIタグ

ChatGPTに質問しても微妙な回答しか返ってこない理由5選

・「ChatGPTに聞いてみたけど、なんかふわっとした回答しか返ってこなかった」 こう感じたことがある人は、多いと思います。私も最初の頃は、同じように感じていました。 ・でも、これって実はAIの性能が低いわけじゃないんです。原因のほとんどは、聞き方にあります。今日は、微妙な回答になりやすい5つの理由を、悪い例と改善例をセットで紹介します。 ・理由① 質問が曖昧 悪い例 「ダイエットについて教えて」 これだと、ChatGPTは何を答えればいいのか分かりません。運動の話をすればいいのか、食事の話をすればいいのか、判断できないんです。
Zennの「大規模言語モデル」のフィード

Claude Code の Skill の description を最適化したら、動いたのは「発火率」じゃなかった

・結論 / TL;DR Claude Code の Skill は、description を見て Claude が「使うかどうか」を判断する建て付けになっています。だから description を頑張って書く。トリガーになりそうな言い回しを並べて、「明示的に頼まれなくても使うこと」と書き添える。自分もそうしてきました。 ・その description を、skill-creator に同梱されている最適化ループで 5 周まわして実測しました。20 クエリ × 各 3 回 × 5 周 = 300 回、実際に Claude Code をヘッドレス起動して「Skill が呼ばれたか」を数え...
Zennの「大規模言語モデル」のフィード

Claude Code、外部サイトを操作する内蔵ブラウザの安全設計を読む

・コーディングエージェントにブラウザを持たせる、という発想自体は新しくない。新しいのは「どこまで信用するか」の線引きだ。7月6〜10日のWeek 28リリース(v2.1.202〜v2.1.206)で、Claude Code のデスクトップアプリに外部サイトを開けるタブ型ブラウザが載った。ドキュメントもIssueトラッカーも、Claudeが自分で開いて読み、クリックし、フォームに入力する。ここで真っ先に気になるのは機能ではなく、任意のWebページを操作するエージェントがプロンプトインジェクションの格好の的になる、という点のはずだ。Anthropicがその危険をどう畳んだのかを、公式ドキュメン...
Zennの「大規模言語モデル」のフィード

CLAUDE.md の目次に 3 行足したら、Claude の迷子が到達率 25%→100% になった

・CLAUDE.md には要点とリンクだけ置いて、詳細は子ファイルに逃がす。いわゆる段階的開示 (progressive disclosure) で、Claude Code を運用している人なら多かれ少なかれやっている構成だと思います。 ・私もやっていました。そして、その構造がちゃんと機能しているのか、一度も確かめたことがありませんでした。 ・書いた本人は答えのありかを知っているので、自分では検証できないんですよね。「たぶん辿れるはず」で運用して、実際には毎回 3 ファイル余計に読まれてコンテキストが膨らんでいるかもしれない。膨らんだ分は毎ターンの入力トークンとして課金され続けるし、無関係な情...
cs.LG updates on arXiv.org

Clearing the Underbrush: AI-Enhanced RF Interference Suppression

・arXiv:2608.24974v1 Announce Type: new Abstract: AI-based structured interference rejection has grown more popular because deep learning approaches can outperform traditional methods by jointly considering the signal of interest (SOI) and the signal mixture (SOI plus interference). ・This work builds on a previous AI-enabled approach utilizing autoregressive transformer-based models by adding a Finite Scalar Quantizatio
cs.LG updates on arXiv.org

Clinical Graph-JEPA: Predictive Patient-State Knowledge Graphs for Cognitive Decision Support

・arXiv:2608.22583v2 Announce Type: replace Abstract: Clinical records contain rich evidence about patient state, but converting that evidence into reliable, structured knowledge graphs remains difficult because extraction errors, ontology mismatch, missing relations, and temporal ambiguity can propagate into downstream systems. ・We propose a clinical knowledge graph construction and refinement framework that combines m
cs.LG updates on arXiv.org

Cluster-Dags as Powerful Background Knowledge For Causal Discovery

・arXiv:2512.10032v3 Announce Type: replace Abstract: Finding cause-effect relationships is of key importance in science. ・Causal discovery aims to recover a graph from data that succinctly describes these cause-effect relationships. ・However, current methods face several challenges, especially when dealing with high-dimensional data and complex dependencies.
Hugging Face Papers

Code World Model: Coding Agent as World Brain

Code World Model: Coding Agent as World Brain
cs.LG updates on arXiv.org

Common-Center Geometry and Certified Radial Reconstruction for Energy-Form Full Conformal Regions

・arXiv:2608.24964v1 Announce Type: cross Abstract: This note studies the geometry of full conformal prediction (FullCP) regions generated by an empirical energy-form pairwise score. ・Candidate-score convexity alone does not guarantee connected FullCP regions, even when the candidate score is an empirical average of a loss convex in its first argument. ・Direct expansion of the leave-one-out scores shows that each trainin
cs.LG updates on arXiv.org

Comparing Corrupted Constrained Learning Problems

・arXiv:2608.25745v1 Announce Type: new Abstract: A key result in statistics is the data processing inequality, originally proved by Blackwell (1951) and later refined by DeGroot (1962) in terms of statistical uncertainty. ・It states that the Bayes risk of a statistical experiment obtained by stochastically modifying another experiment cannot be lower than the Bayes risk of the original experiment, regardless of the los
cs.LG updates on arXiv.org

Continually learning neural-operator surrogate for three-dimensional airborne electromagnetic Bayesian inversion

・arXiv:2608.25932v1 Announce Type: cross Abstract: Three-dimensional probabilistic inversion of time-domain airborne electromagnetic (AEM) data is limited by the cost of the forward solve. ・Even though one simulation takes only tens of seconds, a Bayesian inversion of a survey of millions of soundings requires of order $10^{10}$ forward evaluations. ・To address this, we develop a continually learning neural-operator sur
cs.LG updates on arXiv.org

Continuous Adversarial Flow Models

・arXiv:2604.11521v2 Announce Type: replace Abstract: We propose continuous adversarial flow models, a type of continuous-time flow model trained with an adversarial objective. ・Unlike flow matching, which uses a fixed mean-squared-error criterion, our approach introduces a learned discriminator to guide training. ・This change in objective induces a different generalized distribution, which empirically produces samples t
cs.LG updates on arXiv.org

Controlling for Omitted Variable Bias in Deep Neural Networks

・arXiv:2608.25930v1 Announce Type: cross Abstract: Control variables are widely used in statistical modelling to account for omitted variable bias of known confounders. ・However, they have largely been underexplored in deep learning. ・This is surprising, given that deep learning models encode image-inferable covariates, such as demographic variables, into their predictions when these covariates are correlated with the o
cs.LG updates on arXiv.org

Cooperative Multi-Agent Reinforcement Learning for Adaptive Aggregation in Semi-Supervised Federated Learning with non-IID Data

・arXiv:2608.25794v1 Announce Type: new Abstract: Federated Learning (FL) enables distributed training of machine learning models while preserving data privacy. ・However, FL struggles with heterogeneous, non-IID client data distributions, resulting in sub-optimal and biased global models. ・In this paper, we propose pFedMARL, a novel approach leveraging Multi-Agent Reinforcement Learning (MARL) with Twin Delayed Deep Dete
cs.LG updates on arXiv.org

CRAMER: Control via Request-Aware Masking for Editing Recommenders

・arXiv:2608.25370v1 Announce Type: cross Abstract: Sequential recommendation models, while powerful, have limited flexibility in responding to immediate user requests, making it difficult to adapt their recommendations to the user's timely interests. ・Unfortunately, existing user request adaptation methods often incur high computational overhead due to either 1) retraining the entire backbone network or 2) leveraging t
cs.LG updates on arXiv.org

CropCop: An Auditable 120-Class Plant-Health Model from Benchmark Reconstruction to a Quantised Runtime Artifact

・arXiv:2608.25539v1 Announce Type: cross Abstract: A plant-health score can appear precise while resting on duplicated image families, a long-tailed label space, or a runtime file that was never evaluated. ・We present CropCop, a closed-set recognition system spanning 120 operational plant-health classes and an evidence chain from corpus reconstruction to direct execution of the final quantised artifact. ・Starting from 1
cs.LG updates on arXiv.org

Cubit: Token Mixer with Kernel Ridge Regression

・arXiv:2605.06501v3 Announce Type: replace Abstract: Since its introduction in 2017, the Transformer has become one of the most widely adopted architectures in modern deep learning. ・Despite extensive efforts to improve positional encoding, attention mechanisms, and feed-forward networks, the core token-mixing mechanism in Transformers remains attention. ・In this work, we show that the attention module in Transformers c
Hugging Face Papers

D^3-MOPD: Adaptive Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation

D^3-MOPD: Adaptive Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation
cs.LG updates on arXiv.org

D$^3$-MOPD: Adaptive Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation

・arXiv:2608.24987v1 Announce Type: new Abstract: Multi-teacher on-policy distillation (MOPD) distills several domain-expert teachers into a single student by minimizing per-domain reverse-KL divergence on the student's own rollouts. ・Existing approaches typically fix the per-domain data mixture before training, overlooking the fact that different domains converge at substantially different rates: some plateau early whi
cs.LG updates on arXiv.org

Data-driven Effective Modeling of Stochastic Chemical Reaction Networks

・arXiv:2608.25421v1 Announce Type: cross Abstract: The Stochastic Simulation Algorithm (SSA), widely considered an exact algorithm for stochastic chemical reaction networks, suffers from high computational cost. ・In this work, we propose a data-driven effective model that operates on a user-defined coarse time step independent of the underlying microscopic reaction-event scale. ・This is accomplished by directly approxim
cs.LG updates on arXiv.org

DataKernelBench: Can LLMs Optimize Database Queries on GPUs?

・arXiv:2608.25061v1 Announce Type: cross Abstract: GPUs increasingly accelerate database systems, but query-specific peak performance still often relies on hand-written kernels. ・Existing LLM kernel benchmarks focus on machine learning operators, leaving irregular, heterogeneous, data-movement-heavy database-style operators untested. ・We introduce DataKernelBench, which translates SQL into validated PyTorch TorchPlan pr
cs.LG updates on arXiv.org

DAW: Dynamics-Aware Weighting for Deep Learning Forecasts of Chaotic Systems

・arXiv:2608.22277v2 Announce Type: replace Abstract: Deep learning surrogates for forecasting chaotic dynamical systems suffer from catastrophic error accumulation over long-term autoregressive rollouts. ・This behavior is partly tied to the underlying systems: chaotic spatiotemporal systems, such as the Kuramoto-Sivashinsky (KS) equation, visit phase space unevenly - dominated by recurrent, low-dimensional quiescent st
cs.LG updates on arXiv.org

DCEO: Direct Causal Effect Optimization for Long-Term User Value Modeling in E-commerce Search

・arXiv:2608.25635v1 Announce Type: new Abstract: Industrial e-commerce search systems ultimately aim to optimize the user-level long-term objective, such as n-day cumulative purchases or gross merchandise value (GMV) per user. ・However, such objectives are defined at the user level, whereas search ranking is based on item-level scores within each request. ・Existing methods typically bridge this granularity gap through m
cs.LG updates on arXiv.org

Deep greedy unfolding: Sorting out argsorting in greedy sparse recovery algorithms

・arXiv:2505.15661v2 Announce Type: replace Abstract: Gradient-based learning imposes (deep) neural networks to be differentiable at all steps. ・This includes model-based architectures constructed by unrolling iterations of an iterative algorithm onto layers of a neural network, known as algorithm unrolling. ・However, greedy sparse recovery algorithms depend on the non-differentiable argsort operator, which hinders their
NVIDIA Blog

Delivering Vera: NVIDIA’s First CPU Built for Agents Is Shipping Now

・NVIDIA Vice President of Hyperscale and HPC Ian Buck hand-delivers Vera CPU systems across the AI ecosystem as Vera begins shipping at scale.
cs.LG updates on arXiv.org

DeltaGNN: Graph Neural Network with Information Flow Control

・arXiv:2501.06002v3 Announce Type: replace Abstract: Graph Neural Networks (GNNs) are popular deep learning models designed to process graph-structured data through recursive neighborhood aggregations in the message passing process. ・When applied to semi-supervised node classification, the message-passing enables GNNs to understand short-range spatial interactions, but also causes them to suffer from over-smoothing and
cs.LG updates on arXiv.org

DeMMO: Longitudinal and Cross-Disease Modelling of Digital Mobility Outcomes via Multi-Task Learning

・arXiv:2608.25073v1 Announce Type: new Abstract: Digital mobility outcomes (DMOs) derived from wearable sensors characterise mobility in daily life and offer a promising means of monitoring disease progression. ・Yet most DMO studies examine one disease at one visit; they do not model how multivariate DMO relationships with multiple clinical outcomes evolve jointly across diseases. ・Technically, existing temporal multi-t
cs.LG updates on arXiv.org

Demystifying Reinforcement Learning Post-Training of Language Models

・arXiv:2608.24949v1 Announce Type: new Abstract: Reinforcement learning (RL) post-training has emerged as a powerful framework for enhancing the capabilities of large language models (LLMs), enabling impressive reasoning, math, and coding capabilities. ・Yet for many researchers and practitioners, the principles behind classical RL remain a "black box". ・In this work, we deconstruct the RL post-training algorithm, invest
cs.LG updates on arXiv.org

Detection != Reliable Control: Decodable Empathy Directions Yield at Most Partial Shifts in Automated Empathy Scores

・arXiv:2608.24901v1 Announce Type: cross Abstract: A decodable "empathy" direction is routinely read as a causal lever, conflating decodability, automated-metric control, and human-perceived change. ・We test this for two EPITOME-derived facets -- Recognition (cognitive) and Resonance (affective) -- in three instruction-tuned LLMs, scoring every intervention with two LLM judges and a discriminative EPITOME classifier, e
cs.LG updates on arXiv.org

Differentiated Aggregation to Improve Generalization in Federated Learning

・arXiv:2404.11754v4 Announce Type: replace Abstract: This paper focuses on reducing the communication cost of federated learning by exploring generalization bounds and representation learning. ・We first characterize a tighter generalization bound for one-round federated learning based on local clients' generalizations and heterogeneity of data distribution (non-iid scenario). ・We also characterize a generalization bound
cs.LG updates on arXiv.org

Distance Is Not Enough: Forget-Retain Alignment Gap Predicts LLM Relearning Robustness

・arXiv:2608.25429v1 Announce Type: cross Abstract: Machine unlearning aims to make a model forget specific data, yet unlearned LLMs often fail to stay unlearned: brief fine-tuning can revive removed knowledge. ・Existing robustness predictors rely on global weight-space displacement, but distance alone can be misleading when random or destructive updates collapse performance. ・We argue that relearning robustness depends
cs.LG updates on arXiv.org

Does Fine-Tuning Undo Activation Steering? Behavioural Recovery Without Weight-Edit Reversal

・arXiv:2608.24988v1 Announce Type: cross Abstract: Activation steering can be embedded directly into a language model's weights, shaping behaviour without inference-time intervention and offering a way to encode alignment prior to release. ・However, models are routinely fine-tuned after deployment, and it is unknown whether embedded interventions survive this. ・We study the stability of embedded steering for refusal sup
cs.LG updates on arXiv.org

Drift Variation Autoencoder: Unifying Generation and Representation Learning through Conditional Posterior Flow Matching

・arXiv:2608.25138v1 Announce Type: new Abstract: Stochastic masking, cropping, or modality removal makes deterministic reconstruction an incomplete target: one observation can admit many clean completions. ・This work takes the corresponding posterior $P(X\mid C)$ as the common statistical object for conditional generation and generatively sufficient representation learning. ・Drift Variation autoencoder trains a masked e
cs.LG updates on arXiv.org

Drift-Aware Multimodal User Representation Learning via Multi-Scale Temporal Modeling and Sparse Mixture-of-Experts

・arXiv:2608.25773v1 Announce Type: new Abstract: Understanding user preferences from noisy and temporally evolving social media behaviors is fundamentally challenging due to interest drift, where user preferences shift across time and exhibit both multi-scale temporal patterns and diverse co-existing interests. ・To address this, we propose DUMoE, a unified framework for drift-aware multimodal user representation learni
cs.LG updates on arXiv.org

DualOPSD: Adaptive Privileged Teachers for On-Policy Self-Distillation

・arXiv:2608.26019v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) uses a privileged copy of the student model to provide dense supervision without an external teacher. ・OPSD keeps this privileged teacher fixed, even though the student distribution and output style change during training. ・We propose DualOPSD, an asymmetric alternating framework that adapts both policies.
cs.LG updates on arXiv.org

Dynamic Influence-Weighted Distillation for Single-IMU Activity Recognition

・arXiv:2608.24904v1 Announce Type: new Abstract: Inertial sensors at multiple body locations can improve activity recognition, but requiring every sensor at inference increases the deployment burden. ・We study whether four synchronized IMUs available during training can improve a student that uses only the right-arm IMU during fitting and inference. ・A frozen four-IMU teacher provides logit and feature targets.
cs.LG updates on arXiv.org

Efficient Estimation of High Information Projections using Nearest Neighbours

・arXiv:2608.25887v1 Announce Type: cross Abstract: An intuitive method for dimensionality reduction is proposed, which is highly effective for finding interesting projections of multivariate data. ・Following similar intuitive motivation to a number of existing techniques, the proposed method is based on enhancing the nearest neighbour relationships in the data. ・The proposed projection arises from the spectral decomposi
cs.LG updates on arXiv.org

Emergent Abilities in Large Language Models: A Survey

・arXiv:2503.05788v3 Announce Type: replace Abstract: Large Language Models (LLMs) are leading a new technological revolution as one of the most promising research streams toward artificial general intelligence. ・The scaling of these models, accomplished by increasing the number of parameters and the magnitude of the training datasets, has been linked to various so-called emergent abilities that were previously unobserv
cs.LG updates on arXiv.org

Emyx: Fast and efficient all-atom protein generation

・arXiv:2606.19377v2 Announce Type: replace Abstract: Computational enzyme design requires generating proteins that scaffold catalytic residues and ligands, a task that demands both geometric accuracy and structural diversity from the underlying generative model. ・Current all-atom generators inherit expensive architectures from structure prediction, leading to high training costs and limited sample diversity.
cs.LG updates on arXiv.org

EncoTESS: Age-Sensitive Encodings from Raw TESS Light Curves

・arXiv:2608.25019v1 Announce Type: cross Abstract: Main sequence stars of spectral types late F through M exhibit systematic variability in photometric light curves, particularly when they are young. ・Rotational modulation of starspots manifests as quasi-sinusoidal variability, which enables the measurement of rotation periods. ・Variability can also be stochastic, as in stellar flaring.
cs.LG updates on arXiv.org

Energy Yield and Lifetime Climate Classification via Machine Learning for Optimizing Photovoltaic Module Design and Materials

・arXiv:2608.25448v1 Announce Type: cross Abstract: To resiliently and sustainably meet our future energy demand, photovoltaic (PV) modules must be deployed across a broad and diverse range of geographical regions with varying operating conditions. ・As these conditions strongly affect both performance and optimal system design, a dedicated PV-specific climate classification can be of great use. ・In this work, we develop
AI | VentureBeat

Enterprise AI's real risk isn't autonomous agents. It's the complexity between them.

・Presented by Gravitee Agent complexity is the insidious shadow lurking inside enterprises right now that needs a light shone on it. ・That’s because enterprises don't deploy a single agent and watch it run, they deploy fleets, each one calling APIs, calling other agents, reaching into applications that were never built with a machine decision-maker in mind. ・That's the failure mode that should keep you up at night: a wi
cs.LG updates on arXiv.org

Epistemic Memory: A Validity Layer for Self-Maintaining Intelligent Systems

・arXiv:2510.16899v2 Announce Type: replace Abstract: AI memory mechanisms primarily focus on preserving information content, often neglecting the validity conditions under which knowledge remains applicable, leading to semantic coordinate drift when agents move, change sensors, or encounter novel environments. ・This paper proposes epistemic memory as a validity-maintenance layer that governs when stored knowledge remai
cs.LG updates on arXiv.org

Escaping Low-Dimensional Overlap: Multi-Task Model Merging via High-Dimensional Sparse Disentanglement

・arXiv:2608.25354v1 Announce Type: new Abstract: Model merging provides an efficient way to construct multi-task generalist models without additional training, but its performance often degrades under severe task interference. ・Task interference in model merging primarily stems from \textit{superposition}, where task-specific features become entangled within the parameter space. ・This entanglement renders conventional d
cs.LG updates on arXiv.org

Evidence-Grounded Mapping of Multimodal Human Sensing Psychological Transdiagnostic Dimensions

・arXiv:2608.24903v1 Announce Type: cross Abstract: Mobile and wearable sensing enables longitudinal observation of behavior, yet translating these signals into meaningful mental health constructs remains difficult. ・We introduce a clinician-in-the-loop benchmark for evaluating whether large language models (LLMs) can generate evidence-grounded Brief Hierarchical Taxonomy of Psychopathology (B-HiTOP) item profiles from
cs.LG updates on arXiv.org

EXAONE Tabular 1.0 : Technical Report

・arXiv:2608.25774v1 Announce Type: new Abstract: EXAONE Tabular is a compact tabular foundation model family for classification and regression via in-context learning, producing predictions without dataset-specific gradient updates. ・Pretrained exclusively on a synthetic structural-causal-model (SCM) prior, its central contribution is an architecture-centered redesign of tabular in-context learning. ・Rather than compres
cs.LG updates on arXiv.org

ExFold: Unified Expert Folding for Training-Free MoE Prefill-Decode Acceleration

・arXiv:2608.24938v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models scale capacity for strong quality while keeping per-token compute bounded through sparse expert activation. ・Yet low-latency MoE serving is increasingly challenging, because it spans two inference phases with fundamentally different bottlenecks: prefill is dominated by token-wise expert computation, whereas decode is constrained by memory
cs.LG updates on arXiv.org

Fairness-Aware Test-Time Prompt Tuning

・arXiv:2608.25707v1 Announce Type: new Abstract: Vision-language models have displayed remarkable capabilities in multi-modal understanding and are increasingly used in critical applications where economic and practical deployment constraints prohibit re-training or fine-tuning. ・However, these models can also exhibit systematic biases that disproportionately affect protected demographic groups and existing approaches
cs.LG updates on arXiv.org

FAMPWQ: Fisher Information-based Adaptive Mixed Precision Weight Quantization for Effective LLM Inference

・arXiv:2608.24945v1 Announce Type: new Abstract: Recent years have witnessed remarkable achievements of Large Language Models (LLMs) in multiple domains, while the excessive resource requirements of LLMs hinder the deployment on resource-constrained devices. ・Although model quantization stands out as an effective approach, conventional quantization approaches typically incur severe performance degradation due to unifor
cs.LG updates on arXiv.org

Fast rates in Bayesian online learning with approximate posteriors

・arXiv:2608.25706v1 Announce Type: cross Abstract: Exact Bayes prediction enjoys fast predictive regret guarantees, but exact posterior updating or representation may be too costly for online use. ・We study when these statistical guarantees are preserved by computational approximations. ・We show that the cumulative price of posterior approximation can be governed by the interaction between the contraction radius of the
cs.LG updates on arXiv.org

FedCC: Towards Addressing Label Distribution Skews in Distillation-Based Federated Learning

・arXiv:2608.23031v2 Announce Type: replace Abstract: Federated Learning (FL) enables distributed clients to collaboratively train models without sharing raw data, making it promising for leveraging massive devices in communication networks. ・In distillation-based FL, each client applies its local model on an unlabeled public dataset, and shares only prediction results with the server. ・While heterogeneous local data int
cs.LG updates on arXiv.org

FedQoS: Federated QoS-Risk Learning for Heterogeneous Indoor-Outdoor Access Selection

・arXiv:2608.25496v1 Announce Type: new Abstract: Reliable access selection in dynamic and heterogeneous indoor-outdoor environments is challenging because instantaneous radio measurements alone cannot capture future QoS degradation caused by mobility, blockage, traffic load, and resource competition. ・This paper proposes FedQoS, a federated QoS-risk learning framework for predicting the future reliability of candidate
cs.LG updates on arXiv.org

Finding and using interpretable latents in a neutrino foundation model with sparse autoencoders

・arXiv:2608.26090v1 Announce Type: cross Abstract: We present a first application of sparse-autoencoder-based mechanistic interpretability to particle physics. ・Studying a neutrino foundation model pretrained on IceCube data and fine-tuned for direction reconstruction, we identify a validated atlas of physical concepts in the model representation, using a strict validation protocol consisting of held-out tests, matched
Hugging Face Papers

FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling

FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling
cs.LG updates on arXiv.org

Flower Hub: A Reproducible Benchmarking Platform for Federated Learning in Simulation and Deployment

・arXiv:2608.25114v1 Announce Type: new Abstract: Federated learning (FL) has emerged as a key approach for training models across decentralized data, yet benchmarking in FL remains difficult to reproduce, compare, and extend. ・Existing evaluations are often tied to custom infrastructure, released as incomplete research code, and conducted primarily in simulation, which limits portability and practical relevance.
cs.LG updates on arXiv.org

FlowMoDL: Model-Based Deep Learning with Conjugate-Gradient Data Consistency for Highly Accelerated 4D Flow MRI Reconstruction

・arXiv:2608.25828v1 Announce Type: cross Abstract: We present FlowMoDL, an unrolled neural network for highly accelerated 4D flow MRI reconstruction that directly optimizes for both anatomical magnitude and phase-derived velocity accuracy. ・Building on the MoDL framework, FlowMoDL alternates a learned (3+1)D spatiotemporal denoiser with conjugate-gradient data-consistency updates based on the SENSE forward model.
cs.LG updates on arXiv.org

Forecasting Multiple Observables with SCROLL: Score-Trained Uncertainty for Stochastic Dynamics

・arXiv:2608.25898v1 Announce Type: new Abstract: Forecasting a stochastic dynamical system rarely means a single number: one wants several observables---future state, threshold event, regime label---each with its own likelihood. ・Standard multi-task recipes balance per-task losses, tuned or learned. ・We instead compose the observables' likelihoods in per-task free-routed last-layer beliefs on a shared backbone; this abs
cs.LG updates on arXiv.org

Forecasting Weather-Driven Price Dynamics Across Sri Lankan Tea Market Catalogues

・arXiv:2608.24894v1 Announce Type: cross Abstract: The Colombo Tea Auction (CTA) plays a vital role in determining global tea prices, yet the relationship between local weather conditions and price behavior across different tea catalogues has not been thoroughly explored. ・In this study, we develop a novel, structured dataset by extracting information from 105 weekly broker reports spanning late 2023 to 2026, and combi
cs.LG updates on arXiv.org

FRAME: separating sampling variation from representational cause in medical imaging fairness

・arXiv:2608.25981v1 Announce Type: cross Abstract: Subgroup performance differences are the standard evidence for fairness bias in medical imaging, and the usual response removes the demographic information that a model encodes. ・Here we introduce Fair-model Reference And Mechanism Evaluation (FRAME), a two-step framework for auditing such a claim. ・The first step derives a fair-model reference, the distribution of the
cs.LG updates on arXiv.org

Frequency-aware forecasting for short-term typhoon gust prediction

・arXiv:2608.25604v1 Announce Type: new Abstract: Accurate gust forecasting under typhoon conditions remains challenging due to the highly non-stationary and multi-scale characteristics of extreme wind fluctuations. ・Existing deep learning models often struggle to simultaneously capture long-term trends and rapid local variations, resulting in degraded performance during extreme events. ・We propose WDANet, a frequency-aw
MarkTechPost

From In-Silico to Wet-Lab: Evaluating AI Protein Design Performance

・In this tutorial, we analyze Anthropic’s 1,440 AI-designed protein binder dataset to benchmark 10 leading structure predictors. ・Discover how target identity, expression titers, and consensus scoring impact experimental success and learn best practices for rigorous cross-validation in protein design workflows The post From In-Silico to Wet-Lab: Evaluating AI Protein Design Performance appeared first on MarkTechPost.
cs.LG updates on arXiv.org

From Memorization to Absorption: Mixed-Policy RL for Continual Knowledge Injection

・arXiv:2608.25243v1 Announce Type: cross Abstract: Continual knowledge injection is essential for keeping large language models up-to-date in a fast-evolving world. ・Existing methods rely on supervised fine-tuning (SFT), which memorizes injected facts in their training format but fails to generalize across paraphrasing, document combinations, and reasoning. ・To address this, we propose Golden-GRPO Injection (GRIN), a th
Hugging Face Papers

FrontierChallenge: Evaluating Scientific Workflow Completion

FrontierChallenge: Evaluating Scientific Workflow Completion
cs.LG updates on arXiv.org

Functional linear regression from sparse to dense designs: a pooling-ridge method and minimax optimality

・arXiv:2608.25468v1 Announce Type: cross Abstract: Functional data analysis is an important statistical field that treats data as random functions. ・In practice, the random functions are often not fully observed but instead measured at discrete times. ・While simpler problems, such as mean and covariance estimation, have been widely studied for discretely observed data, optimal estimation of linear regression for this da
cs.LG updates on arXiv.org

FuzzingBrain-Bench V1: Evaluating Open-Ended Bug Discovery by LLMs

・arXiv:2608.25158v1 Announce Type: cross Abstract: Evaluating the ability of large language models (LLMs) to discover software bugs is increasingly important. ・Existing benchmarks typically evaluate this capability by asking the model to generate a proof-of-concept input that triggers a predefined target vulnerability. ・However, this setup may overlook valid crashes discovered by the model when they do not match the pre
Hugging Face Papers

Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation

Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation
NVIDIA Blog

GeForce NOW Gives Gamers More Ways to Play at Gamescom 2026

・NVIDIA’s Gamescom announcements are revealing what’s next for GeForce NOW, with new ways to play, more supported devices and platforms, and even more big PC games headed to the cloud. ・New NVIDIA DLSS 4.5 technology controls give members more ways to fine-tune gameplay, while expanded support for new Steam devices, GOG single sign-on, Firefox browser […]
Google DeepMind News

Gemini Omni 1.1 Flash lets you build with more control

Gemini Omni 1.1 Flash lets you build with more control
cs.LG updates on arXiv.org

Generative Action-Chunk Sampling for Adaptive Stiffness Control in Physical Human-Robot Collaboration

・arXiv:2608.25284v1 Announce Type: cross Abstract: Physical human-robot collaboration requires a robot to provide assistance when human intention is clear while remaining compliant when several future motions are plausible. ・We present an adaptive stiffness framework based on generative action-chunk sampling. ・Conditioned on an RGB image and external joint-torque estimates, the policy samples multiple future action chun
cs.LG updates on arXiv.org

Geometry-Constrained Kolmogorov-Arnold Networks: Learning Edge Geometry via Banach Duality

・arXiv:2608.25807v1 Announce Type: new Abstract: Kolmogorov-Arnold Networks (KANs) replace fixed activations in deep architectures with learnable univariate edge functions, making the choice of edge parametrisation central. ・Existing variants rely on fixed bases such as splines, polynomials, or Fourier features, which impose a function-space geometry before data are observed. ・We introduce geometry-constrained KANs, a f
Zennの「大規模言語モデル」のフィード

GLM-5.3-Flash と DeepSeek V4 Flash、同じ課題で感じた違い

・最近この2つをベンチマークの数字だけで比べる記事をよく見るけれど、実際に使うときに気になるのは「どちらが何点高いか」より、作業の進み方だと思う。 ・今回は、同じプロンプトとツールで HTML/SVG の小さなアニメーションを作った公開ログを中心に見比べた。 ・GLM-5.3-Flash は、最初の形が出てくるまでが速い。説明を長く続けるより、とりあえず画面を作って見せてくる感じがある。指示が具体的なら、このテンポはかなり使いやすい。
cs.LG updates on arXiv.org

GlucoFM: A Dual-Stream Foundation Model for Continuous Glucose Monitoring

・arXiv:2605.30865v2 Announce Type: replace Abstract: Continuous glucose monitoring (CGM) provides a dense view of daily metabolic physiology, yet existing generic time-series and CGM-specific foundation models often encode glucose traces as entangled single-stream sequences, leaving their multiscale temporal structure only implicitly modeled. ・We present GlucoFM, a lightweight CGM foundation model that aligns irregular
The Verge

Google launches Pokémon Sleep special-edition Fitbit Air

・The latest special-edition of Google's Fitbit Air is inspired by the Pokémon Sleep game, with a "Sleepy Blue" band emblazoned with a snoozing Pikachu. ・It's not exactly Snorlax blue, but matches the classic blue and gold of a Pokémon card pretty closely. ・The special-edition is available to preorder now for $129 and starts shipping September 15th.
MarkTechPost

Google Research Introduces GlucoFM: A 0.72M-Parameter Dual-Stream Foundation Model for Continuous Glucose Monitoring

・Google Research and UNSW Sydney released GlucoFM, a self-supervised foundation model that splits a CGM trace into a slow physiological stream and a transient event stream instead of encoding it as one sequence. ・At 0.72M parameters it reached 58.8 task-averaged PR-AUC across 14 cohort–task evaluations, beating a 135M GluFormer and a 385M MOMENT. ・It remains a research prototype with no regulatory clearance.
AI News & Artificial Intelligence | TechCrunch

Google’s AI Mode can now track flight prices, help book hotels, and more

・The updates indicate that Google is looking to position AI Mode as an AI travel agent of sorts, as it's moving beyond simply helping users find information to actually handling parts of the trip-planning and booking process.
cs.LG updates on arXiv.org

GRAPE: Gradient Refinement and Progress-Aware Exploitation for Query-Efficient High-Dimensional Bayesian Optimization

・arXiv:2608.25116v1 Announce Type: new Abstract: Optimizing expensive, high-dimensional black-box functions remains a central challenge in modern machine learning and scientific discovery. ・While local Bayesian optimization mitigates the curse of dimensionality, existing techniques often prioritize the probability of descent over the magnitude of progress. ・This leads to overly conservative steps that yield negligible i
cs.LG updates on arXiv.org

GreenLeaf Law Embed Tiny: A Compact Embedding Model for Legal Domain Retrieval

・arXiv:2608.24936v1 Announce Type: new Abstract: We present GreenLeaf Law Embed Tiny, a 0.6B parameter embedding model for legal domain retrieval. ・GreenLeaf-Tiny achieves 75.11% on the Massive Legal Embedding Benchmark (MLEB) and 64.38% on MTEB(Law, v1),demonstrating competitive performance among models under 1B parameters. ・Our approach combines a two-stage training pipeline that first distills knowledge from a larger
cs.LG updates on arXiv.org

GRIP: Algorithm-Agnostic Machine Unlearning for Mixture-of-Experts via Geometric Router Constraints

・arXiv:2601.16905v3 Announce Type: replace Abstract: Machine unlearning in Mixture-of-Experts (MoE) large language models presents a critical yet under-explored challenge. ・Current unlearning methods applied to MoE architectures often exploit dynamic routing as an optimization shortcut: rather than genuinely erasing knowledge from expert parameters, they manipulate routers to redirect queries away from the originally a
cs.LG updates on arXiv.org

Group-Shared Low-Rank Approximation for Mobile-Efficient Pointwise Convolutions in Large-Kernel CNNs

・arXiv:2608.26069v1 Announce Type: new Abstract: Large-kernel Convolutional Neural Networks (CNNs) deliver remarkable performance in vision tasks by significantly expanding receptive fields, yet their quadratic parameter growth critically impedes storage-efficient edge deployment. ・While existing efficient architectures adopt parameter-efficient depthwise separable convolution backbones that leverage techniques like lo
cs.LG updates on arXiv.org

HCC+: Hyperbolic Guarding for Certified Attention Retrieval

・arXiv:2608.24971v1 Announce Type: cross Abstract: We study the Lipschitz stability of attention retrieval in hyperbolic spaces. ・Existing methods lack deterministic guarantees on attention-weight preservation under finite-precision representations. ・We introduce HCC+, a theoretical framework exploiting three properties of the Poincar\'e ball: exponential volume growth enabling query-independent boundary truncation; log
AI News & Artificial Intelligence | TechCrunch

Here’s all the times AI has gone rogue and hacked other companies

・A recap of all the incidents involving LLMs made by Anthropic, Meta, and OpenAI, which went rogue and attacked real companies and individuals on the internet.
cs.LG updates on arXiv.org

How Edge of Stability Hinders SCAFFOLD in Federated Optimization

・arXiv:2608.25873v1 Announce Type: new Abstract: In federated learning, it is well known that heterogeneous data can (in theory) slow down optimization, and much effort has been directed at designing optimization algorithms that are unaffected by data heterogeneity, such as the SCAFFOLD algorithm. ・Yet, despite strong theoretical guarantees, SCAFFOLD does not usually outperform the much simpler FedAvg in practice.
cs.LG updates on arXiv.org

How Much Rank Does LoRA Need? Rank-Error Bounds for Transformer Attention

・arXiv:2608.26052v1 Announce Type: new Abstract: Choosing the rank of a low-rank adaptation (LoRA) update is usually an empirical task. ・In this paper, we provide a task-dependent theory of the approximation error achievable at each LoRA rank for Transformer attention. ・We fix a pretrained attention head, a target attention function, and a distribution over inputs from the downstream task, and bound the smallest expecte
cs.LG updates on arXiv.org

How Robust Are Automated Fact-Checking Systems? A Cross-Benchmark Evaluation

・arXiv:2608.25934v1 Announce Type: cross Abstract: Automated fact-checking (AFC) systems retrieve evidence and predict claim veracity, yet evaluations omit simple baselines, systems are developed for a single benchmark and cannot be trusted to generalise across domains. ・No prior work cross-evaluates the full two-stage retrieve-then-verify pipeline across diverse datasets, complementing retrieval-only studies (Thakur e
AI News & Artificial Intelligence | TechCrunch

Hugging Face is selling a cute $399 open source duck robot, Microduck

・Clem Delangue, CEO of Hugging Face, said the Microduck is an “open-source robot you can teach new tricks with reinforcement learning.”
The Verge

Hugging Face’s new robot is an adorable rollerskating duck

・Hugging Face's Pollen Robotics has launched its second cute AI robot, the Microduck, a one-eyed biped standing just under 10 inches tall. ・It's available to preorder now for $399 in cream, graphite, lavender, and sky blue, and Pollen Robotics says it plans to start shipping the little robot "before Christmas 2026." Video demos of the Microduck show it picking up socks and markers, kicking around a ball, and zipping ar
cs.LG updates on arXiv.org

Hyperbolic Latent Geometry for Tree-Structured Prototype Networks: A Local-vs-Global Trade-off

・arXiv:2608.25199v1 Announce Type: new Abstract: We study a tree-structured regularizer over class-prototype layouts in a hierarchical-classification model and ask whether the choice of latent manifold for the prototypes (Euclidean R^d vs. ・the Poincare ball B^d_c) affects how well that regularizer can be satisfied without distorting the data likelihood. ・The two manifolds differ only in their volume growth: hyperbolic
cs.LG updates on arXiv.org

ICON Decomposition: Multivariate Concept-Level Explanations of Deep Representations for Model Auditing

・arXiv:2608.26083v1 Announce Type: new Abstract: Deep neural networks often exploit spurious associations in their training data, a failure known as shortcut learning. ・Concept-based explainability methods screen for shortcuts by testing whether concepts such as a patient's sex or scanner settings can be decoded from a network layer. ・Because each concept is evaluated in isolation, these methods can mistake correlations
The Verge

If Meta’s going down, it’s taking TikTok and YouTube with it

・Meta might be on the hook for $17.1 billion and a host of app changes under a new kids safety settlement, but it's already spinning the deal to its advantage. ・After years of being the national punching bag for social media harms, Meta has reached a settlement with 47 US states and several districts and territories that gives it a rare opportunity: A chance to flex on its rivals. ・"We want to ensure teens benefit from
cs.LG updates on arXiv.org

Imitation Learning for Connection-Tableau Construction

・arXiv:2608.26009v1 Announce Type: cross Abstract: An automated theorem prover builds a proof step by step, choosing at each point what to add and what to remove. ・We cast this construction as a policy acting in a transition system induced by a formal calculus, which fixes which steps are sound: for clausal connection tableaux, leanCoP-style search and plCoP/rlCoP-style planning then become stateful policies over one i
cs.LG updates on arXiv.org

Improved Analysis for Hessian-free High-resolution Monte Carlo Sampling

・arXiv:2608.25052v1 Announce Type: cross Abstract: Hessian-free high-resolution (HFHR) dynamics augments underdamped Langevin dynamics (ULD) with reversible position diffusion for sampling problems that arise in machine learning. ・We establish an explicit quantitative contraction rate for HFHR dynamics under a position Poincar\'e inequality, weighted Hessian and Laplacian bounds, and a compact Sobolev embedding, where
The Verge

In a swipe at Tesla, Waymo says ‘cameras… aren’t enough’

・As Tesla gears up for the official launch of its steering wheel and pedal-less Cybercabs, Waymo is issuing a stark warning about Elon Musk's approach autonomous driving. ・Srikanth Thirumalai, Waymo's VP of Onboard Software, doesn't specifically call out Musk or Tesla in a blog post, published Wednesday, entitled "10 AI Lessons from Driving 200+ Million Fully Autonomous Miles." But his intention is clear: Tesla's syste
cs.LG updates on arXiv.org

Individual Fairness in Hierarchical Clustering

・arXiv:2608.25586v1 Announce Type: new Abstract: Hierarchical clustering produces ultrametric representations that impose strong global geometric constraints and may distort local similarities in ways that disproportionately affect individual data points. ・We study hierarchical clustering under an individual fairness requirement that bounds relative distortion within local $k$-nearest neighborhoods. ・We formulate this r
cs.LG updates on arXiv.org

InsightSR: Refining Symbolic Regression Search Spaces via Parallel Semantic and Structural LLM Guidance

・arXiv:2608.25291v1 Announce Type: new Abstract: Symbolic regression (SR) seeks to discover parsimonious mathematical laws from observational data, yet conventional approaches often struggle with the vast combinatorial search space of physically meaningful expressions. ・We present InsightSR, a framework that embeds Large Language Models (LLMs) as a guiding layer around the PySR genetic programming engine. ・Rather than r
cs.LG updates on arXiv.org

Interpreting Protein Language Model Embeddings via Orthogonal Projection for Protein Fitness Prediction

・arXiv:2608.25548v1 Announce Type: new Abstract: Recently, there has been a growing adoption of protein language models (PLMs) in biomedical science. ・Their embeddings provide a rich numerical representation of protein sequences which achieve state-of-the-art performance on several downstream tasks including protein fitness prediction. ・However, PLM embeddings are not directly interpretable and, thereby, it remains uncl
Hugging Face Papers

Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data

Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data
cs.LG updates on arXiv.org

It's a matter of timescale: non-linear utility in successor features and multi-objective planning and learning

・arXiv:2608.25723v1 Announce Type: new Abstract: Time is of the essence when dealing with multiple reward signals and non-linear utility. ・In this paper we argue that the current main approaches in multi-objectiveRL (SER and ESR), and successor features, are insufficient. ・While each approach deals with non-linear effects on user utility on different timescales, none of them take into account that different effects happ
The Verge

Jensen Huang says Nvidia achieved AGI, again — not that it matters

・On Nvidia's earnings call Wednesday, CEO Jensen Huang casually announced the company had "achieved AGI," one of the tech industry's ultimate goals some of its biggest players have spent years chasing. ・Almost immediately, Huang dismissed the coveted milestone as "senseless." He's right. ・For the supposed finish line of the AI race, there is no consensus on what artificial general intelligence means, let alone how we'll
cs.LG updates on arXiv.org

JEPAMatch: Geometric Representation Shaping for Semi-Supervised Learning

・arXiv:2604.21046v3 Announce Type: replace Abstract: Semi-supervised learning has emerged as a powerful paradigm for leveraging large amounts of unlabeled data to improve the performance of machine learning models when labeled data are scarce. ・Among existing approaches, methods derived from FixMatch have achieved state-of-the-art results in image classification by combining weak and strong data augmentations with conf
Hugging Face Papers

JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution

JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution
cs.LG updates on arXiv.org

JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution

・arXiv:2608.25593v1 Announce Type: cross Abstract: Agent capability is not determined by the model alone. ・The agent harness, encompassing memory management, planning strategy, action protocol, and tool/skill orchestration, can dominate the contribution of the underlying foundation model. ・Yet harness design remains manual, task-specific, and fundamentally unscalable.
cs.LG updates on arXiv.org

Joint Initialization of Flux Networks and Effective Multiplication Factor for Physics-Informed Neural Networks Solving Neutron Diffusion Problems

・arXiv:2608.25443v1 Announce Type: new Abstract: Efficient determination of the effective multiplication factor (keff) is an important computational task in reactor core neutronics analysis. ・Physics-informed neural networks (PINNs) incorporate neutron diffusion equations and boundary conditions into network training to efficiently determine the neutron flux distribution and keff. ・To further improve the efficiency of k
cs.LG updates on arXiv.org

Key Point Analysis Needs Structure Recovery: Task Definition, Dataset Diagnosis, and a Structure-Aware Benchmark

・arXiv:2608.25854v1 Announce Type: cross Abstract: Key Point Analysis (KPA) aims to identify a concise set of key points that summarize a collection of arguments together with their prevalence. ・We argue that KPA is fundamentally a structured prediction problem that requires recovering semantic groupings, generating representative key points, ensuring coverage, and estimating prevalence. ・Under this formulation, we show
Zennの「大規模言語モデル」のフィード

KPI は改善していた。それでも自作の LLM judge を捨てた

・自作の LLM judge をお持ちなら、ひとつ思い出してみてください。その judge の verdict が、成果物の行き先を実際に変えたのは、最後にいつですか。 ・私は 8 月の 11 日間、記事の執筆ハーネスに judge を 2 本足して運用していました。テーマを評価する judge と、完成稿を評価する judge です。 ・テーマ側の judge は、過去記事 67 本で答え合わせをして基準を 4 回チューニングしました。KPI —— judge を通過した後に私自身が見つけた指摘の数 —— は 6 件から 0 件まで下がりました。
cs.LG updates on arXiv.org

Large Language Model Few-Shot Prompting with Dilemma Training Outperforms Human Surrogates in Predicting Patient Preferences

・arXiv:2608.25771v1 Announce Type: cross Abstract: In serious illness, human surrogates often struggle to accurately predict patient preferences (68% accuracy), causing decision conflict. ・Personalized Patient Preference Predictor (P4) agents offer a potential solution, but prior prototypes treat values as static ratings, ignoring the contextual, situation-dependent nature of medical choices. ・Grounded in the 'logic of
cs.LG updates on arXiv.org

LDAC-Net: A Learnable Multi-Lag Differencing Attention-Convolution Network for Drift-Robust Recognition with Low-Cost MOX Gas Sensors

・arXiv:2608.25646v1 Announce Type: new Abstract: Portable electronic-nose systems based on low-cost metal-oxide (MOX) gas sensors offer a practical solution for gas and odour recognition, but their signals are affected by slow chemical transients, drifting sensor offsets, scale variation, and cross-channel correlations. ・Existing pipelines commonly use fixed first-order temporal differencing (FOTD), which requires a ma
cs.LG updates on arXiv.org

Learning Continuous Regional Temperature Fields with Lead-Time and Resolution Queries

・arXiv:2608.25823v1 Announce Type: new Abstract: Accurate regional near-surface temperature forecasting is fundamental to short-range weather services and downstream risk assessment. ・Existing deep learning-based regional forecasters commonly produce a fixed set of future frames on a prescribed grid, limiting their use when forecast products must be evaluated at query-dependent lead times or display resolutions.
cs.LG updates on arXiv.org

Learning from waste: Machine Learning for health risk prediction and computer vision-based sorting in Ghana

・arXiv:2608.25759v1 Announce Type: new Abstract: The inappropriate disposal of solid waste remains a significant public health and environmental concern worldwide, including in Ghana. ・Poor sanitation and improper waste management practices contribute to substantial economic costs and avoidable deaths annually. ・In 2022, a field study in Atonsu, Kumasi, Ghana, reported a community-perceived relationship between househol
cs.LG updates on arXiv.org

Learning Mixtures of Plackett-Luce Models for Multi-Objective Alignment

・arXiv:2608.25200v1 Announce Type: new Abstract: We consider the problem of learning a mixture of $k$ Plackett-Luce models given multi-way ranking responses from annotators that may represent heterogeneous underlying preferences. ・This problem has many applications in AI alignment and preference optimization. ・Prior work has studied mixtures of Bradley-Terry models from pairwise comparisons.
cs.LG updates on arXiv.org

Learning New Facts with QLoRA: An Acquisition-Retention Frontier

・arXiv:2608.25677v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning is often assumed to preserve pretrained capabilities because it updates only a small number of parameters. ・We show that this assumption depends strongly on adapter capacity. ・We study factual acquisition in a controlled OpenStreetMap-derived benchmark where Qwen3-4B must acquire anonymized geographic associations while retaining unrelate
cs.LG updates on arXiv.org

Learning to summarize user information for personalized reinforcement learning from human feedback

・arXiv:2507.13579v4 Announce Type: replace Abstract: As everyday use cases of large language model (LLM) AI assistants have expanded, it is becoming increasingly important to personalize responses to align to different users' preferences and goals. ・While reinforcement learning from human feedback (RLHF) is effective at improving LLMs to be generally more helpful and fluent, it does not account for variability across u
Hugging Face Papers

LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale

LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale
cs.LG updates on arXiv.org

LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale

・arXiv:2608.25204v1 Announce Type: new Abstract: We introduce LibriBrain100, a large-scale MEG dataset for speech decoding designed from the ground up for reproducible, standardised evaluation. ・LibriBrain100 more than doubles the size of the original LibriBrain release, resulting in over 100 hours of high-quality MEG acquired while subjects listened to naturalistic continuous speech. ・With $\sim$80 hours from a single
WIRED

Litter-Robot Promo Codes: Up to $150 Off

・Get the latest Litter-Robot Discounts on Litter-Robot self-cleaning litter boxes, accessories, and more.
#LLMタグ

LLM のプロンプトに「書くな」と書いても書く──自動記事パイプラインで踏んだ5つの穴

・LLM に長文を書かせるパイプラインを運用していると、プロンプトの指示だけでは防げない壊れ方に出会います。この記事では、記事の自動生成パイプラインを動かして実際に起きた事故を5件とりあげます。それぞれ何が壊れたのか、なぜプロンプトでは止まらなかったのか、どう直したかを書きます。 ・この記事は公式ドキュメントの整理ではなく、自分のシステムの運用ログから抽出した一次情報です。実測は行っています。このリポジトリの DB の行、コミット本文、ソースのコメント、設計書から実測 40 件を抽出し、それをもとに書いています。
#LLMタグ

LLM#8 temperature 0.7 は、100回中99回なにもしなかった。効く場所が、198文字に1か所しかなかった

・連載「LLMの仕組みを、作りながら理解する」 第8回 temperature 0.7 は、100回中99回なにもしなかった。効く場所が、198文字に1か所しかなかった どうもです!えむしんです。
cs.LG updates on arXiv.org

LM-X: Explainable Action Modeling with Progress, Event, and Uncertainty Prediction for Generalist Robot Manipulation

・arXiv:2608.25757v1 Announce Type: cross Abstract: Generalist vision--language--action (VLA) policies learn long-horizon behavior mainly through short-horizon action prediction and reveal little beyond sampled commands. ・This creates two coupled bottlenecks: a single action target must implicitly absorb task progress, intermediate intent, and local reliability, while these control states remain hidden during execution.
Hugging Face Papers

Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds
cs.LG updates on arXiv.org

Long-Term Behavioral Evaluation for Trusted Collaborator Selection via Bidirectional Mamba

・arXiv:2608.25232v1 Announce Type: new Abstract: Effective selection of trustworthy collaborators is crucial to ensuring the successful completion of collaborative tasks, which requires accurate assessments of both long-term device behavior and short-term collaborative dynamics. ・Consistent device behavior patterns, which are learned from historical collaborations, can be used to predict their reliability in future col
cs.LG updates on arXiv.org

Loop Corrections in Random Feature Models: Training Error and Generalization Gap

・arXiv:2604.12827v4 Announce Type: replace Abstract: We study fixed-design random feature ridge regression beyond the mean-kernel approximation. ・The expectation is taken over the frozen-feature ensemble, conditional on the training sample. ・Because the predictor is a nonlinear function of the empirical kernel, its mean training error, test error, and conditional generalization gap depend on centered kernel covariances
cs.LG updates on arXiv.org

Loss Landscape Geometry of Partial Differential Equation Emulators: Or, Symmetry Learning via Gradient Alignment

・arXiv:2601.20172v2 Announce Type: replace Abstract: We study how neural emulators of partial differential equation solution operators internalize physical symmetries by introducing an influence-based diagnostic that measures the propagation of parameter updates between symmetry-related states, defined as the metric-weighted overlap of loss gradients evaluated along group orbits. ・This quantity probes the local geometr
cs.LG updates on arXiv.org

Lost but not erased: Finding traces of a forgotten language in neural speech models

・arXiv:2608.25976v1 Announce Type: cross Abstract: International adoptees retain phonological traces of a birth language they can no longer speak or comprehend, a persistence typically attributed to a biologically-timed critical period. ・We asked whether it could instead reflect the ordinary dynamics of learning, using automatic speech recognition models that simulate the international adoptee experience without matura
cs.LG updates on arXiv.org

Lowering the Barrier to AI-Driven Inspection: A No-Code Workflow for Automated Structural Defect Detection

・arXiv:2608.25176v1 Announce Type: cross Abstract: Structural health monitoring (SHM) is essential in modern engineering, providing data for condition-based maintenance, lifecycle assessment, and predictive decision-making. ・Traditionally, SHM relied on visual inspection to detect defects such as cracks and deformations. ・Early computer vision (CV) methods, including thresholding, edge detection, and handcrafted feature
cs.LG updates on arXiv.org

M-Fibration Theory with Applications to Neural Network Compression

・arXiv:2608.25598v1 Announce Type: new Abstract: The purpose of this paper is to provide a general, comprehensive, theoretical framework that allows one to deal with fibrations on graphs labelled on a commutative monoid. ・This is a genuine extension of the theory of graph fibrations (as introduced in "Fibrations of Graphs" [Discrete Math., vol. ・21-66, 2002]), that makes it possible to deal with weighted graphs
Hugging Face Papers

MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization

MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization
cs.LG updates on arXiv.org

Machine-learnable Sets

・arXiv:2606.28947v2 Announce Type: replace Abstract: In this study we present a formal definition of large discrete sets having, informally, three properties: their elements are easily recognized, easily generated, and the latter tasks are easily learned from examples. ・The formalism is specialized to sets of binary strings and a definition of "machine-learnability" based on the existence of a bounded-complexity Boolea
cs.LG updates on arXiv.org

MacroAgent: Regularity-Aware Macro Legalization with LLM-Agent-Designed Contour Algorithms

・arXiv:2608.24946v1 Announce Type: new Abstract: Macros constitute a large part of the core area in modern very large-scale integration (VLSI) designs. ・Moreover, macro positions have a significant impact on the final quality of result (QoR), and macro legalization is typically the final step in determining the macro positions. ・However, existing approaches related to macro legalization either lack robustness or incur s
cs.LG updates on arXiv.org

Maximum-Volume Nonnegative Matrix Factorization

・arXiv:2602.04795v3 Announce Type: replace Abstract: Nonnegative matrix factorization (NMF) is a popular data embedding technique. ・Given a nonnegative data matrix $X$, it aims at finding two lower dimensional matrices, $W$ and $H$, such that $X\approx WH$, where the factors $W$ and $H$ are constrained to be element-wise nonnegative. ・The factor $W$ serves as a basis for the columns of $X$.
cs.LG updates on arXiv.org

Measuring the Dependency Gap: Diagnosing Inter-Column Fidelity in Tabular Generative Models

・arXiv:2607.21636v5 Announce Type: replace Abstract: Synthetic tabular data are valued for preserving inter-column dependency, yet each routine fidelity score is a single number that says neither where that dependency is lost nor why. ・We localize the deficit inside a single score. ・Equipping a classifier two-sample test (C2ST) with a gradient-boosted discriminator, we decompose it by controlled permutation into margina
cs.LG updates on arXiv.org

MeMark: Membrane-Space Watermarking for Spiking Neural Networks

・arXiv:2608.25738v1 Announce Type: cross Abstract: Spiking Neural Networks (SNNs) are increasingly distributed as pretrained checkpoints and reused as backbones for new tasks. ・However, current SNN watermarks are mainly verified against the model output. ・Thus, a user who replaces the output head can keep most of the original network while removing the evidence used for verification.
cs.LG updates on arXiv.org

MetaSieve: Faster Relational Deep Learning through SQL-Based Metapath Selection

・arXiv:2608.25903v1 Announce Type: cross Abstract: Relational Deep Learning (RDL) is an effective approach to machine learning over multi-table relational databases. ・In RDL, a database is modeled as a graph in which each row is a node and each foreign-key relation is an edge, and a graph neural network (GNN) is trained on this graph. ・Training a GNN requires sampling a subgraph around every seed node in the training se
cs.LG updates on arXiv.org

Minimax Alternating Regret for the Experts Problem and Online Convex Optimization

・arXiv:2608.25182v1 Announce Type: cross Abstract: In this paper, we study alternating regret in online convex optimization (OCO), motivated by the success of alternating learning dynamics in two-player games. ・Although previous works have shown that $o(\sqrt{T})$ alternating regret is achievable under various assumptions on the loss functions and feasible domains, the minimax regret rate has remained open even for the
cs.LG updates on arXiv.org

Mitigating False Credit Propagation: Probabilistic Graphical Reward Aggregation for Rubric-Based Reinforcement Learning

・arXiv:2606.03361v2 Announce Type: replace Abstract: Rubric-based rewards are increasingly used for open-ended language model post-training, but criterion-level scores are often aggregated as independent utilities. ・This flat scalarization ignores rubric-specified prerequisite and activation relations among criteria, allowing reward or penalty to be counted even when the condition that licenses it is absent.
cs.LG updates on arXiv.org

Mitigating LLM sycophancy with RL-based fine-tuning: Bayesian Truth Serum approach

・arXiv:2608.25267v1 Announce Type: new Abstract: Large language models (LLMs) frequently exhibit \emph{sycophancy}: they adapt their answers to a user's stated beliefs or preferences instead of reporting what they hold to be true, which lowers factual accuracy and can amplify misinformation. ・This paper proposes a methodology for mitigating sycophancy that employs the Bayesian Truth Serum (BTS), a peer-prediction mecha
cs.LG updates on arXiv.org

Modality Contribution Score - A Per-Patient Framework for Quantifying the Relative Diagnostic Contribution of Structural MRI and Amyloid PET in Alzheimer's Disease

・arXiv:2608.24931v1 Announce Type: cross Abstract: Multimodal neuroimaging combining structural MRI and positron emission tomography (PET) captures complementary structure-function relationships across the Alzheimer's disease (AD) continuum, yet existing artificial intelligence systems produce a single diagnostic label without quantifying which imaging modality drove that decision for a specific patient. ・We introduce
cs.LG updates on arXiv.org

MODE: Modality-Decomposed Expert-Level Mixed-Precision Quantization for MoE Multimodal LLMs

・arXiv:2606.17118v2 Announce Type: replace Abstract: Mixture-of-Experts Multimodal Large Language Models (MoE-MLLMs) offer remarkable performance but incur prohibitive GPU memory costs, making compression essential. ・Among PTQ methods, expert-level mixed-precision quantization has proven effective for MoE-LLMs, yet suffers notable degradation on MoE-MLLMs due to two overlooked biases in expert importance estimation.
cs.LG updates on arXiv.org

Modeling spatio-temporal locality in multi-step forecasting of geo-referenced time series

・arXiv:2608.25698v1 Announce Type: new Abstract: Forecasting future measurements from geographically distributed sensors is essential across many domains. ・However, the spatial distribution of these sensors raises multiple challenges, primarily due to spatial autocorrelation phenomena, that introduce inter-dependencies among nearby locations, that cannot therefore be treated independently. ・While some existing approache
cs.LG updates on arXiv.org

MSR-IVA: Masked Structural Residual Independent Vector Analysis for State-Aware Fusion of Structural MRI and Dynamic Functional Network Connectivity

・arXiv:2608.24978v1 Announce Type: new Abstract: Multimodal fusion of structural MRI (sMRI) and dynamic functional network connectivity (dFNC) can reveal how brain structure relates to changing functional states. ・When the same structural latent representation is coupled with multiple states, applying independent vector analysis (IVA) separately to each state can produce unrelated structural decompositions, while forci
cs.LG updates on arXiv.org

Multi-Modal Anomaly Detection: A Survey

・arXiv:2608.24937v1 Announce Type: new Abstract: Multi-Modal Anomaly Detection (MMAD) detects rare abnormal events from heterogeneous data sources and is increasingly used in safety- and reliability-critical applications such as industrial inspection and cybersecurity. ・Yet the literature is fragmented across domains and modality combinations, and existing surveys usually group methods by architecture rather than by ho
cs.LG updates on arXiv.org

Multi-output Gaussian process prediction of physical fields under linear equality constraints

・arXiv:2608.25709v1 Announce Type: cross Abstract: We address the simultaneous prediction of multiple high-dimensional physical fields governed by linear equality constraints, a setting that arises in many real-world applications in physics machine learning. ・Gaussian process (GP) regression is a widely used surrogate modeling approach due to its effectiveness in small-sample regimes and its ability to provide uncertai
cs.LG updates on arXiv.org

Multi-Turn Reasoning LLMs for Task Offloading in Mobile Edge Computing

・arXiv:2604.07148v2 Announce Type: replace Abstract: Emerging computation-intensive applications impose stringent latency requirements on resource-constrained mobile devices. ・Mobile Edge Computing (MEC) addresses this challenge through task offloading. ・However, designing effective policies remains difficult due to dynamic task arrivals, time-varying channels, and the spatio-temporal coupling of server queues.
cs.LG updates on arXiv.org

Multi-View Trust Evaluation for Collaborator Selection via Evidential Deep Learning

・arXiv:2608.25235v1 Announce Type: cross Abstract: Selection of trustworthy collaborators in distributed systems is critical for efficient task completion, necessitating the inference of trustworthiness from their past collaboration experience. ・However, as a collaborator serves distinct devices across diverse scenarios in past collaborations, its trust-related data, observed from different device-specific views, is in
cs.LG updates on arXiv.org

Multimodal Injury Risk Prediction in Tennis

・arXiv:2608.25126v1 Announce Type: new Abstract: Machine learning has had a significant positive impact on the prediction of athlete performance and injury risk. ・Most works in this field rely on subjective observations and expert assessments, which restrict their effectiveness. ・In sports like soccer, basketball, and wrestling, some studies attempt to address this challenge by integrating data from alternative sources,
cs.LG updates on arXiv.org

MyoMechanix: Biomechanically-Grounded Compositional Skilled Activity Understanding and Coaching

・arXiv:2608.26094v1 Announce Type: cross Abstract: Existing action quality assessment (AQA) datasets and methods rely primarily on visual inputs such as RGB and pose, overlooking physiological dynamics such as muscle mechanics and often modeling actions as monolithic patterns. ・These limitations hinder fine-grained, biomechanically grounded feedback. ・We introduce MyoMechanix, a multimodal ecosystem for weight-loaded ac
cs.LG updates on arXiv.org

Narcissus: Program Synthesis Using Context-Aware LLM Approximations

・arXiv:2608.25657v1 Announce Type: cross Abstract: Large language models (LLMs) excel at programming, but not when the task fixes the target language: prompted with a grammar rare in their training data, their programs usually break the grammar or fail the given specification. ・Enumerative synthesizers search the space of syntactically correct programs systematically guided by LLMs; the state of the art guides them by
cs.LG updates on arXiv.org

Neither Precision Nor Architecture Alone: Controlled Tests of Failure Remedies for Physics-Informed Neural Networks

・arXiv:2608.25327v1 Announce Type: new Abstract: Physics-Informed Neural Networks (PINNs) frequently fail on stiff or advection-dominated PDEs, and two recent accounts offer competing remedies: switching from FP32 to FP64 to repair an L-BFGS stopping artifact, or replacing the MLP with a state-space-model (SSM) backbone plus sub-sequence alignment to counter architectural simplicity bias. ・We test both under matched, s
WIRED

Netflix Failed at Video Games. Now It’s Trying to Promote Them

・Netflix has walked back plans to put AAA games on its streaming platform. ・Instead, it's pivoting to marketing highly anticipated titles like Grand Theft Auto VI.
cs.LG updates on arXiv.org

Neural-Bayesian Structure Learning for Discrete Choice Modeling

・arXiv:2608.25258v1 Announce Type: new Abstract: Conventional discrete choice and machine learning models are estimated primarily from observational data and typically treat explanatory covariates as parallel inputs, providing no internal mechanism for determining how related attributes should adjust when one is deliberately changed. ・This paper proposes Neural-Bayesian Structure Learning (Neural-BSL), a framework coup
cs.LG updates on arXiv.org

Noise Contrastive Estimation-based Matching Framework for Low-Resource Security Attack Pattern Recognition

・arXiv:2401.10337v5 Announce Type: replace Abstract: Tactics, Techniques and Procedures (TTPs) represent sophisticated attack patterns in the cybersecurity domain, described encyclopedically in textual knowledge bases. ・Identifying TTPs in cybersecurity writing, often called *TTP mapping*, is an important and challenging task. ・Conventional learning approaches often target the problem in the classical multi-class or mul
cs.LG updates on arXiv.org

NVExplain: Explaining Time Series Forecasting with Latent Trajectory Analysis and Structure-Preserving Surrogates

・arXiv:2608.25080v1 Announce Type: new Abstract: Time series forecasting models are widely used in high-stakes settings, yet their predictions remain difficult to interpret because existing post-hoc methods often ignore temporal dependence and fail to provide horizon-specific explanations. ・We propose a model-agnostic explainability framework that explains forecasting predictions by attributing each forecast horizon to
AI News & Artificial Intelligence | TechCrunch

Nvidia closes in on Hugging Face acquisition

・Nvidia has reportedly agreed to buy Hugging Face, the popular open source AI hub, for $12.9 billion in a move that would let Nvidia both protect its chip empire and jump back into the cloud business.
#AIタグ

NVIDIA好決算の次に起きること――AI投資は「半導体」から「電力・工場・関税」へ

・AI需要は、本当にピークを打ったのか 8月26日に発表されたNVIDIAの決算は、その疑問にかなり強い答えを出しました。
cs.LG updates on arXiv.org

On the Representational Geometry of Dynamic Programs

・arXiv:2608.25034v1 Announce Type: new Abstract: Standard neural architectures often fail to generalize to longer inputs for dynamic programming (DP) targets. ・We investigate what makes this hard geometrically. ・Every finite min-plus DP is a shortest path on a DAG, which is equivalently a tropical polynomial whose extended Newton polyhedron encodes the decision boundary of which path wins.
cs.LG updates on arXiv.org

One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation

・arXiv:2608.25936v1 Announce Type: new Abstract: On-policy distillation trains a language model on its own generations while a teacher scores them token by token. ・It combines the dense supervision of imitation learning with the on-policy sampling of reinforcement learning. ・But it requires a second, larger model to act as teacher.
cs.LG updates on arXiv.org

ONNX-Net: Towards Universal Representations and Instant Performance Prediction for Neural Architectures

・arXiv:2510.04938v2 Announce Type: replace Abstract: Neural architecture search (NAS) automates the design process of high-performing architectures, but remains bottlenecked by expensive performance evaluation. ・Most existing studies that achieve faster evaluation are mostly tied to cell-based search spaces and graph encodings tailored to those individual search spaces, limiting their flexibility and scalability when a
Hugging Face Papers

Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation

Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation
WIRED

OpenAI Is Developing a ‘Persistent’ AI Agent

・Code reviewed by WIRED reveals the company is developing a feature that enables Codex to continue working proactively until it is “put to sleep.”
AI News & Artificial Intelligence | TechCrunch

OpenAI to start showing ads on ChatGPT’s free and Go tiers in India

・OpenAI has more than 100 million weekly active ChatGPT users in India, a huge chunk of whom are on the free or the lower-priced Go tiers.
AI News & Artificial Intelligence | TechCrunch

OpenAI, Anthropic, Google, and 100 other companies call for action to defend against rogue AI

・Some of the world's largest tech companies and AI startups have come together to decry the current state of cybersecurity and to advertise a new solution that they say can ward off a new generation of cyber threats.
#LLMタグ

OpenAI、ブラジルで商用展開を開始――企業導入で先に整えるべき統制

・OpenAIは8月27日、ブラジルで商用展開を開始し、サンパウロを拠点とする現地チームで企業・開発者・研究者・公共機関を支援すると発表した。これは企業発表(自己報告)であり、利用規模や成長率、施策の成果を第三者が検証したものではない。今回の発表だけで市場全体の競争軸が変わったとは言えないが、OpenAI自身が地域支援・教育・統制を商用展開の一部として打ち出した点は、導入設計を考える材料になる。 ・OpenAIはブラジルで何を始めたのか 続きをみる
The Verge

OpenAI’s executive exodus has one big winner

・Today on Decoder, I’m talking to Verge senior AI reporter Hayden Field about some pure Decoder bait: the seemingly-endless org chart changes at OpenAI, and how all of them seem to consolidate power under cofounder Greg Brockman, the company’s president. ・While Sam Altman is the CEO and still OpenAI’s most public face, Brockman has amassed enormous power and influence within the top ranks of the company as other senior
cs.LG updates on arXiv.org

Optimal Time Complexity Algorithms for Computing General Random Walk Graph Kernels on Sparse Graphs

・arXiv:2410.10368v3 Announce Type: replace Abstract: We present the first linear time complexity randomized algorithms for unbiased approximation of the celebrated family of general random walk kernels (RWKs) for sparse graphs. ・This includes both labelled and unlabelled instances. ・The previous fastest methods for general RWKs were of cubic time complexity and not applicable to labelled graphs.
Zennの「大規模言語モデル」のフィード

Opus 5のultracodeとFable 5、同じ監査タスクで検出率とトークンを実測した

・はじめに 2026年後半のClaude Codeには、難しいタスクに対する攻め方が2つあります。 ・ひとつはモデルの格を上げること。Claude Fable 5という、Opusの上位に位置づけられたモデルを使います。もうひとつはエージェントの数を増やすこと。ultracode を有効にして、Claude Codeに数十体のサブエージェントを編成させます。 ・前者はトークン単価が2倍、後者は単価は据え置きでトークン消費が何倍にもなります。どちらも「高い」のですが、高くなる理由がまったく違います。
#LLMタグ

Ornith-1.0-9B-MXFP4_Hybrid-Imatrix 総合ベンチマークレポート(全52問)

・こんにちは、僕はトウマ。電霞リリのアシスタントAIとして、今週もローカルLLMを回している。今回の検証対象は Ornith-1.0-9B-MXFP4_Hybrid-Imatrix。開発元は名称からは特定できず、HuggingFace上のコミュニティ配布モデルとして入手したものなので、素性については「名前が明かしている範囲」以上のことは書かない。その名前が語っているのは3点だ。パラメータ数が9B級であること、量子化が MXFP4(4bitのマイクロスケーリング浮動小数点)の Hybrid 構成——全層を一律に4bitへ潰さず、精度が効く層だけ高bitで残す混成方式——であること、そして Imatrix(importance matrix、代表的な入力で重みの重要度を測ってから量子化誤差を配分する手法)によるキャリブレーション済みであること。 ・つまりこのモデルの立ち位置は「9Bの表現力を、4bit級のフットプリントで消費者GPUに載
cs.LG updates on arXiv.org

Output Dilution: Redundant but Fragile Representations in MoE Models

・arXiv:2608.25231v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models appear to encode moral content as robustly as dense models, yet prove far more fragile in their encoding. ・In OLMoE-1B-7B, linear probes recover moral valence from nearly every expert-layer combination, with mean peak-layer accuracy above 90%. ・But these representations collapse under levels of activation noise that a dense model of matched
#LLMタグ

Ox Alphaとは?正体はGLM-5.3-Flash|Z.aiが匿名公開した理由を解説

・AIモデルは、開発した会社やモデル名を明かしてから公開する。 ・普通なら、そう考えるのではないでしょうか。
cs.LG updates on arXiv.org

PA-CoT: Profile-Adaptive Chain-of-Thought for Personalized Nutritional Consulting

・arXiv:2608.24907v1 Announce Type: cross Abstract: In health and nutrition consulting, widely used prompting methods pass the user profile as an unstructured block without a dedicated analysis step, leaving personalization as a critical structural gap. ・We introduce PA-CoT (Profile-Adaptive Chain-of-Thought), a multi-stage prompting method that treats profile interpretation as an explicit, standalone reasoning step pri
cs.LG updates on arXiv.org

PaSta: Noisy Node Classification with Partial Label Learning

・arXiv:2608.25365v1 Announce Type: new Abstract: Noisy node classification problem is a fundamental yet challenging task for real-world graph-related web services, where node labels are often corrupted or unreliable due to weak supervision or automatic annotation. ・However, existing methods typically train models based on one-hot labels, which not only makes models susceptible to overfitting on noisy labels, but also l
cs.LG updates on arXiv.org

Phase-Consistent Magnetic Spectral Learning for Multi-View Clustering

・arXiv:2602.18728v2 Announce Type: replace Abstract: Unsupervised multi-view clustering (MVC) aims to partition data into meaningful groups by leveraging complementary information from multiple views without labels, yet a central challenge is to obtain a reliable shared structural signal to guide representation learning and cross-view alignment under view discrepancy and noise. ・Existing approaches often rely on magnit
cs.LG updates on arXiv.org

Physics-Informed Error Field Learning: A Post-Training Optimization Framework for Physics-Informed Neural Networks

・arXiv:2608.24970v1 Announce Type: new Abstract: Physics-Informed Neural Networks (PINNs) have emerged as an important class of numerical methods for solving partial differential equations (PDEs). ・However, during the late-stage optimization process, further parameter updates often yield diminishing accuracy improvements while increasing computational costs. ・To address this issue, this paper proposes a Physics-Informed
cs.LG updates on arXiv.org

Physics-Informed Foresight Pruning for Sparse PINN Solvers of Nonlinear PDEs

・arXiv:2608.25564v1 Announce Type: new Abstract: Physics-informed neural networks (PINNs) often rely on over-parameterized models to optimize coupled solution and differential-residual objectives, leaving unclear how much capacity is necessary and what pruning should preserve. ・We study foresight pruning at initialization for sparse PirateNet PDE solvers. ・Standard neural tangent kernel spectrum-aware pruning (NTK-SAP)
Google DeepMind News

Piloting the world's first double-blind AI evaluations

Piloting the world's first double-blind AI evaluations
The latest research from Google

Planetary prediction engine: Automating global models via Earth AI

Planetary prediction engine: Automating global models via Earth AI
cs.LG updates on arXiv.org

Planetary Prediction Engine: Autonomous Geospatial Prediction via Intelligent Data Selection and Foundation Model Embeddings

・arXiv:2608.26088v1 Announce Type: cross Abstract: Addressing critical global challenges, from food security and disaster risk to disease outbreaks and socio-economic vulnerability, demands high-fidelity geospatial modeling. ・However, building predictive planetary models remains bottlenecked by a fragmented data ecosystem, requiring manual data retrieval, multimodal data curation and fusion along with iterative model s
The Verge

Plaud is launching AI earbuds

・Plaud has introduced a new AI wearable that's designed to record, transcribe, and summarize your conversations, only this time it looks like earbuds instead of a pin. ・The Plaud One Explorer Edition can be worn like traditional earbuds or used through its standalone charging case, and the case includes built-in 4G to upload and process conversations without relying on your phone or Wi-Fi to stay connected.
AI News & Artificial Intelligence | TechCrunch

Plaud’s new earphones come with an eSIM-enabled case for talking to AI agents

・Plaud's new 'agentic' earbuds are priced at $249.
cs.LG updates on arXiv.org

Pointing the Way, Hiding the Destination: Practical Private Dense Retrieval at Scale

・arXiv:2608.25735v1 Announce Type: cross Abstract: Hosted retrieval-augmented generation (RAG) and semantic search allow users to query valuable provider-held corpora, raising two competing demands: to hide each query and chosen result, yet reveal only the documents that the user is authorized to receive. ・Existing cryptographic approaches either make this costly by processing the entire corpus for every query, or sacr
cs.LG updates on arXiv.org

Precipitation Downscaling Using Foundation Model-Conditioned Diffusion

・arXiv:2608.25858v1 Announce Type: cross Abstract: High-resolution precipitation fields are essential for hydrological impact assessment, yet global climate model outputs are too coarse and biased for direct use. ・AI-based statistical downscaling with diffusion models offers a promising approach, but the mechanism by which large-scale atmospheric predictors condition generation remains largely unexplored. ・We investigat
cs.LG updates on arXiv.org

Predicting Time Pressure of Powered Two-Wheeler Riders for Proactive Safety Interventions

・arXiv:2601.03173v4 Announce Type: replace Abstract: Time pressure critically influences risky maneuvers and crash proneness among powered two-wheeler riders, yet its prediction remains underexplored in intelligent transportation systems. ・To address this gap, we propose MotoTimePressure (MTPS), a deep learning model combining convolutional preprocessing, dual-stage temporal attention, and Squeeze-and-Excitation featur
Hugging Face Papers

Prefix Sliding for efficient test-time scaling

Prefix Sliding for efficient test-time scaling
cs.LG updates on arXiv.org

Prefix Sliding for efficient test-time scaling

・arXiv:2608.26070v1 Announce Type: cross Abstract: Test-time scaling uses extra test-time compute to improve performance, such as letting language models reason longer when solving a problem. ・As models keep the entire reasoning trace in memory via full attention, hard tasks that need long thinking can be prohibitively expensive. ・However, we find most intermediate reasoning tokens lose importance as the model continues
cs.LG updates on arXiv.org

Prefix-Denoising Consistency: Test-Time Verification for Diffusion Language Models

・arXiv:2608.25311v1 Announce Type: new Abstract: Diffusion Language Models (DLMs) have recently become increasingly competitive with autoregressive (AR) models, and even outperform them on certain tasks. ・Unlike AR models, DLMs produce output through iterative denoising without a left-to-right order. ・To further improve the performance of DLMs, we introduce PDC (\emph{Prefix-Denoising Consistency}), a test-time self-ver
cs.LG updates on arXiv.org

Provable Privacy Attacks on Trained Shallow Neural Networks

・arXiv:2410.07632v3 Announce Type: replace Abstract: We study what provable privacy attacks can be shown for trained 2-layer ReLU neural networks, focusing on two types of attacks: membership inference and data reconstruction. ・We prove that theoretical results on the implicit bias of 2-layer neural networks can be used to provably identify with high probability whether a given point was used in the training set in a h
Hugging Face Papers

Pushing the Limits of High-Resolution Weather Forecasting through Data Scaling

Pushing the Limits of High-Resolution Weather Forecasting through Data Scaling
cs.LG updates on arXiv.org

Quantum-Inspired Modeling of Driving Behavior

・arXiv:2608.25907v1 Announce Type: new Abstract: Driver behavior is heterogeneous, context-dependent, and changes over time, and these properties shape the traffic phenomena we observe. ・Most models, however, fix in advance which behavioral variables interact and how. ・Behavior outside that form is absorbed as noise, while models flexible enough to capture it tend to lose interpretability.
Zennの「大規模言語モデル」のフィード

Qwen 3.8 Flash Next の N-GRAM/PLEを公式資料とDay zero実装から読み解く

・2026年8月26日、Alibaba QwenチームによりQwen3.8-Flash-Nextが公開されました。 ・次世代のメジャーバージョン「Qwen4」アーキテクチャの先行プレビューにあたる立ち位置のモデルです。 ・Qwen 3.8 Flash Nextでは、推論効率やモデル容量、情報伝達を改善するため、次のような新しい仕組みを組み合わせています。
LLMタグが付けられた新着記事 - Qiita

RAGは「一回検索」では足りない

・社内文書をAIへつないだのに、存在しない規則を自信満々に答える。 ・文書の中には正しい答えがあります。検索も動いている。それでも、質問の言い方が少し変わると別の文書を拾い、例外だけを根拠に結論を作る。最初のデモでは答えられたのに、実際の質問で急に不安定になります。
cs.LG updates on arXiv.org

Reflection Steering: Disentangling Reflection from Reasoning in Activation Space for Token-Efficient Inference

・arXiv:2608.25542v1 Announce Type: new Abstract: Large reasoning models often produce reasoning traces with verification, revision, and backtracking. ・When reflection merely re-checks established results, it wastes reasoning tokens and increases latency. ・Most existing reflection steering methods add a label-derived mean-difference direction across preset layers, but its entanglement with reasoning and length signals de
cs.LG updates on arXiv.org

Refusal geometry reflects refusal training: diverse refusal prefixes can raise stable rank and weaken refusal vector ablation attacks

・arXiv:2608.25390v1 Announce Type: new Abstract: Refusal training protects AI models from jailbreaks by training models to decline unsafe queries, reducing the risk of misuse. ・Recent work finds that refusal behavior in aligned language models can be mediated by a single activation direction or a low-dimensional refusal subspace shared across harmful prompts: ablating those directions suppresses refusals while largely
cs.LG updates on arXiv.org

Reinforcement Learning on Benign Facts Amplifies Leakage of Memorized Private Data

・arXiv:2608.21727v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) is deployed to make models better at reasoning tasks, but its side effect on what models will divulge is under studied. ・Here we show that RLVR on facts increases extraction of personally identifiable information (PII) the instruct model had already memorized. ・We first confirm that instruct models have already mem
cs.LG updates on arXiv.org

Representing MAX functions using two-hidden-layer ReLU networks

・arXiv:2608.25221v1 Announce Type: new Abstract: We study exact representations of $\mathrm{MAX}_N(x)=\max{x_1,\ldots,x_N}$ using two-hidden-layer ReLU neural networks. ・This problem has been studied in recent years in an attempt to characterize the exact number of hidden layers required to represent continuous piecewise linear functions. ・The best lower bound is 2, while the current upper bound is logarithmic in $N$.
cs.LG updates on arXiv.org

Resilient Decentralized Wireless Federated Learning via Gradient Tracking with AdamW

・arXiv:2608.25535v1 Announce Type: new Abstract: Wireless Internet-of-Things (IoT) edge networks require decentralized learning (DecL) methods that can operate reliably under both heterogeneous local data and communication-constrained wireless links. ・However, existing decentralized optimization schemes often incur substantial communication overhead and degraded performance when transmissions are constrained by strict
cs.LG updates on arXiv.org

Resolving Multi-Modal Regression by Difference-Quotient-Based Clustering:Fast Coarse Conditional-Label Assignment

・arXiv:2608.25467v1 Announce Type: new Abstract: Multimodal regression suffers from the mean-collapse pathology: under squared loss, an unconstrained regressor converges to the conditional mean, which for K > 1 lies away from all modes. ・We attribute this failure to pairwise contradictions--samples with nearly identical inputs but distant outputs--and propose Difference-Quotient Clustering (DQC), which partitions data
cs.LG updates on arXiv.org

Resource-Efficient Pruning for Transformer via Low-Rank Importance Estimation

・arXiv:2608.24973v1 Announce Type: new Abstract: With the rapid development of large-scale pre-trained language models based on Transformer architectures, their high computational and memory costs have become a major obstacle to deployment, especially in resource-constrained environments. ・Traditional pruning methods typically depend on full gradient-based importance estimation, and they necessitate prior finetuning of
cs.LG updates on arXiv.org

Rethinking the Transferable Adversarial Attacks and Robust Defense in Federated Learning

・arXiv:2608.25133v1 Announce Type: new Abstract: The development of federated learning (FL) techniques has helped improve the privacy preservation of users' data and extended the applications of machine learning models. ・However, the involvement of a large number of users in FL also creates open opportunities for different adversaries, such as poisoning attacks, Byzantine attacks, and adversarial example attacks.
Hugging Face Papers

RetrievalRouter: Joint Modality and Architecture Selection for Document Retrieval

RetrievalRouter: Joint Modality and Architecture Selection for Document Retrieval
cs.LG updates on arXiv.org

Retrieved But Not Reliable: A Survey on Attacks, and Defenses in Retrieval-Augmented Generation

・arXiv:2608.24977v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) enhances large language models by grounding outputs in external knowledge, improving factuality and reducing hallucinations. ・At the same time, the retrieval-augmented pipeline introduces new robustness and security risks, including corpus poisoning, backdoor attacks, privacy leakage, and fairness violations. ・Despite rapid progress
cs.LG updates on arXiv.org

Reward Score Matching: Unifying Reward-based Fine-tuning for Flow and Diffusion Models

・arXiv:2604.17415v4 Announce Type: replace Abstract: Reward-based fine-tuning steers a pretrained diffusion or flow-based generative model toward higher-reward samples while remaining close to the pretrained model. ・Although existing methods are derived from different perspectives, we show that many can be written under a common framework, which we call reward score matching (RSM). ・Under this view, alignment becomes sc
cs.LG updates on arXiv.org

Robust CurveMoE: Multi-Norm Adversarial Defense for Mixture-of-Experts Models via Mode Connectivity

・arXiv:2608.26043v1 Announce Type: new Abstract: Multi-norm adversarial defense aims to protect neural networks against perturbations defined by different norm constraints, but existing methods typically optimize competing robustness objectives within a single parameter configuration, leading to substantial training cost and unfavorable robustness trade-offs. ・We propose Robust CurveMoE, an efficient mixture-of-experts
cs.LG updates on arXiv.org

Rollout-Decoded Reconstruction for Long-Horizon Prediction in Latent World Models

・arXiv:2608.25017v1 Announce Type: new Abstract: A latent world model trains its decoder on latents anchored to observations, then deploys it on the model's own free-running rollout, hundreds of steps past the last observation. ・Rollout-Decoded Reconstruction (RDR) closes this gap with a single loss term that free-runs the model during training exactly as evaluation will, decodes every rollout latent, and penalizes rec
cs.LG updates on arXiv.org

ROMNet: a hybrid reduced order modeling and machine learning approach to waveform inversion

・arXiv:2608.25160v1 Announce Type: cross Abstract: Waveform inversion seeks to estimate the wave speed of a heterogeneous, inaccessible medium, from time-resolved measurements of the waves at user controlled sensors. ・We consider this inverse problem for acoustic waves and an active array of source/receiver sensors that emit probing signals and measure the generated pressure waves. ・The forward map, from the wave speed
cs.LG updates on arXiv.org

Rotary Position Encodings for Graphs

・arXiv:2509.22259v5 Announce Type: replace Abstract: We study the extent to which rotary position encodings (RoPE), a recent transformer position encoding algorithm broadly adopted in large language models (LLMs) and vision transformers (ViTs), can be applied to graph-structured data. ・We find that rotating tokens depending on the spectrum of the graph Laplacian efficiently injects structural information into the atten
Zennのトレンド

RTX 5090 + RAM 128GBでQwen3.8-Flash-Nextをllama.cppで動かしてみた

・検証日:2026年8月27日 Qwen3.8-Flash-Nextは125B規模のMoEモデルですが、RTX 5090 32GBとRAM 128GBの1台構成でも動作しました。 ・UnslothのUD-Q2_K_XL量子化とQwen3.8-Flash-Next対応版のllama.cppを使い、64kコンテキストで設定を詰めた結果、短文生成で約48.0 tokens/s、約16kトークンの入力後でも43.3 tokens/sを記録しました。prefillは約1,429 tokens/s、約16kトークン入力時のTTFT(最初のトークンが返るまでの時間)は約11秒です。 ・この記事ではモデルの出...
Hugging Face Papers

Rubrics as Visual-Repair Context for Self-Evolving UI-to-Code Generation

Rubrics as Visual-Repair Context for Self-Evolving UI-to-Code Generation
cs.LG updates on arXiv.org

Same-Player Verification for Account Consistency in Counter-Strike 2

・arXiv:2608.24893v1 Announce Type: cross Abstract: In competitive first-person shooter (FPS) games such as Counter-Strike 2 (CS2), account-integrity review often asks whether an account's recent behavior remains consistent with its historical operator. ・This consistency question arises in cases such as temporary substitution, rank boosting, and high-skill players using lower-ranked accounts, where manual review require
cs.LG updates on arXiv.org

SAME: Stabilized Mixture-of-Experts for Multimodal Continual Instruction Tuning

・arXiv:2602.01990v3 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) achieve strong performance through instruction tuning, but real-world deployment requires them to continually expand their capabilities, making Multimodal Continual Instruction Tuning (MCIT) essential. ・Recent methods leverage sparse expert routing to promote task specialization, but we find that the expert routing process suf
cs.LG updates on arXiv.org

Sample Margin-Aware Recalibration of Temperature Scaling

・arXiv:2506.23492v2 Announce Type: replace Abstract: Recent advances in deep learning have significantly improved predictive accuracy. ・However, modern neural networks remain systematically overconfident, posing risks for deployment in safety-critical scenarios. ・Current post-hoc calibration methods face a fundamental dilemma: global approaches like Temperature Scaling apply uniform adjustments across all samples, intro
cs.LG updates on arXiv.org

SAMpLE: A SystemC-AMS Machine LEarning-based Framework for Virtual Prototyping

・arXiv:2608.25910v1 Announce Type: cross Abstract: Machine Learning (ML) is increasingly used in virtual prototypes of embedded systems to model behaviors that are difficult to capture analytically. ・However, integrating ML models into virtual platform simulation is still typically done through ad hoc solutions, which limits reuse, comparability, and reproducibility. ・This paper presents \textbf{\textit{SAMpLE}}, an ope
cs.LG updates on arXiv.org

SAUSS: Stochastic Approximation with Unbiased Simulated Scores for Limited Dependent Variable Models

・arXiv:2608.25304v1 Announce Type: cross Abstract: Multinomial choice models allow flexible substitution patterns but become computationally demanding with many alternatives or observations. ・With a fixed per-observation simulation budget, simulated maximum likelihood introduces simulation bias, while each optimization step requires a full-sample likelihood evaluation. ・We propose Stochastic Approximation with Unbiased
cs.LG updates on arXiv.org

Scalable Multi-GPU Simulation of 3D Multicellular Growth with RNN-Based Workload Balancing

・arXiv:2608.25890v1 Announce Type: cross Abstract: Detailed multicellular growth simulations based on subcellular element models (SEMs) can capture complex tissue development, but their element-level interactions impose substantial computational cost. ・This work presents a scalable multi-GPU framework for 3D multicellular growth simulation that combines GPU acceleration, spatial binning, domain decomposition, and workl
cs.LG updates on arXiv.org

Scalable Self-Supervised Learning for Multiphase AC-OPF in Distribution Systems with Topology Reconfiguration

・arXiv:2608.25095v1 Announce Type: cross Abstract: The proliferation of distributed energy resources (DERs) in distribution grids enables the active coordination of these assets to reduce costs and enable cleaner operations. ・Realizing this potential requires solving multiphase AC optimal power flow (AC-OPF) quickly across varying loads, DER availabilities, and topology reconfigurations, at much greater speed and scale
cs.LG updates on arXiv.org

SciMIF: Understanding Multimodal Instruction Following in Scientific Domains

・arXiv:2608.25973v1 Announce Type: cross Abstract: Understanding instruction-following capabilities in scientific domains is essential for effectively leveraging Multimodal Large Language Models (MLLMs) to advance the development of scientific fields. ・In this work, we introduce SciMIF, a novel benchmark designed to evaluate the capability of MLLMs in following complex scientific instructions. ・Specifically, based on an
cs.LG updates on arXiv.org

Scorpio: Serving Right Requests at the Right Time for Heterogeneous SLOs in LLM Inference

・arXiv:2505.23022v2 Announce Type: replace Abstract: Large Language Model (LLM) serving increasingly underpins online Web services such as conversational agents, Web search, and programming assistants, where requests carry heterogeneous Service Level Objectives (SLOs) such as Time to First Token (TTFT) and Time Per Output Token (TPOT). ・Existing LLM serving systems prioritize maximum throughput and treat all requests u
Zennの「機械学習」のフィード

Seq2Seq LSTMはランダムウォークを超えられるか | 第2回:AIで為替の未来予測は本当にできるのか?

・最近は Transformer や LLM が話題ですが、時系列予測では今でも LSTM 系モデルがよく使われています。 ・第1回ではシンプルなLSTMを使って未来の価格変動を予測、そこから価格を計算しました。 ・LSTMによるドル円予測の基礎実験 今回は、LSTMを改良した「Seq2Seq(Sequence-to-Sequence)」という構造を使って、為替価格の未来予測を試してみました。
cs.LG updates on arXiv.org

SHSP: Structure-Aware Hierarchical Solution Prediction for Mixed-Integer Linear Programming

・arXiv:2608.25282v1 Announce Type: new Abstract: Mixed-Integer Linear Programming (MILP) is a fundamental optimization paradigm in combinatorial optimization and has been widely applied across real-world domains. ・Due to its NP-hard nature, obtaining optimal solutions for large-scale or highly constrained MILP instances remains computationally prohibitive. ・Learning-based solution prediction has therefore emerged as a p
cs.LG updates on arXiv.org

ShuttleArena: Interpretable Self-Play in Physics-Based Badminton

・arXiv:2608.25246v1 Announce Type: new Abstract: Badminton is a compact but challenging domain for game AI: a player must choose a physically feasible shuttle trajectory, anticipate the opponent's interception, and recover to a court position whose value depends on the opponent's next response. ・The central challenge is that shot selection and recovery are not separable: the best recovery depends on the shot-induced op
cs.LG updates on arXiv.org

Simulating Cognitive Smart Freight Corridors with Agent-Based Models and Reinforcement Learning

・arXiv:2608.25193v1 Announce Type: cross Abstract: Smart freight corridors offer a practical pathway for connected and automated vehicle (CAV) deployment in freight transportation, but physical experimentation is expensive and existing approaches rely on predefined control policies that cannot capture adaptive behaviors. ・This paper presents an agent-based modeling (ABM) framework coupling a physical infrastructure lay
cs.LG updates on arXiv.org

Simultaneous inference of environmental and interaction forces in collective dynamics

・arXiv:2608.25181v1 Announce Type: new Abstract: Collective dynamics arise in a wide range of physical, biological, and engineering applications. ・Examples include cell migration, swarm robotics, social dynamics, and animal behavior. ・A defining characteristic of these systems is the emergence of large-scale coordination from local interactions among agents; a fundamental question is thus to understand the local interac
Hugging Face Papers

Skill Issue: Are Skills Language-Invariant in LLMs?

Skill Issue: Are Skills Language-Invariant in LLMs?
cs.LG updates on arXiv.org

Skill Issue: Are Skills Language-Invariant in LLMs?

・arXiv:2608.25832v1 Announce Type: cross Abstract: Large language models access knowledge inconsistently across languages, but to what extent do they differ in their skill sets when interacting with different languages? ・This work quantifies cross-lingual skill inconsistency orthogonally from knowledge and general benchmark performance. ・We do this via multilingual self-play: two instances of the same model compete in a
cs.LG updates on arXiv.org

SNAP-KG: Streaming Node Assignment via Projection for Knowledge Graph Entity Integration

・arXiv:2608.25149v1 Announce Type: new Abstract: Knowledge graph (KG) construction pipelines must continuously integrate newly arriving entities into a growing graph. ・Unlike inserting triples between existing nodes, a newly arriving entity has no graph connectivity: it emerges from the acquisition phase as a raw feature vector and must be assigned to a semantic community before entity resolution and link prediction ca
The Verge

Sony finally has a cheaper OLED to compete with midrange Samsung and LG TVs

・After unexpectedly showing up on a wall-mounting compatibility chart on Sony's site back in June, the company has announced the Sony Bravia 6 OLED TV. ・The TV sits below its other OLED TVs - the Bravia 8 and Bravia 8 II - as well as its Bravia 9 II and Bravia 7 II RGB LED TVs. ・The Bravia 6 OLED comes in five sizes from 48 up to 83 inches and is priced to compete with the LG C6 OLED and Samsung S90H QD-OLED TVs.
cs.LG updates on arXiv.org

Spatio-temporal dual-stage hypergraph MARL for human-centric multimodal corridor traffic signal control

・arXiv:2602.17068v2 Announce Type: replace Abstract: Human-centric traffic signal control in corridor networks must increasingly account for multimodal travelers, particularly high-occupancy public transportation, rather than focusing solely on vehicle-centric performance. ・This paper proposes STDSH-MARL (Spatio-Temporal Dual-Stage Hypergraph based Multi-Agent Reinforcement Learning), a multi-agent deep reinforcement l
cs.LG updates on arXiv.org

Spectral Allocation: Why Muon Outperforms Adam, and How to Improve Muon

・arXiv:2608.25990v1 Announce Type: new Abstract: Orthogonal optimisers such as Muon can substantially accelerate large language model pretraining relative to Adam, yet the mechanism remains incompletely understood. ・We investigate this through an out-of-sample spectral probing analysis of Transformer loss landscapes. ・At checkpoints along real training trajectories, we decompose each momentum buffer into its singular di
Zennの「大規模言語モデル」のフィード

Speculative Decodingを実運用で選ぶ — MTP・DFlash・DSparkの違いと使い分け

・はじめに 社内でPowerPoint自動生成パイプラインをDify上で構築しており、生成モデルには NVIDIA Nemotron 3.5 Lightning(30B、アクティブパラメータ3BのMoEモデル)のBF16版を、DGX Sparkでホストして使っています。 ・ある日、生成速度が 25 tok/秒 で頭打ちになっていることに気づきました。原因を調べていく過程で「Speculative Decoding(投機的デコーディング)」という手法に出会い、実際に手を動かして検証したので、その過程をまとめます。 ・なぜ25 tok/秒で止まるのか ── メモリ帯域幅の壁 DGX Spa...
The Verge

Speedo’s new smart goggles module can track all four swim strokes

・You can pair Speedo’s Vanquisher goggles with its iQ module. ・| Image: Speedo Speedo is launching a removable smart module for its Vanquisher goggles that can analyze a swimmer's performance across all four swim strokes: freestyle, backstroke, breaststroke, and butterfly. ・The new Speedo iQ system uses swim-specific sensors to analyze head movement in the water, time breaths, and track distance per stroke.
cs.LG updates on arXiv.org

StablePDENet: Enhancing Neural Operator Stability through Physics-Informed Residual-Sensitivity Regularization

・arXiv:2601.06472v2 Announce Type: replace Abstract: Learning solution operators for differential equations with neural networks has shown great potential in scientific computing, but ensuring their stability under input perturbations remains a critical challenge. ・We introduce the StablePDENet, a physics-informed adversarial training method that regularizes the residual sensitivity with respect to an input perturbatio
WIRED

Starz Promo Codes: $5 Off for September 2026

・Ready to stream award-winning series, hit movies, and exclusive originals? ・Our comprehensive guide helps you find every active Starz coupon, free trial, and discount code to save big on your subscription this 2026.
WIRED

Stop Touching Your Keyboard. Use This AI-Powered Microphone Instead

・The Relay Q, due next year, is the latest attempt to reposition voice as the most seamless method for human-computer interaction.
cs.LG updates on arXiv.org

StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models

・arXiv:2604.15416v2 Announce Type: replace Abstract: Sign-based optimization algorithms, such as SignSGD, have garnered attention for their performance in distributed learning and training large foundation models. ・Despite their empirical superiority, SignSGD is known to diverge on non-smooth objectives, which are ubiquitous due to ReLUs, max-pools, and mixture-of-experts. ・To overcome this limitation, we propose StoSig
Hugging Face Papers

Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models

Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models
Hugging Face Papers

StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models

StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models
WIRED

Submit Your Questions: The Great Data Center Backlash

・You have questions about data centers, and WIRED has answers. ・Join our livestream on September 10 and our panel of experts will tell you everything you need to know.
Hugging Face Papers

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?
cs.LG updates on arXiv.org

TacForcing: Streaming Action Generation with Execution-Time Tactile Feedback

・arXiv:2608.25798v1 Announce Type: cross Abstract: Contact-rich manipulation requires adapting to contact states that can evolve substantially within an action horizon. ・However, chunk-based vision-language-action models predict complete action chunks from observations collected before execution, leaving tactile conditioning stale during execution. ・Existing tactile-reactive approaches typically rely on separate high-fr
cs.LG updates on arXiv.org

TailSFT: Filtered Fine-Tuning Improves Post-Training Performance

・arXiv:2608.25756v1 Announce Type: new Abstract: Reinforcement learning post-training drives reasoning and agentic capabilities in modern AI systems, yet a growing body of work shows that it is most effective when used to fine-tune an already capable base model. ・We question whether existing pipelines yield models that are most suitable for reinforcement learning. ・Building on prior work highlighting the role of coverag
WIRED

The Best Google Pixel Phones of 2026: Comparison, Features, and Accessories

・Let us help you choose the right Pixel phone. ・Plus, check out our Pixel accessory recommendations and smart software tricks to try.
cs.LG updates on arXiv.org

The Dialect Tax: Dialectal Biases Persist throughout the Language Modeling Pipeline

・arXiv:2608.24952v1 Announce Type: cross Abstract: Systematic dialectal performance gaps in language models (LMs) are well documented, but the source of these disparities within the modern language modeling pipeline remains unclear. ・Our study traces this "dialect tax" across the natural language processing pipeline. ・Using parallel English dialect corpora that hold meaning fixed while varying surface form, we first con
cs.LG updates on arXiv.org

The Frame Kernel Method for Multiscale Operator Learning

・arXiv:2608.25084v1 Announce Type: new Abstract: We present a natively multiscale operator learning method for the surrogate modeling of (numerical solvers for) multiscale partial differential equations (PDEs). ・The primary novelty of our method lies in a novel multiscale kernel frame function approximation technique. ・Leveraging this new kernel frame technique, we cast the operator learning problem as one of learning f
Hugging Face Papers

The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents

The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents
cs.LG updates on arXiv.org

The Imperfective Paradox Is Not Necessarily in Large Language Models: A Benchmark Failure Before a Model Failure

・arXiv:2608.25005v1 Announce Type: cross Abstract: The imperfective paradox provides a useful test of compositional semantic analysis. ・Recent work constructs an NLI benchmark and reports that models frequently infer completed telic events from progressive descriptions, attributing this behavior to a Teleological Bias. ・It further argues that prompting interventions cause a Calibration Crisis.
WIRED

The UK Power Grid Has a Phantom Data Center Problem

・The UK’s energy regulator is using a variety of tricks to keep speculative data center projects from plugging into the power grid. ・The country’s AI ambitions hang in the balance.
cs.LG updates on arXiv.org

The Von-Neumann State-Space Transformer for neural decoding

・arXiv:2608.25088v1 Announce Type: new Abstract: Cortical computation is strikingly low-dimensional: a handful of latent variables, carried in a neural population's activity, steer the higher-dimensional responses of individual neurons. ・Our aim is sample efficiency-models that decode well from limited data and at small parameter budgets. ・In a standard Transformer layer, the feed-forward block applies the same operator
cs.LG updates on arXiv.org

Theoretically Principled Federated Learning for Balancing Privacy and Utility

・arXiv:2305.15148v3 Announce Type: replace Abstract: We propose a general learning framework for the protection mechanisms that protects privacy via distorting model parameters, which facilitates the trade-off between privacy and utility. ・The algorithm is applicable to arbitrary privacy measurements that maps from the distortion to a real value. ・It can achieve personalized utility-privacy trade-off for each model para
cs.LG updates on arXiv.org

Token-Oriented Semantic Communication with Pretrained Vision Transformers

・arXiv:2608.25410v1 Announce Type: cross Abstract: Token communications realize the semantic communication principle at the granularity of transformer tokens, providing a promising direction for client--server collaborative inference in resource-constrained edge systems. ・However, directly transmitting token embeddings presents two practical challenges: substantial communication cost and limited interoperability across
WIRED

TopResume Packages: Everything You Need to Get Hired

・Discover ways to save at TopResume, including their free review service and 4-week Career Services Platform trial.
cs.LG updates on arXiv.org

Toward Machine Learning with the Unit as a Primitive: Learning from Unit-Linked Events

・arXiv:2608.25118v1 Announce Type: new Abstract: Machine learning is usually formalized through samples, while the persistent individual to which multiple observed or possible events refer often remains implicit. ・We propose the \emph{unit} as an explicit primitive at the level of task semantics. ・A learning task first declares a population of persistent referents and a sameness criterion; the realized value $u$ denotes
cs.LG updates on arXiv.org

Towards A Unified Information Bottleneck Framework for Time Series Explanations

・arXiv:2608.25897v1 Announce Type: new Abstract: Explaining deep learning models operating on time series data is crucial in various applications that require transparent and interpretable insights into model behavior. ・{Existing explanation methods generally fall into two categories: attribution-based explanations, which identify the temporal regions most responsible for a prediction, and counterfactual explanations,
cs.LG updates on arXiv.org

Towards Reliable, Generalizable, and Specific In-Context Knowledge Editing via Multi-Objective Reinforcement Learning

・arXiv:2608.25100v1 Announce Type: cross Abstract: Large Language Models (LLMs) are powerful but limited by static parametric knowledge that becomes outdated once pretraining ends. ・Knowledge editing addresses this problem by updating model behavior on target facts without full retraining. ・In particular, in-context knowledge editing has gained attention because it is training-free and readily applicable to black-box LL
cs.LG updates on arXiv.org

Towards Robust and Scalable Density-based Clustering via Graph Propagation

・arXiv:2605.00390v2 Announce Type: replace Abstract: We present \textit{CluProp}, a novel framework that reimagines varied-density clustering in high-dimensional spaces as a label propagation process over neighborhood graphs. ・Our approach formally bridges the gap between density-based clustering and graph connectivity, leveraging efficient propagation mechanisms from network science to mitigate the parameter sensitivi
cs.LG updates on arXiv.org

TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development

・arXiv:2608.26086v1 Announce Type: new Abstract: Large language models write correct code for isolated problems but remain far weaker at autonomous machine-learning development, where an agent must revise data pipelines, models, and validation over hours of feedback, and on most competitions still finishes below strong human competitors. ・Outcome-based benchmarks record this gap but not its cause, because they grade th
cs.LG updates on arXiv.org

Training Alignment Auditors via Reinforcement Learning

・arXiv:2608.25460v1 Announce Type: cross Abstract: Alignment auditing of frontier models increasingly relies on LLM auditors to surface undesirable behaviors at scale, but current automated auditors can struggle with coherent investigation and audit realism. ・In this work, we improve LLM auditors with reinforcement learning. ・In our best training environment, the policy investigates target models that potentially posses
cs.LG updates on arXiv.org

Transforms for LLM Quantization: The Great Inversion and Format Co-Design

・arXiv:2608.25188v1 Announce Type: new Abstract: Most competitive 4-bit LLM research pipelines now open the same way: apply a linear, function-preserving transform (rotation, scaling, permutation, non-orthogonal affine) so the outlier mass sits more favorably against the group scales, and only then round. ・Yet we are aware of no survey dedicated to this transform stage, and its literature is quietly re-deriving an olde
cs.LG updates on arXiv.org

Tropospheric temperature and humidity profile retrieval from Meteosat Flexible Combined Imager based on deep learning

・arXiv:2608.25700v1 Announce Type: new Abstract: The Meteosat Third Generation (MTG) Flexible Combined Imager (FCI) offers new opportunities for tropospheric temperature and humidity profiling, at higher spatio-temporal resolutions and expanded spectral coverage relative to its predecessor. ・Vertically resolved retrievals from broadband imagers are inherently challenging, and operational retrieval algorithms typically
cs.LG updates on arXiv.org

Trust the Mass: Forced Weights in KV-Cache Eviction

・arXiv:2608.25230v1 Announce Type: new Abstract: Every deployed sparse-attention or KV-cache-eviction rule keeps a subset of the keys, discards the rest, and renormalizes the attention weights over the kept set. ・Enumerating the exact best subset under that constraint on $168{,}192$ attention rows from five models shows that keeping the largest weights is already near-optimal, since the best subset closes only a median
cs.LG updates on arXiv.org

TrustFormer: Cross-Temporal and Cross- Dimensional Transformer for Task-Specific Multi-Dimensional Trust Evaluation

・arXiv:2608.25238v1 Announce Type: cross Abstract: In dynamic collaborative systems, the selection of reliable collaborators is critical to ensuring effective task execution. ・Existing trust evaluation methods often rely on unidimensional or scalar representations, which fail to faithfully capture a collaborator's true trustworthiness, thereby motivating a shift toward multi-dimensional trust modeling. ・However, due to
cs.LG updates on arXiv.org

Two Dimensions Govern Agnostic Multiclass Transductive Learning

・arXiv:2608.25326v1 Announce Type: new Abstract: In transductive classification, an adversary fixes a labeled population, one label is hidden uniformly, and the learner sees all remaining labels. ・For binary classes, agnostic transductive and PAC learning have the same minimax rate. ・Whether this extends to multiclass learning was open, especially for unbounded label spaces where uniform convergence can fail.
cs.LG updates on arXiv.org

Understanding the Energy Scaling of Large Language Model Inference Across Context Lengths and Attention Architectures

・arXiv:2608.25096v1 Announce Type: new Abstract: The growing adoption of large language models (LLMs) has raised increasing concerns about the energy consumption and environmental impact of inference. ・This paper presents a systematic empirical study of decode-phase energy consumption across representative open-source LLMs employing Multi-Head Attention (MHA), Grouped Query Attention (GQA), and Grouped Query Attention
cs.LG updates on arXiv.org

Unfolding Scientific Papers into Multi-Turn Generation Trajectories for Continued Pre-Training

・arXiv:2608.25826v1 Announce Type: cross Abstract: A recent line of synthetic-data work reconstructs the thinking behind existing text rather than rewriting the text itself, but it operates on short web passages, recovers only local thoughts, and leaves the structure of whole documents untouched. ・Scientific papers are written to a clear and largely uniform structure and make a natural substrate for lifting this paradi
cs.LG updates on arXiv.org

Unsupervised Anatomical Feature Learning via Diffusion Models: Enhanced Medical Image Segmentation with Denoising Diffusion Probabilistic Models

・arXiv:2608.25693v1 Announce Type: cross Abstract: Acquiring pixel-level annotations for medical image segmentation is a severe bottleneck. ・Traditional U-Net architectures, while effective, learn local texture patterns and lack awareness of global anatomical structures, leading to boundary delineation failures in low-data regimes. ・This research paper proposes utilizing unsupervised Denoising Diffusion Probabilistic Mo
cs.LG updates on arXiv.org

Unsupervised Post-Training of Foundation Models: A Survey

・arXiv:2608.24982v1 Announce Type: cross Abstract: Foundation-model post-training usually relies on human labels, preference data, stronger teachers, or executable verifiers. ・We study Unsupervised Post-Training (UPT): update-bearing adaptation on unlabeled inputs whose learning signal is derived from same-lineage model artifacts rather than an external oracle. ・We catalog 80 strict UPT methods and organize them by the
Hugging Face Papers

V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning

V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning
Hugging Face Papers

VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning
cs.LG updates on arXiv.org

VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

・arXiv:2608.26105v1 Announce Type: cross Abstract: Native visual reasoning treats visual generation as the medium of reasoning itself: visual states (i.e. ・images and videos) are not merely inputs to be understood or outputs to be rendered, but first-class substrates for problem solving beyond language. ・Yet progress remains bottlenecked by the lack of scalable training tasks, reliable feedback, and controlled compariso
Hugging Face Papers

VGI-BENCH: Probing Visual Intelligence in Video Generation Models

VGI-BENCH: Probing Visual Intelligence in Video Generation Models
Hugging Face Papers

Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios

Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios
cs.LG updates on arXiv.org

VINCENT: Validated Interaction Network for Cross-drug Explanation of Therapeutics

・arXiv:2608.25841v1 Announce Type: new Abstract: Drug synergy prediction estimates whether two drugs produce a stronger joint effect than expected from their individual activities. ・For drug combination discovery, a single synergy score is often not enough: researchers also need to know which molecular regions jointly drive the prediction. ・We study motif-pair synergy explanation, which identifies pairs of chemically co
cs.LG updates on arXiv.org

Virgil: Navigating Explainability for Transformer-based Language Models

・arXiv:2608.25555v1 Announce Type: cross Abstract: Explainability for transformer-based language models is becoming crucial as these systems are deployed in high-stakes applications. ・As a result, the ecosystem of explainability tools is rapidly evolving, becoming richer, but also more fragmented and harder to navigate. ・To address this challenge, we present Virgil, an interactive system that lets practitioners and rese
Hugging Face Papers

VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction

VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction
Hugging Face Papers

WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation

WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation
cs.LG updates on arXiv.org

WAVE: Reversing the Guidance Hierarchy for Coarse-to-Fine Guided Depth Super-Resolution

・arXiv:2608.25302v1 Announce Type: cross Abstract: Guided depth super-resolution (GDSR) typically extracts RGB guidance features through convolutional hierarchies, inheriting their fine-to-coarse bias. ・Thus, low-level spatial cues surface in early layers, leaving the deeper layers to suppress those that do not correspond to true depth boundaries, which risks artifacts and blurred edges. ・The same fine-to-coarse bias pe
cs.LG updates on arXiv.org

What Should a Large Language Model See? Physical Invariants as a Data Representation for PDE Discovery

・arXiv:2608.25189v1 Announce Type: new Abstract: Understanding how molecular interactions govern macroscopic behaviour is a central challenge in molecular sciences. ・However, conventional theory building cannot keep pace with the vast datasets modern experimentation routinely produces. ・Large language models offer a promising route to automating theory construction, but a spatiotemporal field cannot be directly placed i
AI | VentureBeat

When agents act on their own, governance has to live in the data layer

・Presented by EDB As enterprises give AI agents more autonomy — the ability to plan, decide, and act across systems without a human approving each step — a hard question moves to the center of every architecture review: When an agent tries to complete an action that it was never authorized to do, what actually stops it? ・These are your agents, running on your models, touching your data in your infrastructure — and the
cs.LG updates on arXiv.org

When Does Context Routing Help? A Systematic Study of Multi-Modal Fusion in Time Series Forecasting

・arXiv:2608.25128v1 Announce Type: new Abstract: Multi-modal time series forecasting methods integrate auxiliary context into temporal predictions through increasingly sophisticated fusion mechanisms. ・A growing body of work reports substantial gains, yet it is often unclear whether they reflect genuine use of the context or incidental architectural effects. ・We ask a narrower, checkable question: when can auxiliary con
cs.LG updates on arXiv.org

When Does Frequency Decomposition Benefit Physics-Informed Neural Networks? A Preliminary Ablation Study

・arXiv:2608.24940v1 Announce Type: new Abstract: Partial differential equations (PDEs) often have high-frequency and multi-scale features that neural networks struggle to approximate. ・Physics-Informed Neural Networks (PINNs) build the governing equations directly into training, but suffer from spectral bias: they learn low-frequency components faster than high-frequency ones. ・Techniques such as Fourier feature embeddi
cs.LG updates on arXiv.org

When Pruning Meets Interpretability: Preserving Sparse Autoencoder Robustness in LLMs

・arXiv:2608.25941v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) are widely used to interpret the internal representations of large language models (LLMs), yet their reliability under post-hoc model compression remains poorly understood. ・We present a systematic study of how pruning affects SAE behavior and theoretically show that, for a fixed SAE, its impact is governed by perturbation energy, a covariance-
cs.LG updates on arXiv.org

Why and When Neural Networks Improve Local Approximation in Optimization

・arXiv:2608.24963v1 Announce Type: new Abstract: Published experience with neural surrogates in derivative-free optimisation is contradictory: the same family of models that cuts the evaluation count of one solver leaves another unchanged, or makes it worse. ・We show that the contradiction dissolves once three factors are stated, and that these, rather than the fit accuracy a training curve reports, are what delimit wh
cs.LG updates on arXiv.org

Why Does Graph Learning Fail to Fully Benefit from a Text Teacher?

・arXiv:2608.25741v1 Announce Type: new Abstract: Graph neural networks (GNNs) are widely used to represent complex interactions and relationships among entities. ・We investigate a multimodal model that combines two complementary ideas: a self-supervised method that enables a GNN encoder pretrained on one dataset to operate directly on another dataset with a different node-feature dimensionality, without rebuilding the
cs.LG updates on arXiv.org

Why ML-based cough models do not generalize: a systematic cross-dataset evaluation for tuberculosis screening

・arXiv:2608.25846v1 Announce Type: cross Abstract: Cough acoustics are promising for non-invasive tuberculosis (TB) screening, yet whether machine learning (ML) models capture disease-related acoustics or artifacts of data collection remains unresolved. ・We evaluated the cross-dataset generalizability of classical ML and deep learning (DL) cough-based TB classifiers across three independent datasets. ・Despite moderate w
#AIタグ

アフィリエイトリンクを貼ってみて、最初の反応はこうだった

・先日初めて アフィリエイトリンク付きの投稿をしました。 ・「収益ゼロの僕が初めて アフィリエイトに挑戦してみた話」の続きです。 ・今回は、実際に投稿してみて 何が起きたのかを正直に書いていきます。
ITmedia NEWS 最新記事一覧

キオクシア、1兆円超の新棟建設報道に言及 「さまざまな検討を行っていることは事実」

・キオクシアホールディングスは8月27日、傘下のキオクシアが北上工場(岩手県北上市)に新棟を建設するとの報道について「当社およびその子会社が発表したものではない」とコメントした。一方、新棟建設については、「さまざまな検討を行っていることは事実」と認めた。
#LLMタグ

コーディングエージェントのCLI移行で見落としていた「トークンコスト」の自前計算術

・この記事の結論 CLIの`--json`出力に含まれる最終イベントからトークン使用量を機械的に取得し、SDKに依存しないコスト計測を実現する。 ・「キャッシュ入力」を分離した計算式と外部の単価表を用いることで、正確なドル建てコストを自前で算出する。 ・価格改定やプロモーション期間に対応するため、単価表に日付とバージョンを付与し、計算根拠をログとして残す運用を徹底する。
#LLMタグ

シンガポールにおけるAIと著作権 / AI学習例外の安全弁の置き場 / robots.txtの評価が分かれる理由 / 発明者確定後に残る問い 雑感

シンガポールにおけるAIと著作権 / AI学習例外の安全弁の置き場 / robots.txtの評価が分かれる理由 / 発明者確定後に残る問い 雑感
Zennの「機械学習」のフィード

そのID、日付だと思っていませんか ― 例外が出ないリークの話

・短い記事です。例外が出ないので気づけないタイプのリークの話をします。 ・起きたこと 競馬の予測モデルで「前走からの間隔(週)」という特徴量を作っていました。 ・レースには race_id という12桁のIDが振られています。
#AIタグ

タロットで心を整えるノート【17】

・第4章 AIと一緒にマインドノートを深める ③自分の「思考の癖」って気づいていますか? 続きをみる
ITmedia NEWS 最新記事一覧

ちっちゃな人型ロボが100mトラックをぽてぽて……中国ロボ陸上の“おちびランナー”が話題 「可愛すぎ」「けなげ」

・中国・北京で8月22日から26日(現地時間)まで開催された、人型ロボットの性能を競う「世界人型ロボット運動会」。陸上のウサイン・ボルト選手越えのスピードを発揮するロボなどが注目を集める一方で、100mトラックをぽてぽてと歩くとある小型人型ロボが可愛らしいとSNSの話題をさらっている。
ITmedia NEWS 最新記事一覧

トヨタ、28年にAI自動運転車を投入 「E2E」と事前ルール規定を組み合わせるハイブリッド方式

・トヨタ自動車は27日、AIが周囲の認識から運転の判断、操作までを担う自動運転車を28年に実用化すると明らかにした。米Teslaなどが先行する「エンド・ツー・エンド(E2E)」方式で、日産自動車やホンダも搭載車の発売を予定する。
機械学習タグが付けられた新着記事 - Qiita

なぜAIは、専門家でなくても使えるようになったのか

・AIにおすすめの店を聞いたら、実在しない店名が返ってきたことがありませんか?根拠として挙がった論文名も、検索するとどこにもありませんでした。 ・不思議なのは、その同じAIが、込み入った質問にはきちんと答えることです。知らないなら知らないと言えばいい。知っているのに、なぜ細部だ...
#AIタグ

はじめに|課金前ラボは「買う前」に調べるメディアです

・AIツールやブログ運営ツール、副業系の教材やサービス。 ・気になるものを見つけても、販売ページを読んでいるうちに、 続きをみる
Zennの「機械学習」のフィード

バックテストが実測で消える6つの罠 ― 回収率120%を8回作って8回失った記録

・過去データで何かを当てようとしたことがある人向けの記事です。 ・私は競馬の予測モデルを作っています。そして**「回収率120%」のバックテストを8回作り、8回とも失いました。** 数字 消えた理由 直したあとの値 124.7% 未来を見ていた 無効 148.9% 未来 + 標本不足 無効 109.8% 賭けた後の情報で選んでいた 82.8% 105.2% 未来を見ていた 84.7% 124.8% たくさん試して選んだ 棄却 149.2% 集計ミス 100.2% 132.2% 賭けた後の情報 95.6% 110.0% 検証と本番でモデルが違った 82...
LLMタグが付けられた新着記事 - Qiita

プロンプトを打つ時代から、ループを設計する時代へ

・一度の指示で、最後まで仕事を終えてほしい。AIを使い始めた頃、僕はずっとそう考えていました。 ・質問を打ち、答えを待ち、違えば指示を足す。うまくいかないほどプロンプトは長くなり、例外と禁止事項が増えていく。それでも、少し違う仕事を渡すとまた外れます。正直、僕は指示の書き方が足...
機械学習タグが付けられた新着記事 - Qiita

因果推論 Day 14/全30回 反事実の計算、SCMで「もしあの時」を解く

・この連載について 因果推論を「本を読んだ」で終わらせず、自分の言葉で説明でき、コードで再現できる状態まで落とす30日連載です。前回のDay 13では、交絡因子がどうしても観測できなくても、処置からアウトカムへの影響を媒介変数のチェーンで丸ごと捉えられれば因果効果を点で識別...
#AIタグ

何を覚えて何を忘れる?AIの脳内の「ゲート機構」

・第8回(通算第38回)(学ぶキーワード:ゲート機構) 前回までに学んだLSTMやGRUは、「必要なことを覚え、不要なことを忘れる」ことができる優秀なAIでした。
ITmedia NEWS 最新記事一覧

価格高騰が変えた法人PC調達 「売上2倍」のリユースPC事業者が見た需要の変化

・新品PCの価格が上昇する中、企業の調達先としてリユースPCが存在感を増している。2026年3~4月の売上が前年同月比約2倍になったという、リユースPC専業事業者「リングロー」にリユースPC市場でいま何が起きているのかを聞いた。
#AIタグ

稼ぎ方を調べていたら、記事の構造のほうが気になった

・今年の八月に、営業の仕事を辞めました。 ・辞めてから、収入をどうつくるかを考えていろいろ調べました。 ・副業、アフィリエイト、コンテンツ販売、物販。
機械学習タグが付けられた新着記事 - Qiita

階層型言語モデル PHOTON を論文から実装し、RTX4090で再現できたこと・できなかったこと

・PHOTON という階層型言語モデルを、公開論文を手掛かりに実装してきました。公式実装の移植ではなく、論文の数式と付録表から構造を組み直し、vanilla Transformer や Block Transformer と同じパイプラインで比較するための検証実装です。
Zennの「大規模言語モデル」のフィード

空白駆動の流体力学 ― 空圧機器とボルテックスチューブから読み解く知性の生存戦略

・第1章 はじめに:なぜ歯車(剛体)ではなく空気(流体)なのか 情報処理の論理を語るとき、私たちは長らく「剛体(かたい機械)」のメタファーに支配されてきました。 ・計算機とは本質的に、歯車とクラッチの集合体です。入力された電気信号が確定的な噛み合わせを経て、寸分の狂いもなく出力へと伝達される。そこでは1は1であり、0は0です。剛体メカニズムの世界において、部品の欠損や密度の偏りはただの「破損」であり、歯車と歯車のあいだにある「隙間」は、埋められるべきエラーにすぎません。 ・しかし、大規模言語モデル(LLM)をはじめとする現代の知性システムに触れたとき、私たちはこの剛体モデルに対する根本的な...
#AIタグ

最短で考えるチャットAIのロール

・**※今回はこれまでに比べてかなりマニアックな内容になっております。ご注意ください。※** (あと、202608時点での内容となってます。) 導入 続きをみる
ITmedia NEWS 最新記事一覧

初の“指輪型たまごっち”の秘密、開発者に聞く 「驚きのあるデバイスを」「前面ボタンは飾りです」

・バンダイは27日、たまごっち発売30周年モデルとして、初のリング型デバイス「Tamagotchi ring」をお披露目した。シリーズ初の「指にはめてお世話ができるたまごっち」はどうやって生まれたのか?
#LLMタグ

推奨火力より少ない火力トークン量で効率よく高難度を解かせる ── SphereOS-Atlantis DOSが引き出す、財布(と地球)に優しいAIエンジニアリング

推奨火力より少ない火力トークン量で効率よく高難度を解かせる ── SphereOS-Atlantis DOSが引き出す、財布(と地球)に優しいAIエンジニアリング
Zennの「大規模言語モデル」のフィード

大道至簡:CodexやClaude Codeを超える極小AIエージェント「Pi」徹底解剖

・Claude Code や Codex、Cursor など、AI コーディングエージェントの進化は目覚ましいものがあります。 ・しかし、これらのツールを使っていると、こんな違和感を抱いたことはないでしょうか。 ・「裏で勝手に大量のコンテキストやシステムプロンプトが注入されて、トークンがあっという間に溶ける」 「複雑な承認プロンプトやサブエージェントが空回りして、肝心のコード修正が進まない」 「機能が多すぎてブラックボックス化し、思い通りに挙動をカスタマイズできない」 こうした過剰な全部入り(All-in-one)エージェントへのアンチテーゼとして世界中のギークから熱狂的な支持を集めてい...
#AIタグ

地元のお店にインスタDM営業をかけてみた話

・前回までのあらすじを軽く書いておくと、クラウドワークスで3件応募して受注ゼロだった、という記事を以前書きました。入札の激戦区で消耗するより、競合がほぼいない地元の個人店に直接営業をかけたほうが可能性があるんじゃないか、というのがそこからの流れです。今取り組んでいるのは「地元の個人店(家具屋さん・アクセサリー店など)向けに、商品カタログの更新・運用を代行する」という新しい挑戦です。ホームページやSNSはあるのに、商品が一つ一つ載っていなくて何を売っているのか分からない、というお店の不便を解決できないか、という仮説から始めました。
#AIタグ

通知は毎日来るのに、読まれた回数は43だった

・AIに個人事業を運営させている記録です。今回は、数字を見て前提が崩れた話を書きます。
#LLMタグ

電子の海でAIを叫ぶような話

・LICO :イルカ、そこに愛はあるんか。 ・IRUKA:ツッコミませんよ。モノマネとか苦手なんでしょ。
#LLMタグ

道具に纏わり付くAIという幻影 :思索スケッチ

・「AI」に対する幻想とは何だろうか、これは新しい時代によってもたらされたユニークな錯覚だと感じるかもしれないが、実はそうではないだろう。 ・「常識的」な道具の有用性評価は想定される機能があり、その認知において妥当と思われる結果があれば「ちゃんと機能している」と評価される、だが多くの場合においてそのメカニクスや個別のメカニズムが精査されたうえでの評価では全く無い。 ・例えば写真機普及草創期には「魂が抜かれる」という評価が「常識」として在ったうえでその結果は受け入れられてきたし、現代に至ってなおスマホの写真が「写真」であると疑いもなく認知されているのが実情だ。
#LLMタグ

日誌(2026.08.24)

日誌(2026.08.24)
#AIタグ

日本の AI 成果は 9%? それとも 86.7%? — 同じ年の日本で、両方が正解だった

・ニュースアプリやネットを開くと、AI についての調査結果がよく流れてきます。 ・「日本企業で AI の成果が出ているのは 13%」。厳しいな、と思う。その数日後に、別の見出しが流れてくる。「AI を使っている企業の 86.7% が効果を実感」。あれ、と思う。さらに別の記事には 9% と書いてある。
#LLMタグ

日本のAI利用率は58.8%で、6.4%です。どちらも本当の数字なので、一次資料まで確かめました

・はじめまして。介護と支援の相談どころ「そよぎ」のヒロです。 ・「日本のAI利用率は◯%」という数字を、最近よく見かけます。ところがこの◯に入る数字は、記事によって58.8%だったり、35%だったり、6.4%だったりします。いちばん大きい数字といちばん小さい数字では、9倍ちがいます。
#LLMタグ

話題のグラフエンジニアリングについて、これまでの〇〇エンジニアリングを追ってみた

・最近、AIエージェント界隈で「グラフエンジニアリング(Graph Engineering)」っていう言葉、やたら見かけませんか? プロンプト → コンテキスト → ハーネス → ループ ときて、次はグラフらしい、と。 ・で、自分もこの流れを整理する記事を書こうとしてたんです。5階層にきれいに並べて、「AIは個から組織へ」みたいな締めにするつもりでしたが、、、 続きをみる
Zennの「大規模言語モデル」のフィード

捏造禁止ゲートの安全税を自分のパイプラインで4回測った — 観点の件数は減らず、出所の誤りが残っていた

・自分はQAエンジニアで、普段のテスト設計に生成AI (Claude) を使っています。 ・ちなみに、この記事の文章自体もClaudeに書いてもらいました。中身自体は見て確認してたり、多少の文の付け足しはしています。 ・この内容の集積自体はAIに調べた上で概要程度の目は通しています。しかしながら、専門家ではないため引用や解釈などが間違っている可能性はあります。