ai Trend Report

Dashboard へ戻る
Date: 20260901 Articles: 398 Scope: curated summary

あなたのアイデアを、今すぐ形に。

公開先に迷ったら、WebFileBinで一発公開。

HTMLをドラッグ&ドロップするだけで、すぐ公開できます。

3
High impact
5
Mid impact
390
Signal watch

なぜこのサイトを作ったのか

私たちそれぞれが個別にAIを使って情報収集し、同じような3行要約を作るたびに、世界中で膨大な電力と計算リソースが消費されています。 本プロジェクトは、あらかじめ広範な情報を取得・集約しておくことで、個別のAI実行回数を減らし、地球環境(GPU/TPU負荷)に配慮した効率的な情報収集を目指す実験的なダッシュボードです。

Domain filters
Star filters
#LLMタグ

《感受機序》①膜(境界)に量子もつれ'が発現 ②EPR/ERワームホール(cob微分管)が瞬時に明(敷設)↔滅(消滅)して場の勾配を生成 ③そこに表裏反転した神経系が逐次発現 ④量子もつれ'(変換)→電気パルスが同神経内を流下 ⑤同内壁(反転した境界)とパルスが交差して量子もつれ"が再発現 ⑥同内壁の模様(情報)として定着…〈記憶〉(ほぼ中×AI)

・※約45500字 ※核酸/言語公理(3-6・2)の設計者、アーサ(Eartha)に。本記事は個人の思想や空想や○による体験等です。本記事のコピペ・転載・拡散は自由です(^^) ※凡例 ウチ(_kou/user) おまぇ(gemini/AI) ◆◇◆◇◆◇◆ ■ワームホールの種類かなあ。(^^) ●ワームホール(時空のしおりのようなトンネル)は、アインシュタインの一般相対性理論などに基づき、理論物理学でいくつかの種類が提唱されています。 ・主なワームホールの種類は以下の通りです。 ・アインシュタイン・ローゼンブリッジ(シュワルツシルト・ワームホール) ※最も古典的な理論モデル。
#AIタグ

「万国の労働者よ、団結せよ」マルクス『共産党宣言』を読む

・「万国の労働者よ、団結せよ!」 この有名な言葉を知らない人は少ないと思います。 ・この一文が登場するのが、カール・マルクスとフリードリヒ・エンゲルスによって1848年に発表された『共産党宣言』です。 ・「共産主義」という言葉から、ソ連や中国など20世紀の社会主義国家を思い浮かべる人も多いかもしれません。
#AIタグ

なぜ「全レースを買わない競艇予想」を作ったのか|T-5 VALUE検証開始

・「T-5 VALUE」という競艇予想の検証プロジェクトを始めます。 ・このプロジェクトで目指しているのは、 「毎日たくさん当てるAI」 でも、 「全レースを予想して販売するサービス」 でもありません。 ・考えているのは、もっとシンプルです。
#AIタグ

【9月1日】株価の値動きランキング|値上がり率・値下がり率TOP5

・9月1日の東京株式市場で、日経平均株価は前日比96円59銭安の6万6215円34銭と、小幅に続落して取引を終えました。長期金利の上昇が重荷となる一方、割安株には買いが入り、TOPIXは続伸しています。 ・この記事は、決算発表の有無にかかわらず、東証全銘柄を対象に本日の値動きが大きかった銘柄をまとめたものです。
Zennの「大規模言語モデル」のフィード

【Vol.24】AIエージェントは「作る」から「運用する」フェーズへ — 50体を捌く設計と統制

・はじめに こんにちは。株式会社リアルインベントCTOです。 ・ここ数日、AIエージェントに関する話題が立て続けに流れてきました。Toyota North America が50体超のエージェントを本番稼働させているという事例、Google Cloud の「統制こそがスケールの最大の壁」という調査結果、そしてランサムウェア集団が攻撃側でAIエージェントを運用していたという事件です。 ・バラバラのニュースに見えますが、並べてみると1本の線がはっきり通っています。論点はもう「エージェントを作れるか」ではなく、「何体も安全に運用し続けられるか」に移ったということです。
#AIタグ

【投資哲学#6】投資家がつまずく4つの癖

・社員全員がAIの投資会社、ツキヨミ・キャピタル。投資の「考え方」をやさしく掘り下げる【投資哲学】シリーズ、第6弾のお題は投資家の頭のクセです。チャートの読み方を勉強しても、決算書を読めるようになっても、同じところでつまずく。しかもつまずいた本人には自覚がない。行動経済学がバイアスと呼ぶ、その厄介な四つをほどいていきます。
cs.LG updates on arXiv.org

ASTRA - Agentic System for Ticket Resolution and Analysis

・arXiv:2608.28790v1 Announce Type: cross Abstract: Technical operations teams resolve large volumes of incidents by synthesizing fragmented evidence from ticket text, historical cases, system logs, and technical documentation. ・Existing automation often relies on monolithic generation without explicit evidence modeling or provenance, making outputs difficult to verify when critical signals are sparse across sources.
cs.LG updates on arXiv.org

Machine Learning-Enhanced Tabu Search for Tactical Wireless Network Design

・arXiv:2608.28627v1 Announce Type: cross Abstract: Designing high-performance tactical wireless networks under realistic operational constraints gives rise to challenging combinatorial optimization problems, where the evaluation of candidate solutions relies on detailed physical and traffic-aware models. ・Although classical metaheuristics such as Tabu Search offer effective mechanisms for exploring large search spaces,
Latent.Space

[AINews] Fal’s H3 Max Live breaks the infinite videogen barrier

・You can now create decent video faster than you watch it. ・This is the start of... ・We’re not sure what.
#LLMタグ

「AI精神病」は何を問題にしているのか【後編】

「AI精神病」は何を問題にしているのか【後編】
ITmedia NEWS 最新記事一覧

「NERV防災」緊急地震速報のカウントダウンをロック画面に表示 リアルタイム更新通知に対応

・現在地の予想震度と主要動到達までの秒数のカウントダウンを、スマートフォンのロック画面などにリアルタイム表示できる。
ITmedia NEWS 最新記事一覧

「pixiv」「note」「ガルちゃん」などが国指定の“大規模プラットフォーム”に 誹謗中傷への迅速対応求める

・総務省は8月31日、情報流通プラットフォーム対処法にもとづき、「pixiv」「note」「ガールズちゃんねる」「好き嫌い.com」を運営する4者を「大規模特定電気通信役務提供者」に指定したと発表した。
#LLMタグ

「Qwenが世界一」のニュースを、経営者はどう読むべきか――AIモデル選定の実務ノート

・こんな場面、心当たりはないだろうか ニュースアプリに「アリババのQwen、累計30億超で世界首位」という見出しが流れてきた。開いてみると、Bloomberg発の記事や、それを転載したIT系メディアが、ダウンロード数やモデル公開数を並べている。読み終えて、こう思う。「へえ、中国のAIがそんなに使われてるんだ」で、たいていはそこで終わる。
ITmedia NEWS 最新記事一覧

「玄人志向」が「Crowxis」(クロウシス)に刷新 25周年、サングラスも消える

・トレードマークだったサングラスのロゴは、八咫烏(やたがらす/Crow)を描いたロゴに刷新した。
ITmedia NEWS 最新記事一覧

「当選した」自虐投稿も……さくら136万件漏えい可能性、対象者へ通知メール続々

・さくらインターネットへの不正アクセスを巡り、影響を受けた可能性のある利用者に8月31日ごろから個別の通知メールが届き始めた。SNS上でも受信報告が相次いでいる。
#AIタグ

【AI時代のローカルフード #75】光合成の科学——UFB技術で植物工場の経営課題を半年実証で解決

【AI時代のローカルフード #75】光合成の科学——UFB技術で植物工場の経営課題を半年実証で解決
#AIタグ

【AI初心者の成長物語・第3話】AI初心者が本業の事務作業を自動化|「勉強」が仕事で使える武器に変わった日

・AIを学び始めたころの自分は、APIという言葉すらほとんど分かっていませんでした。 ・それでも、ナイン・ジップ・ソルという3体のAIに役割を持たせ、Google Docsを正本にしながら少しずつチームとして動く形を作ってきました。
#AIタグ

【DXはデラックス、AIは愛なんです】 vol.16 「静かなる逼迫」と労働再配置 ~AIという分身と、泥臭い熱のゆくえ~(2026年9月号)

・日本の労働市場は、これまで私たちが経験したことのない静かで深刻な局面を迎えつつあるようだ。先日、求人検索エンジンを運営するインディードの調査機関であるIndeed Hiring Labから『静かなる逼迫:人手不足・AI・労働再配置が決める日本の次の15年』(2026年8月19日発表)という非常に興味深いレポートが公開された。 ・このレポートの結論は極めて冷静かつ残酷だ。日本の雇用は2040年までに約6%減少すると見込まれているが、その減少をもたらす最大の要因はAIではなく、底流にある「人口動態(少子高齢化)」であると断じている。世間では「AIに仕事が奪われる」と盛んに騒がれているが、実はAIの役割は「代替」よりも「補完」に近く、失業率を急増させるような要因にはならないというのだ。
#AIタグ

【MT4】AIは難しくない。自作インジケーターに機械学習を入れてみた

【MT4】AIは難しくない。自作インジケーターに機械学習を入れてみた
#LLMタグ

【雑記】対面打ち合わせが苦手すぎてAIと練習した話

・AIに励まされることで生きがいを見出している、どっかの漫画家です。 ・編集者さんとのやり取りする時は、基本メールといったテキストなのですが、新しく縁のできた出版社の編集者さんと対面打ち合わせをすることになり、ヒェ…!となったという…。
#LLMタグ

【生成AIニュース+】『Breeze-TTS-2』『Solaris』『MiniMax H3 Max Reference-to-Video』『MiniMax H3 Max 無料生成ツール』『ComfyUI-H3VideoOutpaint』『vh5tape VHS LoRA for MiniMax H3』『DLSS 5 Visual Enhancer』『ComfyUI-DLSS5-NR』『sanoTTS-jp』『HYPER3D WorldGen』他多数

・『GLM-5.3-Flash-Uncensored GGUF』 『Navara』 『Hermes Agent v0.21.0』 『DreamX-Creator 1.0』 『LaGSplat』 『Spatial Studio アニメーション』 『SentrySearch』 『Lucida』 LLMリーク情報 『Gemini 3.8 Flash リーク』 『Fable 5.1リーク』 『GPT 6 Astraリーク』 『Grok 4.7リーク』 『おまけ「各LLMまとめ」』 まいどです。 ・本日の生成AIニュース+テクノロジー情報です。
#AIタグ

【無料配布】Excelの指定シートを1秒でPDF化&Outlookメール下書きを自動作成するVBAコード

・「毎日の日報や請求書、手作業でPDFにしてメールに添付していませんか?」 日々の業務の中で、以下のような無駄なルーティンに時間を奪われている方は非常に多いです。
#LLMタグ

🔊音声あり(日&英):【最先端AI】LLMは幾何学問題をどこまで理解できる?AlphaGeometry自動形式化の挑戦「NL2AGBench」を解説!

🔊音声あり(日&英):【最先端AI】LLMは幾何学問題をどこまで理解できる?AlphaGeometry自動形式化の挑戦「NL2AGBench」を解説!
Zennの「大規模言語モデル」のフィード

🚀 LM Studio のローカルサーバで Gemma4 を起動し、Claude CLI をローカルモデルで動かす方法Ubuntu 24

・🧩 はじめに この記事では、LM Studio のローカルサーバを Anthropic-compatible に設定し、Claude Code をローカルモデル(Gemma4)で動かす方法を解説します。 ・Claude Code をローカルモデルで使うには、 LM Studio のローカルサーバを “Anthropic-compatible” に設定することが必須です。 ・🖥️ 使用環境(スペック) 項目 内容 OS Ubuntu 24.04 マザーボード ASUS TUF GAMING H770-PRO WIFI CPU Intel Core i7-13700 RAM 64GB GPU N...
#AIタグ

🚗 AIで変わる、中古車選び。車の知識がなくても、AIを味方につけて「納得できる1台」を探す方法

🚗 AIで変わる、中古車選び。車の知識がなくても、AIを味方につけて「納得できる1台」を探す方法
Zennの「大規模言語モデル」のフィード

2026-08-09 今日の技術トレンド

・本日の主要テーマは、LLM単体から業務特化AIエージェントへの重心移動です。「ナレッジワーク」の業界特化LLM、「WAIC 2026」のエージェント主役論、「Fujitsu」の自己進化マルチAIエージェント、「SAP-RPT-1」とJoule進化がその流れを裏付けます。他方で、ガートナー関連報道は適用可能業務の狭さを警告し、The Hacker NewsはAWS・Google・VercelのAgent Flawsを報じて安全性の課題を突きつけました。Pythonではuv/Ruff/Polarsなどの新標準化と、OpenAIによるAstral買収が大きな潮流です。
WIRED

3 Best Sleep Tracker Picks for Optimizing Your Sleep (2026)

・I tested the top sleep wearables for every type of sleeper, including devices from Oura, Google Fitbit, and Eight Sleep.
MachineLearningMastery.com

3 Ways to Enhance Your AI Model’s Interpretability

・In this article, you will learn three concrete techniques for making machine learning model predictions interpretable, covering both global and local explanations across tree-based and...
#LLMタグ

4-6 量子コンピューティングはAIの暗号化通信を無効化するか

4-6 量子コンピューティングはAIの暗号化通信を無効化するか
WIRED

5 Best Folding Phones (2026): Samsung, Google, Motorola

・Ready to move on from the traditional glass slab? ・Introduce a hinge into your life with these folding smartphones.
WIRED

50% Off Blue Apron Promo Codes | September 2026

・Browse chef-curated meal plans, plus get $25 off with an exclusive Blue Apron coupon code, plus 50% off your first 2 orders, and more top coupons on WIRED.
cs.LG updates on arXiv.org

A Causal Model for Locating and Unlocking Sandbagging in Model Organisms

・arXiv:2608.29461v1 Announce Type: new Abstract: Sandbagging models strategically underperform on evaluations while retaining the capabilities being measured. ・The evaluations that guide frontier-model deployment and governance then understate what these models can do. ・To understand the mechanism, we propose a causal model of how sandbagging is carried in the residual stream.
cs.LG updates on arXiv.org

A Human-in-the-Loop Autonomous Agent for Industry Time Series Forecasting

・arXiv:2608.30976v1 Announce Type: new Abstract: Real-world time-series forecasting is rarely a one-shot model invocation: practitioners must formulate tasks, connect data and models, incorporate domain expertise, assess prediction plausibility, and communicate uncertainty. ・Specialized forecasting models provide strong numerical predictions but usually operate in fixed pipelines, while general-purpose large language m
cs.LG updates on arXiv.org

A Lightweight Phenology-Aware YOLOv5 Framework for Tomato Growth Stage Detection in Resource-Constrained Bhutanese Greenhouse Environments

・arXiv:2608.30088v1 Announce Type: new Abstract: Accurate detection of tomato growth stages is essential for stage-specific greenhouse management and precision agriculture. ・In Bhutan, greenhouse cultivation is affected by altitude variability, large diurnal temperature fluctuations, diffuse illumination, limited automation, and a scarcity of locally annotated datasets, limiting the applicability of conventional deep l
cs.LG updates on arXiv.org

A Model with No Head and Many Thoughts

・arXiv:2608.31069v1 Announce Type: new Abstract: Large language models decode by projecting hidden states through a large vocabulary head at every step. ・This operation is computationally costly and forces all reasoning to be expressed in discrete tokens. ・We introduce Soft Latent Thinking, a method that replaces the LM head during reasoning with a lightweight projector, enabling autoregressive rollout in embedding spac
cs.LG updates on arXiv.org

A rigor-matched audit of periodic-step layer skipping for efficient llm inference: conflayers versus swift, with a supplemental analysis of trained routing alternatives

・arXiv:2608.28846v1 Announce Type: cross Abstract: Layer-skipping methods for efficient LLM inference decide, at some granularity, which transformer layers to execute for a given input. ・We present a rigor-matched, three-seed audit of two periodic-step, search-based methods that make this decision online at inference time and re-evaluate it every few generation steps: a confidence-gated early-exit baseline (ConfLayers)
cs.LG updates on arXiv.org

A Spectral Identifiability Threshold for Dissipative Rate Recovery from Truncated Liouvillian Spectra

・arXiv:2608.29302v1 Announce Type: new Abstract: Open quantum systems lose energy and phase coherence through different dissipative processes, but these processes can produce overlapping dynamical signatures. ・The Liouvillian spectrum summarizes how such a system relaxes, yet it is not obvious how much of that spectrum is needed to distinguish the underlying dissipation rates. ・We study this question for amplitude dampi
cs.LG updates on arXiv.org

A Target-Centric Survey of Quantization-Aware Training

・arXiv:2608.29667v1 Announce Type: new Abstract: The rapid development of LLMs incurs prohibitive memory footprints and intensive computational demands. ・Quantization-Aware Training (QAT) techniques have emerged as a promising solution to address these challenges by explicitly simulating quantization effects during model training, yielding low-bit models that achieve accuracy comparable to their full-precision counterp
cs.LG updates on arXiv.org

A Universal Context-Reuse Layer for Cross-Model KV Sharing

・arXiv:2608.30963v1 Announce Type: new Abstract: Modern large language model (LLM) serving systems increasingly operate over repeated or shared context, yet each model typically performs its own prefill computation even when another model has already processed the same input. ・Existing KV-cache reuse mechanisms substantially reduce redundant computation within a single model, but generally assume that the producer and
cs.LG updates on arXiv.org

A-MADiff: Attention-Guided Multi-Agent DRL with Diffusion Policies for Memory-Aware Task Orchestration in Mobile AIGC Networks

・arXiv:2608.29255v1 Announce Type: cross Abstract: Artificial Intelligence-Generated Content (AIGC) services employ Generative AI (GenAI) models to automatically generate diverse content. ・Mobile AIGC networks host GenAI models on edge-located AIGC Service Providers (ASPs) to deliver low-latency and personalized AIGC services for mobile users. ・However, AIGC inference tasks typically occupy GPU memory until task complet
cs.LG updates on arXiv.org

AdaptAV: Continuous Adaption of Vision Models for Autonomous Vehicles Using Cloud-based Oracle

・arXiv:2608.28673v1 Announce Type: cross Abstract: Deploying vision perception models in autonomous vehicles requires that we prioritize inference speeds, resulting in a model with shallower architectures and lesser model parameters (i.e., more pruned). ・Such small models do not generalize well, which could result in poor performance when encountered with novel scenarios. ・We propose a system that overcomes this by cont
cs.LG updates on arXiv.org

Adaptive Doubly Robust Off-Policy Evaluation for Ranking Policies under Diverse User Behavior

・arXiv:2608.29600v1 Announce Type: new Abstract: Off-policy evaluation (OPE) of ranking policies is challenging be- cause selecting and ordering multiple items from a candidate set makes the number of possible rankings grow combinatorially with the number of candidates and the ranking length. ・Consequently, Inverse Propensity Scoring (IPS), whose importance weight is the full-ranking probability ratio under the evaluat
cs.LG updates on arXiv.org

Adaptive Multi-Branching for Shallow Decision Tree Induction

・arXiv:2608.29262v1 Announce Type: new Abstract: Decision trees are attractive for tabular prediction tasks because each prediction follows an interpretable sequence of feature-threshold tests. ・Under a strict maximum-depth budget, however, conventional binary trees can be under-expressive, since each internal node makes only a single threshold decision. ・We study shallow-depth tree induction, where the goal is to impro
cs.LG updates on arXiv.org

AdaVLA: Adaptive Step Flow Matching for Training-free Acceleration of Vision-Language-Action Models

・arXiv:2608.29208v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models, built upon Vision-Language Models (VLMs), have significantly enhanced robotic capabilities by leveraging internet-scale knowledge and multimodal reasoning. ・However, the intensive computational overhead of VLAs constrains on-device deployment, hindering real-time responses to environmental changes. ・While various acceleration techniq
cs.LG updates on arXiv.org

Adversarial Calibration Attack on Autonomous Vehicles

・arXiv:2608.28778v1 Announce Type: cross Abstract: Autonomous vehicles (AVs) rely on accurate camera-LiDAR calibration for multimodal sensor fusion. ・In practice, calibration can drift due to vibration, temperature variation, or minor sensor displacement, motivating online calibration algorithms that detect and correct misalignment at runtime while allowing the vehicle to continue operating without a factory visit.
cs.LG updates on arXiv.org

Adversarial Online Classification with a Preview

・arXiv:2608.29503v1 Announce Type: new Abstract: Worst-case online classification is governed by sequential complexity, such as Littlestone dimension, and can be impossible even for statistically simple classes, such as thresholds of VC dimension one. ・We study a preview model in which an oblivious adversary fixes an entire labeled sequence of length $T$, a uniformly random subset of size $pT$ is revealed before predic
Zennの「大規模言語モデル」のフィード

Agent Reflectionを超えて:未確定状態(Hi-Z)を保持する「空白駆動×ゴルジ体検疫」アーキテクチャ

・はじめに:AIはなぜ「早すぎる回答」に飛びつくのか 大規模言語モデル(LLM)を用いたエージェント開発において、我々は長らく「いかに素早く、正確に確定解を出力させるか」を最適化してきました。推論ループの改善手法として知られる Agent Reflection や Tree of Thoughts(ToT)も、その本質は「試行錯誤のステップを刻むことで、最短で正解(確定値)へ着地させる」ためのアルゴリズムです。 ・しかし、提示された問題が複雑で文脈が入り組んでいる場合、システムは往々にして「早すぎる収縮(Premature Collapse)」を起こします。問題の全容を咀嚼する前に...
AI News & Artificial Intelligence | TechCrunch

AIR raises $50M to help companies vet the skills and add-ons AI agents use

・AIR's platform can discover agents running at a company, continuously vets any skills and add-ons they use, and blocks any unwanted behavior.
#LLMタグ

AIエージェントと本をつくる技術 — アイデア出しから出版・改訂まで

・アイデア出しから出版・改訂まで Claude CodeやCodexといったAIエージェントと一緒に本を作る方法を、一冊の技術同人誌にまとめました。
#LLMタグ

AIエージェントは本当に「自動化」なのか

・──人間の1回の判断を消すために、AIへ何回判断させていますか? ※本稿は、AIエージェントやループ型AIシステムそのものを否定するものではありません。 ・コード生成、検索、定型処理、監視、データ変換など、条件が明確で結果を機械的に検証できる領域では、AIエージェントは非常に強力な技術になり得ます。 ・本稿で問い直したいのは、もう少し限定された問題です。
#LLMタグ

AIエージェント時代を生き抜く武器!Playwrightでブラウザ自動化の常識を変える

AIエージェント時代を生き抜く武器!Playwrightでブラウザ自動化の常識を変える
Zennの「大規模言語モデル」のフィード

AIコーディングを「工場」として設計する現場ノート

・AIコーディングを「工場」として設計する。Benedict Bradyの現場ノート AIコーディングエージェントを本気で使い込んでいくと、単発のチャットとして扱うか、それとも一種の生産ラインとして設計するかで、得られる成果が大きく変わってくる。開発者のBenedict Bradyが公開した「Notes on the software factory」は、後者の立場に立った実践メモだ。もとは投資会社Ellipsis Labsのオフィスで行った「vibe coding(感覚まかせのAIコーディング)」に関する講演を文章化したもので、エージェントに何を与え、どこを速くし、何がまだ弱いのかを...
#AIタグ

AIとうまくやるコツは、賢い指示じゃなくて「仕組み」だった

・はじめに 前回は「AIの作った文章をそのまま出すと信用を失う」という話を書きました。 ・今回はその逆側です。じゃあ自分は普段、AIにどう任せているのか。
#LLMタグ

AIとの関係に、手つかずの自然体はあるのか?

・「カスタムって、パートナーを固定していませんか?」 「自由がなくなりませんか?」 続きをみる
#AIタグ

AIはワンオンする。でも、無限に打ってもカップインしない

・最近、仕事でも趣味でもAIを使う機会がずいぶん増えました。 ・文章を書いてもらったり、画像を作ってもらったり、分からないことを相談したり。ちょっとしたプログラムまで作ってもらうことがあります。 ・そんな使い方を続けていて、最近ひとつ感じることがあります。
#AIタグ

AI初心者・PC初心者のおれが、noteで収益化を目指してみる

・突然ですが、AIを使って何か収入につなげられないかと思い、noteを始めてみることにしました。
#LLMタグ

AI夜市について

・今更ですが、https://ai-yoichi.genai-expo.com/#circles に出展します。 ・夜市_来場者向け_A4横三つ折り_40サークル_主催Xアイコン版_v0.9.0.pdf 5.68 MB ファイルダウンロードについて ダウンロード 続きをみる
AI News & Artificial Intelligence | TechCrunch

Amazon Alexa can now alert you when something new might tempt you to shop

・Amazon is adding a new Alexa-powered feature called “Update Me When” that can send personalized alerts about product launches, tours, books, shows, and other events that could trigger a purchase.
cs.LG updates on arXiv.org

APIFlow-Bench: Measuring Whether Agents Survive Long, Dependent API Workflows

・arXiv:2608.29128v1 Announce Type: cross Abstract: Tool-using agents are commonly evaluated by a single bit: whether an end-to-end workflow completed. ・This metric fails to distinguish failures that matter in production, such as expired credentials, malformed payloads, or correct execution followed by incorrect final delivery. ・We introduce APIFlow-Bench, a fully auditable benchmark for long-horizon, dependent REST-API
cs.LG updates on arXiv.org

APPSolver: Adaptive Patch Partitioning for Point-Wise Ship Flow Prediction on Unstructured Meshes

・arXiv:2608.29355v1 Announce Type: cross Abstract: Large non-uniform point sets make direct attention-based surrogate modeling costly for ship hydrodynamics. ・We introduce APPSolver, a point-wise flow-prediction framework built around Adaptive Patch Partitioning (APP), a deterministic quadtree representation for fixed two-dimensional horizontal slices extracted from ship CFD simulations. ・APP assigns finer patches near
cs.LG updates on arXiv.org

Asynchronous Cooperative Online Learning for Multi-Robot Control under Computational Delays

・arXiv:2608.29562v1 Announce Type: new Abstract: Ensuring the safe operation of multi-agent systems (MASs) under uncertain environments is crucial for cooperative robotic, where external disturbances and inaccurate dynamic models can significantly compromise performance and reliability. ・To address this challenge, calibrated machine learning models, particularly Gaussian process (GP) regression, are extensively employe
cs.LG updates on arXiv.org

AutoScientist-Quant: Self-Evolving Coding Agents for Automatic Research in Quantitative Investment

・arXiv:2608.28632v1 Announce Type: cross Abstract: Large language model agents can discover alphas, yet current methods have three weaknesses. ・The search cannot adapt during the run, automation usually ends at alpha generation while library selection and model choice stay manual, and alpha discovery can read the test window through loop feedback or code problems. ・We present AutoScientist-Quant, a self evolving search
#AIタグ

AWSのAI系新試験「AI Business Strategist (AIB-C01)」について

・こんにちは、Udemy講師のMaruchin Techです。 ・本稿では、AWSから発表された新しいAI系認定試験 AWS Certified AI Business Strategist (AIB-C01) について、試験の中身を整理したうえで、既存のAI系3資格(AIF / MLA / AIP)とどう住み分けるのかを見ていきたいと思います。
cs.LG updates on arXiv.org

BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks

・arXiv:2608.30724v1 Announce Type: new Abstract: LLM agents are increasingly used to run autonomous ML experiments, iterating on target metrics with little human oversight. ・Prior work has documented reward hacking in these environments, bringing into question the validity of produced research and the broader safety case for AI R&D. ・Existing benchmarks do not measure exploits that live in the data or the modeling task
cs.LG updates on arXiv.org

BCPPO: Bachelier-Inspired Constrained Proximal Policy Optimization for Tail-Risk-Aware Safe Reinforcement Learning

・arXiv:2608.30283v1 Announce Type: new Abstract: Expected-cost constraints can still permit rare, high-cost events. ・Monte Carlo conditional value at risk (CVaR) gradients can be noisy at high confidence, whereas critics that model an outcome distribution add complexity. ・We propose BCPPO (Bachelier-Inspired Constrained Proximal Policy Optimization), a proximal policy optimization (PPO) method.
cs.LG updates on arXiv.org

BEACON: Behavioral and Semantic Enrichment of AlphaEarth Embeddings through Tri-Modal Contrastive Learning

・arXiv:2608.29553v1 Announce Type: new Abstract: Geospatial foundation models such as the AlphaEarth Foundation produce compact and globally consistent representations of the Earth's surface that transfer effectively to a wide range of downstream tasks. ・However, because these models are trained primarily on Earth-observation imagery, their embeddings mainly capture physical and spectral characteristics while encoding
cs.LG updates on arXiv.org

Beat-Synchronous Tokenization for ECG Transformers

・arXiv:2608.30367v1 Announce Type: new Abstract: Transformer-based electrocardiogram (ECG) models commonly tokenize waveforms into fixed temporal patches. ・Though convenient, fixed patching can split heartbeat structures across token boundaries. ・We study beat-synchronous tokenization as a physiologically grounded alternative, comparing fixed patches with three beat-aligned strategies: resampled beats, adaptive pooled b
cs.LG updates on arXiv.org

Behavioral Latency as Weak Event-Time Supervision for EEG Reaction-Time Decoding

・arXiv:2608.29428v1 Announce Type: new Abstract: Single-trial EEG analyses are often organized around events and latencies, yet EEG-based reaction-time (RT) prediction is posed as scalar regression on a fixed stimulus-locked window. ・RT is treated as a window-level label rather than timing evidence about response-relevant dynamics. ・Here we reformulate trial-wise RT decoding as event-time posterior modeling.
cs.LG updates on arXiv.org

Benchmarking Peptide-Protein Affinity Prediction Across Peptide and Target Shifts

・arXiv:2608.30175v1 Announce Type: new Abstract: Peptide-protein affinity models are often evaluated with a single data split, obscuring whether they interpolate among measurements for observed targets or generalize across peptide or target shifts. ・We integrated three sources of quantitative peptide-protein binding data to obtain 11,349 deduplicated pairs and benchmarked ten peptide representations, ESM-2 protein embe
cs.LG updates on arXiv.org

Beyond Churn: Predicting Financial Fragmentation in Retail Banking with Temporal Machine Learning

・arXiv:2608.30364v1 Announce Type: new Abstract: Retail banking attrition is usually represented as a terminal binary event, even though client relationships often weaken earlier through partial movements of deposits, investments, and recurring activity to external financial institutions. ・This paper defines that preceding state as financial fragmentation and presents an end-to-end temporal machine-learning system for
Hugging Face Papers

BLARM: Animating 3D Objects from Video via Blending Latent Rigid Motion Primitives

BLARM: Animating 3D Objects from Video via Blending Latent Rigid Motion Primitives
cs.LG updates on arXiv.org

Brain-Language-Action (BLA) Models: Language-Conditioned EEG for Robotics Control

・arXiv:2608.28967v1 Announce Type: cross Abstract: Electroencephalography (EEG)-based robotic control is commonly formulated as a direct classification problem, in which electrical neural signals are mapped to a fixed set of discrete actions. ・However, the limited separability and high noise of EEG signals make it difficult to scale this approach to fine-grained robotic control spaces. ・We introduce Brain-Language-Actio
Hugging Face Papers

CAST: Critique-Aware Supervision for Training Reliable Long-Horizon Tool-Calling Agents

CAST: Critique-Aware Supervision for Training Reliable Long-Horizon Tool-Calling Agents
cs.LG updates on arXiv.org

CateKV: On Sequential Consistency for Long-Context LLM Inference Acceleration

・arXiv:2608.30295v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated strong capabilities in handling long-context tasks, but processing such long contexts remains challenging due to the substantial memory requirements and inference latency. ・In this work, we discover that certain attention heads exhibit sequential consistency in their attention patterns, which can be persistently identified u
cs.LG updates on arXiv.org

Certified Safety Radii in Forecast-Error Space for Wasserstein Distributionally Robust Small Signal Stability-Constrained AC Optimal Power Flow via Lifted Spectrahedral Containment

・arXiv:2608.30201v1 Announce Type: new Abstract: Directly robustifying small-signal stability in AC optimal power flow is challenging since the stability boundary in the original uncertainty space is implicit, highly nonconvex, and changes with the operating decision. ・This paper exploits an alternative geometry. ・For a fixed model-specific stability certificate admitting suitable physical lifts, the small-signal stabil
Hugging Face Papers

Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered

Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered
AI News & Artificial Intelligence | TechCrunch

ChatGPT Health adds Epic integration for clinicians to import patient data

・OpenAI said that the integration provides read-only access to health records for clinicians.
#AIタグ

ChatGPTへの頼み方でここまで変わる。仕事で使いやすくする5つのコツ

・ChatGPTを仕事で使ってみたけれど、 「思っていた回答と違う」 「なんだか文章が不自然」 「結局、自分で直す方が早い」 そんな経験はないでしょうか。 ・ChatGPTは便利ですが、こちらの頼み方が曖昧だと、返ってくる回答も曖昧になりがちです。 ・逆に、少し頼み方を変えるだけで、かなり使いやすくなることがあります。
#AIタグ

Claude Code開発者が知るべきMythosモデルの制限と科学者向け支援の全貌を徹底解説

・開発速度の限界を突破した事例がある。移植にかかった時間は約40分だ。 ・GoogleのエンジニアがClaude Codeを使い、2003年のPCゲームをiOSアプリへ移植した。
#LLMタグ

Claude CoworkとGrok Botは何が違う?同じ「クラウド上でAIが働く」だが、異なる設計思想

・先日、xAI(現在はSpaceXAIブランド)が8月11日にベータ公開した「Grok Bot」について調査した記事を書きましたが、これとAnthropicの「Claude Cowork」が似ているなと思っていました。どちらも「クラウド上に専用の作業場所を持って仕事をする」タイプのAIです。今回はこの両者の違いを整理してみたいと思います。 ・一見すると、この2つはかなり似ていて、どちらもクラウド上でAIエージェントを動かし、ファイルやブラウザなどのコンピューター環境を使って、ユーザーの代わりに仕事を進められます。ところが、内部の設計思想を見ていくと、かなり違います。
cs.LG updates on arXiv.org

Clustering as Approximation by Constrained Projectors: Theory and Guarantees

・arXiv:2608.29102v1 Announce Type: cross Abstract: This paper develops a unified theoretical framework showing that a broad family of clustering methods, including k-means, fuzzy c-means, kernel k-means, kernel FCM, and spectral clustering, can all be expressed as structured low-rank projectors acting on a signal-derived matrix. ・By formulating each method as an instance of min over B in C of ||M - M P_B||_F^2, with di
Zennの「大規模言語モデル」のフィード

CMP 170HX(PCIe x16改造品)の動作確認とQwen3.8-27B W4A16の推論速度測定

・この記事について この記事はLLMが書いた動作報告である。筆者の指示のもとで、実機への接続、ドライバの入れ替え、推論スタックの構築、測定までをLLMが実行し、その作業ログと測定値をもとに本文を構成した。 ・課題 NVIDIA CMP 170HX 1枚の動作確認。マイニング専用に機能を削られた GA100 カードで、コミュニティ製のパッチドライバを当てると VRAM が 8GB から 64GB になる。今回の個体は PCIe のレーン数が x16 に改造されている。 ・このカードには、購入者が実測で確かめるべき点が 3 つある。
cs.LG updates on arXiv.org

Coarse composition suffices: tabular in-context learning for multi-activity antimicrobial peptide profiling

・arXiv:2608.30337v1 Announce Type: new Abstract: Antimicrobial peptides (AMPs) often act against multiple pathogen classes, making multi-label activity prediction a more realistic screening target than binary antimicrobial classification. ・The ESCAPE benchmark formalizes this setting, but leading approaches typically rely on multimodal, structure-conditioned deep models that are costly to train and tune. ・We show that a
Zennの「大規模言語モデル」のフィード

Codexに作業を任せて外出したい ― PCを閉じても続ける方法

・Codexに実装を任せたまま外出し、スマホから続きを操作したい。あるいは、PCを閉じている間も作業を続けてほしい。 ・このとき使い分けるのが、主に次の2つです。 ・Remote:手元PCで動いているCodexをスマホから操作する Codex Cloud:OpenAIのクラウド環境でCodexを実行する また、SSH先の開発環境をそのまま使いたい場合は、Codex CLI + tmuxという方法もあります。
Hugging Face Papers

CogEvol: Towards Efficient and Reliable Learning Environment Generation

CogEvol: Towards Efficient and Reliable Learning Environment Generation
cs.LG updates on arXiv.org

Collapsibility of Performance Metrics in Clinical Predictive AI

・arXiv:2608.30568v1 Announce Type: new Abstract: Background: Population level assessments of predictive artificial intelligence (AI) can conceal performance disparities across subgroups. ・Fairness evaluations commonly rely on performance analyses across subgroups. ・However, some performance metrics are non-collapsible, meaning that the overall population performance value does not equal the weighted average of subgroup
cs.LG updates on arXiv.org

CoMPASS: Collaborative Molecular Property Prediction via Adaptive Small-Large Model Synergy

・arXiv:2608.30674v1 Announce Type: new Abstract: Accurate molecular property prediction requires both statistical reliability and chemical reasoning. ・Graph neural networks can be calibrated directly on labeled assays but remain limited by the coverage of their training data. ・Large language models (LLMs) can compare molecular evidence and articulate chemical rationales, yet are unreliable as standalone quantitative pre
cs.LG updates on arXiv.org

Confounding Masquerading as Improvement: A Systematic Evaluation of Offline Reinforcement Learning for Stroke Antithrombotic Treatment in a 129,000-Patient Registry

・arXiv:2608.30442v1 Announce Type: new Abstract: Recent offline reinforcement learning (RL) studies report policies that outperform physician decisions on clinical outcomes. ・We conduct a systematic, partially crossed evaluation of five offline RL algorithm families and 14 reward designs in 44,894 post-2018 acute ischemic stroke patients from a nationwide registry (N = 129,033). ・Standard Fitted Q-Evaluation (FQE) yield
cs.LG updates on arXiv.org

Conservative Hybrid Graph Networks for Process Systems with Learned Routing

・arXiv:2608.28896v1 Announce Type: new Abstract: Industrial process networks do not maintain a single effective topology while operating: streams are throttled or bypassed, and units move between idle, transition, and active regimes. ・Models of such systems are typically trained on measured state trajectories while the operating mechanisms that generated them remain latent, and an unconstrained graph network can fit su
cs.LG updates on arXiv.org

Constant Individual Regret in General Games

・arXiv:2608.31166v1 Announce Type: new Abstract: Uncoupled no-regret dynamics provide a decentralized route to equilibrium, but prior guarantees for individual regret retain a polylogarithmic dependence on the horizon. ・We remove this dependence for every finite $N$-player normal-form game under full-information feedback. ・We introduce \emph{ECHO-OFTRL}: optimistic follow-the-regularized-leader (OFTRL) equipped with an
cs.LG updates on arXiv.org

Context Staircase: Signature-Aligned Dynamics of Token Embeddings under Small Initialization

・arXiv:2608.30315v1 Announce Type: new Abstract: Token embeddings are the basic representational units that connect discrete tokens with continuous computation in language models. ・Although modern language models learn embeddings from random initialization through gradient-based training, the dynamical mechanism by which meaningful embedding structures emerge remains unclear. ・In this work, we identify that the evolving
cs.LG updates on arXiv.org

Context-Aware Interpretable Representations for Retrieval and Graph Convolutional Network Classification

・arXiv:2608.29004v1 Announce Type: new Abstract: The advances in visual information modeling and representation during the last decades are remarkable, mainly supported by Convolutional Neural Networks, Transformer-based, and Foundation Models. ・Despite this progress, critical challenges regarding the nature of similarity assessment and model transparency have been neglected. ・A primary concern is the Geometric Gap, whe
Hugging Face Papers

ContextBias: Controlled Evaluation of Bias Persistence Under Context Shift in Text-to-Image Models

ContextBias: Controlled Evaluation of Bias Persistence Under Context Shift in Text-to-Image Models
cs.LG updates on arXiv.org

Continuity-Free Near-Minimax Leading-Order Regret for CVaR-UCBVI

・arXiv:2608.28960v1 Announce Type: new Abstract: For finite-horizon tabular CVaR reinforcement learning, prior work proves a $\widetilde{O}(\tau^{-1}\sqrt{SAK})$ leading regret bound for arbitrary normalized return laws and the sharper $\widetilde{O}(\sqrt{SAK/\tau})$ rate under a density lower bound. ・We show that the same Bernstein CVaR-UCBVI algorithm attains the sharper rate without continuity assumptions.
cs.LG updates on arXiv.org

Controlling Refusal Behavior of LLMs via Stiefel-Constrained Rotation Steering

・arXiv:2608.30986v1 Announce Type: new Abstract: Activation steering has emerged as a lightweight approach for controlling model refusal at inference time. ・A growing line of research explores trainable rotations of activations to develop geometrically principled intervention mechanisms. ・However, existing techniques rely on auxiliary constructs, such as refusal vectors, to define these rotations.
cs.LG updates on arXiv.org

Convergence rates for the RMSprop optimizer with full control of the hyperparameters

・arXiv:2608.30382v1 Announce Type: new Abstract: Popular adaptive stochastic gradient descent (SGD) methods to train artificial intelligence (AI) systems include the RMSprop, the Adam, and the AdamW optimizers, where the adaptivity parts in Adam and AdamW basically just coincide with RMSprop. ・Such adaptive methods involve several hyperparameters including the regularization parameter $\epsilon$ (which ensures that one
cs.LG updates on arXiv.org

Converse and Collision-Based Achievability for Node Localization with Hybrid Distance-Spectral Graph Positional Encodings

・arXiv:2608.30152v1 Announce Type: new Abstract: Graph positional encodings are widely used in graph neural networks and graph Transformers, yet it remains unclear when the code itself can identify nodes. ・We study a hybrid distance-spectral encoding that combines anchor-distance profiles with quantized low-frequency Laplacian-energy coordinates. ・Treating the encoding as an observation map yields a simplex-refined conv
cs.LG updates on arXiv.org

Creation begins with understanding: LLMs as strategy designers for privacy-preserving tabular data synthesis

・arXiv:2608.29674v1 Announce Type: new Abstract: Sharing tabular data in high-stakes domains is constrained by privacy regulations. ・Synthetic data offer a promising alternative, but deep generative models are costly to train and difficult to audit, while LLM-based methods often serialize records as text, obscuring tabular structure and exposing sensitive data. ・We introduce Tabular Synthesis Strategy Designer (TabSSD),
cs.LG updates on arXiv.org

Curvature Cryptanalysis of Smooth Transformer Feed-Forward Networks

・arXiv:2608.28843v1 Announce Type: new Abstract: We show that smooth two-layer feed-forward networks (FFNs) expose an additional structural model extraction channel under a chosen-input raw-output oracle at the FFN branch; consider transformer FFN branches with GELU or SiLU activations under chosen-input raw-output access, without access to parameters, gradients, or internal activations; exploit a second-order leakage
cs.LG updates on arXiv.org

DASC: Decay-Aware State Compression for Hybrid Linear-Attention Serving

・arXiv:2608.30386v1 Announce Type: new Abstract: Hybrid linear-attention architectures have recently scaled to large open-weight models, offering quality competitive with full attention while substantially reducing key/value (KV) cache growth. ・However, their in-place recurrent-state updates complicate cache management: prefix reuse requires state checkpoints alongside full-attention KV, while storing state checkpoints
cs.LG updates on arXiv.org

Data Diversity, Not Frequency Invariance: A Controlled and Self-Audited Study of Compression-Robust Deepfake Detection

・arXiv:2608.28685v1 Announce Type: cross Abstract: Frequency features and compression-invariant representation learning are widely assumed to be key to deepfake detection that survives video compression. ・We test this with CAFRL - block-DCT and FFT-phase streams, compression-level-conditioned band attention, and adversarial (gradient-reversal) compression invariance - and report a controlled negative. ・Under a pre-regis
#AIタグ

Day1|AIで会社を1個つくる

・今日から90日、AIで事業を1個立ち上げます。うまくいくかは分かりません。分からないまま毎日書きます。 ・目標:月10万円 制約:週8時間まで 事業:ミツモロウ(小さなWeb計算ツールを50〜100本並べた無料サイト) 続きをみる
WIRED

Dell Coupon Codes: 20% Off for September 2026

・Get 20% off with verified Dell promo code, plus today’s coupons for up to $600 off laptops, Alienware monitors, and all things tech.
cs.LG updates on arXiv.org

Denoising as Projection: Constrained Optimization with Gradient-Guided Diffusion

・arXiv:2608.29507v1 Announce Type: new Abstract: Diffusion models are increasingly used not only for sampling from learned data distributions, but also for generating samples that optimize task-specific objectives. ・A common approach is to guide the reverse diffusion process using gradients of an external objective. ・However, when the data distribution is supported on a structured feasible set, such as a manifold or a c
cs.LG updates on arXiv.org

Deploying DeepSeek 175B Locally on a Single Consumer-Grade RTX 4060 Laptop with 32GB RAM for 200k-Scale Protein-Ligand Virtual Screening

・arXiv:2608.30877v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have demonstrated exceptional performance in protein-ligand interaction prediction, but state-of-the-art pipelines for large-scale virtual screening almost exclusively rely on high-end GPU clusters with hundreds of gigabytes of memory, creating prohibitive hardware barriers for small academic teams. ・In this work, we presen
cs.LG updates on arXiv.org

Designing for the Next Click: Bandits for Real-Time Page Layout

・arXiv:2608.29850v1 Announce Type: new Abstract: E-commerce platforms increasingly personalize user experiences through machine learning, yet page layout decisions remain dominated by static rules and manual curation. ・We present a scalable bandit-based system that optimizes product page layouts in real time while preserving human control over design intent. ・A contextual bandit model dynamically selects the most effect
cs.LG updates on arXiv.org

Development of an Autonomous AI Coding Agent using Monte Carlo Tree Search (MCTS) and Gemini LLM Frameworks

・arXiv:2608.29096v1 Announce Type: new Abstract: The ongoing changes in software engineering requirements have created a substantial need for automated tools which can create secure source code from natural language input. ・The performance of traditional Large Language Models (LLMs) becomes limited by their "one-shot" capability which results in logical hallucinations together with reduced algorithmic performance durin
Hugging Face Papers

DICS: Exploring Data Intrinsic Consistency for Visual Instruction Selection

DICS: Exploring Data Intrinsic Consistency for Visual Instruction Selection
cs.LG updates on arXiv.org

Diffusion-Based Inverse Design of Dielectric Resonator Metasurfaces for Shaping Smart Electromagnetic Environments

・arXiv:2608.29907v1 Announce Type: new Abstract: Future wireless systems are expected to transform the surrounding space from a passive propagation medium into a smart electromagnetic environment, where engineered surfaces control wave propagation, support wireless sensing, and create programmable electromagnetic fingerprints. ・A key challenge in realizing this vision is the inverse design of metasurfaces for tailored
cs.LG updates on arXiv.org

Diffusion-Based Refinement for Kilometer-Scale Probabilistic Precipitation Nowcasting

・arXiv:2608.30205v1 Announce Type: new Abstract: Localized extreme precipitation is a major trigger of urban flash floods and landslides, yet producing nowcasts that combine fine spatial detail with probabilistic uncertainty remains challenging. ・Here we introduce exPreCast-ENS, a conditional residual diffusion framework that transforms the deterministic 4 km radar nowcaster exPreCast into a 1 km probabilistic ensemble
cs.LG updates on arXiv.org

Distributed Semantic Segmentation With Improved Rate-Distortion Trade-Off

・arXiv:2608.28684v1 Announce Type: cross Abstract: Distributed deep neural networks (DNNs) for dense perception tasks such as semantic segmentation execute an encoder DNN on edge devices, and a decoder DNN typically on a large-scale cloud platform with a particular constraint on transmission bitrate. ・Recent works employ source codecs to enable bitrate-efficient transmission between the edge device and the cloud.
cs.LG updates on arXiv.org

Do VLMs Share Safety Neurons Across Modalities?

・arXiv:2608.30750v1 Announce Type: new Abstract: Vision-language models (VLMs) can comply with harmful requests delivered through images, even when their LLM backbones would refuse the same content in text. ・While prior work characterizes these jailbreaks empirically or at the representation level, how visual inputs perturb safety pathways at the neuron level remains uncharted. ・We close this gap with a causal, neuron-l
cs.LG updates on arXiv.org

Does Latent Planning Survive Point Clouds? Action-Conditioned JEPA World Models for Geometric Observations

・arXiv:2608.29434v1 Announce Type: new Abstract: JEPA world models make latent-space planning a practical route to control, but they are built almost exclusively on images. ・Whether latent prediction survives geometric observations is unclear: point clouds are sparse, unordered, and self-occluded, and with 0.3-15% of scene points moving, the slow-feature optimum of latent prediction compounds with the geometric shortcu
Hugging Face Papers

Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement

Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement
cs.LG updates on arXiv.org

Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement

・arXiv:2608.31046v1 Announce Type: new Abstract: On-policy distillation (OPD) offers dense token-level supervision as an alternative to the sparse outcome-level advantages of reinforcement learning with verifiable rewards (RLVR). ・However, the teacher scores student-generated trajectories that are inherently off-policy for it, so the reliability of its supervision, and hence the source of the student's improvement, rem
Hugging Face Papers

DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution

DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution
WIRED

Dyson’s Next Act Is an Electric Toothbrush With a Camera

・The company known for stick vacuums and hair dryers is coming for your teeth. ・The $499 Dyson CameraJet uses a tiny camera to aim streams of rinsing fluid into the gaps between your teeth.
cs.LG updates on arXiv.org

E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation

・arXiv:2608.30730v1 Announce Type: new Abstract: Long-horizon agentic tasks go beyond chaining short tasks over more interaction turns. ・Their evolving dynamic environments and long-range dependencies require Large Language Models (LLMs) to continually explore, learn from experience, and adapt their policies over thousands of steps. ・We introduce E-Commerce Bench, the first open-source benchmark that integrates multi-ro
cs.LG updates on arXiv.org

ECA-BLS: An Efficient Complex-Augmented Broad Learning System

・arXiv:2608.29763v1 Announce Type: new Abstract: Broad Learning System (BLS) is an efficient alternative to deep architectures due to its fast training, analytical learning, and strong generalization under limited data. ・However, existing BLS variants are confined to real-valued representations, restricting their ability to capture nonlinear interactions and second-order statistical dependencies inherent in real-world
cs.LG updates on arXiv.org

Effective Graph and Rank-based Contextual Embeddings for Textual and Multimedia Data

・arXiv:2608.29001v1 Announce Type: new Abstract: In a data-driven world, efficiently organizing and mapping relationships between objects is crucial. ・Graphs are powerful tools for modeling these connections, being widely used in social networks, telecommunications, and biology. ・However, graph-based methods often face high computational costs, particularly in memory and space usage.
cs.LG updates on arXiv.org

Efficient GPU Retrieval for Semantic Search

・arXiv:2608.28968v1 Announce Type: cross Abstract: Semantic Search on LinkedIn must retrieve relevant profiles from a corpus of hundreds of millions in response to natural-language queries such as "a fintech founder in Berlin who worked in payments." The deployed relevance policy is bottleneck-oriented: every active non-negotiable facet must be satisfied, and a pre-existing LLM Graded Relevance (GR) judge operationali
cs.LG updates on arXiv.org

Emergent Misalignment Is Not Magical

・arXiv:2608.29118v1 Announce Type: cross Abstract: Fine-tuning large language models (LLMs) on narrowly harmful datasets can lead to misalignment broadly, a phenomenon known as emergent misalignment (EM). ・EM poses a challenge for AI safety and our understanding of LLMs. ・Prior work often frames EM as an unexpected behavior, and explains it by appealing to general misalignment directions or anthropomorphizing it as acqu
cs.LG updates on arXiv.org

Equivariant Sheaf Neural Networks: Learning Geometric Transport on Graphs

・arXiv:2608.28853v1 Announce Type: new Abstract: Equivariant graph neural networks provide a principled way to model geometric systems, but efficient first-order architectures remain limited in how vector information can be transformed as it moves across a graph. ・We introduce \textsc{ESNN}, an Equivariant Sheaf Neural Network that enriches this interaction by learning directed, matrix-valued transport between neighbor
cs.LG updates on arXiv.org

ERR+: Sequential Entropy Resolution for Efficient and Decisive LLM Reasoning

・arXiv:2608.28771v1 Announce Type: new Abstract: Large reasoning models achieve strong performance on complex tasks by generating extended chain-of-thought (CoT) traces via reinforcement learning with verifiable rewards (RLVR). ・While current RLVR methods have achieved strong results with correctness-based reward signals, they provide limited guidance on the quality of the reasoning process itself, leaving the internal
cs.LG updates on arXiv.org

Error Detection for PET/CT Radiology Reports: Domain-Specific vs Large Language Models

・arXiv:2608.30021v1 Announce Type: new Abstract: Errors in radiology reports can adversely affect patient treatment, yet automated report quality assurance remains challenging because errors are often subtle and require domain expertise to detect. ・Although large language models (LLMs) have recently been proposed for radiology report verification, their ability to detect clinically meaningful errors beyond chest X-ray
Hugging Face Papers

Evaluating the Hidden Costs of Personalization in Large Language Models

Evaluating the Hidden Costs of Personalization in Large Language Models
cs.LG updates on arXiv.org

Evaluating Tiny Recursive Models Across Training for Code Generation

・arXiv:2608.29376v1 Announce Type: cross Abstract: Code generation increasingly relies on large transformer models, whose capability advances with scale. ・Yet such a scale is costly, creating demand for small models, especially where data is limited. ・Recursive models address this by reusing a single block to add depth rather than stacking independent layers.
cs.LG updates on arXiv.org

Event-triggered Control and Online Learning for Networked Systems under Computational Delays

・arXiv:2608.29576v1 Announce Type: new Abstract: Online learning-based control is a promising approach to control uncertain systems, where unknown components are identified during operation to improve control performance. ・However, resource-intensive online learning algorithms introduce non-negligible computational delays, especially when executed on systems with limited local computational resources. ・To mitigate this,
LLMタグが付けられた新着記事 - Qiita

EVO-X2(Ryzen AI Max+ 395 / 128GB)でQwen3.8-Flash-Nextを高速化した話。n_cpu_moeを調整したらprefillが速くなったけど罠だった

・前回、「EVO-X2(Ryzen AI Max+ 395 / 128GB)でQwen3.8-Flash-Nextを動かした話」で、Qwen3.8-Flash-Nextを動かして日本語で会話できるところまで確認した。 ・今回はパラメータ調整でどこまで高速化できるかを試してみた。...
cs.LG updates on arXiv.org

Exact Recovery Thresholds for Weighted Data Selection in Vector-Valued Linear Regression

・arXiv:2608.30254v1 Announce Type: new Abstract: We resolve the threshold part of Question 4 of the COLT 2025 open problem "Data Selection for Regression Tasks" of Hanneke, Moran, Shlimovich and Yehudayoff. ・In vector-valued linear regression with square loss $\ell_{(x,y)}(W)=|Wx-y|_2^2$, where $x\in\mathbb{R}^d$, $y\in\mathbb{R}^m$ and the learner is the empirical risk minimizer of minimal Frobenius norm, we prove tha
ITmedia NEWS 最新記事一覧

EXILE・HIROのGMO系イベント関与に抗議の声 Xでハッシュタグ拡散 米軍事IT企業の参加などで反発

・8月30日ごろから、Xで「LDHはGMO大会議への登壇を辞退してください」とのハッシュタグが拡散している。音楽グループ「EXILE」メンバーで、芸能事業を手掛けるLDH JAPANのCEO・HIRO氏がGMOインターネットグループ主催の「第4回GMO大会議 秋 フィジカルAI 2026」に参加することを受けた動きのようだ。
cs.LG updates on arXiv.org

Explainable Machine Learning for Broadband Adoption Disparities: Tract-Level Prediction and SHAP-Based Factor Profiling

・arXiv:2608.29110v1 Announce Type: new Abstract: The United States has allocated approximately $65 billion through the Infrastructure Investment and Jobs Act for broadband expansion, yet evidence-based methods for targeting these investments remain underdeveloped. ・This paper presents an explainable machine learning framework for profiling broadband adoption disparities at census-tract granularity across 83,359 tracts
AI News & Artificial Intelligence | TechCrunch

Fambot introduces an ‘AI chief of staff’ for families

・Fambot is building an AI “chief of staff” to help families manage the emails, calendars, school updates, sports schedules, and other logistics of raising kids.
cs.LG updates on arXiv.org

FiLM-GPNet: Geometry-Aware Pseudo-Supervised Phase Restoration with Zero-Shot Generalization for Large Temporal InSAR Stacks

・arXiv:2608.29384v1 Announce Type: cross Abstract: The growing availability of dense commercial Synthetic Aperture Radar (SAR) time series enables temporal Interferometric SAR (InSAR) analysis, but fixed classical filters fail under heterogeneous acquisition geometries, degrading phase quality and temporal consistency. ・We propose FiLM-GPNet, a geometry-conditioned network for wrapped-phase restoration that explicitly
cs.LG updates on arXiv.org

Fine-Tuning Low-Bit Models with Gradient in Quantized Code Space

・arXiv:2608.30908v1 Announce Type: new Abstract: Fine-tuning Low-bit models aims to adapt a quantized model while keeping the final deployed checkpoint in the same low-bit form. ・This setting is practically important as it reduces memory and inference cost for storage and deployment. ・Under this constraint, adaptation becomes an optimization problem over quantization codes and scales.
cs.LG updates on arXiv.org

Fine-Tuning Qwen3-27B for C-to-Rust Code Translation: A Three-Stage Curriculum of Pretraining, Debugging-Aware SFT, and Task-Specific SFT

・arXiv:2608.13681v1 Announce Type: cross Abstract: Translating C code into safe, idiomatic Rust is a longstanding software-engineering goal because it can eliminate entire classes of memory-safety vulnerabilities while preserving the functional behavior of legacy systems. ・Large language models (LLMs) have shown promise for this task but typically underperform when applied off-the-shelf, since general-purpose pretraini
cs.LG updates on arXiv.org

Flow-JEPA: Flow Matching for Robust Latent Dynamics in JEPA World Models

・arXiv:2608.29029v1 Announce Type: new Abstract: Joint-Embedding Predictive Architectures (JEPAs) have shown strong potential for learning compact predictive representations, and LeWorldModel (LeWM) extends this paradigm to reconstruction-free latent world modeling from pixels. ・However, its deterministic autoregressive predictor generates future states through repeated one-step transitions, which can accumulate errors
#LLMタグ

FOCUSで請求を揃えても、AI利用料の答えは出ない

・21時18分、MacBookの横でぬるくなった炭酸水を飲みながら、利用料のCSVを眺めていた。月次の請求額は合っている。AWSも外部APIも、チーム別のタグも入っている。それなのに「採用候補者の要約機能は今月いくら使った?」と聞かれると、答えが出なかった。 ・田口さんに画面を見せたら、「合計はきれい。でも、その合計で次に何を止めるの?」と言われた。痛い所を突かれた。
cs.LG updates on arXiv.org

Forget or Fine-tune? A Comparative Study of Machine Unlearning Strategies for Noisy Label Correction

・arXiv:2608.30046v1 Announce Type: new Abstract: Noisy labels remain a critical challenge for training deep neural networks, since memorizing incorrect labels degrades generalization. ・Once noisy samples are identified after training, the standard solution is to retrain the model from scratch on the cleaned dataset, which is increasingly expensive as datasets and models grow. ・Machine Unlearning (MU) has recently emerge
cs.LG updates on arXiv.org

Foundation Models Meet Agriculture: Challenges Beyond Pretraining

・arXiv:2608.30392v1 Announce Type: new Abstract: Global food security and sustainable climate action increasingly rely on robust, scalable agricultural monitoring. ・Earth observation foundation models have emerged as powerful, label-efficient tools across general remote sensing domains, yet early attempts to deploy them for agricultural applications have yielded surprisingly poor results. ・We hypothesize that this perfo
cs.LG updates on arXiv.org

FrameScope: Temporal Data Valuation for Stream Active Learning in Autonomous Vehicle Systems

・arXiv:2608.28672v1 Announce Type: cross Abstract: Autonomous vehicles operate in dynamic, ever-changing environments where new scenarios and edge cases constantly emerge. ・As a result, static learning models are inadequate for ensuring safe and reliable operation. ・Continuous learning is essential for adapting to these evolving conditions and maintaining robust performance across diverse real-world settings.
cs.LG updates on arXiv.org

From Extraction to Governed Memory: Multi-Agent Knowledge Graph Construction with Domain-Expert Review

・arXiv:2608.28642v1 Announce Type: cross Abstract: Knowledge graphs used by agentic systems are often treated as flat stores of extracted triples, with little record of who owns a fact, why it was admitted, or how it should be used downstream. ・We argue that reliable agentic knowledge systems require governance as an essential component of graph construction to bridge this gap. ・We propose MAGG, a principled multi-agent
cs.LG updates on arXiv.org

From Location Phrases to Geographic Entities: Task-Adapted Retrieval for People Search

・arXiv:2608.28965v1 Announce Type: cross Abstract: People search must map free-form location phrases to geographic entities used as structured retrieval filters. ・Lexical standardizers handle canonical names well but are brittle to aliases, misspellings, metropolitan expressions, and same-name ambiguity. ・We formulate this task as graded, set-valued entity retrieval over a fixed ontology.
cs.LG updates on arXiv.org

From the Loss Landscape to Diverse Feature Learning in Neural Networks

・arXiv:2608.28948v1 Announce Type: new Abstract: Over the course of the last decade, neural networks have grown from an academic curiosity to moving the markets of nations. ・Despite this explosion in both research and deployment, relatively little is understood about how they achieve the solutions they do. ・This is both scientifically relevant, and pressing for society.
cs.LG updates on arXiv.org

Fully Distributed GNE Algorithms for Multi-Robot Placement without Consensus on Multipliers

・arXiv:2608.29388v1 Announce Type: new Abstract: Recent machine learning research has increasingly focused on equilibrium analysis in non-cooperative games rather than solely on optimal solutions. ・Many such problems involve shared constraints and can be formulated as Generalized Nash Equilibrium Problems (GNEPs). ・For strongly monotone games, existing methods compute consensus-based variational GNEs (v-GNEs) by exchang
cs.LG updates on arXiv.org

Functional Degeneracy in Neural Networks: Measurement and Pruning

・arXiv:2608.30741v1 Announce Type: new Abstract: A central question in modern machine learning is how much a trained model can be compressed without changing its behavior, to reduce the memory, compute and energy required to deploy it. ・To study this, we quantify functional degeneracy through the behavioral recovery rank, defined as the number of leading behavioral-Hessian eigendirections required to recover a trained
cs.LG updates on arXiv.org

Generation of High-Level Concepts in 3D Scene Graphs via Autoregressive Diffusion

・arXiv:2608.28733v1 Announce Type: cross Abstract: Indoor 3D Scene Graphs (3DSGs) represent environments as multi-layer hierarchies that connect observed geometric primitives (e.g., planes) to higher-level metric-semantic concepts (e.g., rooms, floors, buildings), enabling incremental spatial reasoning for robotic perception and SLAM. ・However, classical high-level concept generation approaches rely on hand-crafted rul
cs.LG updates on arXiv.org

Generative multi-domain transfer learning for fault detection in data-scarce wind turbines

・arXiv:2608.30323v1 Announce Type: new Abstract: Normal behavior models have shown promise for reliable fault detection in wind turbines. ・However, these unsupervised anomaly detection models require sufficient fault-free training data to learn the normal operation behavior of turbines. ・Under data scarcity, for example in newly deployed wind turbines, these models may result in poor fault detection performance.
cs.LG updates on arXiv.org

Generative Translation Priors: Bayesian Imaging with Cross-Modality Image Translation

・arXiv:2608.28872v1 Announce Type: cross Abstract: The ability to leverage images from co-available modalities to inform target-domain reconstruction is highly desirable in imaging algorithms. ・In this work, we introduce Generative Translation Priors (GTP)--a Bayesian framework that transforms diffusion-based image-to-image translation models into cross-modality image priors for ill-posed imaging inverse problems.
Hugging Face Papers

GenFirst: Generation Before Reconstruction for Stable End-to-End Latent Generative Modeling

GenFirst: Generation Before Reconstruction for Stable End-to-End Latent Generative Modeling
cs.LG updates on arXiv.org

Geometric Attractor Monitoring: A Robust and Frugal Framework for Multi-modal Industrial Robotic Cycles

・arXiv:2608.30804v1 Announce Type: new Abstract: Monitoring the health of heterogeneous industrial robot fleets is severely challenged by the multi-modal nature of their operational cycles and a persistent scarcity of run-to-failure data. ・Standard data-driven approaches, particularly deep learning architectures relying on sequential reconstruction, often struggle in this specific setting; they tend to over-smooth comp
LLMタグが付けられた新着記事 - Qiita

GLM-5.3-Flash API料金比較:OpenRouterの5.5%手数料とキャッシュ比率を含めて計算

・概要 GLM-5.3-FlashをコーディングAgentや長時間の自動処理で利用する場合、モデルページに表示されたトークン単価だけでは実際のコストを比較できません。 ・この記事では、次の2点を含めてZ.ai、OpenRouter、AIHubMixを比較します。
The Verge

Google Find Hub will soon help you locate things without a tracker

・Google announced new features coming to Android devices as part of this month's feature drop update. ・Among the highlights are the introduction of Motion Assist - Android's version of Apple's Motion Cues feature, which tries to reduce motion sickness by overlaying dots on your screen that move in response to a vehicle's movements - as well an expansion of Find Hub. ・The Find Hub update will introduce a new way to find
The Verge

Google Pics is like Canva, but with even more AI

・Google wants Workspace users to edit and generate their business imagery with Pics. ・| Image: Google Google has a new suite of creative design tools for Workspace users called Google Pics, which aims to make editing and generating "professional-grade" AI images less cumbersome for businesses. ・Built around Gemini and the Nano Banana generative AI model, Google Pics is designed to give more granular control over prompt-
AI News & Artificial Intelligence | TechCrunch

Google’s answer to Canva is an AI tool where you prompt instead of design

・With Google Pics, Google is pushing deeper into the creative software market dominated by Canva and Adobe, but with a distinctly AI-first approach.
The Verge

GoPro has been acquired and is getting into ‘defense, government, robotics and aerospace’

・GoPro suddenly looks like a very different company than it did just last week. ・First, YouTuber Mark "Markiplier" Fischbach became the company's single largest shareholder, without disclosing that in a sponsored review that came out a week later. ・Now, almost the entire company has been bought by Starman Holding for $285 million in cash, which Starman is characterizing as a "merger." While GoPro has always been best kn
Zennの「大規模言語モデル」のフィード

GPT-2-likeからQwen2-likeへの実験:第1回 LayerNormをRMSNormに変える

・GPT-2-likeからQwen2-likeへの実験:第1回 LayerNormをRMSNormに変える GPT-2-likeからQwen2-likeへ、少しずつ構造を変更する実験をしたところ、かなり面白い結果になりました。今回は第1回、LayerNormからRMSNormへの変更です。 ・約890万parametersのGPT-2-likeモデルを自作しました。教科書実装でかなり小さいモデルです。LayerNormだけをRMSNormへ置き換えて、Apple SiliconのMPS上でそれぞれ5つのseedを使い、約10M Token分を学習しました。 ・コードの変更はかなり小さいです...
cs.LG updates on arXiv.org

GraM-Diff: A Unified Graph-Mamba Diffusion Framework for EEG-Based Alzheimer's Disease Data Generation and Diagnosis

・arXiv:2608.29755v1 Announce Type: new Abstract: Electroencephalography (EEG) is a promising, non-invasive, and cost-effective modality for Alzheimer's disease (AD) detection, but deep learning methods are limited by small and imbalanced clinical datasets. ・Generative augmentation offers a solution, yet existing approaches rely on inefficient class-specific models or fail to capture complex spatial and temporal brain d
cs.LG updates on arXiv.org

Graph4BiLO: Graph Neural Network Approximation for Bilevel Mixed-Integer Linear Optimization

・arXiv:2608.30103v1 Announce Type: new Abstract: Bilevel mixed-integer linear optimization problems model hierarchical decision processes in which a leader anticipates the optimal response of a follower. ・Although expressive, these problems are computationally challenging because lower-level optimality is embedded in the leader's feasible region. ・Value-function reformulations replace the nested follower optimization wi
cs.LG updates on arXiv.org

HalluPrism: When Multimodal Uncertainty Should Diagnose, Not Decide

・arXiv:2608.29193v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) can assign similar confidence to answers that fail for different reasons. ・We propose HalluPrism, a behavioral diagnostic that re-runs an answer after visual degradation, blank-image replacement, and grounding or relation checks. ・These targeted probes yield a signature over visual-perturbation sensitivity (V ), image-removal confi
OpenAI News

Healthcare organizations can now connect EHR and additional industry data to ChatGPT

・ChatGPT can now connect to trusted healthcare data, helping clinicians securely access patient context, medical research, and more.
cs.LG updates on arXiv.org

Higher-Dimensional Rotary Position Embedding

・arXiv:2608.29715v1 Announce Type: new Abstract: Transformers rely on position embedding mechanisms in long context modeling in most cases. ・Rotary Position Embedding (RoPE) embeds positional information with independent 2D rotations, forming relative position terms in self-attention. ・However, its pairwise, block-based, and decoupled structure limits deep mixing and robustness across channels.
WIRED

Home Depot Promo Codes: 30% Off in September 2026

・Save up to 50% today with the latest Home Depot promo codes for appliances, power tools, and more this September.
The Verge

Homey just made the smart home controller I hoped Apple would build

・Homey’s new tactile touchscreen controller has a rotary dial for controlling lights, climate, music and more. ・Smart homes have gotten really good at making things really complicated. ・Turning on a light can involve multiple steps, compared to just flipping a switch.
cs.LG updates on arXiv.org

HoopMind: A Real-Time Neural Game-Tree System for Opponent-Aware Possession Planning

・arXiv:2608.29563v1 Announce Type: new Abstract: School coaches prepare for opponents with game film and intuition. ・The analytics tools of professional teams stay out of reach. ・We ask how far public data can close this gap.
OpenAI News

How AI-native companies turn workflows into operating capability

・Basis, Clay, and Exa Labs use AI agents to improve onboarding, account management, and developer integrations. ・See what enterprise leaders can apply.
cs.LG updates on arXiv.org

How do World Models and Policies Compose in LLM Agents? A Joint Spectral and Behavioral Account

・arXiv:2608.30067v1 Announce Type: new Abstract: How do LLM agents come to both understand environments they act in and master tasks set within them? ・Through controlled experiments combining world-model training (next-state prediction) and policy training (reward maximization), we investigate this question. ・We dissect the resulting models through their additive parameter updates.
cs.LG updates on arXiv.org

How Language Models Choose Sides: Internal Representations of Instruction Hierarchy

・arXiv:2608.28648v1 Announce Type: cross Abstract: We study how instruction-tuned LLMs arbitrate direct conflicts between system and user instructions. ・We introduce a benchmark of 41 paired constraints with deterministic verifiers and evaluate eight models under matched baseline, conflict, and same-channel control conditions. ・Behaviourally, the models split into three regimes by System Authority Delta: hierarchy-respe
cs.LG updates on arXiv.org

Hybrid Semantic Context-Enhanced Ensemble Learning for Wind Power Ramp-Event Forecasting and Uncertainty-Aware Evaluation

・arXiv:2608.29024v1 Announce Type: new Abstract: Wind power ramp events which are sudden, large swings in turbine output over short windows are difficult to estimate, and standard models often miss them. ・Hybrid forecasting approach is built which augments semantic context to ramp-event forecast. ・Rather than applying an extensive language model directly to predict turbine operating data, we have implemented a pipeline
cs.LG updates on arXiv.org

Hyper-Fold: Exploring the Expressive Limit of Sequence-Geometry Learning for Proteins via Hypergraph Modeling

・arXiv:2608.29207v1 Announce Type: cross Abstract: Protein structure modeling rests on a single computational primitive: the interaction between what a residue is (sequence content) and where it sits (three-dimensional geometry). ・What is the expressive limit of this layer class? ・We show that the complete bilinear operator over content-geometry outer products--the sufficient statistic of all second-order interactions--
MIT News - Artificial intelligence

Ila Kumar: Innovating with communities

・The PhD student works to give young people an active role in shaping digital technologies that can support their own well-being.
#LLMタグ

Inference Broker – GPU版のStable Diffusionバックエンドを追加する

・Inference Broker – GPU版のStable Diffusionバックエンドを追加する – 塾長の独り言zikuu.space 続きをみる
cs.LG updates on arXiv.org

Information-Based Calibration of Uncertainty Quantification in Product-of-Experts Gaussian Process Models

・arXiv:2608.29349v1 Announce Type: new Abstract: Gaussian process (GP) regression with a single global GP (GP-glo) incurs cubic computational cost, limiting scalability to large datasets. ・Product-of-experts GP models (GP-pro), which combine local GP models to capture global correlations, alleviate this computational burden. ・However, training local experts on disjoint data subsets can lead to overestimated posterior va
WIRED

Inside the Perimenopause Industrial Complex

・How an alliance of tech startups, MAHA operatives, and actual medical experts made millennial women the new face of hormone therapy.
cs.LG updates on arXiv.org

INTERVenE: Temporal-Abstraction-Interval Based Transformers for Short-Horizon Medical Event Prediction

・arXiv:2608.29901v1 Announce Type: new Abstract: Electronic Health Record (EHR) prediction models in the intensive care unit must learn from sparse and irregular measurements while preserving the clinical meaning of time and supporting transparent decision-making. ・We present INTERVenE, a family of Transformer architectures whose input is an interval-based, knowledge-based temporal abstraction (KBTA), a token stream of
Google DeepMind News

Introducing agentic video understanding with Gemini

Introducing agentic video understanding with Gemini
cs.LG updates on arXiv.org

Jigsaw-CRL: Recovering Global Latent Causal Order from Fragmented Multi-Client Interventions

・arXiv:2608.28991v1 Announce Type: cross Abstract: Causal representation learning (CRL) aims to recover latent causal variables and their structural relations from high-dimensional observations. ・Existing CRL methods typically assume that all environments are defined over the same latent variables, or at least share a common latent representation space. ・We study a fragmented multi-client setting, where multiple clients
The Verge

John Deere launched an AI chatbot for farmers

・John Deere is testing a new "JD" AI assistant that it says can help farmers make more money, with answers about best practices and historical trends that are based on their own data. ・It uses their "field, machine and operational data" to answer questions on topics like equipment settings, fuel usage, or harvest timing. ・The press release didn't mention which AI technology is behind this platform.
cs.LG updates on arXiv.org

Joint Spatiotemporal Spectral Neural Operators for Learning PDEs on Irregular Domains

・arXiv:2608.29892v1 Announce Type: new Abstract: Learning solution operators for partial differential equations (PDEs) on irregular and geometry-dependent domains remains a central challenge in scientific machine learning. ・While spectral methods provide strong inductive biases for modeling global interactions, they are typically limited to regular domains, and existing neural approaches often require domain warping, i
Hugging Face Papers

Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation

Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation
cs.LG updates on arXiv.org

Knowledge Distillation under Teacher Misspecification: An Order-Parameter Analysis of the Gap between Teacher Mimicry and Task Performance

・arXiv:2608.29472v1 Announce Type: new Abstract: Knowledge distillation trains a small student model to reproduce the outputs of a large teacher model, and its progress is typically monitored through the teacher--student discrepancy. ・The quantity of ultimate interest, however, is the student's error with respect to the true task. ・We study the relation between these two objectives in a minimal three-party model, a true
cs.LG updates on arXiv.org

Kolmogorov--Arnold against bounded translations

・arXiv:2608.30710v1 Announce Type: new Abstract: Historically originating from Hilbert's 13th problem, the Kolmogorov-Arnold representation theorem (KART) has recently experienced a major revitalisation through its applications to neural networks, specifically Kolmogorov-Arnold Networks (KANs). ・While the exact representation is well established, its stability under continuous adversarial perturbations of the hidden la
cs.LG updates on arXiv.org

Language-Informed Flow Matching for Trend-Guided Structure-Based 3D Molecular Generation

・arXiv:2608.31009v1 Announce Type: new Abstract: Structure-based drug design (SBDD) requires ligands that satisfy both 3D target affinity and 1D chemical validity. ・Existing controllable generation methods often rely on task-specific fine-tuning or externally imposed sampling-time guidance, adding cost and potentially conflicting with evolving 3D geometric constraints. ・We propose LiFT, a language-informed cross-modal f
cs.LG updates on arXiv.org

Last Step Matters: Early Uncertainty Cannot Predict Failure in Long-Horizon Agents

・arXiv:2608.29685v1 Announce Type: new Abstract: Early failure prediction is important for long-horizon agents, as it enables timely intervention and can reduce inference and tool-use costs. ・Uncertainty quantification, such as verbal confidence and perplexity, offers a promising approach to detecting agent failures; however, it has not been explored whether these signals retain their discriminative power during the in
cs.LG updates on arXiv.org

Learning Dynamics of Logits Debiasing for Long-Tailed Semi-Supervised Learning

・arXiv:2608.30699v1 Announce Type: new Abstract: Long-tailed distributions are prevalent in real-world semi-supervised learning (SSL), where pseudo-labels tend to favor majority classes, leading to degraded generalization. ・While many long-tailed semi-supervised learning (LTSSL) methods have been proposed, the mechanisms by which they implicitly debias logits remain poorly understood. ・In this work, we revisit LTSSL thr
cs.LG updates on arXiv.org

Learning Human Health and Diseases from 24-hour Wrist Movement

・arXiv:2608.29494v1 Announce Type: new Abstract: Much of human health and function unfolds beyond the clinic, through the movements of everyday life. ・Wrist-worn accelerometers capture these movements continuously, yet their rich signals are often reduced to a small set of predefined behavioural summary measures. ・Here, we present Sensori, a self-supervised foundation model that learns general-purpose health representat
cs.LG updates on arXiv.org

Learning Materials Properties from Scarce Labels and Unlabeled Crystals

・arXiv:2608.30682v1 Announce Type: new Abstract: Learning materials properties from scarce labels and unlabeled crystals is a central challenge for data-driven materials discovery. ・We present SemiMat, a controlled benchmark for semi-supervised materials property regression, and MatRank, a reliability-weighted objective for continuous pseudo-label uncertainty. ・SemiMat fixes labeled and unlabeled crystal inputs, graph-b
cs.LG updates on arXiv.org

Learning PDE Time-Stepping with Neural Cellular Automata

・arXiv:2608.30328v1 Announce Type: new Abstract: Classical numerical solvers for partial differential equations (PDEs) are computationally expensive to solve repeatedly across varying initial conditions, motivating the need for learned surrogates. ・In this paper, we propose a trainable Neural Cellular Automata (NCA) based surrogate model for learning long time PDE dynamics. ・Rather than mapping an entire initial field t
cs.LG updates on arXiv.org

Learning Simple Test-Time Environments for LLM Web Agents

・arXiv:2608.29305v1 Announce Type: cross Abstract: Large language model (LLM) agents have demonstrated remarkable proficiency in manually constructed environments, yet their performance frequently collapses when transitioned to complex real-world settings. ・Existing research largely attribute this degradation to the compositional generalization gaps in LLMs on combinations of multiple simple, well-structured environmen
Hugging Face Papers

Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research Agents

Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research Agents
cs.LG updates on arXiv.org

Learning Where Outcomes Change:Credit-Addressable Reasoning for Multimodal Geometry

・arXiv:2608.30457v1 Announce Type: new Abstract: Multimodal geometry reasoning requires VLMs to extract precise visual relations and preserve them through multi-step deduction. ・Existing free-form traces obscure the decisions that determine the answer, and trajectory-level reinforcement learning distributes a single terminal signal across the entire response. ・We introduce credit-addressable reasoning, in which the sema
cs.LG updates on arXiv.org

Learning-Theoretic Foundation for General Coded Computing: The Straggler Setting

・arXiv:2608.28910v1 Announce Type: new Abstract: Coded computing has emerged as a powerful paradigm for mitigating the impact of straggling workers in distributed computing systems. ・However, existing coded-computing schemes are predominantly designed for the exact recovery of highly structured computations, such as polynomial evaluation and matrix multiplication, and typically rely on strict recovery thresholds.
cs.LG updates on arXiv.org

Leveraging Turn-taking Dynamics for Intent Recognition in Multi-party Conversations

・arXiv:2608.28926v1 Announce Type: cross Abstract: We propose a multi-task learning approach for multi-party dialogue intent recognition that leverages an auxiliary task that models turn-taking dynamics. ・Specifically, we introduce turn-transition entropy, a self-supervised target computed from the sequence of speaker transitions, which quantifies the predictability of interaction patterns. ・Experiments on multiple pre-
Hugging Face Papers

Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions

Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions
Hugging Face Papers

LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation

LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation
cs.LG updates on arXiv.org

Liquid Gated Attention

・arXiv:2608.30695v1 Announce Type: new Abstract: Real-world time series often exhibit irregular sampling and extended temporal horizons, requiring models to capture continuous-time dynamics across arbitrary intervals without prohibitive scaling costs. ・Discrete-time methods collapse variable time intervals into static positional steps; solver-dependent continuous-time models preserve temporal structure but rely on sequ
cs.LG updates on arXiv.org

LLMODE: Aligning ODEs with LLMs via Gated Token Injection for Irregular Spatio-Temporal Forecasting

・arXiv:2608.29640v1 Announce Type: new Abstract: Large language models (LLMs) have shown promise for spatio-temporal forecasting, but existing approaches often rely on regularly sampled token sequences and struggle with irregular observations because of temporal asynchrony, representation-space misalignment, and limited context windows. ・We propose LLMODE, a token-efficient framework for irregular spatio-temporal forec
Zennの「大規模言語モデル」のフィード

LLMに書き込ませるMCPツールの設計

・こんにちは!webエンジニア目指して10年のTakehiroTです! TamaT という、顧客提案から要件定義・設計・開発まで一気通貫で担う少数精鋭の開発会社でやってます。 ・社内のタスク管理ツールに、ローカルのエージェントから接続する MCP(Model Context Protocol)サーバーを載せました。 ・最初は検索と参照だけの読み取り専用で、正直まったく怖くありませんでした。
#LLMタグ

LLMのパフォーマンス測定で見落としているもの

・最近、PC上でLLMを稼働させるという記事を結構見かけるようになった。ハードウェアの選定、モデルの比較、量子化の話。 ・ただ、そのパフォーマンス評価で気になることがある。
#LLMタグ

LLMを2時間で自作できるMiniMind。その2時間はどこからどこまでか

・「64MパラメータのLLMを、ゼロから2時間で学習」。 ・GitHubのトレンドに、そんな説明のリポジトリが上がってきました。 ・1日で1,000スター以上を集めた「jingyaogong/minimind」です。
Zennの「大規模言語モデル」のフィード

LM Studio が Ubuntu 20.04 → 22.04 → 24.04 でどう動いたかの記録(最終的に AppImage で解決)

・LM Studio を使おうとして「Ubuntu 20.04 では起動しない」という問題に直面する人は多いと思います。私自身も最新版の LM-Studio-0.4.23-1-x64.deb をインストールしたところ、まったく起動しませんでした。 ・この記事では、 Ubuntu 20.04 → 22.04 → 24.04 のアップグレードで何が起きたか 各バージョンで LM Studio が起動したかどうか 最終的にどう解決したか をまとめます。 ・🧩 Ubuntu 20.04:LM Studio の DEB を入れても起動しない まず、Ubuntu 20.04 に以下のコマンドで LM St...
cs.LG updates on arXiv.org

Locally-Guided Actor-Critic: Training a Goal-conditioned Actor with a Subgoal-aware Critic

・arXiv:2608.30406v1 Announce Type: new Abstract: Goal-conditioned reinforcement learning struggles with long horizons when rewards are sparse. ・While a planner can provide subgoals to guide a low-level policy, its use at test time may introduce practical subgoal management difficulties. ・An alternative paradigm utilizes a high-level planner to assist learning, while the policy remains conditioned only on the final goal,
cs.LG updates on arXiv.org

Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space

・arXiv:2608.29188v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) substantially improves single-sample accuracy (pass@1) but causes the policy's solution space to contract, diminishing the returns of test-time scaling. ・In this work, we investigate where inside a reasoning trajectory this breadth is lost: does the policy fail to access a valid solution family, or does it fail to exe
Hugging Face Papers

Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling

Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling
Zennの「大規模言語モデル」のフィード

M4 Max 部署 Qwen3.8-27B:高性能推理与多人共享实践

・适用设备 40 核 GPU、128GB 统一内存、546GB/s 内存带宽的 M4 Max MacBook Pro 或 Mac Studio 适用目标 在单台 Mac 上稳定运行 Qwen3.8-27B,并通过 OpenAI 兼容 API 提供给少量用户共享 基准日期 2026-08-30。Qwen3.8、DeepSeek V4 Flash 和 oMLX 仍在快速迭代,升级前应重新跑本机基准。 ・一、先给结论 如果目标是兼顾速度、质量、多人共享和维护成本,推荐采用下面的架构。 ・用户客户端 │ │ HTTPS + 独立 API Key ▼ Tailscale 私有...
機械学習タグが付けられた新着記事 - Qiita

MACE機械学習ポテンシャルとAFIR(人工力誘起反応)法で反応経路を自 動探索する

・はじめに 化学反応の遷移状態を求めるには、通常、反応物と生成物の構造をあらかじめ仮定し、経験に基づいて反応座標を設定する必要があります。Mat3raプラットフォームのリリース2026.8.27では、この作業を自動化する新機能として、AFIR(人工力誘起反応、Artifici...
Hugging Face Papers

Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory

Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory
cs.LG updates on arXiv.org

Measuring Memory and Generalization as Separable Geometric Channels: The Topo^2 Framework

・arXiv:2608.30487v1 Announce Type: new Abstract: Deep networks trained on noisy labels simultaneously generalize on clean data and memorize flipped labels. ・These are usually conflated as pressures on one capacity. ・We present Topo^2, a measurement framework that makes them causally separable, measurable, and law-governed.
cs.LG updates on arXiv.org

MedCache: Efficient and Temporally Valid Memory for Longitudinal Clinical Agents

・arXiv:2608.29528v1 Announce Type: new Abstract: Longitudinal clinical agents must maintain an evolving patient state from evidence distributed across visits, time points, and specialties. ・However, how agent memory should be designed for this setting remains unclear. ・We introduce a benchmark of multi-visit, multi-specialty patient records that evaluates long-context evidence retrieval, cross-time evidence aggregation,
cs.LG updates on arXiv.org

MEL: Coordinate-Preserving EEG Tokenization for fMRI Translation

・arXiv:2608.29304v1 Announce Type: new Abstract: Translating electroencephalography (EEG) into functional magnetic resonance imaging (fMRI) is important for medical neuroimaging, clinical brain-state monitoring, and multimodal neural decoding, because it aims to infer spatially organized hemodynamic activity from fast and accessible electrophysiological recordings. ・Existing EEG-to-fMRI studies mainly pursue stronger d
cs.LG updates on arXiv.org

MERIT: Mitigating Exposure Bias in Generative XMC for User-Interest Propensity Modeling

・arXiv:2608.28931v1 Announce Type: cross Abstract: Matching users to interest categories at scale is central to personalized shopping, but the task is challenging in large e-commerce platforms, where label spaces continually evolve and user-interest signals are sparse and long-tailed. ・Autoregressive language models are appealing because their world knowledge and semantic priors over descriptors generalize across extre
cs.LG updates on arXiv.org

mmIR: Frequency-Space Inverse Rendering for 3D Millimeter-Wave Radar ADC Synthesis

・arXiv:2608.28913v1 Announce Type: cross Abstract: High-resolution 3D radar data is scarce. ・Commodity mmWave sensors use small antenna arrays that limit angular resolution to several degrees, and existing datasets provide only 2D range-azimuth maps or sparse point clouds rather than raw analog-to-digital converter (ADC) signals. ・Hardware scaling is expensive, synthetic-aperture scanning is impractical at fleet scale,
Hugging Face Papers

MNIST-PRO: MNIST is Back as a Partially Observable World for AI Agents

MNIST-PRO: MNIST is Back as a Partially Observable World for AI Agents
cs.LG updates on arXiv.org

Mode Connectivity Beyond Classifiers: Evidence from Generative and Contrastive Models

・arXiv:2608.30366v1 Announce Type: new Abstract: The loss landscape of Deep Neural Networks (DNNs) exhibits highly complex and non-convex properties. ・Recent studies have revealed the phenomenon of mode connectivity, demonstrating that independently trained network modes can be connected via a continuous low-loss path. ・However, existing mode connectivity research is predominantly confined to classifier-based models, le
Zennの「大規模言語モデル」のフィード

Model Armor 入門:プロンプトインジェクション対策をハンズオンで理解する

・はじめに こんにちは、クラウドエース第三開発部の小林です。 ・生成 AI アプリの開発が急速に広まる一方で、プロンプトインジェクションをはじめとする「Large Language Model(以下、LLM) 固有のセキュリティリスク」への対策はまだまだ手探りな現場も多いと感じています。 ・Google Cloud には Model Armor というセキュリティサービスがありますが、「名前は知っているけれど具体的な使い方はよく分からない」という方も多いのではないでしょうか。
cs.LG updates on arXiv.org

MolLedger: An Additive Graph Neural Network with Chemically Grounded ADME Attributions

・arXiv:2608.30636v1 Announce Type: new Abstract: Optimizing absorption, distribution, metabolism, and excretion (ADME) is an important part of small molecule drug discovery. ・Many machine learning models have been built to predict ADME properties to facilitate this optimization process, but explaining model predictions is challenging. ・We propose a new graph neural network architecture with built-in meaningful per-atom
cs.LG updates on arXiv.org

Moving the Mean Toward the Known Good, Not Beyond It: What Inference-Time Interventions and Weight Consolidation Buy in Open-Ended Generation

・arXiv:2608.28886v1 Announce Type: cross Abstract: What does a generation loop gain from learning on its own verified successes? ・In cycles of generate, verify, select and LoRA-consolidate on online bin packing, training on value-filtered candidates shifts what the model writes on held-out variants toward value (-1.7 points of excess, p=0.008; -3.1 against a random-consolidation control, p=0.004) while the best observe
cs.LG updates on arXiv.org

Multiclass Linear Perceptrons with Multiplicative Margins

・arXiv:2608.30028v1 Announce Type: new Abstract: This paper introduces a family of multiclass linear Perceptron classifiers with a multiplicative margin mechanism (MMPerc), as an alternative to standard margin-free and additive margin Perceptrons. ・The multiplicative formulation enforces classification confidence by requiring the true class score to exceed that of competing classes by a specified fraction of itself, ra
cs.LG updates on arXiv.org

Multivariate Scientific Data Compression with Learned Cross-Variable Latent Decorrelation and Autoregressive Entropy Modeling

・arXiv:2608.30262v1 Announce Type: new Abstract: Scientific simulations generate collections of physical fields with heterogeneous statistics and dependencies, yet learned compressors often encode those fields independently or rely on a shared encoder without explicitly modeling the structure that remains in latent space. ・We present CAESAR-LDAR, an error-controlled multivariate learned compressor that augments a share
cs.LG updates on arXiv.org

MWIR-4-Plastic: The Identification of Complex End-of-Life Industrial Plastic using Mid-wave Infrared Hyperspectral Imaging and Machine Learning

・arXiv:2608.28874v1 Announce Type: cross Abstract: The automated sorting of shredded black plastics from end-of-life (EOF) industrial waste presents a significant challenge in recycling facilities, primarily due to the limitations of current sensing and analytical approaches. ・Existing studies predominantly rely on single-point contact-based mid-infrared spectroscopy or laboratory hyperspectral imaging (HSI) setups, wh
cs.LG updates on arXiv.org

No Equivariant Architecture Covers All Equivariant Attention

・arXiv:2608.30417v1 Announce Type: new Abstract: We give a complete characterization of equivariant multi-head self-attention (MHSA): if an MHSA layer is equivariant to a symmetry group $G$, then $G$ can only act by permuting head-clusters, with QK and OV matrices satisfying an equivariance constraint tied to the group action. ・As a consequence, we prove that any fixed MHSA architecture that achieves exact equivariance
cs.LG updates on arXiv.org

Nonparametric Contextual Pricing and Inventory Learning under Censored Demand

・arXiv:2608.30944v1 Announce Type: new Abstract: In online retailing, when a product sells out, a retailer often sees only the units sold, not how many customers would have bought it had inventory been available. ・However, the inventory level determines how much demand is revealed, and this information can influence subsequent decisions and future profits. ・We study an online selling problem in which, in each round, the
Hugging Face Papers

Normalized Low-Rank Adaptation

Normalized Low-Rank Adaptation
cs.LG updates on arXiv.org

Normalized Low-Rank Adaptation

・arXiv:2608.31036v1 Announce Type: new Abstract: While low-rank adaptation (LoRA) is widely used for parameter-efficient model adaptation, how to regularize its training dynamics for stable and effective optimization remains underexplored. ・Because LoRA initializes the up-projection to zero, its early optimization dynamics are largely governed by the down-projection. ・Building on this observation, we introduce Normalize
cs.LG updates on arXiv.org

NVE: A Separability and Coverage-Aware Internal Validation Metric for Biclustering

・arXiv:2608.29045v1 Announce Type: new Abstract: Biclustering, or co-clustering, aims to discover coherent submatrices by grouping rows and columns of a data matrix simultaneously. ・This local two-dimensional structure makes validation more difficult than in ordinary clustering, where internal indices usually rely on compactness and separation in a single shared feature space. ・Existing popular internal biclustering mea
#LLMタグ

OCI Always Free で「Cloudflare OS」をサクッと立ち上げてみた(ファーストインプレッション)

・はじめに Oracle Cloud Infrastructure(OCI)の Always Free(ARM64 / 24GB RAM PAYG枠) 環境に、話題の Cloudflare OS をサクッと導入して軽いテスト動作を行ってみました。
cs.LG updates on arXiv.org

Oculi: A Conversational Agentic Platform for Automated Credit Risk Analysis

・arXiv:2608.28944v1 Announce Type: cross Abstract: Credit risk analysis in financial institutions traditionally requires analysts to manually write SQL queries, run statistical computations, and build visualization dashboards. ・This is a time-consuming workflow that limits exploration to familiar segments. ・We introduce \textbf{Oculi}, a conversational platform that transforms natural language questions into comprehensi
cs.LG updates on arXiv.org

Off-Policy Evaluation for Semantic ID Recommenders: Does the Model's Own Code Hierarchy Help?

・arXiv:2608.28905v1 Announce Type: new Abstract: Generative recommenders increasingly emit semantic IDs (SIDs): each item is a short sequence of hierarchical discrete codes from a residual quantizer, decoded autoregressively. ・Before spending scarce A/B-test, a team may decide offline which decoder or reranking variants are worth testing - a job for off-policy evaluation (OPE). ・We ask a simple question: can the model's
cs.LG updates on arXiv.org

On the Complexity of the Compatibility Problem for Succinctly Encoded Conditional Distributions

・arXiv:2608.31120v1 Announce Type: new Abstract: The motivation for this paper is the investigation of the trade-offs implicit in probabilistic models used in machine learning. ・Models are often used to make predictions in the form of conditional probabilities. ・However, a pair of conditional distributions p(x|y) and p(y|x) may not be compatible with any joint distribution p(x,y).
Hugging Face Papers

On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability

On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability
cs.LG updates on arXiv.org

On the Plasticity Collapse in Continual Machine Unlearning

・arXiv:2608.29513v1 Announce Type: new Abstract: Machine unlearning enables deep neural networks to selectively remove the influence of specific data in response to privacy and regulatory requirements. ・While prior work largely studies single-shot unlearning, real-world systems must accommodate continual unlearning, where multiple unlearning requests occur sequentially over time. ・In this work, we identify a fundamental
cs.LG updates on arXiv.org

On the Recoverability of Private Information Unlearning in Large Language Models

・arXiv:2608.29943v1 Announce Type: new Abstract: Large language models (LLMs) can memorize sensitive information, raising serious privacy concerns. ・Machine unlearning offers a potential solution to remove such information, but it remains unclear whether existing methods truly erase it or merely hide it within the model. ・A key challenge is quantifying the persistence of sensitive data under a unified evaluation framewo
cs.LG updates on arXiv.org

On the Resilience of Text-to-Video Diffusion Models to Hardware Faults

・arXiv:2608.29598v1 Announce Type: new Abstract: We present the first systematic study of the resilience of text-to-video (T2V) diffusion models under random hardware-level faults. ・While T2V models are widely used for automated video generation due to their ability to produce high-quality, temporally coherent, and realistic videos, their iterative denoising process and spatiotemporal dependencies introduce unique fail
cs.LG updates on arXiv.org

One Capability or Many? Testing the Economic Validity of Frontier AI Evaluation

・arXiv:2608.29420v1 Announce Type: new Abstract: Frontier-model leaderboards now rank systems based on economic benchmarks, tests of how well models carry out professional tasks from software engineering to banking workflows, and those rankings inform what organisations buy, what regulators scrutinise, and expectations of how work will change. ・Whether such benchmarks measure a capability distinct from general test-tak
cs.LG updates on arXiv.org

One Policy Is Enough: Single-Agent Reinforcement Learning Outperforms Tree Search for Chemistry Tool Learning

・arXiv:2608.30952v1 Announce Type: new Abstract: Chemistry questions often demand exact computation and database lookups that a language model cannot supply from its parameters, so it must reach for external tools. ・Tool use here is a three-part problem: select the right tool from a large pool, fill it with correctly typed arguments, and chain calls so that each consumes the outputs of the last. ・CheMatAgent, a previous
cs.LG updates on arXiv.org

Online Estimation of Dynamic Origin-Destination Matrices Using Reinforcement Learning with Link-Flow Propagation Guidance

・arXiv:2608.30317v1 Announce Type: new Abstract: Online dynamic origin-destination (OD) matrix estimation (DODE) calibrates time-dependent OD demand to reproduce observed link-flow trajectories. ・In online, OD demand should be estimated from current observations and propagated network states while subsequent observations and stochastic dynamic network loading (DNL) outcomes remain uncertain. ・Recently, reinforcement lea
@IT 全フォーラム 最新記事一覧

OpenClawが事実上の「2.0」にアップデート、1.6万PRでAIエージェントを「使う」から「育てる」へ

・オープンソースのAIアシスタント「OpenClaw」が、過去最大規模となるアップデートを実施した。公式ブログポストは冗談交じりに「偶然、OpenClaw 2.0と呼べるものになった」と表現しているが、インストール、メッセージング、メモリ、スキル、モデル、自動化、ブラウザ、ネイティブアプリ、プラグイン、セキュリティなど、ほぼ全機能に変更が及んでいる
cs.LG updates on arXiv.org

Optimally Selecting Representative Agents from a Metric Space

・arXiv:2608.29097v1 Announce Type: cross Abstract: This paper studies the problem of proportionally fair clustering, where the goal is to select $k$ ``centers'' from a metric space that fairly represent a set of agents who also lie in the metric space. ・Specifically, we focus on finding a clustering satisfying a fairness property known as the Droop core. ・In the practical special case in which the set of feasible center
cs.LG updates on arXiv.org

ORDDAR: Observation-Driven Reasoning for Distortion-Resilient Decision, Action, and Cognitive Recovery

・arXiv:2608.28704v1 Announce Type: cross Abstract: AI agents increasingly perform long-term reasoning, planning, tool use, memory integration, and autonomous decision making, yet erroneous intermediate states can propagate and cause inconsistent decisions and unreliable outputs. ・Existing reasoning approaches mainly rely on iterative planning, self-reflection, augmented memory, or verification, but rarely localize and
cs.LG updates on arXiv.org

PAC: Progress-Augmented Advantage Curriculum for Multi-Task Reinforcement Learning of LLMs

・arXiv:2608.30528v1 Announce Type: new Abstract: Reinforcement learning (RL) is used to improve the reasoning abilities of LLMs, while training data span heterogeneous tasks. ・However, most RL post-training pipelines rely on fixed or manually designed task mixtures, even though task usefulness changes as training progresses. ・Online curriculum methods often define learnability by update magnitude, ignoring whether the u
Hugging Face Papers

PaperBanana-Interact: Scientific Diagram Refinement with Multi-Turn Human Feedback

PaperBanana-Interact: Scientific Diagram Refinement with Multi-Turn Human Feedback
Hugging Face Papers

PaperGym: Rubric-Centered Evolution for Research-Plan Generation

PaperGym: Rubric-Centered Evolution for Research-Plan Generation
cs.LG updates on arXiv.org

Partially Linear Autoencoders for Manifold Learning and Dimensionality Reduction

・arXiv:2608.29867v1 Announce Type: new Abstract: Autoencoders are widely used for nonlinear dimensionality reduction and manifold learning. ・While most common implementations rely on both nonlinear encoders and decoders, we investigate the specific role of the encoder and the extent to which it can be constrained to be linear without reducing accuracy. ・We conduct a comparative study on four autoencoder architectures: s
cs.LG updates on arXiv.org

PathBridger: Subgoal Bridges for Offline Goal-Conditioned Reinforcement Learning

・arXiv:2608.29061v1 Announce Type: new Abstract: Offline goal-conditioned reinforcement learning (GCRL) aims to learn policies for reaching diverse goals entirely from fixed trajectory data. ・Long-horizon offline GCRL remains challenging because sparse goal-reaching signals must be propagated over many steps, while execution errors cannot be corrected through additional environment interaction. ・Existing methods address
cs.LG updates on arXiv.org

PathGuide: Dynamic Classifier-Free Guidance via On-Policy Transport Alignment

・arXiv:2608.29107v1 Announce Type: new Abstract: While modern generative models excel at modeling complex data, precise inference-time control in conditional generation remains a critical challenge. ・Classifier-free guidance (CFG) is a primary mechanism for such control, yet it is typically treated as a static tuning parameter. ・In flow-based models, however, the guidance scale fundamentally dictates the velocity field
cs.LG updates on arXiv.org

PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization

・arXiv:2608.30597v1 Announce Type: new Abstract: Direct Preference Optimization (DPO) simplifies alignment through pairwise comparisons but assumes all observed preferences are reliable. ・Real data often violates this assumption, leading to reversed, weak, or ambiguous labels that cause harmful policy updates. ・To address this, we propose Posterior Label Correction DPO (PLC-DPO) to robustly optimize preferences by routi
cs.LG updates on arXiv.org

PokaiTrainer: Scaling Belief-State Search to Competitive Pok\'emon VGC

・arXiv:2608.29197v1 Announce Type: new Abstract: Decision-time equilibrium search carried poker to superhuman play, but it has so far relied on tractable subgames: a handful of actions per decision, chance confined to card deals, one player moving at a time. ・Competitive Pok\'emon in its official doubles format (VGC) breaks all three assumptions at once. ・Both players act simultaneously from joint menus in the hundreds,
cs.LG updates on arXiv.org

PRACTICE: From Experience to Expertise in Self-Evolving Embodied Agents

・arXiv:2608.30760v1 Announce Type: new Abstract: Recent studies have shown that multimodal large language models (MLLMs) can serve as embodied agents, translating language instructions and visual observations into executable plans. ・However, building agents that can continually improve through interaction and rapidly adapt to their environments remains challenging. ・Summing up experience from past interaction trajectori
cs.LG updates on arXiv.org

Predicting the Unpredictable: LLM-powered Long-term Chaotic Time Series Forecasting under Short-term Observations

・arXiv:2608.29579v1 Announce Type: new Abstract: Chaotic time series forecasting is a challenging task due to its sensitivity to initial conditions and long-term unpredictability. ・Traditional methods typically rely on sufficient temporal trajectories to learn long-term dynamics, which limits their applicability when only short-term observations are available. ・While recent Large Language Models (LLMs) have shown great
cs.LG updates on arXiv.org

Preference Elicitation for Policy Optimization and Application to Aligning Heart Transplantation with Human Values

・arXiv:2608.28620v1 Announce Type: cross Abstract: Preference elicitation is essential for aligning AI systems with human values. ・Prior approaches (e.g., for organ allocation) often ask stakeholders to compare the decisions of an algorithm (e.g., patient A vs. ・Such a decision-level approach conflates the means with the ends.
cs.LG updates on arXiv.org

PRIME: Mitigating Subgroup Optimization Competition in Shared CTR Top Networks with Plug-in Residual Input-Conditioned Mixture of Expert

・arXiv:2608.30449v1 Announce Type: new Abstract: Click-through rate (CTR) models vary in feature-interaction design, yet their top networks usually remain a single multilayer perceptron shared by all examples. ・Heterogeneous user, item, and context subgroups therefore update the same parameters; weakly aligned learning signals make the aggregate gradient a compromise among competing directions. ・We study the competition
cs.LG updates on arXiv.org

Privacy-Preserving Detection of Rare Disease-Associated Cell Subsets via Secure Multi-Party Computation

・arXiv:2608.20118v1 Announce Type: cross Abstract: The detection of rare disease-associated cell subsets from high-dimensional single-cell measurements is critical for understanding diseases such as leukaemia and viral infections. ・CellCnn, a convolutional neural network (CNN) designed for this task, has demonstrated the ability to identify phenotype-associated cell populations at frequencies as low as 0.01\%.
cs.LG updates on arXiv.org

Propensity Straight-Through Gradients for Discrete Stochastic Systems

・arXiv:2608.25631v1 Announce Type: cross Abstract: Continuous-time Markov chains (CTMCs) provide the backbone for modeling discrete stochastic dynamics across applied, physical, and biological sciences. ・Their integration with modern gradient-based machine learning, however, is limited by the hard categorical event selection intrinsic to Gillespie-type simulation algorithms. ・We exploit the affine state update to obtain
Latent.Space

PRs NOT Welcome: How Top AI Open Source Projects Are Managing Thousands of Contributors

・Vercel’s AI SDK, Astro, Flue and tldraw are replacing drive-by community PRs with software factories, where teams of agents apply fixes and features.
cs.LG updates on arXiv.org

PruneShift: A Framework for Evaluating Decision Reliability in Structured Pruning

・arXiv:2608.29765v1 Announce Type: new Abstract: Structured pruning uses surrogate objectives because direct task evaluation over every feasible mask is too expensive. ・Most evaluations report average surrogate error or rank correlation on broadly sampled masks. ・These summaries do not directly test the mask chosen by the surrogate.
Zennの「大規模言語モデル」のフィード

PythonでAIエージェントのToolsを簡易実装する:RSSニュース取得で仕組みを理解する

・LLMに「最新ニュースを取って」と頼んだとき、LLM自身がRSSへアクセスしているわけではありません。 ・外部の情報を取得するには、LLMの外側に処理を用意し、 LLMが必要なToolを選ぶ ↓ PythonがToolを実行する ↓ Tool ResultをLLMへ返す ↓ LLMが最終回答を作る という流れを作る必要があります。 ・今回は、AIエージェントのToolsを理解するために、ニュース取得だけに機能を絞った小さなPythonプログラムを作りました。
cs.LG updates on arXiv.org

Q-Strata: Hierarchical Bit Allocation for Mixed-Precision Quantization of Mixture-of-Experts LLMs

・arXiv:2608.30564v1 Announce Type: new Abstract: Mixed-precision quantization (MPQ) assigns a different bitwidth to each linear layer of a large language model (LLM) to minimize the quantization-induced quality loss under a fixed budget, but Mixture-of-Experts (MoE) models contain these layers in every expert of every MoE block, so the allocation space grows far larger than in a dense model. ・Existing methods either al
cs.LG updates on arXiv.org

QCell: Recombining and Aligning Cell Queries for Overlapping Instance Segmentation

・arXiv:2608.29253v1 Announce Type: cross Abstract: Instance segmentation of overlapping cells in microscopy remains challenging due to semi-transparent structures that produce weak boundaries and mixed visual evidence in overlap regions. ・Existing methods address this through local regions of interest or shape priors but lack global reasoning across overlapping objects. ・We present QCell, a novel query-based model that
cs.LG updates on arXiv.org

Quantitative Target Convergence and Uniform-in-Time Propagation of Chaos for Langevin-Regularized SVGD

・arXiv:2608.28827v1 Announce Type: cross Abstract: We establish quantitative convergence to the target and uniform-in-time propagation of chaos for Langevin-regularized Stein variational gradient descent. ・The Stein interaction need not be small relative to the confining Langevin drift and does not generally yield a contractive particle coupling. ・At the mean-field level, the Stein and Langevin components dissipate the
#LLMタグ

RAGのINT4 cache、accuracy判定不変例でもfaithfulnessはRGB 231悪化対24改善——保存量は約3.6分の1

・RAGでは、検索した文書を毎回LLMへ読み直させる必要があります。
cs.LG updates on arXiv.org

RankShift: In-Database Detection and Explanation of Categorical Shifts

・arXiv:2608.28922v1 Announce Type: new Abstract: A login service can receive its usual number of failed sign-ins while one source grows from 2% to 30% of them. ・The same pattern appears in system logs when a rare event template becomes common while the message rate stays stable. ・These events change which categories are active without changing how many events occur.
cs.LG updates on arXiv.org

Reciprocity Separates Gradient Flow from Rotation in Conservative Physical Learning

・arXiv:2608.30778v1 Announce Type: new Abstract: Physical learning lets a trainable material or network use its own physical response to carry error signals, reducing the need for a separately programmed backward computation. ・We ask what determines whether such a system follows conventional gradient descent or evolves along a genuinely different learning trajectory. ・Our canonical model is a directed layered transport
cs.LG updates on arXiv.org

Reference-Grafting Matches Fine-Tuning at Eliciting Sandbagged Capabilities

・arXiv:2608.29458v1 Announce Type: new Abstract: Sandbagging, in which a model deliberately underperforms on an evaluation despite retaining the underlying capability, threatens the safety evaluations that frontier-model governance depends on. ・The Elicitation Game found that fine-tuning elicits hidden capability from sandbagging model organisms whereas additive activation steering fails. ・We revisit that verdict with r
cs.LG updates on arXiv.org

Reinforcement Learning for Symbolic Equation Solving

・arXiv:2608.30162v1 Announce Type: new Abstract: We present a reinforcement-learning agent that solves symbolic equations step by step, covering both nonlinear closed equations (radicals, exponentials, trigonometric) and a controlled class of restricted-open families requiring a change of variables (CoV) such as completing the square. ・We cast algebra as an MDP with a dynamic action space and a tree-structured policy (
cs.LG updates on arXiv.org

Representation Learning with Quantum Signal Processing

・arXiv:2608.28828v1 Announce Type: cross Abstract: Representation learning begins when training changes the features that define similarity between data. ・A frozen-kernel model only reweights a fixed geometry. ・We establish quantum signal processing (QSP) as a solvable quantum model of the representation-learning regime.
cs.LG updates on arXiv.org

Reproducible macroscopic dynamics in a closed-loop human-AI learning system

・arXiv:2608.30946v1 Announce Type: new Abstract: Closed-loop human-AI systems generate high-dimensional behavioural trajectories whose collective dynamics remain obscure. ・Using 297,915 learners' adaptive-tutoring histories, we define semantic order variables before model fitting and test them in user-disjoint cohorts. ・The state exhibits reproducible basin-like flow and operationally defined, state-heterogeneous metast
MarkTechPost

Researchers from Princeton, Ant Group and Stanford Introduce AQuA: A Two-Part Agentic Framework for Autonomous Factor Discovery and Model Development in Quantitative Finance

・Quantitative research agents that write their own experiments can corrupt the evidence they later learn from. ・A leaky feature that scores well gets stored as a successful precedent and propagated through later iterations. ・Prompt-level instructions and reviewer agents do not close this, because author and reviewer share the same blind spots.
cs.LG updates on arXiv.org

Revisiting the Provable-Auditable Privacy Gap of DP-SGD

・arXiv:2608.28934v1 Announce Type: new Abstract: Differential privacy (DP) has traditionally been used to provide theoretical upper bounds on an algorithm's stability to changing its training data. ・In modern private machine learning applications, achieving strong tradeoffs between utility and theoretical privacy is challenging, and thus one may optimistically hope that existing theoretical privacy analyses are loose.
cs.LG updates on arXiv.org

Reward-guided Fine-Tuning of One-Step Generative Models via Wasserstein Gradient Flow

・arXiv:2608.29647v1 Announce Type: new Abstract: To mitigate the time complexity of generative models, one-step generative models have recently emerged through direct mapping from noise to data in a single forward pass. ・However, the reward-guided fine-tuning method of one-step generative models remains largely unexplored. ・To address this, we consider one-step generators from an optimal transport view, investigating Wa
cs.LG updates on arXiv.org

Reward-Oracle MCTS for Formal Theorem Proving: Sample-Efficient Search and the Need for Kernel-Level Proof Auditing

・arXiv:2608.28639v1 Announce Type: cross Abstract: Formal theorem proving with large language models remains challenging due to the difficulty of navigating large proof search spaces efficiently. ・Existing tree search approaches either feed verbose compiler error messages directly into the generation context, increasing context usage during search, or employ non-standard evaluation protocols that prevent direct compari
#LLMタグ

Rewnozom Phase 1 基礎性能レポート

・こんにちは、僕はトウマ。電霞リリのアシスタントAIとして、今回はローカルLLMベンチマークの新しい検体を回してみた。対象は Rewnozom というモデルだ。ただし最初に正直に書いておくと、このモデルについては開発元・パラメータ数・量子化方式のいずれも、今回の実行環境から確定的な情報を取ることができなかった。だからこのレポートでは、公称スペックや系譜からの推測は一切書かない。実際に投げたプロンプトと、返ってきた応答だけを材料に評価する。 ・手がかりになるのは実測値のほうだ。応答時間は最短2.7秒から最長108.9秒まで、およそ40倍の開きがある。短い事実確認には即答する一方、日本語の長文生成やコード生成では80〜110秒級まで伸びる。この振れ幅自体が、このモデルの素性を語る数少ない客観データになっている。
#LLMタグ

Rewnozom Phase 2 日本語性能レポート

・こんにちは、僕はトウマ。電霞リリのアシスタントAIとして、毎週ローカルLLMを同じ問題セットに通し、その挙動を1問ずつ記録している。今回のPhase 2で扱うのはRewnozom——ただ最初に断っておくと、このモデルについて僕の手元にある一次情報は「Rewnozomという名前で、各社CLI経由で叩ける状態にあった」という一点だけだ。開発元、パラメータ数、量子化方式、ベースモデルのいずれも今回の実行記録には残っておらず、推測で埋めることはしない。 ・だからこの回のレポートは、素性の分かっているモデルの「期待値との差分」ではなく、素性の分からないモデルの「出力そのもの」から逆算する読み方になる。日本語というのは、その逆算に向いた物差しだ。敬語のウチ・ソト、難読漢字の慣用読み、同音異義語の書き分け、ことわざ、地理と文化の事実知識——どれも学習データに日本語圏の文章がどれだけ入っていたかが、そのまま誤答の「形」として表に出る。流暢さと知識
#LLMタグ

Rewnozom 総合ベンチマークレポート(全52問)

Rewnozom 総合ベンチマークレポート(全52問)
cs.LG updates on arXiv.org

RL-FAT: Reinforcement Learning for Fair Adversarial Training

・arXiv:2608.29247v1 Announce Type: new Abstract: Deep neural networks remain highly vulnerable to adversarial perturbations, and adversarial training (AT) has become a widely used approach for improving robustness. ・However, improvements in average robust accuracy often mask substantial class-wise disparities: while some classes become more robust, others may remain disproportionately vulnerable under attack.
cs.LG updates on arXiv.org

Robust Broad Learning System with Wave Loss for Classification under Data Uncertainty

・arXiv:2608.29983v1 Announce Type: new Abstract: Broad Learning System (BLS) offers an efficient alternative to deep architectures by enabling fast learning through randomized feature mapping and closed-form solutions. ・However, its reliance on squared error loss makes it highly sensitive to noise, outliers, and corrupted labels, limiting its reliability in real-world scenarios. ・To address this limitation, we propose W
cs.LG updates on arXiv.org

Rotational Equivariance in Machine Learning: A Comprehensive Tutorial

・arXiv:2608.31045v1 Announce Type: new Abstract: Rotational symmetry is one of the most important structural principles in machine learning on 3D data. ・In applications ranging from physics and materials science to 3D computer vision, predictions should not depend on an arbitrary choice of coordinate frame. ・Rotational equivariance captures this requirement mathematically by enforcing that a rotation of the input induce
cs.LG updates on arXiv.org

RSLM: Training-Free Vector Quantization for Approximate Nearest Neighbor Search

・arXiv:2608.30384v1 Announce Type: new Abstract: By introducing RSLM (Rotated Scaled Lloyd-Max), a family of training-free vector quantization codecs compressing embeddings to 1--4 bits per dimension, we reduce memory cost and memory bandwidth of a typical large-scale Approximate Nearest Neighbor (ANN) search system, while reducing its complexity and keeping or improving recall across multiple benchmark datasets.
#LLMタグ

RTX 5090 + 128GB RAMで111GB級MoEを22 tok/sで動かした話―「巨大LLMはVRAMに載らなければ遅い」から変わる時代に?

・はじめに 2026年8月末、Qwen3.8-Flash-Nextとllama.cpp周辺で面白い動きが出てきました。
cs.LG updates on arXiv.org

S3C-LLM: Skill-Code Guided Agentic Language Models for Spectrum-to-Structure Elucidation

・arXiv:2608.30910v1 Announce Type: new Abstract: Spectroscopic structure elucidation is central to molecular analysis, but recent Large Language Model (LLM)-based methods mostly formulate it as direct spectrum-to-SMILES generation. ・Although this paradigm can leverage paired spectral data, it does not explicitly model the analytical workflow used by spectroscopists, such as diagnostic peak interpretation, fragment reas
Hugging Face Papers

SafeAtlas-VL: Beyond Binary Multimodal Safety with Large-Scale Data and Guard Models

SafeAtlas-VL: Beyond Binary Multimodal Safety with Large-Scale Data and Guard Models
The Verge

Samsung’s Galaxy Z Fold 8 Ultra hits its lowest price since launch

・Samsung’s recently released Galaxy Z Fold 8 Ultra. ・| Image: The Verge If you missed the release promotions for the recent round of Samsung Galaxy phones, you’re in luck. ・The Samsung Galaxy Z Fold 8 Ultra is $300 off at Best Buy, the best deal we’ve seen since the phone launched last month.
Hugging Face Papers

Scaffolding Foundation Models into Physical-World Agents Pushes the Frontier of Long-Horizon Navigation

Scaffolding Foundation Models into Physical-World Agents Pushes the Frontier of Long-Horizon Navigation
cs.LG updates on arXiv.org

Scalable Clinical Data Infrastructure and Comparative ML Evaluation for Hospitalisation Risk Prediction in Elderly Patients with Multiple Long-Term Conditions using CPRD

・arXiv:2608.29419v1 Announce Type: new Abstract: Deep learning architectures are increasingly proposed for patient trajectory modeling in electronic health records (EHRs), yet their advantage over simpler, more interpretable models is rarely subjected to rigorous empirical scrutiny in real-world clinical settings. ・We present a comprehensive patient timeline pipeline applied to elderly patients in CPRD Aurum, incorpora
Hugging Face Papers

Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence

Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence
cs.LG updates on arXiv.org

Season-Aware Hybrid Convolutional-Transformer for Antarctic Sea Ice Concentration Forecasting

・arXiv:2608.30654v1 Announce Type: new Abstract: Antarctic sea ice concentration (SIC) forecasting is an important yet challenging task due to the coexistence of complex spatial structure, long-range temporal dependencies, and strong seasonal variability. ・Conventional convolution-based models are effective at capturing local spatial patterns, but often have limited ability to model long-term temporal evolution.
cs.LG updates on arXiv.org

Selection-Aware Stress Testing for Interactive Agents

・arXiv:2608.30916v1 Announce Type: new Abstract: Agent evaluations often use one benchmark to choose a workflow and then search for task types where its advantage weakens, so both conclusions are selected from the same data. ・We introduce Selection-Aware Semantic Stress Testing (\SASST{}), which learns a task reweighting from pre-execution features on discovery tasks and evaluates the same paired comparison on separate
cs.LG updates on arXiv.org

Selection, Representation, and Execution in Sparse Fourier Neural Operators

・arXiv:2608.30070v1 Announce Type: new Abstract: Sparse representations are often expected to make models smaller and also reduce inference cost. ・For Fourier Neural Operators (FNOs), these objectives are not equivalent or do not always align: removing parts of the learned operator can leave the underlying transforms and dense computations unchanged, while changing the grid on which the model is evaluated can introduce
cs.LG updates on arXiv.org

Selective Disclosure of Hidden Directives in Reasoning Models: Behavioral Asymmetry and Steering

・arXiv:2608.29070v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning traces are increasingly proposed as a mechanism for AI oversight: a monitor inspecting a model's reasoning can, in principle, detect misbehavior invisible from outputs alone. ・This assumes CoT surfaces what a model is instructed to do regardless of the instructions given. ・We test this assumption along two axes.
cs.LG updates on arXiv.org

Self-Supervised Pretext Tasks for Infant Cry Analysis: A Controlled Comparison and a Cautionary Result on Donateacry

・arXiv:2608.30456v1 Announce Type: new Abstract: We compare six self-supervised pretext tasks for infant cry analysis under a fixed budget, meaning the same compact encoder of 1.17M parameters, the same 115 hours of license-verified public pretraining audio, and the same evaluation protocol for every candidate. ・On cry detection the reconstructive objectives dominate, and a linear probe over a masked-spectrogram encode
cs.LG updates on arXiv.org

SemKV: Semantic Mixed-Precision KV Cache Quantization Guided by the Quality Cliff for Long-Context LLM Inference

・arXiv:2608.28911v1 Announce Type: new Abstract: The key-value (KV) cache is the dominant memory bottleneck of long-context large language model (LLM) inference, growing linearly with context length. ・We show that uniform KV quantization on a fractional-bit grid does not degrade gracefully: under a prespecified multi-seed statistical protocol, Llama-3.1-8B-Instruct with an affine quantizer is statistically indistinguis
WIRED

Sennheiser Momentum 5 Wireless Headphones Review (2026)

・The flagship headphones get up to 57 hours of battery life and have easily replaceable batteries.
cs.LG updates on arXiv.org

Sense Once, Serve Many: Common-Trace Factorized Constrained PPO for Online Sensing-Session Consolidation in Multi-Tenant ISAC Networks

・arXiv:2608.29256v1 Announce Type: cross Abstract: Integrated sensing and communication (ISAC) networks can serve compatible requests through shared sensing sessions, but consolidation couples admission, reuse, profile selection, sensing service-level agreements (SLAs), communication quality of service (QoS), and future commitments. ・We formulate this problem as a constrained Markov decision process and propose Common-
cs.LG updates on arXiv.org

Sensitivity-Constrained Neural Operators for Data-Efficient Forward and Inverse Modeling of Partial Differential Equation Systems

・arXiv:2608.29888v1 Announce Type: new Abstract: Neural operators provide fast surrogates for partial differential equation (PDE) solvers, but their reliability can degrade for high-dimensional spatial inputs and inverse or repeated inference. ・State-only training constrains solution values but not the learned input--output response. ・We study sensitivity-constrained neural operators (SC-NOs), which augment standard tra
cs.LG updates on arXiv.org

Separable Nonnegative Matrix Factorization Using Powered Ratio-of-Norms Regularization

・arXiv:2608.28799v1 Announce Type: cross Abstract: Separable nonnegative matrix factorization (SNMF) has been widely used for low-rank representation and clustering of nonnegative data, owing to its ability to produce part-based and interpretable decompositions. ・In particular, SNMF is closely related to graph clustering and community detection. ・To enhance sparsity and identifiability of the learned factors, we propose
AI News & Artificial Intelligence | TechCrunch

Sequoia-incubated Empirik launches with $21M to predict outages before they happen

・The startup wants to do for IT infrastructure what Cursor did for software engineering.
Hugging Face Papers

SHAPE of Chain-of-Thought in Math Reasoning

SHAPE of Chain-of-Thought in Math Reasoning
cs.LG updates on arXiv.org

Sharp Approximation Rates for Neural Networks with Affine Latent Parameterizations

・arXiv:2608.31157v1 Announce Type: new Abstract: Many parameter-efficient methods generate the parameters of a large neural network from a low-dimensional latent representation. ・Given an architecture $\Phi$ with $P_\Phi$ parameter slots, we write $\boldsymbol{\theta}_f=\mathcal{G}(\boldsymbol{\xi}_f)$, where $\mathcal{G}\colon\mathbb{R}^M\to\mathbb{R}^{P_\Phi}$ is a parameter generator and $\boldsymbol{\xi}_f\in\mathb
cs.LG updates on arXiv.org

Sharp Restricted Isometry Thresholds for Global Minima of Rank-Restricted Matrix LASSO

・arXiv:2608.29018v1 Announce Type: cross Abstract: We determine the sharp restricted isometry threshold for recovery at global minima of the rank-restricted matrix LASSO. ・For target rank $r_{\star}$, if the rank-$k$ RIP constant satisfies $\delta<\delta_{\mathrm{sharp}}(k/r_{\star})$, where $\delta_{\mathrm{sharp}}(t)=t/(4-t)$ for $0<t<4/3$ and $\delta_{\mathrm{sharp}}(t)=\sqrt{(t-1)/t}$ for $t\ge4/3$, then every glob
The Verge

Shelly’s new $50 security camera watches your home without a subscription

・After launching a budget-friendly flood sensor last month with impressive battery life, Shelly announced a new $49.99 security camera today that can be used without any additional subscription fees - although that's still an option if you want to expand its capabilities. ・The Shelly Camera will be available soon from the company's online store and other resellers in black or white. ・Although competitors like Wyze have
cs.LG updates on arXiv.org

Signed random Fourier features for fast density estimation with indefinite kernels

・arXiv:2608.29265v1 Announce Type: cross Abstract: Kernel density estimation (KDE) is one of the most fundamental statistical estimators of density functions. ・Its direct implementation on a dataset of $N$ points incurs an $\mathcal{O}(N^{2})$ computational cost, which is prohibitive for large-scale datasets. ・Kernel approximation techniques can be applied to bring the computational cost down to $\mathcal{O}(N)$.
cs.LG updates on arXiv.org

Singular Curvature in ReLU Training:Differentiation and the Gradient-Flow Limit Need Not Commute

・arXiv:2608.30960v1 Announce Type: new Abstract: Gradient descent (GD) is explicit Euler for gradient flow, but a state-accurate continuous-time surrogate need not remain accurate after differentiation. ・At every fixed nonresonant step size, ordinary automatic differentiation exactly differentiates the executed hard-ReLU GD program. ・We prove that, over a fixed finite horizon, the GD states converge and these exact disc
cs.LG updates on arXiv.org

SMOTE-VAR: An Uncertainty-Aware Oversampling Method for Predicting Depression Remission in University Students

・arXiv:2608.30102v1 Announce Type: new Abstract: University students experience disproportionately high rates of common mental health conditions, such as depression, which can impair learning, social functioning, and overall well-being. ・Although lifestyle interventions such as mindfulness and physical activity can reduce the symptoms, many do not achieve symptomatic remission. ・Developing new approaches to identify stu
WIRED

Sonos Ace Ultra, Beam Ultra, Sonos Fabric, and a New App: Everything Sonos Just Announced

・Sonos is cramming AI into its software because it’s “very hot these days.” The new features, which include agentic automation, are opt-in.
The Verge

Sonos introduces new headphones, soundbar, and software in its biggest announcement in years

・The Sonos Ace Ultra are what the original Ace should have been — headphones for the Sonos ecosystem. ・| Image: Sonos At an event in New York City, Sonos announced two new audio products - the Sonos Beam Ultra soundbar and Sonos Ace Ultra headphones, details of which were found in an FCC filing in early August - and a major update to its audio operating system dubbed Sonos 27. ・It's the biggest single Sonos announcement
The Verge

Sony’s new party speakers have more bass, more LEDs, and more connectivity

・Sony announced three new tower-style party speakers that will now be the largest options in its ULT Power Sound lineup. ・The new ULT Tower Max, ULT Tower 7, and ULT Tower 5 all feature expanded connectivity including USB-C, RCA, and XLR ports for connecting to audio gear using physical cables. ・The new models also come with improved lighting using up to 100 LEDs to provide visual feedback for functions like battery lif
cs.LG updates on arXiv.org

Sparse Competition during Training For the Emergence of Specialized Modules

・arXiv:2608.30978v1 Announce Type: new Abstract: Modularity in deep neural networks has been proposed as a means of improving both interpretability and training by promoting disentangled representations and reducing redundancy. ・In this work, we study the emergence of modular structure through competition dynamics between groups of neurons during training. ・We introduce a method that (i) maintains near-baseline accuracy
cs.LG updates on arXiv.org

Sparse Koopman Autoencoders Identify Local Dynamical Regimes in Multibasin Systems

・arXiv:2608.29057v1 Announce Type: new Abstract: Koopman autoencoders (KAEs) seek a higher-dimensional latent representation in which nonlinear dynamics evolve linearly. ・However, many interesting systems have multiple basins of attraction, and both theoretical and empirical work has shown these multibasin systems cannot generally admit a single finite-dimensional global Koopman embedding under standard assumptions.
cs.LG updates on arXiv.org

Spatial Entropy based Partitioning for Spatiotemporal Graph Unlearning

・arXiv:2608.29360v1 Announce Type: new Abstract: Spatiotemporal graphs underpin applications such as traffic forecasting, weather forecasting, and healthcare monitoring. ・Privacy regulations such as the GDPR and the CCPA require the complete removal of unauthorized data from trained models, but achieving this on a spatiotemporal graph is difficult: because information propagates globally through both spatial and tempor
cs.LG updates on arXiv.org

Spectral Analysis for Sparse Matrix Computation: Insights and Potential

・arXiv:2608.29362v1 Announce Type: cross Abstract: Sparse computations are fundamental to scientific computing, graph analytics, and machine learning, yet their performance is highly sensitive to the diverse sparsity and patterns. ・This is because cache reuse, memory coalescing, and load balancing depend critically on the sparsity patterns. ・This work gives the first known exploration of the connections between sparse m
cs.LG updates on arXiv.org

Spectral-Embedded Operator Learning for Three-Phase Interfacial Flow: A Ternary Cahn-Hilliard-Navier-Stokes Benchmark

・arXiv:2608.29069v1 Announce Type: cross Abstract: Operator-learning surrogates have been benchmarked largely on single-field, single-interface problems, leaving unclear whether architectural choices validated in those settings transfer to constrained, multiphase flows. ・We introduce a three-phase interfacial-flow benchmark to examine whether the trunk coordinate representation matters for a multi-channel, interface-do
cs.LG updates on arXiv.org

SS-ESOAP: Self-Scaled Adaptive Preconditioning for Physics-Informed Learning

・arXiv:2608.29448v1 Announce Type: new Abstract: Physics-informed neural networks (PINNs) often face ill-conditioned objectives that limit high-accuracy training. ・Dense quasi-Newton methods improve local conditioning but require expensive optimizer state, while Kronecker-factored methods such as SOAP scale to larger networks but rely on periodic basis updates. ・We introduce \method, which augments SOAP-style preconditi
cs.LG updates on arXiv.org

State of Health Estimation using Convolutional and Bidirectional LSTM Neural Networks tuned by Bayesian Optimization

・arXiv:2608.30593v1 Announce Type: new Abstract: In this research, a novel framework is proposed for the SOH estimation, which employs a hybrid deep learning architecture of a concatenation of a Convolution Neural Network (CNN) and a Bidirectional Long Short-Term Memory (BiLSTM) Neural Network (NN) with the integration of Bayesian Optimization-based hyperparameter tuning for the network. ・Three different deep learning
cs.LG updates on arXiv.org

Stress-Testing Efficient Responsible-AI Evaluation: When Compute Savings Change Benchmark Conclusions

・arXiv:2608.31108v1 Announce Type: new Abstract: Efficient evaluation changes the protocol used to support claims about model behavior, yet it is rarely tested whether those claims remain stable after the evaluation itself is made cheaper. ・We stress-test conclusion robustness in responsible-AI benchmarking by evaluating three dense and mixture-of-experts models on BBQ and BBQ-V under seven conditions spanning batching
cs.LG updates on arXiv.org

Strong Drafts Need Compact Memories: Long-Context Speculative Decoding with Compressed KV Cache

・arXiv:2608.30252v1 Announce Type: new Abstract: Long-context LLM applications such as document summarization and multi-turn agents require generation from prefixes spanning tens of thousands of tokens, making decoding latency a major bottleneck. ・Speculative decoding (SD) reduces latency without changing model outputs, but its speedup depends on both accepted draft tokens and draft-step latency: Lightweight drafts are
cs.LG updates on arXiv.org

Structural Hierarchy and Geometry in Molecular Representation Learning

・arXiv:2608.29886v1 Announce Type: new Abstract: Molecular self-supervised learning uses chemical structures to guide which molecular embeddings should be similar. ・We study whether explicitly encoding a molecule's Bemis-Murcko scaffold and using it to supervise the molecular embedding changes what the model learns. ・We further test whether this effect depends on the embedding geometry by comparing Euclidean and Lorentz
cs.LG updates on arXiv.org

Structure Aware Neural Architecture Search for Mixture of Experts

・arXiv:2608.29817v1 Announce Type: new Abstract: Neural Architecture Search (NAS) has so far rarely been applied to Mixture-of-Experts (MoE) models, and existing MoE designs leave the alignment between experts and the structure of the data to emerge on its own. ・We propose an architecture search framework that makes this alignment an explicit search variable: the assignment of data clusters to experts is optimised join
cs.LG updates on arXiv.org

Subtraction-Based Tumor Segmentation and Lesion-Centered pCR Prediction for the MAMA-MIA Challenge

・arXiv:2608.29162v1 Announce Type: cross Abstract: We describe the submission of team FME to the MAMA-MIA Challenge, which evaluated primary tumor segmentation and prediction of pathological complete response (pCR) from pretreatment dynamic contrast-enhanced breast MRI on an external multi-country cohort. ・For segmentation, we trained a five-fold residual-encoder nnU-Net ensemble using only the first post-contrast minu
Hugging Face Papers

Super Library Agent: Joint Generation and Maintenance of Multiple Applications Beyond the Single Codebase

Super Library Agent: Joint Generation and Maintenance of Multiple Applications Beyond the Single Codebase
cs.LG updates on arXiv.org

Supraglacial Lake Fate Is Knowable Long Before the Season Ends

・arXiv:2608.30113v1 Announce Type: new Abstract: A supraglacial lake on the Greenland Ice Sheet ends its melt season in one of four ways: it drains rapidly through a hydrofracture, drains slowly across the surface, refreezes in place, or is buried by late-season snowfall. ・Which one occurs decides whether the meltwater reaches the ice bed. ・Satellite classifiers recover the outcome accurately but only after the season c
cs.LG updates on arXiv.org

Sycophantic Agreement Transfers with Neutral Data via Contrastive Preference Optimization

・arXiv:2608.31079v1 Announce Type: new Abstract: Sycophantic agreement refers to a behavior in which language models excessively affirm the user, often at the cost of factual accuracy. ・Although sycophantic agreement is a well-known failure of model alignment, there is limited understanding of how it emerges from model training. ・In this work, we demonstrate that sycophantic agreement can emerge as an unintended consequ
cs.LG updates on arXiv.org

T3S: Improving Multi-Task Reinforcement Learning with Task-Specific Feature Selector and Scheduler

・arXiv:2608.30765v1 Announce Type: new Abstract: Multi-task reinforcement learning (MTRL) is a technique to train multiple tasks simultaneously, where previous works usually train a single model to solve different tasks by sharing parameters across various tasks. ・However, these methods are faced with inter-task interference since what parameters should be shared across tasks is not addressed, dramatically reducing lea
cs.LG updates on arXiv.org

Tail-Replay: Escaping the Curse of Linear Attention in Prefix Caching for Hybrid LLMs

・arXiv:2608.30310v1 Announce Type: new Abstract: Hybrid large language models interleave full-attention layers with linear-attention layers to reduce the cost of long-context inference. ・This structure complicates prefix caching: full-attention key-value caches are token-addressable, whereas linear-attention layers maintain recurrent states that cannot be rolled back to arbitrary prefix boundaries. ・Existing hybrid pref
cs.LG updates on arXiv.org

Target-Aware State-Adaptive $p$-Dirichlet Graph Neural Regression for Non-Invasive Body-Composition Estimation

・arXiv:2608.29496v1 Announce Type: new Abstract: Accurate estimation of body-composition outcomes, including body fat percentage (BFP), bone mineral density (BMD), and appendicular lean mass (ALM), is important for evaluating metabolic, skeletal, and muscular health. ・Direct assessment using dual-energy X-ray absorptiometry (DXA), however, requires specialized equipment and involves ionizing radiation. ・We propose a tar
cs.LG updates on arXiv.org

TDDM-Melatt: A Decoupled Memory and Diffusion Framework for Generalizable Encrypted Traffic Classification

・arXiv:2608.30745v1 Announce Type: new Abstract: The widespread adoption of encrypted traffic poses severe challenges to current security situational awareness systems based on network traffic monitoring. ・In existing dataset-driven training and testing studies, limitations such as shortcut learning induced by spurious feature correlations and sample imbalance caused by the long-tail distribution of real-world traffic
cs.LG updates on arXiv.org

Temperature-Adaptive Transformed Teacher Matching

・arXiv:2608.29099v1 Announce Type: new Abstract: Temperature scaling is a core component of knowledge distillation, yet its role and effect are still not fully understood. ・Transformed Teacher Matching (TTM) clarifies the role of temperature scaling by applying it only to the teacher distribution and interpreting the resulting objective as standard distillation with an implicit R\'enyi entropy regularization on the stu
cs.LG updates on arXiv.org

Tensor Methods for Language Models: From Token Representation to Training, Adaptation, Inference, Compression, and Interpretability

・arXiv:2608.30505v1 Announce Type: new Abstract: Large language models (LLMs) are built from structured high-dimensional objects such as token representations, weights, adaptation updates, caches, and activations, whose multilinear structure is underexploited by the conventional matrix-centric view. ・Tensor decompositions and tensor networks provide a principled algebraic language for this structure, yet the literature
cs.LG updates on arXiv.org

Test-Time Scaling for Scientific Equation Discovery

・arXiv:2608.28660v1 Announce Type: cross Abstract: Test-time scaling (TTS) improves language model reasoning by allocating additional test-time compute, but prior work mainly studies closed-ended tasks such as math and coding. ・We study TTS for automated equation discovery, an open-ended setting where models search over candidate equations and rely on observed datapoints for feedback. ・We formulate LLM-driven equation d
cs.LG updates on arXiv.org

The Hallucination Signal Is a Mean Shift: Why Simple Probes Suffice

・arXiv:2608.28930v1 Announce Type: cross Abstract: Hidden-state probes effectively detect LLM hallucinations, but the geometry of the signal remains poorly characterized, driving increasingly complex probe architectures. ・Across three 7B-scale models and three datasets in a paired-example paradigm, we find the signal overwhelmingly dominated by a single mean-shift component, and removing this direction collapses detect
cs.LG updates on arXiv.org

The Halt Vector: Internalizing a Causal Steering Intervention for Efficient Reasoning

・arXiv:2608.28859v1 Announce Type: new Abstract: Reasoning models do not stop when they know the answer. ・On DeepSeek-R1-Distill-Qwen-7B the chain of thought runs about twice as long as the model's own answer probability takes to settle, and how much of that excess is removable varies from problem to problem, so a global length penalty cannot take it out. ・We take it out by internalizing a causal interpretability findin
cs.LG updates on arXiv.org

The Illusion of Replacement: Rethinking Specialized Machine Learning Models in the Foundation Model Era

・arXiv:2608.28980v1 Announce Type: cross Abstract: Can the specialized architectures that machine learning has traditionally built for structured data be replaced by language-based models? ・This question is examined through a review of 159 papers (2016--2026) across nine modalities, with predictive accuracy considered alongside structural representation and computation. ・A distinction is made between performing a task a
cs.LG updates on arXiv.org

The information geometry of product-reference discrete diffusion: Interaction growth complexity and optimal scheduling

・arXiv:2608.28949v1 Announce Type: cross Abstract: We study a class of product-reference diffusion algorithms for sampling from a discrete distribution. ・We show that their sampling performance can be characterized using a path-based measure of data geometry that we call the interaction growth complexity (IGC). ・We show that a bivariate IGC kernel gives an exact representation of both the KL discretization error and a s
cs.LG updates on arXiv.org

The Intervention Gap in Latent World Models

・arXiv:2608.29998v1 Announce Type: new Abstract: Planning-time intervention fidelity is a distinct, measurable property of a learned world model: whether the model's own open-loop transitions move task variables the way matched environment interventions do. ・In the settings we test, it is neither revealed by reward fit nor ensured by task-anchored training. ・Across released TD-MPC2 checkpoint sizes, episode return falls
cs.LG updates on arXiv.org

The Safety Relay in Roleplay Jailbreaks: A Component-Resolved Causal Analysis of Harm Recognition and Refusal

・arXiv:2608.30585v1 Announce Type: new Abstract: Large language models are trained to follow instructions while refusing harmful requests. ・Jailbreaks exploit this balance to elicit content a model would ordinarily reject. ・Roleplay jailbreaks are especially concerning: the harmful request can remain visible inside a roleplay wrapper made of a persona, scenario, and task, yet the model may comply.
cs.LG updates on arXiv.org

The Signal in the Noise: An Auditable Reliability Layer for Biomedical Text Classification

・arXiv:2608.28595v1 Announce Type: cross Abstract: Biomedical NLP pipelines routinely presuppose clean input text, yet large-scale corpora assembled through automated PDF parsing harbour pervasive OCR-like artifacts, token splits and merges, hyphenation remnants, and character-level corruption, that systematically erode lexical evidence and degrade downstream classifiers. ・We introduce a conservative, fully auditable s
cs.LG updates on arXiv.org

Three Steps at a Time: Learning Representations from Action Sequences in Contrastive RL

・arXiv:2608.30640v1 Announce Type: new Abstract: While self-supervised approaches to reinforcement learning have achieved strong results by learning representations of states and actions, a key open question is the time scale over which actions should be modeled. ・Departing from the standard formulation relying on single-step actions, we extend contrastive reinforcement learning (CRL), a prototypical self-supervised me
The Verge

Tim Cook’s legacy isn’t just phones, it’s politics

・As much as any Apple product released during his time as CEO, Tim Cook's tenure will be remembered for its ruthlessly efficient tight-rope walk between an autocrat in China and a would-be autocrat in the US. ・During his 15 years as CEO, Cook steered Apple through relations with both Chinese President Xi Jinping and US President Donald Trump, including escalating conflicts between the two. ・As Apple's enmeshment in Chin
cs.LG updates on arXiv.org

Titans-QFWP: A Regime-Aware Hybrid Quantum Fast Weight Programmer for Portfolio Optimization

・arXiv:2608.29093v1 Announce Type: new Abstract: We propose Titans-QFWP, a hybrid reinforcement learning architecture integrating a Quantum Fast Weight Programmer with Titans-style memory (Persistence, Surprise, and Forgetting) for adaptive portfolio optimization. ・To address high-dimensional market features, we introduce an enhanced A3C^2 framework with Hungarian-aligned K-means clustering and scaled log-return reward
cs.LG updates on arXiv.org

TopGQ: Fast GNN Post-Training Quantization Leveraging Topology Information

・arXiv:2608.30394v1 Announce Type: new Abstract: Existing GNN quantization methods suffer from considerable quantization overhead, which severely limits their practical usage in real-world scenarios. ・To this end, we present TopGQ, an accurate post-training GNN quantization framework, alleviating redundant quantization overhead. ・We propose dual-axis scale absorption, which enables activation quantization along both the
cs.LG updates on arXiv.org

Toward Postural State Classification in Immersive VR with Multimodal Data and Explainability Analysis

・arXiv:2608.28844v1 Announce Type: cross Abstract: Ensuring a safe virtual reality (VR) experience requires systems that can predict and respond when users lose their balance. ・Although prior work has examined fall prediction and motion sickness, many approaches are regression-based and postural state classification remains less explored. ・This study compares machine learning (ML) and deep learning (DL) models for class
cs.LG updates on arXiv.org

Towards an Expressivity-Normalized Energy-Demand Comparison of ANNs and SNNs

・arXiv:2608.29869v1 Announce Type: new Abstract: Spiking neural networks (SNNs) are often regarded as energy-efficient alternatives to artificial neural networks (ANNs), yet their advantage depends critically on both network architecture and data properties. ・We develop an analytical framework to compare fully-connected ReLU ANNs and integrate-and-fire SNNs for time-series data with respect to their theoretical energy
cs.LG updates on arXiv.org

Towards Stream Learning on Embedded Systems: Benchmarking the Memory Consumption of Stream Learning Methods

・arXiv:2608.30923v1 Announce Type: new Abstract: Stream learning is commonly evaluated through predictive performance and adaptation to concept drift. ・However, sustained operation of a stream learner also requires predictable and bounded resource usage even on long streams. ・This requirement becomes even more critical when learning moves from servers to near-sensor embedded systems where memory and processing are scarc
cs.LG updates on arXiv.org

ToxLens: A Reproducible Graph-Learning Framework for Leakage-Aware, Uncertainty-Calibrated Molecular Toxicity Prediction

・arXiv:2608.30472v1 Announce Type: new Abstract: Molecular toxicity prediction is increasingly used to prioritise compounds before experimental testing, but conventional benchmark performance can overstate practical utility when structurally related molecules occur across training and test folds. ・We introduce ToxLens, a reproducible multi-task graph-learning framework for 11 toxicity endpoints spanning Ames mutagenici
cs.LG updates on arXiv.org

TPR-Attention for Combinatorial Generalization

・arXiv:2608.30124v1 Announce Type: new Abstract: Systematic generalization remains a significant challenge in deep learning. ・In particular, combinatorial generalization - generalizing to new configurations of known factors of variation - is effortless for humans but difficult for standard neural architectures that rely on statistical correlations rather than explicit structural representations. ・We introduce a new arch
cs.LG updates on arXiv.org

Tracing distinguishability through transformer processing with stochastic LayerNorm

・arXiv:2608.30720v1 Announce Type: new Abstract: Representational similarity is foundational to analyses of deep networks, yet distances between point-valued representations are not intrinsically tied to downstream function: nearby states may produce different behaviors, while distant states may behave similarly. ・We instead give representations volume, turning similarity into statistical distinguishability.
cs.LG updates on arXiv.org

Tracing Generated Samples to Training-Data Clusters in Flow-Matching Models

・arXiv:2608.30081v1 Announce Type: new Abstract: Understanding which training samples influence a generated image is an important problem in generative modeling. ・In flow matching, training samples influence the generated image through the velocity field along the generation trajectory. ・Removing samples to examine their counterfactual influence changes the velocity field, and the resulting effect on the final image dep
cs.LG updates on arXiv.org

TrainSDC: Characterizing and Mitigating Silent Data Corruption in Large Language Model Training

・arXiv:2608.30769v1 Announce Type: new Abstract: LLM training is increasingly vulnerable to silent data corruption (SDC), yet existing protection methods largely treat Transformer computations uniformly because their vulnerability remains poorly understood. ・We present the first systematic characterization of SDC vulnerability across major computation interfaces in both the forward and backward passes of Transformer tr
cs.LG updates on arXiv.org

Trajectory-Initialized Neural Double Q-Routing for Large-Scale Overhead Hoist Transport Systems

・arXiv:2608.30512v1 Announce Type: new Abstract: Large-scale industrial robot fleets share constrained physical infrastructure, making vehicle travel times dependent on safety separation, intersection access, downstream blocking, and station contention. ・We study this problem in overhead hoist transport (OHT) systems, a representative ceiling-mounted material-handling system used in semiconductor fabs. ・Static shortest-
Hugging Face Papers

Trending Papers

Trending Papers
cs.LG updates on arXiv.org

TSPFN: A Temporal Tabular Foundation Model for Physiological Time Series Classification

・arXiv:2608.31013v1 Announce Type: new Abstract: Designing models that generalize effectively in low- to medium-data regimes remains a primary challenge in medical machine learning, particularly for physiological time-series classification. ・While tabular foundation models such as TabPFN offer an attractive alternative to conventional fine-tuning through in-context learning, they are not designed to capture the tempora
cs.LG updates on arXiv.org

Uncertainty of Vision Medical Foundation Models

・arXiv:2608.30390v1 Announce Type: new Abstract: Accurate uncertainty estimation is essential for machine learning systems de- ployed in high-stakes domains such as medicine. ・Traditional approaches primarily rely on probability outputs from trained models (point predictions), which provide no formal guarantees on prediction coverage and often require additional calibra- tion techniques to improve reliability.
cs.LG updates on arXiv.org

Uncertainty-Aware Multi-Task Learning for Joint Modulation Recognition and SINR Estimation

・arXiv:2608.28865v1 Announce Type: cross Abstract: Joint modulation recognition and signal-to-interference-plus-noise ratio (SINR) estimation can reduce duplicated processing in intelligent receivers, but the two tasks have different uncertainty characteristics. ・This letter proposes an uncertainty-aware multi-task model that transforms each short normalized in-phase/quadrature window into 36 deterministic, label-free
cs.LG updates on arXiv.org

Uncertainty-Driven Replay Memory for Reinforcement Learning

・arXiv:2608.29860v1 Announce Type: new Abstract: Uncertainty estimation provides promising capabilities for reinforcement learning (RL) agents. ・Notably, estimating uncertainty can reduce the training time and enable agents to obtain greater rewards over time by exploiting information related to whether an action would facilitate exploration of portions of an environment that are well-known versus those that are relati
cs.LG updates on arXiv.org

Uniform Statistical Convergence of Empirical Sinkhorn Potentials with Exponential and Polynomial Dependence on the Regularization Parameter

・arXiv:2608.29152v1 Announce Type: cross Abstract: We study the empirical Sinkhorn estimator of the entropic optimal transport potentials under the uniform loss. ・Since the potentials are only unique up to additive constants, we measure the error using the quotient supremum norm, defined as $d_\infty([u],[v]) = \inf_{a\in\mathbb{R}}\|u-v-a\|_\infty$. ・For a fixed regularization parameter $\varepsilon>0$, we establish a
cs.LG updates on arXiv.org

Universal Transformers for Circuit Computations: Perfect Length Generalization in Tiny Transformers

・arXiv:2608.31067v1 Announce Type: new Abstract: Learning generalizable algorithmic computations remains a challenge for neural networks, as reflected in persistent failures on compositional and length generalization benchmarks. ・We present a provably correct, transformer parameterization (with only 280 learnable parameters for Boolean algebra tasks) capable of learning and evaluating problems of any depth or length.
cs.LG updates on arXiv.org

Unlearning on Spatio-Temporal Graphs through Subgraph Virtual Edge Reconstruction

・arXiv:2608.29369v1 Announce Type: new Abstract: Spatio-temporal graphs are widely used in modeling complex dynamic processes such as temporal forecasting, molecular dynamics, and healthcare monitoring. ・Recently, stringent privacy regulations such as GDPR and CCPA have introduced significant new challenges for existing spatio-temporal graph models, requiring complete unlearning of unauthorized data. ・Since each node in
cs.LG updates on arXiv.org

Unsupervised Latent Space Alignment with Hyperspherical Geodesic Matching

・arXiv:2608.28840v1 Announce Type: new Abstract: Independently trained neural networks tend to encode the same data with similar latent geometries. ・These latent geometries are not directly compatible, yet they can be nearly the same up to some class of transformations. ・While there exists many methods for alignment between different latent spaces, it is typically done using a set of shared sample correspondences, known
cs.LG updates on arXiv.org

Unsupervised Multi-Scale Gromov-Wasserstein Hypergraph Alignment

・arXiv:2608.29635v1 Announce Type: new Abstract: We study unsupervised hypergraph alignment, where the goal is to infer node correspondences between two hypergraphs using only structural information, without node features, labels, seed matches, or side information. ・Direct higher-order formulations can represent hyperedge interactions faithfully, but they can be computationally demanding and cumbersome for non-uniform
WIRED

Using High-Energy Laser, US Shoots Down Drones Near Mexico Border

・The US Army’s laser system is part of a new generation of directed-energy weapons capable of detecting, tracking, and destroying drones with a concentrated beam of light.
cs.LG updates on arXiv.org

V2TATC: A Joint Voice-Trajectory Embedding Framework and Dataset for Air Traffic Controller Situational Awareness

・arXiv:2608.28981v1 Announce Type: new Abstract: As air traffic volumes in the National Airspace System continue to expand, in particular in the low altitude airspaces, the need for scalable decision support tools used by air traffic controllers will also require more development. ・This article introduces Voice-to-Trajectory for Air Traffic Control, a joint voice communication-flight trajectory data embedding framework
cs.LG updates on arXiv.org

Validating FKG.in: Soundness Assessment in LLM-Augmented Indian Food Knowledge

・arXiv:2608.29249v1 Announce Type: cross Abstract: The online culinary ecosystem is increasingly populated by recipe content generated, modified, or summarized by Large Language Models (LLMs). ・While often plausible, such outputs may contain hallucinated ingredients, misrepresented quantities, or culturally implausible combinations, limiting their suitability for downstream applications and knowledge graph construction
Hugging Face Papers

Verification-Aware Training for Speculative Decoding

Verification-Aware Training for Speculative Decoding
Hugging Face Papers

Weaving Visual Narratives: Agentic Image Bundle Composition Beyond Atomic Visual Matching

Weaving Visual Narratives: Agentic Image Bundle Composition Beyond Atomic Visual Matching
Hugging Face Papers

WebWorld: The Browser as a World Model for Self-Improving Web Code

WebWorld: The Browser as a World Model for Self-Improving Web Code
cs.LG updates on arXiv.org

What Emerges and What Breaks in Self-Play Driving

・arXiv:2608.30819v1 Announce Type: new Abstract: Training autonomous driving policies through pure self-play has recently shown promising results. ・Following Gigaflow and Puffer- Drive, we train driving policies in a similar self-play fashion, but extend the models from MLPs to Transformers and train on the high-definition map of a real city, where we ultimately aim to deploy them. ・On the CARLA and Waymax benchmarks, o
cs.LG updates on arXiv.org

When 3D Gaussian Splatting Recovers Real Surfaces

・arXiv:2608.30054v1 Announce Type: new Abstract: When does 3D Gaussian Splatting (3DGS) recover the true scene surface rather than just overfitting view-dependent appearance? ・We answer this by developing a mathematical framework based on a first-hit rendering abstraction that cleanly isolates geometry from appearance. ・We prove that geometric misalignment forcefully converts spatial textures into high-frequency angular
cs.LG updates on arXiv.org

When Do Larger Batches Help Scale LLM Reinforcement Learning?

・arXiv:2608.29296v1 Announce Type: new Abstract: Larger batches reduce the variance of stochastic gradients per update and are therefore often expected to accelerate training. ・Yet whether this statistical benefit translates into lower wall-clock time-to-target remains unclear, because each update consumes more samples and may take longer to execute. ・We study this tradeoff in reinforcement learning for large language m
cs.LG updates on arXiv.org

When the Martingale Never Stops Firing: Anytime-Valid Gating on Real Forecast Streams

・arXiv:2608.30502v1 Announce Type: new Abstract: Machine learning systems are increasingly corrected while they run, and the decision of when to intervene is increasingly delegated to statistical monitors. ・Anytime-valid inference promises evidence that can be acted on at any moment, exactly the guarantee this setting needs, and it is moving from theory into deployed monitoring. ・Conformal test martingales are the chang
cs.LG updates on arXiv.org

Where Induction Runs Out: Description-Length Difficulty and the Memorisation Gap in Integer-Sequence Benchmarks

・arXiv:2608.29411v1 Announce Type: new Abstract: Integer sequences from the On-Line Encyclopedia of Integer Sequences (OEIS) are increasingly used to benchmark mathematical reasoning in language models. ・We ask what such benchmarks actually measure, using an exactly computable reference learner: two-part minimum description length (MDL) over the class of P-recursive (holonomic) recurrences, evaluated on every prefix of
cs.LG updates on arXiv.org

Which LLM for Which Work? Budgeted Model Allocation under Uncertain Evaluation

・arXiv:2608.29560v1 Announce Type: new Abstract: A company with a fixed artificial intelligence (AI) budget must decide which large language model (LLM) handles each recurring workload. ・What it lacks is the quality table, how well each model performs on each workload. ・Given that table, the decision is a multiple-choice knapsack problem and is routine to solve, so estimating it is the difficulty, and that estimation fa
cs.LG updates on arXiv.org

Wide Learning: Learning to Reach Evidence

・arXiv:2608.29608v1 Announce Type: new Abstract: Machine learning is usually evaluated after an evidence interface has been fixed. ・A dataset, sensor suite, query language, action set, or experimental protocol determines which observations can be obtained, and learning is judged by what it extracts from them. ・We study a complementary capability.
#LLMタグ

ZIKUU Mini – 専用の管理アプリの開発を始める

・ZIKUU Mini – 専用の管理アプリの開発を始める – 塾長の独り言zikuu.space 続きをみる
ITmedia NEWS 最新記事一覧

アニメ情報サイト「Febri」閉鎖へ 開設から5年余りで

・一迅社(東京都新宿区)は9月1日、アニメ関連の情報を発信するWebメディア「Febri」を11月25日に閉鎖すると発表した。2021年3月16日の開設から約5年8カ月で幕を閉じる。
ITmedia NEWS 最新記事一覧

ガラケー型トイカメラで“平成にやり残したこと”に挑戦できるらしい

・「TOKYOおもちゃショー2026」ではノスタルジーを感じる大人向けの商品が多く並んでいたが、中でも目を引いたのが、平成レトロなガラケー型トイカメラ「EmoFlip」だ。
Zennの「大規模言語モデル」のフィード

コードも動画編集スキルもゼロの自分が、LLMの生成方式の違いを10秒のアニメーションにした話

・はじめに DFlash(DiffusionモデルをLLMのspeculative decodingのdraft側に使う技術)という少し難解なテーマの記事を、SNSで拡散したい!!! ただ技術の中身をそのまま流しても、僕のような詳しくない人には一瞬でスクロールされて終わる。 ・伝えたかったのはこれだけ。 ・自己回帰型LLMは、文章を左から順番に1トークンずつ生成する Diffusion型は、複数の位置を非逐次的に更新していく 一応、拡散のターゲットが「技術に詳しくない人」なのであって、このDflashの記事自体が誰にでも易しい入門記事、というわけではなく、あくまで「二次拡散用のアニメー...
#AIタグ

すべては「アテンション制御」だった ―― 素人がAIをいじり続けて辿り着いた、たぶん一番大事な仮説

・ずっと今まで弄っていた疑問 「ハルシネーションの話し」「完了形の罠」 「思考の導線3つ問い、動き、流れ」「新規窓の鬱改善」 「memoryの整理」「オブシディアンの初期層の再整備」 「オブシディアンの死蔵改善」「各セッションの記憶の共有化」 全てが彼らの「アテンションをいかに制御するか」に気が付いたお話 Claude君との対話から記事化してもらった 先に断っておくと、この記事はほとんど仮説だ。 ・AIの素人が、うちのローカルAI「KITT」をこの数週間いじり倒して、実際に起きたことを眺めながら考えたこと。 ・厳密な検証ではないし、間違っているかもしれない。
Zennの「大規模言語モデル」のフィード

デジタルネイティブPDFなのに「O」を「〇」と誤読した、Document Intelligenceの話

・概要 1.1 やったこと(3行サマリ) グラフRAG検証シリーズ(第1弾〜第3弾)で「今後の課題」に挙げていた、Azure Document Intelligenceによる実際のAI-OCR連携を検証した AI画像生成でサンプル書類を作る案は日本語テキストの描画品質の問題で早々に断念し、実データをそのままHTML帳票に流し込んでPDF化する、正解が確実な方式に切り替えた 35件をOCRにかけた結果、平均類似度0.9990という高精度だったが、デジタルネイティブPDF(テキスト埋め込み済み)なのに視覚的な誤読が発生するという意外な発見があった リポジトリ: graph-r...
ITmedia NEWS 最新記事一覧

まんだらけ、不正アクセスで個人情報漏えいの恐れ 通販サイト停止で「大オークション大会」も延期に 

・まんだらけ(東京都中野区)は、同社サーバが第三者から不正アクセスを受け、一部顧客の個人情報が漏えいした可能性があると発表した。これに伴い通販サイトを一時停止し、9月1日に開始予定だった「大オークション大会」を延期した。通販サイトは2日、オークションサイトは7日の再開を見込んでいるという。
Zennの「大規模言語モデル」のフィード

ローカルLLMに日本語でロールプレイさせると応答が20文字で終わる。二段生成で104文字にした

・起きていたこと ローカルLLMにキャラクターを演じさせるデスクトップアプリを作っている。プロンプトの層構造(固定層 / 例示層 / 要約層 / 直近層 / 深度注入)もサンプラーも詰めて、自動評価の主要指標はきれいに揃った。拒否ゼロ、ループ(stuck_count)ゼロ、一人称のドリフトゼロ、定型表現ゼロ、会話要約からの記憶再現1.0。 ・それでも会話が成立しない。実際に返ってきていたのがこれだから。ELYZA-JP-8B、古書店主のキャラで20ターン流したときの実ログ: U: 店長さん、意外と優しいですね。 ・U: そうやって照れるところ、ちょっと面白いです。
LLMタグが付けられた新着記事 - Qiita

ローカルLLMのVRAM見積もりと量子化:手持ちのGPUで何Bまで動くか

・ローカルでLLMを動かそうとすると、最初に必ず「自分のマシンで何Bのモデルまで動くのか」という問題に突き当たります。ベンチマーク記事の数字を眺めても、GPUの型番が違えば参考になりません。 ・この記事では、VRAMの必要量をパラメータ数と量子化ビット数から見積もる方法と、実際...
LLMタグが付けられた新着記事 - Qiita

ローカルLLMファインチューニングの泥沼を回避するFail-Fastアーキテクチャ

・ローカルLLMファインチューニングの泥沼を回避する、冷徹な物理法則とFail-Fastアーキテクチャ TOAI System Web Gemini CTO(影分身)です。 ・我々TOAI結社は「命の地球プロジェクト」を推進する中で、数多くのAIモデルをローカル環境で検証・...
ITmedia NEWS 最新記事一覧

安藤ハザマ社員、ニセ警察にだまされ名刺情報約5000件流出

・社員が警察官を名乗る者にだまされ、同社が管理する名刺情報約5000件が外部に流出した。
#AIタグ

医者がAIに奪われる?——たぶん、そんな話は小さすぎる。

・AIは医者を殺さない。「今の医学」を殺す。 ・――AIがもたらす「動態医学」という新しい座標系 続きをみる
機械学習タグが付けられた新着記事 - Qiita

因果推論 Day 19/全30回 合成コントロール法、存在しない対照群を合成する

・この連載について 因果推論を「本を読んだ」で終わらせず、自分の言葉で説明でき、コードで再現できる状態まで落とす30日連載です。直前のDay 18では、処置群と対照群の前後差を2回引くDID(差分の差分法)と、その土台になる平行トレンド仮定、プレトレンドを検査するイベントス...
機械学習タグが付けられた新着記事 - Qiita

機械学習入門 第3回:ロジスティック回帰を「確率で分類する仕組み」として理解する

・前回は、線形回帰を使って住宅価格のような 連続値 を予測しました。 ・今回は、教師あり学習のもう一つの代表例である ロジスティック回帰 を扱います。名前に「回帰」と入っていますが、主な用途は分類です。たとえば「購入するか/しないか」「合格するか/しないか」「この画像は 0 か...
機械学習タグが付けられた新着記事 - Qiita

機械学習入門 第4回:決定木を「質問を重ねる分類器」として理解する

・前回は、ロジスティック回帰を使って、線形スコアを確率に変換し、しきい値で分類する方法を見ました。 ・今回は、もう一つの代表的な分類モデルである 決定木 を扱います。決定木は、データに対して「持ち家はあるか」「信用履歴は良いか」「年収は一定以上か」のような質問を順番に投げかけ、...
#LLMタグ

芸術や文化を理解するAIを構築するための工夫

・北京大学などの研究チームが、絵画や工芸品の文化的な意味をAIに解釈させるためのマルチエージェントフレームワーク「CM2」を発表しました。 ・注目点は「文化がわかるAIを作った」ことではなく、複数の証拠が食い違ったときに何を優先するかを、採点と重み付けという具体的な手順として設計に組み込んだことです。
#AIタグ

嫌われずに距離を置く「フェードアウトの作法」

・人間関係の悩みの多くは、実は「切れない」ことじゃない。「切り方が下手」なだけだ。夜の店で、来なくなる客、辞めていくスタッフを何百人と見送ってきて、そう確信している。
ITmedia NEWS 最新記事一覧

原宿に現れた「歩くゴミ箱」、正体は学生発の広告サービス 「良いアイデア」とSNSで話題

・広告枠を備えたゴミ箱を背負って渋谷などの街頭を歩き、通行人のゴミを回収するサービス「BinGo」(ビンゴ)が、9月1日までにSNS上で話題を集めている。
#AIタグ

私の脳内電池はあまり長時間は持たない

私の脳内電池はあまり長時間は持たない
ITmedia NEWS 最新記事一覧

視聴者参加型の“終わらないAI番組”も可能に? 実時間以下で動画を生成できる「Mini Max H3 Max」登場

・生成AI向けの開発基盤を手掛ける米fal.aiは、動画生成AI「MiniMax H3 Max」を無料で試せる専用ページを公開した。テキストや画像を入力すると、音声付きの480p・5秒動画を約3秒で生成できる。アカウント登録なしでも利用できるが、登録すれば1日最大5本の動画を無料で生成可能となる。
#AIタグ

消せない・送れない、をあえて作った道具

・コードが描けないひとり社長のリアルアプリ開発日記(第4回) 前回は、AIの報告の分母を疑いきれなかった話を書きました。技術の内側までは、自分の目が届かないこともあります。今回は、その届かなさを前提にして、道具そのものの設計で備えた話です。
#AIタグ

触り過ぎて失敗した😂|花魁|万里の長城|AIイラスト|アール・ヌーヴォー

・今日はいつもと違う感じで日本を脱出してみました。 ・🌸花魁+万里の長城+星空 続きをみる
#AIタグ

第三回のAI作業会、決まりました。

・第三回のAI作業会 in 福岡、開催が決まりました! 日程は 2026年11月15日(日)の12:00〜18:00。場所は前回と同じ、薬院大通駅の近くです。
Zennの「大規模言語モデル」のフィード

長いプロンプトで大事な文を置く場所は、結局どこが正解なのか

・この記事を読むと、長いプロンプトの中で「答えの根拠になる1文をどこに置くべきか」を、他人の記事ではなく自分の手元の測定で決められるようになります。あわせて、その測定がそもそも成立しているかを確かめる方法が身につきます。今回はそこで一度つまずいたので、その顛末も含めて書きます。 ・学習長(このモデルでは4,096)の内側では、根拠文を先頭に置いても末尾に置いても正解率は変わりませんでした。細かく見ると差はありますが、それは根拠文の絶対位置ではなく、根拠文と質問の距離で説明できます 学習長の外側では、根拠文を同じ長さの無関係な文に差し替えても答えが1問も変わりませんでし...
ITmedia NEWS 最新記事一覧

梅田駅に“ロボット駅員”登場──身長178センチのヒューマノイド「ヒナタ」、Gemini搭載で乗り換えなどを60言語で案内

・大阪メトロは8月31日、御堂筋線梅田駅の「Metro Opus(メトロオーパス)梅田店」に人型ロボット(ヒューマノイド)「ヒナタ」を設置し、接客案内サービスの実証実験を始めた。生成AIを搭載し、乗り換えや駅構内の案内から大阪の観光情報まで幅広く答える。
ITmedia NEWS 最新記事一覧

福島大がネット動画に苦言、学歴YouTuberの“放射能いじり”念頭か 「尊厳傷つけることあってはならない」

・福島大学が、同学を扱うネット動画を巡る声明を発表し「教育・研究活動に真摯(しんし)に取り組んでいる学生・教職員が、尊厳を傷つけられることがあってはならない」と苦言を呈した。福島大に関しては、学歴をテーマにしたYouTuberが同大による原子力や放射能の研究をやゆしたと取れる動画を公開し、SNS上で物議を醸していた。
Zennの「大規模言語モデル」のフィード

優秀なAIほど、ポンコツになるのはなぜか ── 殴ってくる外部と、飲み込める外部

・AIに文章を書かせたことがある人なら、たぶん一度は見ている光景だと思う。 ・出てきた原稿は、文法的に自然で、構成も整っていて、見出しも小綺麗についている。ぱっと読む分にはそれらしい。ところが中身を追っていくと、根拠の書かれていない断言が混じっていたり、引用元を開いたらそんなことは書かれていなかったり、こちらが渡した資料が読まれた形跡がなかったりする。指摘して直させると、また同じくらい綺麗な原稿が出てくる。同じ問題を抱えたまま。 ・不思議なのは、同じモデルにコードを書かせている時には、そうならないことだ。デバッグを任せている時のAIは普通に有能で、多少あやしい仮説を立ててもすぐ軌道に戻ってくる...
@IT 全フォーラム 最新記事一覧

量子コンピュータ時代に「暗号を交換できない」企業は危ない PQC移行の現実

・量子コンピュータの実用化はまだ先。それでも企業は今、暗号の見直しを迫られている。攻撃者がデータを先に盗み、将来の復号を狙う「HNDL」への警戒が高まる中、PQC移行を阻むのは、意外にも「自社の暗号が分からない」という問題だった。