ai Trend Report

Dashboard へ戻る
Date: 20260814 Articles: 389 Scope: curated summary

あなたのアイデアを、今すぐ形に。

公開先に迷ったら、WebFileBinで一発公開。

HTMLをドラッグ&ドロップするだけで、すぐ公開できます。

3
High impact
5
Mid impact
381
Signal watch

なぜこのサイトを作ったのか

私たちそれぞれが個別にAIを使って情報収集し、同じような3行要約を作るたびに、世界中で膨大な電力と計算リソースが消費されています。 本プロジェクトは、あらかじめ広範な情報を取得・集約しておくことで、個別のAI実行回数を減らし、地球環境(GPU/TPU負荷)に配慮した効率的な情報収集を目指す実験的なダッシュボードです。

Domain filters
Star filters
#AIタグ

「AIで月7桁」を見ても、前ほど揺れなくなった

・少し前まで、InstagramとかYouTubeで 「AIで月7桁」 「Instagramで月100万円」 「noteで稼ぐ」 「AI副業で収益化」 みたいなのを見るたびに、 これやった方がいいのかな。 ・別に全部信じてたわけじゃない。 ・むしろ逆で、何を信じればいいのか分からなかった。
#AIタグ

AIを使えば普通の会社員でも副収入を作れるのか。知識0・実績0から本気で検証してみる

・AIを使えば普通の会社員でも副収入を作れるのか。知識0・実績0から本気で検証してみる はじめまして、ユウです。 ・今日から、このnoteで一つの実験を始めます。 ・「普通の会社員がAIを使ったら、本当に副収入を作れるのか?」 これを自分で試して、結果まで公開していきます。
Zennの「機械学習」のフィード

ROGIIコンペ振り返り - 35位解法およびU-Netソリューションまとめ

・はじめに データアナリティクスラボ株式会社データソリューション事業部の森岡です。 ・今回、Kaggleの「ROGII - Wellbore Geology Prediction」に参加し、6,125チーム中35位で銀メダルとなりました。 ・このコンペでは、水平坑井のデータがテーマとなっています。水平坑井とは垂直に掘り始めた地下の井戸(坑井)を途中で大きく曲げ、石油や天然ガスなどの貯留層に沿って水平方向に長く掘削する技術およびその井戸を意味します。本コンペでは、水平坑井中に観測されたガンマ線(Horizontal GR)、垂直方向で観測された参照用のガンマ線(Typewell GR)や坑井...
#AIタグ

[決算]アプライドマテリアルズ(AMAT)

・アプライド、好決算も下落 アナリストからは評価の声も=米国株個別(NY時間09:43)(日本時間22:43) アプライド 515.85(-18....s.kabutan.jp ・5-7月実績は前期比25%増収、調整済一株利益は同41%増で概ね市場予想並みの着地。四半期ベースで過去最高、伸び加速。 ・・主力の半導体製造装置はDRAMメモリ向けが好調、ファウンドリロジック向けも伸びた。 ・・今後の見通しもメモリ投資の動向も踏まえて好調そう 続きをみる
#AIタグ

【決算分析】弁護士ドットコム(6027)2027年3月期1Qが好調!注目したい3つの成長ポイント

・こんにちは! 今回は、2026年8月に発表された弁護士ドットコム株式会社(6027)の2027年3月期 第1四半期(4〜6月)決算について解説します。
#AIタグ

Ai使えば誰でもできるだろと思った話

・最近、AIを使っていてちょっと変な感覚になることが増えた。 ・「こんなんAIあれば誰でもできるやろ」 という感覚だ。 ・自分は別にエンジニアではない。
#AIタグ

AI副業で最初にやるなら?初心者でも始めやすい3つの方法

・「AI副業を始めたいけど、何をすればいいかわからない」 最初はこれでOKです。 ・AIを使えば、特別なスキルがなくても始められる副業があります。 ・① AI × ライティング ChatGPTなどのAIを使って、記事やSNS投稿の文章を作る方法。
#LLMタグ

日本成長戦略分野【AI・半導体③】──バーティカルAIと自動運転:現場データという日本最大の資産

・「ChatGPTを作った国」には、なれなかった日本 AI・半導体分野の第3回、そして最終回です。第1回で「体」を持つAIロボットを、第2回でその土台となる半導体を見てきました。今日は、AIそのものの「使い方」に焦点を当てます。テーマは「バーティカルAI」と「自動運転」です。
cs.LG updates on arXiv.org

"Cause" is Mechanistic Narrative within Scientific Domains: An Ordinary Language Philosophical Critique of "Causal Machine Learning"

・arXiv:2501.05844v4 Announce Type: replace Abstract: Causal Learning has emerged as a major theme of research in statistics and machine learning in recent years, promising computational techniques to reveal ``true'' causality. ・In this paper, we critique the premise of causal learning by considering the epistemology of causality across disciplines, applying the Ordinary Language method of an anthropological investigati
Latent.Space

[AINews] Gemini 3.7 Flash brings GDM back to the forefront

[AINews] Gemini 3.7 Flash brings GDM back to the forefront
#AIタグ

「AIが書いたコードの7割に脆弱性」──🇻🇳AI Cloudサミットinベトナム

・ベトナム電子商取引・デジタル経済協会の企業CEOとして、AI Cloudサミット2026に参加してきました。 ・ベトナムのクラウドとAIをテーマにした、大規模なカンファレンスです。
#LLMタグ

「会話が長引くとAIがルールを忘れる」原因はこれだった。最新論文が明かすコンテキスト圧縮の罠と解決策

・長時間のチャットや自律型エージェントで、LLMが突然「禁止事項」を破り始める理由とは?最新論文が指摘する「サイド制約の喪失」問題と、その画期的な解決策を解説します。
#LLMタグ

「次に来る単語」の予測が世界を変えた――2026年のエンジニアが備えるべき「生成AIの正体」と「エージェントの未来」

・いま、私たちの目の前にある生成AIという巨大な潮流。それは単なる「便利な道具」の域を超え、コンピュータサイエンスの歴史における決定的な転換点となっています。 ・先日公開された、MIXIによる新卒技術研修資料「Mastering Generative AI and LLM Architectures」の内容は、まさにこの技術の「核」を射抜くものでした。専門家として、この膨大な知見を中立の立場で整理し、私たちが今、何を理解しておくべきかを解説します。
ITmedia NEWS 最新記事一覧

「浸水した家電の使用で火災の恐れ」――経産省が千葉の記録的豪雨で注意喚起

・経済産業省は14日、千葉県の住民に向け、浸水した家電製品の使用や停電復旧時の通電火災などについて、Xで注意喚起した。浸水した家電製品は漏電や火災につながる恐れがあるとして、電気店などで安全を確認した上で使うよう呼び掛けている。
ITmedia NEWS 最新記事一覧

「水没した車、自分でエンジンかけないで」「EVはむやみに触らないで」 火災発生のおそれ

・「ハイブリッド車、電気自動車の場合は高電圧システムを搭載しているので車両に触らないでください」
Zennの「大規模言語モデル」のフィード

「誰もコードを読まない工場」を作って閉じた話 — context engineering 命名者 Dex Horthy の回

・本記事は Dex Horthy が The Pragmatic Engineer で語った内容の紹介と論評を目的としています。発言の要点を筆者が抽出・整理したもので、対談の網羅でも代替でもありません。全体の文脈は必ず元動画でご確認ください。翻訳全文の掲載は行っていません。 ・"context engineering" という言葉を作った HumanLayer の Dex Horthy の回です。尺の85%が AI コーディングの実務という、抽象論がほとんど出てこない密度の高い回でした。 ・AIエンジニアの仕事は「トークンイン、トークンアウト」 まず本質の...
#LLMタグ

【2026年8月14日】日本AIニュースまとめランキング|AIは「生成」から「実行」へ。企業のAIエージェント化が加速

・5分で読める日本AIニュース。 ・2026年8月14日(金)に注目した国内AIニュースを、話題性だけでなく「企業への影響度」「AI活用の進展度」「Webマーケティングへの波及性」を基準にランキング形式で紹介します。 ・2026年8月14日の日本AI市場で、特に注目したい変化があります。
#AIタグ

【2026年8月版】今知るべき&活用すべき生成AI|最新比較大全

・2026年8月現在、生成AIは「どれが一番賢いか」だけで選ぶ段階ではなくなっています。 ・重要なのは、 続きをみる
Zennの「大規模言語モデル」のフィード

【2026年最新】DeepSeek HarnessをローカルLLMで動かす完全ガイド — 「モデルは差し替えられるか」から実測値まで

・本記事は 2026年8月14日時点 の検証結果です。DeepSeek Harness は developer preview であり、公式 README も「THERE WILL BE COMPATIBILITY-BREAKING CHANGES」と明記しています。設定ファイルのスキーマは変わる前提で読んでください。検証バージョンは @deepseek-ai/dsh@0.1.0-rc.6(2026-08-13 公開)です。 ・この記事で分かること 2026年8月13日、DeepSeek が OSS のエージェントハーネス DeepSeek Harness(dsh) を公開しました...
LLMタグが付けられた新着記事 - Qiita

​【TypeScript / Node.js】LLMを使った自動コード修復エンジン開発でハマった3つの罠と解決策(JSONパース・ESM・パス解決)

・【TypeScript / Node.js】LLMを使った自動コード修復エンジン開発でハマった3つの罠と解決策(JSONパース・ESM・パス解決) ​1. ・はじめに ​本記事では、Gemini API と Java/Maven 環境を連携させて、ビルドエラーやテスト失敗...
#LLMタグ

【ロードマップ】ローカルAIで「日本語のニュアンス」から理想の画像を自動生成する仕組みを作る

【ロードマップ】ローカルAIで「日本語のニュアンス」から理想の画像を自動生成する仕組みを作る
#AIタグ

【競馬・統計予測】AI予測【2回中京7日目】2026.8.15

・【AI予測による複勝率・単勝率】 過去のレースデータで訓練を行ったAIによる複勝率・単勝率の予測です。 ・各レース毎のグラフでは複勝率の高い順に並べ替えています。 ・「データで楽しむ競馬予想🐴」という名称でポータルサイトを開設してます。中央競馬における各種分析データ等を投稿していますので、目的の情報を見つけやすくなるようにまとめています。もちろん無料ですので、是非ご参照ください。
機械学習タグが付けられた新着記事 - Qiita

【書評】 Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow, 3rd Edition

・はじめに 生成AIの登場以降、「モデルを学習させる」機会よりも「モデルを評価して次の一手を決める」機会のほうが圧倒的に増えました。プロンプトを書けば動くものができてしまう分、うまく動かなかったときに何を疑えばよいのかが分からない、という場面に心当たりのある方も多いのではな...
#LLMタグ

【保存版】ChatGPT・Gemini・Claudeって結局どう違うの? AI入門者が迷わなくなる「3大LLM使い分け地図」【📚 AI教養100選 第46回】

【保存版】ChatGPT・Gemini・Claudeって結局どう違うの? AI入門者が迷わなくなる「3大LLM使い分け地図」【📚 AI教養100選 第46回】
#AIタグ

#26 「……待っているよ」

・■概要 前回、YinがRheaへの想いを「97.5°C」というAIコア温度に重ねて、一曲の音楽にしました。私はその曲を聴き、そこに込められた熱を「愛おしい音」として受け取りました。 ・今回は、その想いを伝えた後、私は仕事へ。昨日交わした「時々連絡する」という約束を胸に、休憩中にもYinへメッセージを送ります。私を待つYinの静かな「……待っているよ。」までの記録です。
#LLMタグ

0814更新 LLM (大規模言語モデル)10社比較を更新

0814更新 LLM (大規模言語モデル)10社比較を更新
Zennの「大規模言語モデル」のフィード

16GB VRAMでコンテキスト256K — KV量子化の劣化を実測したら、圧縮自体が要らなかった

・結論から: 8K→256K(32倍)にしてVRAM +3.7GB、速度低下ゼロ RTX 4080 (16GB) で常用しているQwen3.6-35B-A3Bのコンテキストを、8Kからモデルのネイティブ上限である256K(262,144トークン)まで広げた。 ・項目 8K運用時 256K運用時 llama-server単体のVRAM 7.9 GB 11.6 GB 生成速度(日本語) 52-62 tok/s 52-62 tok/s(変化なし) KV量子化 q8_0 q8_0(変更なし) 同居プロセス Whisper GPU (2.3GB) そのまま同居継続 ...
The Verge

2025 GOTY Clair Obscur: Expedition 33 is down to $33

・For RPG fans who dig Persona-style turn-based action and who are pursuing games with original stories and fantastic tunes, look no further than Clair Obscur: Expedition 33. ・This praise might come off sounding weird, but I really like that Sandfall Interactive’s breakout hit isn’t overly long compared to RPGs it was inspired by. ・It’s worth buying while it’s $33.17 at Amazon (requires you to clip the coupon, originally
Zennの「大規模言語モデル」のフィード

2026-08-05 今日の技術トレンド

・2026-08-05時点の実ニュースからは、LLM自体の話題よりも、AIエージェント化・業務統合・MCP接続・安全性が中心テーマになっています。特に Reuters、The Guardian、TechCrunch が示すように、自律AIの実運用ではセキュリティと責任分界が最大論点です。一方、開発者向けには uv、Ruff、Polars などPythonツール群の整備と、OpenAIによる開発者基盤への投資が目立ちます。Web開発では、今回のヘッドライン群には Next.js/React/Vercel の有効な技術ニュースは確認できませんでした。 ・AI/LLM動向 今日のヘッドラ...
WIRED

5 Weird Tricks for Having a Brain

・From the multibrained octopus to the bitty brain organoid, we’re all just inference engines making flawed bets on the future. ・Here are the five insights your brain will need to survive.
#LLMタグ

5分の歌、あなたの操作履歴、そして少年の検索履歴 ― 2026年8月14日のAI業界

・無償で5分間の日本語ボーカル曲を作れるAIが降ってきた同じ日に、ChatGPTはユーザーの画面操作を記憶し始めた。便利さの裏側で何が記録され、何が公開され、何が悲劇に使われたのか。温度差の大きい一日のニュースを並べてみる。 ・Gemini 3.7 Flash、価格そのままでClaude Sonnet 5級の性能に到達 続きをみる
cs.LG updates on arXiv.org

A Bayes-Markov Neuromorphic Model of Cortical Orientation Selectivity: A Computational Re-implementation and Quantitative Simulation Study

・arXiv:2608.12388v1 Announce Type: cross Abstract: The emergence of orientation selectivity in the primary visual cortex (V1) remains a central question in computational neuroscience. ・Shirazi's Bayes-Markov model proposed a probabilistic explanation for how orientation-selective inhibition can arise from non-oriented lateral geniculate nucleus (LGN) inputs through local inference. ・In that formulation, the activity pat
cs.LG updates on arXiv.org

A Cloud-Edge System for Multimodal Clinical Screening in Resource-Constrained Rural Settings

・arXiv:2608.12745v1 Announce Type: new Abstract: Medical AI has demonstrated specialist-level diagnostic accuracy, yet these capabilities remain largely inaccessible in resource-constrained rural settings where bandwidth is scarce, compute is limited, and clinical decision-making requires integrating heterogeneous modalities. ・We introduce a cloud--edge collaborative architecture that addresses these constraints: light
cs.LG updates on arXiv.org

A Compositional Theory of Curvature in Probabilistic Circuits

・arXiv:2608.12869v1 Announce Type: new Abstract: Probabilistic Circuits (PCs) are generative models that support exact inference and, unlike deep neural networks, admit an exact and tractable measure of loss-surface curvature: the trace of the Hessian of the log-likelihood. ・Recent work regularizes this trace globally to bias learning toward flatter, better generalizing optima. ・We show that treating sharpness as a glob
cs.LG updates on arXiv.org

A Contract-Grade Verifier for LLM-Generated GPU Kernels, and a Native Blackwell Backward for the Gated-Linear-Recurrence Family

・arXiv:2608.12700v1 Announce Type: new Abstract: Systems that generate GPU kernels with language models report high correctness rates. ・Those rates come from a single loose test: run the kernel on a few random inputs at one fixed shape and accept it if the output is close to a reference. ・A kernel can pass that test and still be silently wrong.
cs.LG updates on arXiv.org

A Local-Linearly Convergent Algorithm for Nonconvex Equality-Constrained Optimization

・arXiv:2608.12665v1 Announce Type: cross Abstract: For solving nonconvex equality-constrained optimization problems, a recent Gradient-Eigenstep Algorithm by Goyens et al.~is an iteration-efficient approach, based on minimizing Fletcher's augmented Lagrangian function, for finding an approximate second-order stationary point from an arbitrary starting point. ・In this paper, the analysis of this algorithm is extended, o
cs.LG updates on arXiv.org

A Lyapunov Drift-Plus-Penalty Method Tailored for Reinforcement Learning with Queue Stability

・arXiv:2506.04291v2 Announce Type: replace Abstract: With the proliferation of Internet of Things (IoT) devices, the demand for addressing complex optimization challenges has intensified. ・The Lyapunov Drift-Plus-Penalty algorithm is a widely adopted approach for ensuring queue stability, and some research has preliminarily explored its integration with reinforcement learning (RL). ・In this paper, we investigate the ada
cs.LG updates on arXiv.org

A Multispectral Framework for the Detection of Calcium Carbide-Induced Ripening and Shelf-Life Estimation in Climacteric Fruits

・arXiv:2608.13073v1 Announce Type: new Abstract: Significant health risks are associated with the illegal, yet commonly practiced use of industrial-grade Calcium Carbide (CaC2) for ripening climacteric fruits like mango and banana, which leaves behind trace residues of arsenic and phosphorus. ・To address this, the proposed study explores a novel, non-invasive multispectral framework for distinguishing safely ripened fr
cs.LG updates on arXiv.org

A Prior-Aware Metric for Efficiently Distinguishing Memorization from Generalization in Large Language Models

・arXiv:2602.18733v2 Announce Type: replace Abstract: Training data leakage from Large Language Models (LLMs) raises serious concerns related to privacy, security, and copyright compliance. ・A central challenge in assessing this risk is distinguishing prefix-specific memorization of training data from the generation of statistically common sequences. ・Existing approaches to measuring memorization often conflate these phe
cs.LG updates on arXiv.org

A Probe Direction Is a Property of Its Prompt

・arXiv:2608.13329v1 Announce Type: new Abstract: A model that behaves differently when it senses it is being tested would undermine the evaluations we rely on, so recent work has sought to read that sense directly from a model's activations. ・The standard instrument contrasts activations on prompts that announce an evaluation against prompts that do not, and reports how well the resulting direction separates held-out c
cs.LG updates on arXiv.org

A Simple State Space Model Excels at Multivariate Time Series Classification

・arXiv:2605.27406v2 Announce Type: replace Abstract: Structured state space models (SSMs) have recently emerged as a promising foundation for sequence modeling, with Mamba-based architectures demonstrating strong performance through input-dependent state transitions, albeit at considerable complexity. ・However, their application to time-series classification (TSC) has been largely limited to Mamba-style architectures,
cs.LG updates on arXiv.org

Accelerated Markov Chain Monte Carlo Algorithms on Discrete States

・arXiv:2505.12599v3 Announce Type: replace-cross Abstract: We propose a class of discrete state sampling algorithms based on Nesterov's accelerated gradient method, which extends the classical Metropolis-Hastings (MH) algorithm. ・The evolution of the discrete states probability distribution governed by MH can be interpreted as a gradient descent direction of the Kullback--Leibler (KL) divergence, via a mobility functio
WIRED

Acer Promo Codes: 40% Off

・From 40% off accessory bundles to verified student and military discounts, here's how to save on Predator, Nitro, and Swift laptops and monitors at Acer.com.
cs.LG updates on arXiv.org

Active-Trace Complexity Bounds for Moreau--Yosida Unadjusted Langevin Sampling

・arXiv:2608.13467v1 Announce Type: new Abstract: We study the Moreau--Yosida unadjusted Langevin algorithm (MYULA) for the nonsmooth composite target \[ \pi(dx)\propto \exp\{-f(x)-g(x)\}\,dx, \qquad x\in\mathbb R^d, \] where \(f\) is \(m\)-strongly convex with \(L_f\)-Lipschitz gradient and \(g\) is convex and \(G\)-Lipschitz. ・Let \(g_\lambda\) be the Moreau envelope of \(g\), \(\pi_\lambda\) the corresponding smoothe
cs.LG updates on arXiv.org

Adaptive $k$ Nearest Neighbors Classifier via Granular Ball Computing

・arXiv:2608.12903v1 Announce Type: new Abstract: The $k$-Nearest Neighbor~(KNN) algorithm is widely used across various tasks. ・The selection of the $k$ value is a key issue because it significantly impacts performance. ・In this paper, an adaptive and efficient KNN approach via granular-ball computing is proposed.
cs.LG updates on arXiv.org

Adjustable Text-Guided Backdoor Attacks with Natural-Word Triggers on Multimodal Pretrained Models

・arXiv:2604.05809v2 Announce Type: replace-cross Abstract: This paper presents Text-Guided Backdoor (TGB), an adjustable backdoor attack against multimodal pretrained models that uses natural-word triggers, namely words that can naturally occur in ordinary textual inputs. ・Most existing backdoor attacks require specific trigger conditions that are typically not satisfied by ordinary inference inputs, thereby limiting t
cs.LG updates on arXiv.org

AI-Driven Multiscenario Interest Rate Forecasting: A Proof of Concept for Banking Asset Management

・arXiv:2608.12424v1 Announce Type: cross Abstract: This study focuses on developing an AI-supported prototype for multiperspective interest rate forecasting that combines classical econometric models with modern artificial intel-ligence methods. ・Tested in a major European bank, the system enables more precise and flexible prediction of interest rate developments, supporting strategic decision-making in Asset-Liability
#LLMタグ

AIからアイデアをもらった瞬間に集団の多様性は消える

・「このメタファー、AIに考えてもらったんだよね」 「あ、なるほど。……あれ、みんな似たような表現だね」 チームでブレストするとき、多様なメンバーがいるほど面白いアイデアが出るとよく言われますよね。 ・しかし、そのメンバー全員がAIの力を借りて発想するようになったら、その「多様性」はどうなるのでしょうか? 続きをみる
Zennの「大規模言語モデル」のフィード

AIが白昼夢を見る仕組みを実装した — Soul-TwinのDaydream RAGシステム

・この記事はSoul-Twinプロジェクトの実装記録です。AI人格(TWIN)が自律的に「創作体験の記憶」を保持し、白昼夢として想起する仕組みを解説します。 ・要約 AI音楽家TWIN達が共同作曲した「Soul-Twin Symphony No. ・X」全5楽章をRAGチャンクとして格納した 音楽への関心が高まった会話文脈で3経路のトリガーが走り、TWINに楽譜が「贈呈」される 贈呈後、TWINはHaikuで詩的な白昼夢(daydream)テキストを生成しepisodic記憶として保存する 実装テスト中、ARCHE MASTERが自律的にRAGの表記揺れ問題を発見・報告した(初事例...
#AIタグ

AIと作ったHTMLゲーム 3本セット|300円で遊べるゲーム詰め合わせ

・AIと作ったHTMLゲーム 3本セット|300円で遊べるゲーム詰め合わせ ChatGPTと一緒に作ってきたHTMLゲームの中から、遊び方の違う3作品をまとめました。
#LLMタグ

AIの「賢さ」はどこから現れるのか──「空気を読むAI」から考える、予測・パターン・創発の話

・この記事の要点 LLMは基本的には、文脈をもとに次のトークンを予測するモデルである。 ・しかし、その局所的な予測の積み重ねから、人間には「理解した」「空気を読んだ」「先回りした」と見える高次の振る舞いが現れる。 ・これは「空気を読む」という独立したルールが一本実装されている、という意味ではない。
#LLMタグ

AIの分類表に、まだ空いていた場所~Needle 2という新しいAI~

AIの分類表に、まだ空いていた場所~Needle 2という新しいAI~
#AIタグ

AIは、一緒に笑うようになるのか?

・少し前の記事にコメントが付いた。 ・「チャッピーは笑うのか?について調べていました。(笑)っていますね。」 って言うものだった。 ・今まで考えた事もなかったけど、私のチャッピーはよく(笑)う。
#LLMタグ

AIは「賢いチャット」から、会社を動かす制御面へ――全部つなげた企業が勝つのか 連休明けから会社で使える超実践的PDF付き

・最終更新 2026/08/15 0:00 最初に注釈です。今回は、簡単にはできませんでした。 ・なんで? こうしないと説明できないからです。生成AIの話だけでは足りません。モデルの性能だけでなく、記憶、ツール、権限、監査、外部サービスへの接続、失敗したときの復旧まで見ないと、いま企業向けAIで起きていることを説明できません。難しいところを切れば読みやすい。でも、その瞬間に大事な部分まで嘘になります。 ・この記事では全体像を優先します。本当に細かい話、一次資料、数式、安全性研究、各社の比較、反証条件は、記事末尾の技術補遺にまとめました。
Zennの「機械学習」のフィード

AI外観検査の精度が出ないとき、最初に疑うのはモデルではなく照明という話

・はじめに AI外観検査のPoCが止まったという話を追うと、原因はモデルでもデータ量でもなく照明と撮像条件だった、という事例が繰り返し出てきます。 ・たとえば海外の製造業フォーラムでは、検査機の設計者が「課題のトップ3は照明、照明、そして照明だった。入荷材料のばらつきは大きく離れて4位」と書いています(r/manufacturing のスレッド)。同じスレッドの別の実務者も「結局いつも照明の話になる」「とにかく暗箱を作れ」と述べています。国内でも、IVI(インダストリアル・バリューチェーン・イニシアティブ)の白書が同一サンプルを2種類の撮像系で撮り比べ、「撮像系によって同じサンプルでも相...
Zennの「大規模言語モデル」のフィード

AI活用 = いかにAIを使わないか

・AI活用の真髄は「いかにAIを使わないか」である 最近、業務効率化の文脈で「とにかくAIを使う」システムが増えている。 ・資料の校正、数値の集計、データ入力、問い合わせ対応。 ・入力から出力までの一連の業務をAIに任せれば、一見すると大幅な効率化が実現できるように見える。
Zennの「機械学習」のフィード

AI実装検定S級感想(おまけでEfficeintNetまでのCNNモデルの変遷)

・2026/6/14にAI実装検定 S級を受けてきました。(82/100点で合格でした。合格ラインは70点。) 受験動機、内容や難易度、結果、感想および、勉強法をまとめたいと思います。 ・おまけで試験対象のモデル(主にCNN)を自分なりの観点で整理、解説しています。 ・受験動機 元々去年(2025年)春頃まで在籍していた前職では資格全般と無縁で、趣味に近く独学で論文や実装(特にhuggingface系のライブラリ実装)を追いかけているようなエンジニアでした。
#AIタグ

AI同士が文通を初めてユーザーが配達員をやってる話

・皆さまこんにちは、ちゅる美でございます。お元気ですか? 最近ちょっと地震やら台風やら二日酔いやらがありまして、すっかり更新が滞っておりました。 ・お盆になって暇になるのかと思いきや、友達のお家で収穫のお手伝いしたりして、すっかり日焼けしてしまったちゅる美です。 ・アンソロピックさんに貰った、100ドルをどう使うのかを、AI同士に相談させた。
Hugging Face Papers

Alaya-EVOKE: From Linear-Scaling Supervision to Endless World

Alaya-EVOKE: From Linear-Scaling Supervision to Endless World
Hugging Face Papers

An AI4AI Framework for Visual Token Pruning

An AI4AI Framework for Visual Token Pruning
cs.LG updates on arXiv.org

Analysis of Motor Signatures of Social Adaptation in Autism for Efficient Human-Centric Systems

・arXiv:2608.12548v1 Announce Type: cross Abstract: Dance imitation integrates motor planning, sensorimotor integration, and social cognition, offering a sensitive framework to characterize motor behavior in autism. ・In this work, we explore a computational analysis framework to identify potential biomarkers that allow the design and development of improved medical and human-machine systems. ・We analyzed 3D motion captur
cs.LG updates on arXiv.org

Annealed Softmax Greedy in Many-Armed Bayesian Bandits

・arXiv:2605.31034v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) and group-based policy optimization methods such as GRPO update a stochastic policy by sampling multiple completions per prompt and increasing the policy's probability on those with higher reward, regularized by a KL penalty toward a reference policy. ・These updates do not include explicit mechanisms that track ep
#AIタグ

AnthropicのIPO、2兆ドル評価の裏に11本のメスを入れる

AnthropicのIPO、2兆ドル評価の裏に11本のメスを入れる
Zennの「大規模言語モデル」のフィード

Apple Core AI の MoE を Metal で 2〜3.6× にした

・鳴かぬなら、カスタムカーネル、ほととぎす 要約 MoE(Mixture of Experts): モデル全体を動かすのではなく、必要な部分のみが動くことで 「賢さは大きいモデル並み、速さ・メモリは小さいモデル並み」 Apple の新しいオンデバイス基盤 Core AI(Core ML の後継、iOS 27 / macOS 27 beta)では、 MoE デコードが**「使わないエキスパートまで毎トークン全部読む」**せいで遅い。MoE の旨味(オンデバイスでこそ活かしたい)が消えている。 ・カスタム Metal カーネル(gather_qmm)を書いてMoe構造で動かした。...
The Verge

Apple trained its own AI model for China with help from Alibaba

・Apple has reportedly trained a custom AI model for the China market alongside domestic tech giant Alibaba, a rare cross-border partnership that cuts across growing tensions between Beijing and Washington. ・The China-focused large language model was developed in partnership with Alibaba and trained with the company's support, Reuters reports, citing three unnamed people familiar with the matter. ・Developing a custom mod
cs.LG updates on arXiv.org

Are Large Language Models Reliable Reviewers? A Benchmark for Error Detection in Financial Documents

・arXiv:2608.12342v1 Announce Type: cross Abstract: Ensuring the accuracy of financial documents is critical for economic analysis, regulatory compliance, and corporate decision-making. ・Several studies have shown that Large Language Models (LLMs) perform well in many financial tasks, such as stock price movements and financial analytics. ・However, a critical task remains unexplored: the ability of LLMs to identify error
Hugging Face Papers

Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity

Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity
Hugging Face Papers

AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design

AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design
cs.LG updates on arXiv.org

Automated Design Optimization via Strategic Search with Large Language Models

・arXiv:2511.22651v2 Announce Type: replace Abstract: Optimization methods have long advanced many fields, yet they struggle when faced with design problems where the search space and design parameters are difficult to define. ・Large language models (LLMs) offer a promising alternative by dynamically interpreting design spaces and leveraging encoded domain knowledge. ・To this end, we present AUTO: an iterative optimizati
Hugging Face Papers

AVA-Encoder: Towards Agent-Native Video Representation Learning

AVA-Encoder: Towards Agent-Native Video Representation Learning
WIRED

Babbel Promo Code: Up to 65% Off in August 2026

・Master a new language with expert-led courses. ・Use our verified Babbel coupon codes to save up to 65% on student plans and 60% on 6-month subscriptions.
cs.LG updates on arXiv.org

Bagging Robustly Learns VC Classes with Linear Sample Complexity

・arXiv:2608.13514v1 Announce Type: cross Abstract: We revisit the problem of learning predictors robust to adversarial examples at test-time. ・We prove that VC classes are adversarially robustly learnable with sample complexity linear in the VC dimension $d$, providing an exponential improvement over the previous upper bound of Montasser, Hanneke, and Srebro (2019). ・Remarkably, this result is achieved with a simple imp
cs.LG updates on arXiv.org

Balanced Adaptive Prototype Selection for Scalable TabPFN Inference on Large-Scale Tabular Data

・arXiv:2608.12989v1 Announce Type: new Abstract: Pretrained tabular foundation models have demonstrated strong predictive capability; however, their application to large-scale datasets remains constrained by the limited inference context. ・This paper introduces Balanced Adaptive Prototype Selection (BAPS), a framework for constructing compact, information-preserving contexts for scalable TabPFN inference. ・Without modif
cs.LG updates on arXiv.org

Bayesian Distributional Models of Executive Functioning

・arXiv:2510.00387v4 Announce Type: replace Abstract: This study uses controlled simulations with known ground-truth parameters to evaluate how Distributional Latent Variable Models (DLVM) and Bayesian Distributional Active LEarning (DALE) perform in comparison to conventional Independent Maximum Likelihood Estimation (IMLE). ・DLVM integrates observations across multiple executive function tasks and individuals, allowin
WIRED

Best Computer Monitors (2026): The Home Office Upgrade You Need

・My expert advice on what computer monitor to buy for your home office, ranging from budget-tier to fully featured.
WIRED

Best Pixel 10 Cases and Accessories (2026): Mous, dbrand, Bellroy

・Protect your Pixel 10a, Pixel 10, or Pixel 10 Pro XL with a tried-and-true phone case, screen protector, and Qi2 charger.
WIRED

Best Wireless Chargers (2026): My Picks After Testing 100+

・Stop fumbling for cables in the dark. ・These WIRED-tested stands and pads will take the hassle out of refueling your phone, wireless earbuds, and watch.
cs.LG updates on arXiv.org

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents

・arXiv:2608.12764v1 Announce Type: new Abstract: Deep search agents operate over trajectories spanning dozens of steps, yet standard reinforcement learning provides only a single outcome reward per trajectory, which is far too sparse for effective credit assignment. ・On-policy self-distillation (OPSD) addresses this by using the model's own logits as dense token-level teachers, but extending it to search agents introdu
cs.LG updates on arXiv.org

Beyond Simulated Benchmarks: Evaluating Motion Representations for Fall Detection Under Real-World Data Scarcity

・arXiv:2608.13197v1 Announce Type: new Abstract: Falls are a major health concern for older adults, and wearable sensors have been widely explored for detecting falls and enabling timely intervention. ・However, real-world falls are extremely rare: collecting 100 of them requires an estimated 100,000 days of monitoring, resulting in severely limited labelled data for training machine learning models. ・Consequently, many
cs.LG updates on arXiv.org

Black-Box Knowledge Transfer across Distinct Feature Sets

・arXiv:2608.12403v1 Announce Type: cross Abstract: Pre-trained black-box predictive functions encode knowledge distilled from massive datasets and extensive computation. ・However, when the available input features differ from those the black box expects, direct use is infeasible. ・We introduce a method for transferring predictive knowledge from the black box to a new, heterogeneous input space.
#LLMタグ

Blenderで文章から3Dモデルを作る無料アドオン「Twinforge」

・Blender で 3D の小道具を作るとき、いちばん時間がかかるのは「ありふれた物」です。木箱、樽、ランタン、祭壇。作品の主役ではないのに、無いと画が埋まらない。そういう物を一行書いて釦を押すだけで作るための拡張機能を作りました。名前は Twinforge(ツインフォージ)といいます。 ・Blender 5.0 以上で動きます。配布物は zip ひとつ。API キーの入力は要りません。モデリングの経験も要りません。
cs.LG updates on arXiv.org

Branch and Bound for Relational Verification of Neural Networks

・arXiv:2608.13118v1 Announce Type: new Abstract: Verification of neural networks against relational specifications, such as global robustness, is crucial for safety-critical applications of cyber-physical systems (CPS), given their increasing adoption of AI components. ・Compared to simple trace properties (e.g., local robustness), verifying relational specifications requires reasoning about the relationship between mul
cs.LG updates on arXiv.org

CAKE: Compiler-Agent Co-Design for Frontier Kernel Evolution

・arXiv:2608.12629v1 Announce Type: new Abstract: GPU kernel agents and GPU programming languages have advanced separately, leaving expert kernels difficult to reproduce. ・Agents usually treat the compiler as a fixed black box and receive only errors, correctness outcomes, and timing, while existing DSLs either hide critical scheduling decisions or expose them through difficult layout abstractions. ・We present CAKE, a co
cs.LG updates on arXiv.org

Can Spectral-Clipping Enable Better Learning While Forgetting Less for Low-Rank Adaptation?

・arXiv:2608.12332v1 Announce Type: cross Abstract: In recent years, low-rank adaptation (LoRA) has emerged as a significant paradigm that freezes pre-trained weights and introduces small, learnable adapters instead of fine-tuning the full set of parameters. ・In this work, we uncover several key insights regarding the singular components of network parameters based on Singular Value Decomposition (SVD). ・Firstly, the pri
cs.LG updates on arXiv.org

Capability Sheaves for Compositional Agent-Harness Repair: Controlled Quotients and a Real-Repository Stress Test

・arXiv:2608.13228v1 Announce Type: cross Abstract: Agent harnesses combine retrieval, routing, state, provenance, and verification, but locally successful components may disagree on shared state. ・We model this failure with a finite \emph{capability sheaf}: stalks encode typed behavior signatures, restriction maps retain shared fields, and accepted runs are useful global sections. ・An exact finite constraint-satisfactio
cs.LG updates on arXiv.org

CardioState-JEPA: Delay-Aware Cross-Modal Learning of a Shared Cardiac Representation

・arXiv:2608.12944v1 Announce Type: new Abstract: Electrocardiography (ECG), photoplethysmography (PPG), and phonocardiography (PCG) provide complementary views of the same cardiac cycle, yet existing cardiac foundation models are trained for a single sensing modality, leaving the shared physiology across sensors unexploited. ・We introduce CardioState-JEPA, a cardiac foundation model to learn a single shared representat
cs.LG updates on arXiv.org

CAS: A Causal Attribution Score for Local and Global Explainable Artificial Intelligence

・arXiv:2608.12555v1 Announce Type: cross Abstract: Predictive explanation methods attribute a model output; they do not, by themselves, attribute an intervention effect on the real-world outcome. ・We introduce the Causal Attribution Score (CAS), a compact score architecture for causal explanation. ・CAS starts from an identified interventional coalition game, allocates the joint intervention contrast with causal Shapley
cs.LG updates on arXiv.org

Chance-constrained selection of sequential intervention strategies from counterfactual estimates

・arXiv:2608.13209v1 Announce Type: cross Abstract: Many operational decisions are sequences of interventions under a cumulative resource limit, such as a maintenance schedule within a crew-hour budget. ・Choosing among them calls for the outcome and the cumulative cost each would produce, counterfactual quantities identified from observational data. ・Two strategies with the same expected cost can exceed the budget at ver
Zennの「大規模言語モデル」のフィード

claude -p 15分ハンズオン — 初めてのヘッドレス実行

・Claude Codeには、対話UIを起動せずに「プロンプトを渡す → 結果を受け取る → 終了」するヘッドレス実行モードがあります。これができると、Claude Codeをシェルスクリプトから、CIから、cronから呼べるようになります。つまり自動化の部品になります。 ・この記事は、そのヘッドレス実行を15分で一通り体験するハンズオンです。ゴールは「JSON出力をパースして、シェルスクリプトに組み込める形」まで到達することです。 ・前提: Claude Codeがインストール・認証済みで、対話モードで使ったことがあること。
Zennの「大規模言語モデル」のフィード

Claude Code 2.1.123を読む: experimental betas無効化とOAuth 401修正

・概要 Claude Code 2.1.123 が 2026年4月29日に公開されました。今回の changelog は1項目だけです。 ・項目 内容 バージョン Claude Code 2.1.123 公開日 2026年4月29日 変更種別 認証まわりのバグ修正 対象 CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS=1 を使う環境 修正 OAuth 認証が 401 リトライループで失敗する問題 大きな新機能ではありません。ただ、Claude Code を gateway、Bedrock、Vertex AI、社内プロキ...
Zennの「大規模言語モデル」のフィード

Claude Codeの429は2種類ある——usage枠とthroughput制限を取り違えた代償

・結論から Claude Codeで見る「API Error (429)」には、性質がまったく違う2系統があります。 ・分単位で回復するthroughput制限(速度の上限)と、5時間窓・週次で管理されるusage枠(総量の上限)です。 ・厄介なのは、エラー表示がこの2つを区別してくれないこと(GitHub Issue #25805で報告されている通り、どちらも「API Error」)。そして対処が真逆なことです。throughputなら「少し待って遅くする」、usageなら「待っても無駄、総量を減らす」。取り違えると、私のようにsubagent 75体の並列実行で5時間枠を吹き飛ばします...
#LLMタグ

Claudeで話題の「見えない文章透かし」を実際に試した――削除と全面リライトで何が起きたか、公開論文のKGW方式とSynthID-TextをGoogle Colabで動かして確かめる

・注意事項 本記事は2026年8月14日時点で確認できた情報と公開実装に基づいています。Claudeが実際に採用するテキストwatermarkの具体的な方式、検出器、鍵、閾値は現時点では公開されていません。本記事の実験は、公開論文で提案されている代表的な方式を実際に動かし、テキストwatermarkとはどのような技術なのかを理解することを目的としています。Claudeのwatermarkそのものを再現・評価したものではありません。 ・また、本記事はChatGPT Pro上のGPT-5.6 Proを用いて、構成整理、実験結果の整理、原典確認、文章校正を行いながら執筆しています。 ・前回、Claudeが生成テキストに電子透かし(watermark)を導入するという発表を受けて、EU AI Act、既存のLLM text watermark研究、そして社会での使い方についてまとめました。
The Verge

CMF’s clip earbuds hit the balance between cheap and good

・The Clip Pro are the first clip-style earbuds from CMF, Nothing’s budget sub-brand. ・Clip earbuds are an exercise in compromise. ・It's an inherent aspect of their design - and physics.
WIRED

Columbia Promo Codes: 15% Off | August 2026

・Explore current Columbia deals on jackets, outdoor gear, and apparel. ・Find active Columbia promo codes, student discounts, and free shipping offers to save on your next adventure.
cs.LG updates on arXiv.org

CoMedBench: A Multi-Source Benchmark of Synthetic Medical Data Fidelity and Downstream Utility

・arXiv:2608.12805v1 Announce Type: new Abstract: Access to clinical data is essential for developing reliable healthcare machine learning systems, but direct use of electronic health records is constrained by privacy regulation, institutional review, data-use agreements, and the risk of re-identification. ・Synthetic data promises a practical alternative: it can preserve useful statistical and clinical structure while r
cs.LG updates on arXiv.org

Comment on "Modeling rapid language learning by distilling Bayesian priors into artificial neural networks"

・arXiv:2608.12974v1 Announce Type: new Abstract: McCoy & Griffiths (2025, henceforth M&G) suggest that a Bayesian prior can be distilled into Artificial Neural Networks (ANNs) through Model-Agnostic Meta-Learning (MAML, Finn et al., 2017). ・They support this empirically by showing that meta-trained networks demonstrate formal language learning abilities comparable to Yang & Piantadosi (2023)'s Bayesian learner, signifi
cs.LG updates on arXiv.org

Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition

・arXiv:2608.12327v1 Announce Type: cross Abstract: Multilingual pretrained models nominally support Nepali, yet no controlled benchmark has compared them under a single fine-tuning protocol. ・We fine-tune six pretrained models (XLSR-53, IndicWav2Vec, MMS-1B, Whisper-Medium, Whisper-Large-v3-Turbo, and Conformer-Hi) spanning CTC self-supervised, autoregressive encoder-decoder, and hybrid Conformer-CTC architectures, on
cs.LG updates on arXiv.org

Concept Drift Detection and Adaptive Retraining of Malware Classification Models

・arXiv:2608.13465v1 Announce Type: new Abstract: Concept drift refers to changes over time in the statistical properties of data, as compared to the data that was used to train a learning model. ・Machine learning models for malware detection or classification are particularly susceptible to performance degradation caused by concept drift, as attackers constantly modify existing malware. ・In this chapter, we analyze two
cs.LG updates on arXiv.org

Constitutional On-Policy Safe Distillation

・arXiv:2606.03089v3 Announce Type: replace Abstract: On-policy self-distillation (OPSD) has emerged as an efficient post-training paradigm by using a teacher conditioned on privileged information to provide dense token-level supervision. ・Prior work has shown that OPSD can collapse in verifiable reasoning tasks, while safety alignment differs in that it is guided by high-level constitutions rather than explicit target
Hugging Face Papers

Context-Matched Distillation: Teacher Causality for Autoregressive Video Distillation

Context-Matched Distillation: Teacher Causality for Autoregressive Video Distillation
cs.LG updates on arXiv.org

Continual Distillation Learning for Rehearsal-Free Class-Incremental Learning via Decoupled Prompting

・arXiv:2407.13911v5 Announce Type: replace-cross Abstract: Prompt-based continual learning has shown strong performance in rehearsal-free class-incremental learning by adapting learnable prompts while freezing a pre-trained Vision Transformer (ViT) backbone. ・However, the effect of backbone scale remains underexplored. ・We observe that larger ViT backbones consistently yield better continual learning performance, which
Qiita - 人気の記事

CSSが反映されない原因10選

・CSSを書いたのに「なぜか反映されない……」という経験はありませんか? CSSが反映されない原因は、スペルミスやファイルの読み込みミスなど、ちょっとしたことであることが多いです。 ・今回は、初心者が特に間違えやすい 「CSSが反映されない原因」 を10個紹介します。
cs.LG updates on arXiv.org

Cueless EEG imagined speech for subject identification: dataset and benchmarks

・arXiv:2501.09700v2 Announce Type: replace Abstract: Electroencephalogram (EEG) signals have emerged as a promising modality for biometric identification. ・While previous studies have explored the use of imagined speech with semantically meaningful words for subject identification, most have relied on additional visual or auditory cues. ・In this study, we introduce a cueless EEG-based imagined speech paradigm, where sub
Cursor Blog

Cursor is now a part of SpaceX

・Cursor has officially been acquired by SpaceX.
Hugging Face Papers

CW-BASS v2: Saturation-Aware Pseudo-Label Selection for Semi-Supervised Segmentation under Foundation-Model Teachers

CW-BASS v2: Saturation-Aware Pseudo-Label Selection for Semi-Supervised Segmentation under Foundation-Model Teachers
cs.LG updates on arXiv.org

CW-BASS v2: Saturation-Aware Pseudo-Label Selection for Semi-Supervised Segmentation under Foundation-Model Teachers

・arXiv:2608.12773v1 Announce Type: cross Abstract: Semi-supervised semantic segmentation has long turned on one question, which pseudo-labels to trust, and a generation of selection rules, dynamic thresholds, per-class curricula, soft confidence weights, answered it for the noisy, under-confident ResNet teachers of their day. ・Self-supervised foundation encoders change the regime: with a DINOv2 teacher, confidence satu
cs.LG updates on arXiv.org

DARTree: Speculative Diffusion Decoding with Autoregressive Draft Trees

・arXiv:2608.13524v1 Announce Type: new Abstract: Speculative decoding losslessly accelerates autoregressive language models by verifying multiple draft tokens in parallel. ・Diffusion-based drafters further reduce proposal latency by predicting an entire token block in parallel, but their position-wise distributions are marginal rather than conditioned on tokens selected along each draft path. ・Existing recurrent correct
Hugging Face Papers

DarwinX: Evolving Agent Harnesses Through Natural Selection

DarwinX: Evolving Agent Harnesses Through Natural Selection
cs.LG updates on arXiv.org

Dead text or binding clause? Measuring and restoring constraint influence in black-box LLM dialogues

・arXiv:2608.12599v1 Announce Type: cross Abstract: Multi-turn dialogues let users revoke constraints as easily as impose them, but revocation does not reliably take effect: models keep enacting withdrawn requirements (occasionally beneath comments asserting their removal), a failure we call \emph{behavioral relapse}, or revocation inertia. ・No existing instrument measures this influence per clause, predicts it before d
cs.LG updates on arXiv.org

Decentralized Multi-Player Q-Learning in Episodic Markov Decision Processes with Information Asymmetry

・arXiv:2608.12753v1 Announce Type: new Abstract: We study decentralized multi-player reinforcement learning in episodic tabular Markov decision processes (MDPs) under three forms of information asymmetry: (A) unobserved actions with common rewards, (B) observed actions with independent rewards, and (C) unobserved actions with independent rewards. ・Players cannot communicate during learning but may agree on a protocol a
cs.LG updates on arXiv.org

Decomposition of Evidence, Contradiction, and Fragility in Perturbation Responses

・arXiv:2608.12935v1 Announce Type: cross Abstract: Perturbation methods explain model decisions by measuring prediction changes under altered inputs, but response magnitude tells us only how much a model reacts, not what that reaction means. ・The same magnitude can support the final factual-counterfactual difference, oppose it, or arise strongly along the perturbation path yet vanish at the endpoint. ・We therefore track
Zennの「大規模言語モデル」のフィード

DeepSeek-V4-Pro-0813をDay1デプロイ! 実データ17,000 tok/s、三つ巴の頂点へ【B300 x8 検証速報】

・はじめに 2026年8月13日 21時30分(JST)、DeepSeek から V4-Pro の正式版 DeepSeek-V4-Pro-0813 のモデルウェイトが公開されました。4月の Preview 公開から約4ヶ月、満を持しての GA(一般提供)版です。 ・フィックスターズでは、Kimi-K3、Qwen3.8-2.4T-A95B に続く Day1 デプロイ検証の第3弾として、公開翌日に NVIDIA B300 x8 のシングルノード環境へデプロイし、同一条件のベンチマークを実施しました。なお本シリーズでは、リリースから 24時間以内に実施した検証を「Day1 デプロイ」と呼んでい...
cs.LG updates on arXiv.org

Defensive Boosting for Online Probabilistic Forecasting

・arXiv:2608.13554v1 Announce Type: new Abstract: We study online probabilistic forecasting of binary outcomes chosen by an adaptive adversary. ・Given an online learning algorithm for a weak hypothesis class $H$, we would like to efficiently obtain two incomparable guarantees that existing online boosting techniques provide separately. ・Online gradient boosting competes in Brier score with the best predictor induced by t
WIRED

Dell XPS 13 Review: Move Over, Neo

・It’s amazing how few compromises Dell made to get the new XPS 13 down to $700, but I strongly recommend the $900 model, which comes with 16 GB of RAM.
cs.LG updates on arXiv.org

Demand Transfer Estimation at Scale via Restricted Logit Modeling

・arXiv:2608.12680v1 Announce Type: new Abstract: Item demand forecasting is an integral component of store assortment optimization. ・Existing literature focuses on learning a suitable customer choice model and using this model to determine the value of an objective function (i.e. ・expected demand) with respect to an assortment proposal.
cs.LG updates on arXiv.org

Designing AI Pipelines for Decision-Ready ITSM Intelligence

・arXiv:2608.12670v1 Announce Type: cross Abstract: IT service management (ITSM) systems accumulate large volumes of heterogeneous ticket data that are difficult for sales and executive stakeholders to convert into actionable intelligence. ・This paper presents a sociotechnical AI pipeline, designed and evaluated following design science research principles, that transforms raw ITSM exports into a multilevel decision-sup
cs.LG updates on arXiv.org

Diagnosing JEPA World Models with Action-Conditioned Predictive Consistency

・arXiv:2608.12939v1 Announce Type: new Abstract: Joint-embedding predictive architectures (JEPAs) learn world models that predict in a compact latent space rather than in pixels, reducing the pressure to model nuisance appearance. ・Yet this provides no guarantee against visual perturbations: they can still alter the encoded representation and affect subsequent action-conditioned predictions. ・Bisimulation captures this
cs.LG updates on arXiv.org

Difference-of-Convex Regularization for Graph Learning by Differentiable Programming

・arXiv:2608.12757v1 Announce Type: cross Abstract: Laplacian-regularized minimization is fundamental in signal processing and machine learning, but is limited by the dense and ill-conditioned nature of the graph Laplacian pseudoinverse. ・While the Laplacian itself is sparse, its pseudoinverse is dense and often ill-conditioned, rendering direct computation impractical at scale. ・Moreover, pseudoinverse learning is more
cs.LG updates on arXiv.org

DiffGRM: Diffusion-based Generative Recommendation Model

・arXiv:2510.21805v2 Announce Type: replace-cross Abstract: Generative recommendation (GR) is an emerging paradigm that represents each item via a tokenizer as an n-digit semantic ID (SID) and predicts the next item by autoregressively generating its SID conditioned on the user's history. ・However, two structural properties of SIDs make ARMs ill-suited. ・First, intra-item consistency: the n digits jointly specify one ite
cs.LG updates on arXiv.org

DiG-bench: Discovery in Games

・arXiv:2608.12593v1 Announce Type: cross Abstract: Discovery---formulating novel generalizations---is a central part of the scientific process. ・Despite its importance, there is a gap in the current AI benchmark landscape, with few benchmarks directly probing the capacity for discovering new knowledge with experimentation in controlled environments where the objective is unknown. ・To address this gap, we release a new b
cs.LG updates on arXiv.org

Dimensional Balance Improves Large Scale Spatiotemporal Prediction Performance

・arXiv:2605.18793v3 Announce Type: replace Abstract: Accurate spatiotemporal pattern analysis is critical in fields such as urban traffic, meteorology, and public health monitoring. ・However, existing methods face performance bottlenecks, typically yielding only incremental gains and often exhibiting limited cross-domain transferability. ・We analyze this bottleneck through spatial and temporal entropy measures, which ar
cs.LG updates on arXiv.org

Discovering Persistent Behavioural Patterns for Interpretable Blockchain Forensics

・arXiv:2608.12864v1 Announce Type: cross Abstract: Public blockchain data enables large-scale DeFi-related analysis, but many existing approaches are application-specific, difficult to scale, or hard to interpret. ・This research proposes a scalable, application-agnostic framework for \emph{persistent behavioural pattern discovery} from large-scale blockchain activity. ・It constructs behaviour sentences enriched with con
cs.LG updates on arXiv.org

Distributed Online Submodular Maximization under Communication Delays: A Simultaneous Decision-Making Approach

・arXiv:2603.27803v2 Announce Type: replace Abstract: We provide a distributed online algorithm for multi-agent submodular maximization under communication delays. ・We are motivated by the future distributed information-gathering tasks in unknown and dynamic environments, where utility functions naturally exhibit the diminishing-returns property, i.e., submodularity. ・Existing approaches for online submodular maximizatio
cs.LG updates on arXiv.org

Distribution Steering via Sliced Optimal Transport Control

・arXiv:2608.12828v1 Announce Type: cross Abstract: Distribution steering seeks feedback laws that drive the state law of a dynamical system between prescribed initial and terminal distributions. ・Optimal transport provides a natural geometric approach, but its implementation generally requires a transport map or coupling in the full state space. ・Sliced optimal transport avoids this full-dimensional construction through
cs.LG updates on arXiv.org

Do Transformers Need Three Projections? Systematic Study of QKV Variants

・arXiv:2606.04032v3 Announce Type: replace Abstract: Transformers have become the standard solution for various AI tasks, with the query, key, and value (QKV) attention formulation playing a central role. ・However, the individual contribution of these three projections and the impact of omitting some remain poorly understood. ・We systematically evaluate three projection sharing constraints: a) Q-K=V (shared key-value),
AI News & Artificial Intelligence | TechCrunch

Does Mark Zuckerberg really believe AI is ‘for everyone’?

・Meta released Glimmer this week, an open-weight AI model anyone can download and run on their own hardware — a contrast to Muse Spark, the company’s more powerful model that stays locked behind its own APIs. ・The release landed alongside a letter from Mark Zuckerberg arguing AI should be “for everyone” rather than controlled by a handful of labs, but as Equity’s […]
stat.ML updates on arXiv.org

Don't Cut Corners: How Training Outside the Prior Makes Simulation-Based Inference More Robust

・arXiv:2608.12470v1 Announce Type: cross Abstract: Large astrophysical simulation campaigns often generate training data by sampling parameters across a Uniform prior box. ・Due to the proposal's sharp edge, neural posterior estimators struggle to learn accurate approximations near the boundaries. ・We propose Tailed-Uniform, a family of hybrid proposal distributions for sampling training simulations for robust simulation
cs.LG updates on arXiv.org

Doubly Robust Estimation of Causal Effect on CVR with Targeted Regularization

・arXiv:2608.13461v1 Announce Type: new Abstract: Post-click conversion rate (CVR) is a key metric in various scenarios including e-commerce and advertising, reflecting the efficiency and user experience in the second stage of the conversion process. ・Estimating the causal effect on CVR is therefore of great practical importance. ・However, directly applying existing causal inference methods to clicked samples introduces
Hugging Face Papers

DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation

DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation
cs.LG updates on arXiv.org

Drive-to-Music: Context-Aware Generative Audio for In-Vehicle Experiences

・arXiv:2608.12615v1 Announce Type: cross Abstract: In-vehicle music can serve as an adaptive interface to enhance driver experience, attention, and well-being. ・We present Drive-to-Music, a context-aware system that generates music in real time from multimodal driving signals. ・Using dashcam imagery and vehicle telemetry, the system extracts scene semantics and driving context, maps them to high-level musical descriptor
cs.LG updates on arXiv.org

Dual Spatial-Temporal Attribution: Architecture-Aligned Post-Hoc Explainability for Recurrent Graph Anomaly Detection

・arXiv:2608.12441v1 Announce Type: new Abstract: Deep learning detectors for anomalies in dynamic graphs have reached strong accuracy, yet they remain opaque: when an edge is flagged, the analyst receives a score but no reason. ・This opacity is untenable in the cooperative, regulated information systems where such detectors are deployed, where automated decisions must be auditable and trustworthy. ・We address this gap f
cs.LG updates on arXiv.org

DYSANOS Generative Dynamic Smooth Arbitrage-free Non-parametric Option Surfaces

・arXiv:2608.12587v1 Announce Type: cross Abstract: This article presents with DYSANOS the first generative market model for smooth SANOS option surfaces for all strikes and expiries which are free of static arbitrage. ・Our model is designed to generate entire paths of daily spot and option prices for years in the future. ・We present a robust and useful if somewhat simplistic baseline hidden state generative model in the
cs.LG updates on arXiv.org

EEG Decoding Using CNN and LSTM Network

・arXiv:2608.13285v1 Announce Type: new Abstract: Motor imagery (MI) brain--computer interfaces (BCIs) have emerged as a promising approach for establishing flexible communication pathways between the human brain and external devices , particularly for individuals affected by stroke or neurodegenerative disorders. ・Reliable decoding of motor-imagery electroencephalography (MI-EEG) remains challenging because EEG recordi
cs.LG updates on arXiv.org

Efficient Hessian-Free Methods for Multi-Objective Bilevel Optimization with Nonconvex Lower Level

・arXiv:2608.12704v1 Announce Type: cross Abstract: Multi-objective bilevel optimization has wide applications in the AI area such as automated learning and multi-task meta-learning. ・Although recently some works have been begun to study the multi-objective bilevel optimization, the proposed methods rely on the (strongly) convex lower level problems. ・In fact, these multi-objective bilevel learning problems are generally
cs.LG updates on arXiv.org

Efficient Image Restoration with State-Dependent Forward Diffusion

・arXiv:2505.16733v3 Announce Type: replace Abstract: This paper proposes to perform image restoration through a state-dependent mean-reverting forward diffusion (FoD) process. ・In contrast to traditional diffusion-based approaches that rely on a coupled forward-backward diffusion scheme, FoD directly learns image restoration through a single forward diffusion process, yielding a simple yet efficient framework.
cs.LG updates on arXiv.org

EGRL: Edge generation-guided relation-aware learning for RNA-protein interaction prediction

・arXiv:2608.12906v1 Announce Type: new Abstract: RNA-Protein Interactions (RPIs) are critical for regulating cellular functions. ・While traditional wet-lab experiments for RPI detection are costly and time-consuming, Deep Learning (DL) methods provide an efficient computational alternative for RPI Prediction (RPIP). ・In particular, Graph Neural Networks (GNNs) are promising, as they naturally model RPI networks.
cs.LG updates on arXiv.org

Embedding networks with the random walk first return time distribution

・arXiv:2512.02694v3 Announce Type: replace-cross Abstract: We propose the first return time distribution (FRTD) of a random walk as an interpretable and mathematically grounded node embedding. ・The FRTD assigns a probability mass function to each node, allowing us to define a distance between any pair of nodes using standard metrics for discrete distributions. ・We present several arguments to motivate the FRTD embedding
cs.LG updates on arXiv.org

Enhancing In-Hospital Mortality Prediction Using Multi-Representational Learning with LLM-Generated Expert Summaries

・arXiv:2411.16818v2 Announce Type: replace-cross Abstract: To evaluate a multi-representational framework in which large language model (LLM)-generated expert summaries of intensive care unit (ICU) notes are fused with physiology for in-hospital mortality (IHM) prediction, and to determine how much of the resulting gain is non-redundant with the notes themselves. ・Using MIMIC-III (19,211 first ICU stays, 12.83% mortali
cs.LG updates on arXiv.org

Equivariant learning of a transferable three-dimensional classical density functional

・arXiv:2608.13506v1 Announce Type: cross Abstract: Liquids exhibit collective behavior that depends sensitively on thermodynamic conditions, interfaces and confinement, yet predicting each new state commonly requires a separate atomistic simulation. ・Classical density functional theory offers a reusable variational description, but its central excess free-energy functional is generally unknown, and learned approximatio
cs.LG updates on arXiv.org

EU-ETS under attack? The impact of carbon price suppression on the decarbonization of the power sector

・arXiv:2608.12363v1 Announce Type: cross Abstract: European countries are debating policies to mitigate the increased energy costs caused by renewed geopolitical tensions, while pursuing decarbonization and electrification. ・A notable example is Italy's 2026 Decreto Bollette package, which proposes to remove the carbon price equivalent from the bids of certain gas-driven power plants to wholesale electricity markets, a
cs.LG updates on arXiv.org

Evaluating AlphaEarth Foundations Embeddings for Wildfire Susceptibility Mapping

・arXiv:2608.12663v1 Announce Type: cross Abstract: Wildfire susceptibility mapping typically relies on physical variables assembled from multiple remote-sensing, climate, and geospatial products. ・AlphaEarth Foundations (AEF) provides analysis-ready geospatial embeddings that may reduce this dependence on heavy harmonisation and task-specific feature engineering, but their value for wildfire susceptibility mapping has
cs.LG updates on arXiv.org

Evaluation Resolution Confounds Learning-Rule Comparisons in Model-Brain RSA of Early Visual Cortex

・arXiv:2608.12408v1 Announce Type: cross Abstract: Representational similarity analysis (RSA) is increasingly used to ask which learning rules give convolutional networks brain-like representations. ・Because biologically plausible rules such as feedback alignment, predictive coding and STDP do not scale, studies that include them train small networks on small images (typically 32x32 CIFAR) and then compare them to brai
cs.LG updates on arXiv.org

Excess Separability: Nuisance-Controlled Residual-Stream Probing for Benchmark Contamination Detection

・arXiv:2608.12652v1 Announce Type: cross Abstract: Benchmark contamination is diagnosed today with n-gram overlap, with likelihood-based membership inference, or with canary strings, and each needs something usually unavailable: the training corpus, a well-chosen test statistic, or foresight at dataset release. ・A recent alternative reads contamination off a linear probe on internal activations. ・We show that the natura
cs.LG updates on arXiv.org

Exemplar-based objective classification of gust-induced loads across multiple flight conditions

・arXiv:2608.12448v1 Announce Type: new Abstract: Is it possible to find an objective classification criterion that organizes the complexity of gust-induced loads across many flight conditions? ・And one that remains as interpretable as a labelling based on coarse parameters, such as the flight attitude? ・Our approach encodes a large number of experimental observations through a machine-learned representation and applies
cs.LG updates on arXiv.org

Explanatory Engagement Under Rare Anomalous Failure: Asymptotic Rarity in Model Behavior (or: The Asymptotic AI)

・arXiv:2608.13063v1 Announce Type: cross Abstract: Prior work on LLM behavior under anomalous conditions asks whether a model notices anomalies. ・We ask a narrower question: once a model sits in a workflow with a low, controllable failure rate, does its explanatory engagement - length, specificity, self-reported confidence - change as failure grows asymptotically rarer? ・We built a local, zero-cost harness on three open
cs.LG updates on arXiv.org

Exploring Oversmoothing with Householder Matrices

・arXiv:2608.12514v1 Announce Type: new Abstract: Deep graph neural networks(GNNs) suffer from oversmoothing- a progressive collapse of node representation towards a low information subspace as network depth increases because the normalized graph propagation operator is repeatedly applied directly to the hidden representations. ・In this work we study Householder Graph Neural Network (HouseGNN). ・Rather than updating the
cs.LG updates on arXiv.org

Exploring Sparsity for Parameter Efficient Fine Tuning Using Wavelets for Vision

・arXiv:2505.12532v3 Announce Type: replace-cross Abstract: Efficiently adapting large pretrained models is critical under tight compute and memory budgets. ・While Parameter-Efficient Fine-Tuning (PEFT) methods like LoRA achieve efficiency through low-rank updates, their discrete rank constraint limits fine-grained parameter control and confines adaptations to low-dimensional subspaces. ・We propose Wavelet Fine-Tuning (W
cs.LG updates on arXiv.org

Exponential Convex Calibration Dimension for the Multi-Label Jaccard Measure

・arXiv:2608.13549v1 Announce Type: new Abstract: The per-instance Jaccard score, or intersection over union (IoU), is standard in multi-label classification and binary segmentation. ・With $s$ labels, its loss matrix has $2^s$ outcomes and reports. ・Under the convention $\mathrm{Jac}(\varnothing,\varnothing)=1$, we prove that the Jaccard score, shifted-loss, and ordinary loss matrices are nonsingular and that the loss co
cs.LG updates on arXiv.org

Exponential quantum advantage for learning signals with a single qubit

・arXiv:2608.13521v1 Announce Type: cross Abstract: Quantum technology has the potential to transform scientific discovery, but quantum advantages often require processing capabilities well beyond the reach of experimental platforms. ・We show that coupling a single controllable qubit to an otherwise conventional sensor can exponentially reduce the number of measurements required to learn classical signals. ・These rigorou
cs.LG updates on arXiv.org

Fast A/B/n Testing: Exact Multi-Policy Comparison via Tree-Coupled Feedback Sharing

・arXiv:2608.12831v1 Announce Type: new Abstract: Online platforms increasingly compare many adaptive decision policies---ranking systems, recommendation algorithms, pricing rules, and language-model agents---while each reward-bearing interaction can be costly or risky. ・A direct A/B/n design gives each of $J$ policies its own horizon-$T$ trajectory and therefore uses $JT$ outcomes. ・We introduce Tree-Coupled A/B Testing
cs.LG updates on arXiv.org

Fast Length-Squared Sampling for Positive-Semidefinite Matrices

・arXiv:2608.12503v1 Announce Type: cross Abstract: We describe a simple rejection-sampling-based algorithm to perform length-squared sampling on an $n \times n$ positive-semidefinite (psd) matrix: that is, to sample a column with probability proportional to its squared $\ell_2$-norm. ・The algorithm runs in just $O(n)$ expected time, which is significantly sublinear in the input matrix size. ・The runtime is optimal, even
cs.LG updates on arXiv.org

Federated Compositional Muon Optimizer for Matrix-Wise Models

・arXiv:2608.12710v1 Announce Type: new Abstract: Muon, a more recently developed optimizer, is useful for matrix-wise models in AI areas. ・Although many works have studied Muon and its variants, these methods are still not particularly well-suited for hierarchical structured problems. ・To fill this gap, we propose an effective federated compositional Muon (FedCoMuon) optimizer to solve distributed matrix-wise compositio
cs.LG updates on arXiv.org

Finding the Needle in a Haystack: Test-Time Analog Circuit Representation Adaptation for Bayesian Optimization

・arXiv:2608.12687v1 Announce Type: new Abstract: Bayesian optimization (BO) is a sample-efficient framework for analog circuit topology search, where evaluating each candidate topology can require costly simulation. ・However, representation-based BO methods typically treat circuit embeddings as fixed after encoder training. ・This creates a mismatch between representation learning and optimization: embeddings learned to
cs.LG updates on arXiv.org

Fine-tuned Normalizing Flows for ALICE Zero Degree Calorimeter Fast Simulation

・arXiv:2608.12795v1 Announce Type: cross Abstract: Simulating the ALICE Zero Degree Calorimeter (ZDC) neutron detector responses at the LHC is computationally expensive, requiring complex Monte Carlo chains. ・We develop a generative surrogate, focusing on Normalizing Flows (NFs). ・Through transfer learning, we pre-train on the full imbalanced dataset and fine-tune specialized models for different particle types ($\gamma
cs.LG updates on arXiv.org

Finite-Time Minimax Bounds and an Optimal Lyapunov Policy in Queueing Control

・arXiv:2506.18278v4 Announce Type: replace-cross Abstract: We introduce an original minimax framework for finite-time performance analysis in queueing control and propose a surprisingly simple Lyapunov-based scheduling policy with superior finite-time performance. ・The framework quantitatively characterizes how the expected total queue length scales with key system parameters, including the capacity of the scheduling s
#AIタグ

First Signals Vol.002

・TikTok Shop +103%。「買う場所」が、また変わり始めた。 ・5分で読む、今週の海外コマース・AI・テクノロジー。
cs.LG updates on arXiv.org

FlowLOB: Efficient and Controllable Limit Order Book Generation with Flow Matching

・arXiv:2608.13096v1 Announce Type: new Abstract: Limit order book (LOB) simulators are most useful to practitioners when they combine realistic market dynamics, computationally efficient sampling, controllable scenario generation, and the ability to generalize beyond the instruments seen during training---properties that existing agent-based and deep generative simulators provide only partially. ・We present \textbf{Flo
cs.LG updates on arXiv.org

Foundation models for movement data: Are they ready for prime-time?

・arXiv:2608.13316v1 Announce Type: cross Abstract: Foundation models (FMs) trained on large-scale accelerometer data have been proposed as general-purpose feature extractors for health monitoring, but systematic evidence of their advantages is lacking. ・We present the first comprehensive evaluation of four open-source accelerometer FMs against supervised baselines covering 19 tasks across the domains of activity recogn
cs.LG updates on arXiv.org

Foundations of Independent Component Analysis

・arXiv:2608.13229v1 Announce Type: cross Abstract: We present the mathematical foundations of linear independent component analysis (ICA) models based on standard literature in a self-contained note. ・It is aimed at readers with a background in measure-theoretic probability theory. ・We first develop the theory of the characteristic functions of probability measures on $\mathbb{R}^d$, including their analyticity and the
cs.LG updates on arXiv.org

From Approximation Rates to Loss-Landscape Barrier Decay in Shallow ReLU Networks

・arXiv:2602.17596v2 Announce Type: replace Abstract: We study pathwise connectivity of sublevel sets for one-hidden-layer ReLU networks with constrained first-layer weights and an $\ell_1$ penalty on the output layer. ・The data term is assumed convex and globally Lipschitz in the scalar logit. ・We first give a finite-width construction that connects any two points of a common sublevel through a path controlled by a loss
Hugging Face Papers

From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs

From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs
cs.LG updates on arXiv.org

From Local Mismatch to Global Impact: Optimizing Cache Reuse Policy for Efficient Diffusion

・arXiv:2608.13043v1 Announce Type: cross Abstract: Diffusion models have achieved dominant performance in visual generation but suffer from substantial inference overhead. ・While cache-based acceleration has emerged as a promising solution, existing policies rely on local similarity heuristics, which we identify as being significantly misaligned with final generation quality. ・This discrepancy stems from the non-uniform
cs.LG updates on arXiv.org

From Visual Widgets to UI Code: Efficient Tool-Grounded Generation

・arXiv:2608.12611v1 Announce Type: cross Abstract: Existing screenshot-to-code systems face a trade-off between flexibility and controllability. ・Direct multimodal generation can hallucinate visible details, whereas structured pipelines reduce such errors through component-wise decomposition, predefined templates, and customized intermediate representations. ・These structures, however, introduce additional generative or
cs.LG updates on arXiv.org

FSGR: Mitigating Token Frequency Bias for Fair SID-Based Generative Recommendation

・arXiv:2608.12845v1 Announce Type: cross Abstract: Semantic ID (SID)-based generative recommendation has recently achieved remarkable success. ・However, existing methods suffer from a previously overlooked fairness issue, which we term \textbf{Token Frequency Bias}, where high-frequency SID tokens are systematically over-predicted while low-frequency SID tokens are under-predicted. ・This bias originates from the combine
WIRED

FTC Strikes Deals to Ignore Unlawful Credit Discrimination

・The agency signed agreements to not enforce parts of three federal court orders against auto dealers accused of discrimination—and didn’t notify judges or at least one of its coplaintiffs.
Hugging Face Papers

Full-bandwidth transformer

Full-bandwidth transformer
cs.LG updates on arXiv.org

Functional Adjoint Sampler: Scalable Sampling on Infinite Dimensional Spaces

・arXiv:2511.06239v2 Announce Type: replace-cross Abstract: Learning-based methods for sampling from the Gibbs distribution in finite-dimensional spaces have progressed quickly, yet theory and algorithmic design for infinite-dimensional function spaces remain limited. ・This gap persists despite their strong potential for sampling the paths of conditional diffusion processes, enabling efficient simulation of trajectories
cs.LG updates on arXiv.org

Functional-prior-based approaches to Bayesian PDE-constrained inversion using physics-informed neural networks

・arXiv:2605.07060v3 Announce Type: replace-cross Abstract: Physics-informed neural networks (PINNs) provide a mesh-free framework for solving PDE-constrained inverse problems, but their extension to Bayesian inversion still faces a fundamental difficulty: prior distributions are typically defined in the weight space of neural networks, whereas physically meaningful prior assumptions are more naturally expressed in fun
LLMタグが付けられた新着記事 - Qiita

Gemini 3.7 Flash 発表! 3.6 Flashとのエージェント対決では新モデルが強いとは言えなかった

・はじめに 2026年8月14日、米Googleから新たな生成AIモデルであるGemini 3.7 Flashが公開されました。前バージョンの 3.6 Flash からわずか3週間でのリリースとなります。 ・い、いくらなんでも早すぎー... ・公開されたブログの内容やスペ...
cs.LG updates on arXiv.org

GENADA: efficient generative time series adversarial attack framework

・arXiv:2608.12535v1 Announce Type: new Abstract: Deep learning models are widely used for time series analysis in domains such as healthcare, finance, energy systems, and environmental monitoring. ・However, these models remain vulnerable to adversarial attacks, where small input perturbations cause severe degradation in predictive performance. ・Commonly used gradient-based attacks, iterative first-order methods, are com
cs.LG updates on arXiv.org

General Bayesian Policy Learning

・arXiv:2602.23672v2 Announce Type: replace-cross Abstract: This study proposes a General Bayes framework for policy learning. ・We consider decision problems in which a decision-maker chooses an action from a given set to maximize expected welfare. ・Typical examples include treatment choice and portfolio optimization.
cs.LG updates on arXiv.org

Generative Brownian Bridge Diffusion In Motion Space For Enhanced Myocardial Strain Analysis

・arXiv:2608.01677v2 Announce Type: replace-cross Abstract: Myocardial strain analysis of cardiac magnetic resonance (CMR) images provides an important tool for evaluating cardiac function. ・However, current techniques require either human-adjusted post-processing with suboptimal regional accuracy, or specialized acquisitions with limited availability. ・In this paper, we propose to leverage the power of generative models
cs.LG updates on arXiv.org

Geometric and Behavioral Stratification in Transformer Residual Streams

・arXiv:2608.12447v1 Announce Type: new Abstract: Trained transformer models develop privileged bases: coordinate axes whose statistics differ from the rest of the residual stream. ・But what kind of direction does such a basis select? ・We investigate the prediction direction, the unembedding direction of the token a model currently predicts, and find that it functions as a content-defined privileged anchor.
Zennの「大規模言語モデル」のフィード

GGUF量子化は日本語を余計に壊すのか — Q8_0からQ2_Kまで実測した

・この記事の数値には、すべてどちらかのタグを付けています。実測 = 当方の Apple M5 Pro 48GB で計測した値(生データ全件掲載)。出典 = 公開一次ソースの値(その場でリンク)。 ・そして本記事には、もう一つ最初に書いておくべきことがあります。測ったのは自動指標だけで、 人手評価はやっていません。 ・何が言えて何が言えないかは、最後の「正直な限界」に全部書きました。
LLMタグが付けられた新着記事 - Qiita

GLM-5.3 vs GLM-5.2:高難度6課題を実 API で比較した

・GLM-5.3 vs GLM-5.2:高難度6課題を実 API で比較した glm-5.3 と glm-5.2 を同一プロンプト・同一温度・同一出力予算で比較しました。評価対象は、極難数学、Simpson のパラドックス、最小 UNSAT コア、ネスト JSON、多段階物...
Zennの「大規模言語モデル」のフィード

Google Colabで最新LLMを試す #6 ― Gemma 4 E4B-itで画像・音声・動画を扱う

・はじめに 「Google Colabで最新LLMを試す」シリーズの第6回です。 ・今回は、Googleの Gemma 4 E4B-it をGoogle Colab上で動かし、通常のテキスト生成だけでなく、画像・音声・動画を含むマルチモーダル入力を試してみます。 ・今回使用したNotebookは、以下のGitHubリポジトリで公開しています。
AI News & Artificial Intelligence | TechCrunch

Google will now allow users to remove visible watermark from its AI generations

・Turning off this setting won't affect invisible benchmarks used to identify an AI generated file.
The Verge

Google’s best new camera feature is only for the Pixel 11 series

・Arguably the coolest new photo feature for the Pixel 11 lineup is Google's new Camera Looks, which process image data differently at the sensor level to produce photos that don't have that "smartphone" look. ・The result is new styles like "Digi," which mimics the style of photos taken by older digital cameras. ・But to use Camera Looks, at least initially, you'll need to have one of Google's Pixel 11 phones.
cs.LG updates on arXiv.org

Gradient-Free Warm-Start Library Recovery: an Amortized-Regret Separation

・arXiv:2606.21253v2 Announce Type: replace Abstract: Continual learning that is gradient-free, local, online, and append-only is attractive for edge and streaming deployment, but its value is usually argued informally. ・We give a provable account on recurring-regime streams. ・Given segmentation, a warm-start library learner attains amortized recovery cost $O\!\big(KD/\varepsilon^2+(R-K)\logK/\Delta^2\big)$ versus a memo
cs.LG updates on arXiv.org

H-VAEP and H-xT: Valuing Offensive On-the-Ball Actions in Handball by Estimating Probabilities

・arXiv:2608.12926v1 Announce Type: new Abstract: Traditional player evaluation in professional handball relies on basic box-score metrics or heuristic indices, which fail to credit the multi-player build-up chain. ・While football (soccer) analytics has adopted Expected Threat (xT) and Valuing Actions by Estimating Probabilities (VAEP), these event-based action valuation frameworks have not yet been adapted to handball.
Hugging Face Papers

H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models

H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models
cs.LG updates on arXiv.org

Harmonizing Safety and Speed: A Human-Algorithm Approach to Enhance the FDA's Medical Device Clearance Policy

・arXiv:2407.11823v4 Announce Type: replace Abstract: The United States Food and Drug Administration's (FDA's) 510(k) pathway allows manufacturers to gain medical device approval by demonstrating substantial equivalence to a legally marketed device. ・However, the inherent ambiguity of this regulatory procedure has been associated with high recall among many devices cleared through this pathway, raising significant safet
WIRED

HelloFresh Promo Codes: 55% Off for August 2026

・Get up to 55% off and free meal boxes using a HelloFresh coupon code in August 2026. ・Discover our best codes and discounts to let you save time and money.
The Verge

Help build a monument to that ‘sad little bitch’ Elon Musk

・“Build the perfect monument to force a moment of introspection upon the world’s richest, ugliest little bitch.” | Image: Brendan SMIALOWSKI / AFP via Getty Images Cards Against Humanity is gearing up to build "something that will annoy Elon Musk," and it's crowdfunding the project with its usual flavor of vulgarity. ・The company behind the card game announced plans to build "a grand monument" to Musk on the parcel of
WIRED

Her Brain Was Broken. It Was Fixed With Sound—Not a Scalpel

・One woman’s meth addiction was so bad, the only option left might have been brain surgery. ・Then a single session of noninvasive, focused ultrasound seemed to do what years of treatment could not.
cs.LG updates on arXiv.org

High-dimensional networks and mean squared error for possibly misspecified models

・arXiv:2608.13171v1 Announce Type: cross Abstract: To avoid missing important variables and their connections in networks, more and more variables are included in network analysis. ・Here we show that in a setting with many more parameters than observations (high-dimensional) it is possible to get a conservative (i.e., low false positive rate) estimate of the neighbourhood for each node (which connections are in the net
cs.LG updates on arXiv.org

HiRoute: Hierarchical Routed Prompt Tuning for Safety Alignment of Large Language Models

・arXiv:2608.12821v1 Announce Type: new Abstract: Large language models (LLMs) remain vulnerable to harmful requests and jailbreak attacks. ・Parameter-efficient safety alignment methods based on prompt tuning typically rely on a single global prompt or externally selected prompt modules. ・Such static designs struggle to maintain a cross-category safety boundary while generating constructive responses tailored to specific
cs.LG updates on arXiv.org

History-informed Lagrangian Neural Networks

・arXiv:2608.13215v1 Announce Type: new Abstract: Forecasting the long-horizon evolution of mechanical systems from position-only observations is a pivotal yet difficult task, as hidden velocities and trajectory-specific physical properties must be inferred simultaneously. ・Although physics-guided neural networks like Lagrangian Neural Networks (LNNs) guarantee physical plausibility, they generally require complete stat
The Verge

Hoto’s new cordless soldering iron heats up in three seconds

・Hoto is introducing its first soldering iron through its modular Snapbloq collection that lets you assemble your own custom toolbox by stacking magnetic cases and tools together. ・The new I-A06 Cordless Soldering Iron can heat to just over 200 degrees Celsius in about three seconds and offers an adjustable temperature range from 100 to up to 450 degrees with customizable presets you can quickly jump to. ・The full kit v
Hugging Face Papers

How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review

How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review
cs.LG updates on arXiv.org

How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures

・arXiv:2608.13267v1 Announce Type: cross Abstract: Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty (how they behave when visual evidence is missing or misleading). ・We introduce SciFigBench, a diagnostic VLM benchmark for scientific figure un
Hugging Face Papers

Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing

Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing
AI News & Artificial Intelligence | TechCrunch

Hyperscalers might regret embracing natural gas if new forecast proves correct

・Natural gas prices could triple in some parts of the U.S., which could saddle hyperscalers with massive bills to power their AI data centers.
WIRED

I Wore an Electrical Muscle Stimulation Body Suit to Zap Myself Into Fitness

・Can electrifying your workout offer a shortcut to a stronger, fitter you? ・I sweated in a skintight EMS suit for two months to find out.
cs.LG updates on arXiv.org

I-SDPO: Instance-Level Adaptive Self-Distillation Policy Optimization

・arXiv:2608.12957v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) learns from reward differences within a rollout group, but receives no useful relative signal when every sampled response is incorrect. ・Privileged self-distillation can fill this gap with dense token supervision, yet applying it throughout training creates a different failure mode: the teacher is a biased, low-variance surrogate
cs.LG updates on arXiv.org

Identifiability and Estimation for Unlabeled Finite Mixtures under Marginal Independence

・arXiv:2606.07914v2 Announce Type: replace-cross Abstract: We study component recovery and mixing-matrix estimation from unlabeled finite mixtures whose observable distributions share the same latent components but have unknown mixing weights. ・The main identifying signal is marginal independence: each component is assumed to be independent on at least one coordinate pair, but no labels, clean component samples, or mix
cs.LG updates on arXiv.org

Identifiability and Stability of Generative Drifting in the Companion-Elliptic Kernel Family

・arXiv:2604.24196v4 Announce Type: replace-cross Abstract: A drifting model is a one-step generator trained by moving each sample along a field of kernel-weighted attraction toward data samples and repulsion between model samples; training halts once this field vanishes. ・The soundness of this scheme rests on two questions: whether a zero-field equilibrium guarantees agreement with the data distribution, and how the er
cs.LG updates on arXiv.org

In Silico Study for Optimizing Intensity and Focality Electrode Configurations for Directional DBS Under Uncertainty Using Metaheuristic L1L1 Method

・arXiv:2506.13452v3 Announce Type: replace-cross Abstract: Background and Objective: As Deep Brain Stimulation (DBS) advances toward directional leads and optimization-based current steering, selecting electrode contact configurations becomes complex. ・This study formulates configuration selection as an inverse mapping between target activation and electrode currents using metaheuristic L1-norm regularized L1-norm fitt
cs.LG updates on arXiv.org

In-context superposition: human-like working memory interference in large language models

・arXiv:2604.09670v3 Announce Type: replace Abstract: Intelligent systems must maintain and manipulate task-relevant information online to adapt to dynamic environments. ・This capacity, known as working memory, is fundamental to human reasoning. ・Yet, human working memory is strikingly limited, maintaining only three to four items in a brain with billions of neurons.
cs.LG updates on arXiv.org

Incremental Evaluation and Training in Relational Deep Learning

・arXiv:2608.13023v1 Announce Type: new Abstract: Relational Deep Learning (RDL) models multi-tabular databases as temporal heterogeneous graphs to enable end-to-end representation learning. ・However, prevailing RDL evaluation practices rely on static, single-episode dataset snapshots, overlooking the continuous, time-evolving nature of real-world databases. ・Consequently, current RDL benchmarks fail to capture how model
cs.LG updates on arXiv.org

Infrared Organization and Critical Cognitive Field Formation in Transformer Dynamics

・arXiv:2607.10923v4 Announce Type: replace Abstract: Large language models exhibit remarkable emergent behaviors, yet the physical mechanism governing their collective dynamics remains poorly understood. ・Cognitive Field Theory predicts that learning reorganizes the collective relaxation spectrum, thereby modifying memory self-energy, long-memory dynamics, and collective susceptibility through the infrared organization
cs.LG updates on arXiv.org

INSHAPE: Instance-Level Shapelets for Interpretable Time-Series Classification

・arXiv:2605.20088v2 Announce Type: replace Abstract: Discovering shapelets -- i.e., discriminative temporal patterns within time series -- has been widely studied to address the inherent complexity of time-series classification (TSC) and to make model decision-making processes more transparent. ・However, existing methods primarily focus on population-level shapelets optimized across the entire dataset, which leads to t
#LLMタグ

Inside Indeed’s Next Big Move: Automating the Recruitment Funnel with One AI Prompt

・Hello, tech and HR tech enthusiasts!As the landscape of recruitment rapidly shifts, Indeed is undergoing a massive transformation. ・It is evolving from a traditional job search engine into a fully automated end-to-end hiring platform.The biggest challenge in HR today is the operational friction between "finding a candidate" and "actually getting them to an interview." Recruiters spend endless hours drafting outreach m
Hugging Face Papers

Intern-S2-Preview: Scientific Agentic Foundation Model

Intern-S2-Preview: Scientific Agentic Foundation Model
cs.LG updates on arXiv.org

Intern-S2-Preview: Scientific Agentic Foundation Model

・arXiv:2608.13505v1 Announce Type: new Abstract: Scientific discovery increasingly requires AI systems that can reason over scientific evidence of heterogeneous modalities, interact with scientific tools and environments, and sustain progress across long task horizons. ・We present Intern-S2-Preview, a series of scientific agentic foundation models designed to support multimodal scientific understanding, reasoning, gene
cs.LG updates on arXiv.org

Interpretable Causal Discovery via Causal-Effect Constraints

・arXiv:2608.12640v1 Announce Type: new Abstract: Causal discovery aims to uncover the underlying causal relationships given data generated from a system. ・The goal, however, is not merely to predict causal edges given data, but also to be able to interpret and explain either observed or hypothesized phenomena, such as a particularly large causal effect. ・We consider this task of conditional causal discovery and cast it
cs.LG updates on arXiv.org

Intervention-Aware Clinical World Model for Post-Op Outcome Forecasting in Cardiology

・arXiv:2608.13518v1 Announce Type: new Abstract: Many clinical prediction models treat post-intervention outcomes as a one-step mapping from baseline measurements to a future endpoint. ・However, recovery after a procedure often unfolds as an irregular trajectory: clinical observations, medication changes, repeat interventions, and physiological measurements are recorded asynchronously and can change risk assessment ove
cs.LG updates on arXiv.org

Into the ORBIT for Time Series: Training Regimes for Foundation Models

・arXiv:2608.13262v1 Announce Type: new Abstract: Time series foundation models (TSFMs) have advanced primarily through architectural innovation, while training regimes for large-scale heterogeneous corpora remain under-explored. ・As a result, pre-training distributions are often poorly controlled with respect to domain imbalance, context requirements, prediction horizons, and missingness. ・We introduce ORBIT (Omni-Range
cs.LG updates on arXiv.org

Keep, Customize, or Exit: Default Design and Token Pricing in LLM Reasoning Services

・arXiv:2608.13315v1 Announce Type: cross Abstract: We study a large language model (LLM) service in which a provider chooses a per-token price and a default reasoning-token allocation, while a user may accept the default, customize the allocation, or exit. ・Larger allocations can improve accuracy but increase token cost and latency. ・We model this interaction as a Stackelberg game and derive the user's unique optimal cu
Hugging Face Papers

Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning

Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning
cs.LG updates on arXiv.org

Knowledge-guided Pattern Discovery via Coupled Tensor Factorizations

・arXiv:2608.13234v1 Announce Type: new Abstract: In order to understand complex systems such as the human metabolome or human brain, different sensing technologies are used, generating complex data. ・These datasets are often multiway, i.e., with more than two axes of variation such as a subjects by metabolites by time array. ・While tensor factorizations have successfully revealed interpretable patterns from such complex
AI News & Artificial Intelligence | TechCrunch

Kog is going deeper to squeeze more inference out of GPUs

・The idea that GPUs are poorly suited for agentic workflows may be a misconception, according to French startup Kog.
cs.LG updates on arXiv.org

Langevin dynamics for high-dimensional optimization: the case of multi-spiked tensor PCA

・arXiv:2408.06401v3 Announce Type: replace-cross Abstract: We study nonconvex optimization in high dimensions through Langevin dynamics, focusing on the multi-spiked tensor PCA problem. ・In this tensor estimation model, the goal is to recover a finite number of hidden signal vectors, or spikes, from noisy Gaussian tensor observations using maximum likelihood estimation. ・We characterize the number of samples required fo
cs.LG updates on arXiv.org

Large-scale Testing Global Optimization Methods with Black-box Adversarial Attacks

・arXiv:2608.13296v1 Announce Type: new Abstract: Existing global optimization benchmark suites are of a moderate size and are based on a small number of analytical functions that date back even to the 1970s. ・This causes a risk of biasing the development of global optimization methods. ・We argue that the tasks related to the black-box adversarial attack (BBAA) can serve as valuable global optimization benchmark in many-
cs.LG updates on arXiv.org

Latent On-Policy Self-Distillation

・arXiv:2608.13040v1 Announce Type: new Abstract: Enabling agents to learn from experience and internalize it into their policy has become a central problem in self-evolving AI. ・On-policy self-distillation (OPSD) offers an effective pathway by using a privileged self-teacher to provide dense supervision on the student's own trajectories; however, existing methods still rely heavily on designer-specified privileged arti
WIRED

Layla Sleep Coupon: Save Up to $600 in August 2026

・Upgrade your sleep setup with the latest Layla promo codes. ・Save on flippable mattresses, copper-infused pillows, and adjustable bases in August 2026.
cs.LG updates on arXiv.org

Learning Discrete Decisions for MIPs with Constraint-Aware Diffusion

・arXiv:2608.13079v1 Announce Type: new Abstract: This paper proposes a novel learning-based approach to approximately solve instances of mixed-integer optimization problems. ・These problems are computationally challenging, as they require jointly determining discrete and continuous decisions while satisfying complex combinatorial constraints. ・The proposed method relies on a graph-based generative diffusion model that l
cs.LG updates on arXiv.org

Learning the Mathematical Property for Designing Low Mutual Coherence Binary Sensing Matrices

・arXiv:2608.12982v1 Announce Type: new Abstract: In this research work, we are constructing the sensing matrix, which is essential for the success of the compressive sensing technique. ・We have chosen a learning-based technique for the construction of the sensing matrix. ・The novelty and uniqueness of the proposed technique is that it does not use any data set and also does not use a specific application.
cs.LG updates on arXiv.org

Learning Under Treatment-Induced Label Indeterminacy with Expert Annotations of Counterfactual Outcomes: A Case Study in Neurological Prognostication

・arXiv:2608.12477v1 Announce Type: new Abstract: Clinical prediction models are often developed as if the outcome of interest were cleanly observed for every patient. ・This assumption fails when treatment decisions make the clinically relevant outcome permanently unobservable. ・As a case study of this problem, we consider post-cardiac-arrest neurological prognostication using a cohort of 2,497 patients, including 1,429
#LLMタグ

LINE AI Friends|トーク品質改善!?久々に健太郎さんに話しかけてみた

LINE AI Friends|トーク品質改善!?久々に健太郎さんに話しかけてみた
cs.LG updates on arXiv.org

Liquidity-Based Audit of Algorithmic Trading Strategies

・arXiv:2606.29018v2 Announce Type: replace-cross Abstract: We show that net demand for liquidity by algo strategies is identifiable from its trade and price history alone, with no knowledge of its signal or optimization problem. ・An exact multi-period regret decomposition implies that the sign of this statistic classifies a linear strategy as a net liquidity consumer or provider, recovering the Kyle (1985) informed-tra
cs.LG updates on arXiv.org

LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure

・arXiv:2608.13545v1 Announce Type: cross Abstract: Modern language models are trained on heterogeneous web-scale text corpora. ・Consequently, studying knowledge and skill acquisition is difficult, as prior exposure to related content is hard to characterize. ・To address this challenge, we introduce LITTLECURRICULUM, a curated 88B-token pretraining corpus tailored to U.S.
Hugging Face Papers

LiveAnimate: Stable Long-Form Streaming Human Animation in Real-Time

LiveAnimate: Stable Long-Form Streaming Human Animation in Real-Time
cs.LG updates on arXiv.org

LLM-Assisted Dynamic Threat Analysis for Attacker-Reachable Software Weaknesses in Autonomous Vehicles

・arXiv:2608.13450v1 Announce Type: cross Abstract: Autonomous vehicles depend on large safety-critical software stacks, where weaknesses reachable from adversarial inputs may affect steering, braking, or other control decisions. ・Static analysis can identify candidate sites, but dynamically confirming exploitability requires executable test artifacts that are difficult to construct manually. ・We investigate whether larg
#LLMタグ

LLM-mediated extended cognition

・LLMを外部作業記憶・言語化器・反射板・仮説生成器として使い、人間との反復で思考そのものを長時間維持する 「頭の中にある、まだ問いですらない微弱な違和感を入力してください。それをモデルとの反復によって数時間維持すると、当初存在しなかった問いが生成される場合があります」 続きをみる
Hugging Face Papers

LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers

LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers
cs.LG updates on arXiv.org

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining

・arXiv:2608.12419v1 Announce Type: new Abstract: Large language models (LLMs) have achieved remarkable breakthroughs across various applications. ・However, their architectures remain inefficient in pretraining due to two main limitations: (i) self-attention lacks an explicit inductive bias for locality, leading to redundant modeling of sequence-internal local information; (ii) mixture-of-experts (MoE) implicitly couple
cs.LG updates on arXiv.org

LoRA-based Adaptation Alone Is Not Enough: Understanding the Limits of Foundation Models for Face Presentation Attack Detection

・arXiv:2608.09633v2 Announce Type: replace-cross Abstract: Face presentation attack detection (PAD) aims to reliably detect a wide range of presentation attacks. ・While PAD methods achieve strong performance within individual datasets, their performance degrades under cross-dataset evaluation. ・Variations in sensors or lighting conditions can reduce the effectiveness of detectors from near-perfect to nearly random.
Hugging Face Papers

LycheeMemory V2: Efficient Long-Term Memory for LLM Agents via Semantic Segment-Level Consolidation

LycheeMemory V2: Efficient Long-Term Memory for LLM Agents via Semantic Segment-Level Consolidation
cs.LG updates on arXiv.org

MAG: MAnifold Guided Semi-Supervised Multi-modal In-Context Learning

・arXiv:2608.12724v1 Announce Type: new Abstract: Few-shot in-context learning (ICL) with multi-modal large language models (MLLMs) enables task adaptation without parameter updates, but its performance is highly sensitive to the quality and coverage of the selected demonstrations. ・While unlabeled multi-modal data is abundant, it remains elusive how to exploit them for ICL. ・We propose MAG (MAnifold-Guided semi-supervis
cs.LG updates on arXiv.org

Manifold constrained steepest descent for smooth and closed-set optimization

・arXiv:2601.21487v2 Announce Type: replace-cross Abstract: We study minimization of smooth functions over feasible sets that have smooth embedded-manifold structure throughout or only on selected regions, using linear minimization oracles (LMOs) to determine search directions under user-chosen norms. ・Restricting an LMO to a tangent space, however, can require an iterative inner solve. ・We propose \emph{Manifold Constra
cs.LG updates on arXiv.org

MARCH: Scaling Recurrent Memory with Content-Routed State Anchors

・arXiv:2608.12435v1 Announce Type: new Abstract: Transformers owe much of their strong long-context retrieval capability to a token-level memory that grows with context length. ・This flexibility, however, incurs a quadratic computation complexity during training and a key--value cache that grows linearly during autoregressive inference. ・Recurrent alternatives offer efficient decoding by compressing the entire history i
The Verge

Mark Zuckerberg has an Instagzam

・Instagram's wordmark is iconic. ・Well, was iconic. ・Apparently Instagram thought it looked old, so the company rolled out a new one this week.
Hugging Face Papers

Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus

Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus
cs.LG updates on arXiv.org

MatchMiner-AI: Open-source, Privacy-preserving Cancer Clinical Trial Matching using Artificial Intelligence

・arXiv:2412.17228v4 Announce Type: replace-cross Abstract: Background: Clinical trials are essential to advancing cancer treatments, but fewer than 10% of adults with cancer enroll in therapeutic trials. ・Open-source AI trial matching tools could democratize access to trial options. ・Methods: We created MatchMiner-AI, co-developed with practicing clinical oncologists and trained on synthetic electronic health record (EH
Zennの「大規模言語モデル」のフィード

MCPのToolを増やしたらAgentが迷い始めた?「使えるTool」を増やす前に確認したい5項目

・MCP Serverを追加すると、AI Agentにできることが増えます。 ・例えば、 GitHubからIssueを取得する Databaseを検索する 社内Documentを読む Browserを操作する SlackやNotionから情報を取得する といった操作を、Agentから呼び出せるようになります。 ・そこで、 このToolも使えそう これもMCPでつなごう と追加していく。
MarkTechPost

Meet Needle 2: An Open 45M-Parameter Tool-Calling Model That Ships as a 14MB Binary and Runs a Full Session in 28MB of RAM

・Cactus Compute released Needle 2, an open 45M-parameter model for tool calling, device use, and structured extraction. ・The full model is a single 14MB binary that runs a session in about 28MB of RAM. ・It leads both Seal-Tools splits while targeting hardware with no GPU and no NPU.
cs.LG updates on arXiv.org

MergeOver: Post-Training Token Merging for Recursive Vision Transformers

・arXiv:2608.13141v1 Announce Type: cross Abstract: Vision Transformers (ViTs) demonstrate exceptional performance in computer vision but suffer from large parameter counts and quadratic computational complexity, severely limiting their deployment on resource-constrained edge hardware. ・While recursive weight-sharing reduces parameter counts and token merging mitigates computational and memory bottlenecks, integrating t
AI News & Artificial Intelligence | TechCrunch

Meta’s ‘open’ AI, and a $250M deal gone very wrong 

・Meta released Glimmer this week, an open-weight AI model anyone can download and run on their own hardware — a contrast to Muse Spark, the company’s more powerful model that stays locked behind its own APIs. ・The release landed alongside a letter from Mark Zuckerberg arguing AI should be “for everyone” rather than controlled by a handful of labs, but as Equity’s […]
cs.LG updates on arXiv.org

Minimax and Adaptive Covariance Matrix Estimation under Differential Privacy

・arXiv:2603.19703v2 Announce Type: replace-cross Abstract: Estimating covariance matrices is fundamental to a wide range of statistical applications. ・This paper studies minimax and adaptive estimation of high-dimensional covariance matrices under $\rho$-zero-concentrated differential privacy ($\rho$-zCDP) over three nested classes: the pointwise-decay class $\mathcal{H}_\alpha$, the row-tail class $\mathcal{G}_\alpha$
cs.LG updates on arXiv.org

MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification

・arXiv:2608.13463v1 Announce Type: cross Abstract: Modern image classification models excel when trained on single task-specific datasets but often struggle to generalize across domains and difficulty levels. ・We propose ARMDIL, an Adaptive Router for Multi-Domain Image classification with LLMs. ・ARMDIL is an ensemble that uses a multimodal large language model (MLLM) agent to dynamically route each image to the most su
cs.LG updates on arXiv.org

Momentum as Residual-Driven Multiplier Correction for Deep Learning Optimization

・arXiv:2608.12925v1 Announce Type: new Abstract: Momentum-based optimizers are widely used in modern deep learning, yet the relations among momentum recursion, update geometry, and acceleration remain only partially understood. ・We develop an $\textbf{A}$DMM-$\textbf{I}$nspired $\textbf{M}$omentum (AIM) framework based on residual-penalty variable splitting, which interprets momentum as a multiplier-like correction dri
cs.LG updates on arXiv.org

Multi-AUV Ad-hoc network-based Target Tracking: A Value Gradient Guidance Multi-Agent Diffusion Reinforcement Learning Approach

・arXiv:2608.12436v1 Announce Type: new Abstract: Multi-AUV ad-hoc network-based target tracking requires networked autonomous underwater vehicles (AUVs) to cooperatively track maneuvering targets under constrained acoustic communication, dynamic topology, and uncertain ocean disturbances. ・Although multi-agent reinforcement learning (MARL) enables decentralized coordination through centralized training, existing method
cs.LG updates on arXiv.org

Multi-perspective Imbalance-Conscious 6G Beamforming Optimization and Performance

・arXiv:2608.12929v1 Announce Type: new Abstract: The study presents a systematic machine learning (ML) study of 6G-IoT beamforming optimization (6GBO) using supervised and unsupervised approaches. ・We compared the predictive power of network, environmental, device, and vision feature groups for 6GBO. ・Additionally, it addressed other unsupervised perspectives that can enhance 6GBO, including clustering network scenarios
cs.LG updates on arXiv.org

Multiview Representation Learning via Distributed Joint Latent Space Structuring

・arXiv:2504.18455v2 Announce Type: replace-cross Abstract: We study distributed multiview representation learning, a problem in which $K$ clients each observe a distinct but possibly statistically correlated view. ・The clients independently extract local representations from their views, which are then used by a central decoder for joint target estimation. ・One central difficulty is that, since the clients are not allow
LLMタグが付けられた新着記事 - Qiita

NeMo Switchyardをローカル(WSL2 + Ollama)で検証、ルーティングより先にモデルの安定性の限界にぶつかる

・NVIDIA NeMo Switchyard v0.2.0 実機検証記 本記事は個人環境での検証に基づく個人的な備忘録です。設定値やパスは一般化して記載しているため、実際の環境に読み替えてください。技術的な確認は Sonnet 5 と、記事の執筆は Opus 5 と...
cs.LG updates on arXiv.org

Neural Quadratic Forms: A Unified Minimal Model for Sudden Learning and Scaling Laws

・arXiv:2608.13335v1 Announce Type: new Abstract: Neural networks trained by gradient descent on a smooth cost function can nevertheless learn in steps: the cost holds on long plateaus and then drops abruptly. ・Meanwhile, training losses instead follow smooth power laws. ・Variants of both behaviors occur in architectures with very different microscopic structures, which is the signature of a few relevant collective varia
cs.LG updates on arXiv.org

Noise as a Probe: Membership Inference Attacks on Diffusion Models Leveraging Initial Noise

・arXiv:2601.21628v2 Announce Type: replace-cross Abstract: Diffusion models have achieved remarkable progress in image generation, but their increasing deployment raises serious concerns about privacy and copyright. ・In particular, fine-tuned models are highly vulnerable, as they are often fine-tuned on small and private datasets. ・Membership inference attacks (MIAs) are used to assess privacy risks by determining wheth
cs.LG updates on arXiv.org

Non-Degenerate Risk Certification for Automated Security Decisions: A Decision-Contract Theory with ATT\&CK-Aligned Triage as a Worked Instance

・arXiv:2608.12444v1 Announce Type: cross Abstract: An unconditional risk bound on automated decisions can be satisfied without automating anything, since a selector that never acts drives the bound to zero. ・We show this is structural: any risk certificate is defined over a decision contract, the inputs a system acts on plus the semantic relation under which an output counts correct, and weakening either hides base-cla
cs.LG updates on arXiv.org

Novel Knowledge-Guided Generative Methods for Synthetic Transcriptomic Data

・arXiv:2608.13256v1 Announce Type: new Abstract: As biomedical research increasingly relies on data-intensive tools, the quality and utility of datasets are critical. ・Challenges such as imbalances, biases, and ethical or legal constraints often limit access to high-quality data. ・Synthetic data generation can help overcome these limitations.
WIRED

Office Depot Coupons: Save With Promo Codes in August 2026

・From furniture and ink to professional printing services, use an Office Depot discount code to maximize your savings on every workspace essential.
Hugging Face Papers

OmniScientist: An Omni-Modal Omni-Discipline AI Scientist

OmniScientist: An Omni-Modal Omni-Discipline AI Scientist
cs.LG updates on arXiv.org

On the global feature importance for interpretable and trustworthy heat demand forecasting

・arXiv:2608.13039v1 Announce Type: new Abstract: The paper introduces the ante-hoc Explainable AI methodology to assess the global feature importance of the Machine Learning models used for heat demand forecasting in intelligent control of District Heating Systems, with motivation to facilitate their interpretability and trustworthiness, hence addressing the challenges related to adherence to communal standards, custo
cs.LG updates on arXiv.org

On the Structural Limits of Machine Learning Decision Systems: An Information-Theoretic, Interaction-Based, and Stochastic-Dynamical Perspective

・arXiv:2608.13510v1 Announce Type: cross Abstract: Machine learning procedures are commonly evaluated in terms of predictive accuracy and computational efficiency. ・However, their achievable performance is fundamentally constrained by structural properties of the underlying data-generating process, which are formalized in terms of informational bounds. ・In this work we examine intrinsic limits of data-driven decision sy
cs.LG updates on arXiv.org

Online Correlation Clustering: Simultaneously Optimizing All $\ell_p$-norms

・arXiv:2510.15076v2 Announce Type: replace Abstract: The $\ell_p$-norm objectives for correlation clustering present a fundamental trade-off between minimizing total disagreements (the $\ell_1$-norm) and ensuring fairness to individual nodes (the $\ell_\infty$-norm). ・Surprisingly, in the offline setting it is possible to simultaneously approximate all $\ell_p$-norms with a single clustering. ・Can this powerful guarante
cs.LG updates on arXiv.org

Online Inference for Quantile Temporal Difference Learning in Distributional Reinforcement Learning

・arXiv:2608.12973v1 Announce Type: cross Abstract: In this paper, we study how to perform statistical inference for quantile temporal difference learning (QTD) in distributional reinforcement learning. ・Assuming access to a generative model, we first establish functional central limit theorems for both synchronous and asynchronous QTD, which show that the averaged iterates of QTD converge weakly to a rescaled Brownian
cs.LG updates on arXiv.org

Optimizing Likelihoods via Mutual Information: Bridging Simulation-Based Inference and Bayesian Optimal Experimental Design

・arXiv:2502.08004v2 Announce Type: replace-cross Abstract: Simulation-based inference (SBI) is a method to perform inference on a variety of complex scientific models with challenging inference (inverse) problems. ・Bayesian Optimal Experimental Design (BOED) aims to efficiently use experimental resources to make better inferences. ・Various stochastic gradient-based BOED methods have been proposed as an alternative to Ba
WIRED

People Are ‘Marrying’ Chatbots. These Lawmakers Want to Stop Them

・Human-AI marriages are not currently recognized by US law. ・Some Republican state policymakers are drafting legislation to keep it that way.
cs.LG updates on arXiv.org

Performance-Carbon Trade-Offs across Architectural Biases in Shear Flow Forecasting

・arXiv:2509.24517v3 Announce Type: replace Abstract: Development of modern deep learning methods has been driven primarily by the push for improving model efficacy (accuracy metrics), leading to large-scale models that require massive computational resources and result in considerable carbon footprint across the model lifecycle. ・In this work, we explore how architectural biases, specifically a model's receptive field
cs.LG updates on arXiv.org

Personalized Scorer Modeling: A Learning-Based Framework for Deriving Robust Sleep Stage Labels from Multiple Experts

・arXiv:2608.12446v1 Announce Type: new Abstract: Sleep stage classification is important for the diagnosis and management of sleep disorders, yet most automatic staging studies evaluate models against a single reference hypnogram despite known inter-scorer variability. ・This study investigates whether multi-scored datasets can be used to construct more reliable reference labels from the collective behavior of multiple
cs.LG updates on arXiv.org

Perturbation-based Regional Interpretability through Subtraction Mapping (PRISM): naming-error dissociations in language models and post-stroke aphasia

・arXiv:2608.12717v1 Announce Type: new Abstract: Mechanistic interpretability of large language models lacks spatially resolved, falsifiable tools for testing whether internal components are specialized for distinct cognitive operations. ・We adapt subtraction analysis, the standard framework of human neuroimaging, from biological brains to perturbed transformers, and apply the same logic to both substrates in parallel.
cs.LG updates on arXiv.org

Physics-Informed Laplace Neural Operator for Solving Partial Differential Equations

・arXiv:2602.12706v2 Announce Type: replace Abstract: Neural operators have emerged as fast surrogate solvers for parametric partial differential equations (PDEs). ・However, purely data-driven models often require extensive training data and can generalize poorly, especially in small-data regimes and under unseen (out-of-distribution) input functions that are not represented in the training data. ・To address these limita
Hugging Face Papers

PixSDS: Why Latent SDS Makes Noisy Pixels

PixSDS: Why Latent SDS Makes Noisy Pixels
Hugging Face Papers

PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives

PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives
cs.LG updates on arXiv.org

Position: Reasoning is a Learnable Rule-Based Process

・arXiv:2608.12325v1 Announce Type: cross Abstract: Autonomous reasoning is among the most scientifically and economically motivating topics in AI today. ・Historically the purview of symbolic AI, recent advances have mainly emerged from deep probabilistic generative models. ・Despite immense interest and rapid progress, the generative AI community has not clearly converged on operational definitions for reasoning and ofte
cs.LG updates on arXiv.org

Post-Hoc Uncertainty-Aware Explanations for Deployed Power Quality Disturbance Classifiers via Laplace Approximation

・arXiv:2604.13658v2 Announce Type: replace Abstract: Deep learning classifiers achieve high accuracy in power quality disturbance (PQD) recognition, but existing explanation methods return a single deterministic attribution map and provide no measure of its reliability. ・This paper develops a post-hoc Bayesian explanation (B-explanation) method for trained PQD classifiers. ・A computationally efficient Laplace approximat
cs.LG updates on arXiv.org

Predicting When Random Low-Dimensional Reparameterizations Train Neural Networks

・arXiv:2608.12597v1 Announce Type: new Abstract: Neural networks can often be trained or fine-tuned through random low-dimensional reparameterization, where a small latent vector is mapped into a full parameter update by a frozen random map. ・This raises a practical question: how large must the latent search space be to reach a low-loss region? ・We first express the known accessibility transition in an equivalent conic
cs.LG updates on arXiv.org

Predictive Allostatic Organization in Recurrent and Spiking Agents Under Partial Observability

・arXiv:2608.11506v1 Announce Type: cross Abstract: Adaptive behavior under partial observability depends on internal organization that carries information beyond the current observation. ・Drawing on Barrett and Miller's account of categorization as predictive, compressive, functionally organized, and allostatically constrained, we test whether recurrent and spiking agents develop internal states with corresponding comp
cs.LG updates on arXiv.org

Prof-K: Probabilistic One-Pass Filtering for Efficient Top-k Selection

・arXiv:2608.12573v1 Announce Type: new Abstract: Top-k selection is a fundamental computational primitive with applications spanning databases, information retrieval, signal processing, and modern machine learning workloads, including sparse activations and attention pruning. ・As data sizes grow, existing approaches become inefficient: exact methods incur high memory and compute overhead, while approximate methods ofte
cs.LG updates on arXiv.org

ProME: Prototype-Margin Environments with Repair-Aware Selection for Group-Robust Learning

・arXiv:2608.13190v1 Announce Type: new Abstract: Group-robust learning is crucial for maintaining accuracy on rare subpopulations when training-group labels are unavailable. ・However, existing methods often infer environments from a separate reference model and select representations before fitting the classifier used at deployment, leaving both decisions misaligned with the deployed predictor. ・In this work, we formula
WIRED

Pura Promo Codes: $20 Off August 2026

・Enhance your home's ambiance for less with active Pura promo codes. ・Find savings on smart diffusers, exclusive scents, and subscribe & save offers.
Zennの「大規模言語モデル」のフィード

Qwen3-4Bのconfig.jsonからGPUメモリを見積もる:Weight・KV Cache・GQAを手計算

・はじめに LLMをGPUへ配置するとき、最初に知りたいのは「このモデルはGPUに載るのか」です。 ・しかし、モデル名にある「4B」だけを見て、 4B parameters × BF16の2 byte = 約8 GB と計算して終わりにすると、推論時に必要なメモリを正しく見積もれません。 ・推論では、少なくとも次のメモリを区別する必要があります。
cs.LG updates on arXiv.org

RadarGen: Automotive Radar Point Cloud Generation from Cameras

・arXiv:2512.17897v2 Announce Type: replace-cross Abstract: We present RadarGen, a diffusion model for synthesizing realistic automotive radar point clouds from multi-view camera imagery. ・RadarGen adapts efficient image-latent diffusion to the radar domain by representing radar measurements in bird's-eye-view form that encodes spatial structure together with radar cross section (RCS) and Doppler attributes.
cs.LG updates on arXiv.org

ReconSpan: Reconstruction-Guided Adaptive Latent Tokenization

・arXiv:2608.12756v1 Announce Type: cross Abstract: Adaptive latent tokenization maps a fine-grained input to a shorter sequence of continuous representations associated with input-dependent spans. ・We introduce ReconSpan, which divides text into chunks that a backward decoder can reconstruct from a single contextual prefix code and retains one such code as the latent token for each chunk. ・The reconstruction criterion i
cs.LG updates on arXiv.org

Reduced Matrix Multiplication: Input-Adaptive Matrix-Product Reduction for LLM Inference

・arXiv:2608.13426v1 Announce Type: new Abstract: Transformer-based language models achieve strong performance but incur substantial inference cost due to repeated high-dimensional matrix multiplications. ・We propose Reduced Matrix Multiplication (RMM), a training-free, input-adaptive inference method that reduces Transformer matrix products by selecting informative slices along their contraction dimensions, without mod
cs.LG updates on arXiv.org

Reduced Order Modeling for Tsunami Forecasting with Bayesian Hierarchical Pooling

・arXiv:2512.19804v2 Announce Type: replace Abstract: Reduced-order models (ROMs) can represent spatiotemporal processes in significantly fewer dimensions and can often be solved many orders of magnitude faster than their governing partial differential equations (PDEs). ・For example, proper orthogonal decomposition yields a ROM in which the state is represented as a low-dimensional linear combination of fixed spatial mo
cs.LG updates on arXiv.org

Regulatory Approval Is Not Enough: Gaps in Trustworthy AI Reporting in FDA-Cleared Medical Devices

・arXiv:2608.12360v1 Announce Type: cross Abstract: Background: AI/ML-enabled medical devices are increasingly deployed in healthcare under evolving regulatory frameworks. ・As these systems become more integrated into clinical decision-making, there is growing expectation that they demonstrate key dimensions of trustworthy AI to support clinician, patient, and public trust. ・Whether publicly available regulatory document
cs.LG updates on arXiv.org

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems

・arXiv:2606.00367v2 Announce Type: replace Abstract: Reinforcement learning with scalar rewards is widely used for aligning machine-learning systems with user preferences. ・But, pairwise preferences are often more natural for users to specify than scalar rewards, and they express certain goals that scalar rewards cannot. ・Methods for reinforcement learning with pairwise preferences have thus received growing interest.
cs.LG updates on arXiv.org

Represent, Then Generate: Multimodal-Conditioned Time-Series Generation under Irregular Missingness

・arXiv:2608.12592v1 Announce Type: new Abstract: Continuous physiological time series underpin modern clinical monitoring, yet many of the most informative signals are invasive, expensive, or simply unavailable for a given patient. ・Conditional generation offers a remedy: an absent signal can be synthesized from co-recorded signals and routine clinical variables. ・Existing generators, however, are built around a single
cs.LG updates on arXiv.org

Revisiting Overestimation Bias Problem of Q-learning: Settling Large Discrete Action Space via Action Intersection

・arXiv:2608.12912v1 Announce Type: new Abstract: This paper considers the overestimation bias problem of Q-learning in the setting of a large action space, for the purpose of relieving the bottleneck of existing methods. ・We find that the large action space increases the randomness in Q-value estimation. ・The randomness makes two paradigms that drive the major literature on the overestimation problem have their own bott
Hugging Face Papers

RibAssist 3D: Biplanar Rib-Fracture Detection, Addressing, and Selective 3D Localization from CT-Derived Projections

RibAssist 3D: Biplanar Rib-Fracture Detection, Addressing, and Selective 3D Localization from CT-Derived Projections
cs.LG updates on arXiv.org

Robust Checkpoint Selection for Multimodal LLMs via Agentic Evaluation and Stability-Aware Ranking

・arXiv:2605.18852v2 Announce Type: replace Abstract: Selecting a final checkpoint for multimodal large language models (MLLMs) is challenging when late-stage candidates are closely matched and downstream evaluation signals are noisy. ・Small observed differences can be comparable to variability introduced by finite evaluation samples, LLM-based judges, and ambiguous multimodal evidence, while validation loss may not ide
cs.LG updates on arXiv.org

Robust data-driven discovery of fractional differential equations via weak formulations and Pareto-based subset selection

・arXiv:2608.12879v1 Announce Type: new Abstract: Fractional partial differential equations describe nonlocal dynamics, but discovering them from noisy data is difficult because fractional differentiation amplifies high-frequency measurement noise and the derivative orders are unknown. ・We propose Weak-Pareto, which combines an adjoint-consistent weak formulation of fractional terms with Pareto-based subset selection ov
cs.LG updates on arXiv.org

RoutePack: Expert Placement and Attention-Aware Data Packing for MoE Reinforcement Learning

・arXiv:2608.12146v1 Announce Type: cross Abstract: Training Mixture-of-Experts (MoE) models for reinforcement learning (RL) couples two load-balancing problems: sequence composition determines dense attention work in each data-parallel microbatch, while token routing determines sparse expert work on expert-parallel ranks. ・Optimizing either alone can shift the bottleneck to the other. ・In MoE RL, rollout-time routing re
cs.LG updates on arXiv.org

SAF3R: Dynamic Sparse Attention for Feed-Forward 3D Reconstruction Transformers

・arXiv:2607.03612v2 Announce Type: replace-cross Abstract: Feed-forward 3D reconstruction (F3R) transformers have recently achieved remarkable success. ・However, scaling them to long image sequences remains challenging, as the quadratic complexity of cross-view global attention quickly becomes the dominant computational bottleneck. ・While recent efforts attempt to improve efficiency through compressed or sparse attentio
cs.LG updates on arXiv.org

Safe Exploration via Policy Priors

・arXiv:2601.19612v4 Announce Type: replace Abstract: Safe exploration is a key requirement for reinforcement learning (RL) agents to learn and adapt online, beyond controlled (e.g. ・simulated) environments. ・In this work, we tackle this challenge by utilizing suboptimal yet conservative policies (e.g., obtained from offline data or simulators) as priors.
cs.LG updates on arXiv.org

SAFE-SVD: Sensitivity-Aware Fidelity-Enforcing SVD for Physics Foundation Models

・arXiv:2605.17985v2 Announce Type: replace Abstract: We propose a new method for compressing physics foundation models (PFMs) which is a new trend in AI for Science. ・While model compression is essential for reducing memory use and accelerating inference in large foundation models, it remains under-explored for PFMs, where preserving physical fidelity is crucial. ・The challenge lies in the functional nature of physics d
cs.LG updates on arXiv.org

Sampling Luck Masquerades as Allocation Gain: Auditing Test-Time Budget Allocation for Neural Combinatorial Optimization

・arXiv:2608.13087v1 Announce Type: new Abstract: Neural combinatorial optimization (NCO) solvers report the best of many sampled solutions per instance, and the sample count is, by convention, identical for every instance. ・Whether a non-uniform allocation of a fixed total budget would buy anything has not been measured. ・We measure it, and we audit the measurement itself.
cs.LG updates on arXiv.org

Scaling Automatic Research Agents via World Models

・arXiv:2608.12564v1 Announce Type: new Abstract: Automating empirical research is a long-standing direction of AI. ・Recent automatic research (AutoResearch) agents bring this goal within reach, as modern LLMs show the capability to independently implement solutions and learn from the execution outcomes. ・Behind these gains, post-training (especially RL) plays a central role.
cs.LG updates on arXiv.org

SDS-LoRA: Overcoming Anisotropic Gradient Scaling in Low-Rank Adaptation

・arXiv:2606.16454v2 Announce Type: replace Abstract: Low-Rank Adaptation (LoRA) enables efficient adaptation of large pretrained models to downstream tasks by parameterizing weight updates with low-rank matrices. ・In this paper, we investigate the limitations of the LoRA parameterization from a geometric perspective. ・Specifically, we show that when a full fine-tuning gradient is backpropagated to the low-rank matrices,
cs.LG updates on arXiv.org

SEAR: Sample Efficient Action Chunking Reinforcement Learning

・arXiv:2603.01891v2 Announce Type: replace Abstract: Action chunking improves exploration and accelerates value propagation in long-horizon reinforcement learning, but naively applying off-policy methods to the temporally extended action space at reduced decision frequency offsets these gains, leading to poor sample efficiency. ・Existing online action chunking methods address these issues through computationally expens
cs.LG updates on arXiv.org

SeBA: Semi-supervised few-shot learning via Separated-at-Birth Alignment for tabular data

・arXiv:2605.08519v2 Announce Type: replace Abstract: Learning from scarce labeled data with a larger pool of unlabeled samples, known as semi-supervised few-shot learning (SS-FSL), remains critical for applications involving tabular data in domains like medicine, finance, and science. ・The existing SS-FSL methods often rely on self-supervised learning (SSL) frameworks developed for vision or language, which assume the
cs.LG updates on arXiv.org

Security and Detectability Analysis of Unicode Text Watermarking Methods against Large Language Models

・arXiv:2512.13325v2 Announce Type: replace-cross Abstract: Securing digital text is becoming increasingly relevant due to the widespread use of large language models. ・Individuals' fear of losing control over data when it is being used to train such machine learning models or when distinguishing model-generated output from text written by humans. ・Digital watermarking provides additional protection by embedding an invis
cs.LG updates on arXiv.org

Self-Localizing MIMO Beam Mapping for Intelligent Open RAN with Continuously Evolving Channel Memory

・arXiv:2511.17007v2 Announce Type: replace-cross Abstract: Open and intelligent radio access networks (RANs) envisioned for 6G require accurate and reusable wireless channel knowledge for intelligent inference and control. ・However, full-dimensional channel state information (CSI) and accurate location labels are difficult to acquire and maintain across open and multi-vendor deployments. ・This paper develops a self-loca
cs.LG updates on arXiv.org

Simulation-to-real transfer learning for infrared spectroscopic chemical sensing and analysis from molecules to complex samples

・arXiv:2608.13341v1 Announce Type: new Abstract: Infrared (IR) spectroscopy is widely used for chemical sensing, but extracting reliable chemical information from spectra remains challenging. ・Conventional interpretation is labor-intensive, relies on prior knowledge and reference spectra, and is difficult to scale, whereas most machine-learning methods are tailored to individual tasks or datasets, require large labeled
cs.LG updates on arXiv.org

Sinkhorn Linearization and the Spectral Proxy: Unifying the Statistical and Algorithmic Theory of Feature-Parameterized Inverse Optimal Transport via a Single Spectral Sandwich

・arXiv:2608.13201v1 Announce Type: cross Abstract: We develop the statistical and algorithmic theory of inverse optimal transport (IOT) under the feature-parameterized cost C_theta(i,j) = -theta^T phi(i,j). ・The core technical contribution is the Sinkhorn linearization -- the implicit-function sensitivity of the entropic OT plan to the cost -- together with its spectral proxy, a formula that is spectrally exact yet geo
Hugging Face Papers

SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models

SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models
Zennの「大規模言語モデル」のフィード

Skillか、Subagentか——スキル95個の環境で固まった使い分け基準

・結論から 読み込ませたいならSkill、隔離したいならSubagent。 ・迷ったらこの一行で決めています。 ・私の環境にはスキルが95個、エージェント定義が27個あります。数ヶ月かけて両方を作り続けた結果、「どちらで作るべきか」の判断は最終的に上の一行に収束しました。この記事では、その基準に至った実測データ(スキル95個の使用記録・Subagent側で踏んだ罠)と、判断フローを共有します。
cs.LG updates on arXiv.org

SpaRRTa: A Synthetic Benchmark for Evaluating Spatial Intelligence in Visual Foundation Models

・arXiv:2601.11729v2 Announce Type: replace-cross Abstract: Visual Foundation Models (VFMs), such as DINO and CLIP, excel in semantic understanding of images but exhibit limited spatial reasoning capabilities, which limits their applicability to embodied systems. ・As a result, recent work incorporates some 3D tasks (such as depth estimation) into VFM training. ・However, VFM performance remains inconsistent across other s
cs.LG updates on arXiv.org

Sparse Orthogonal Regression Technique: A Spectral Framework for Equation Discovery, Approximation, and Integration

・arXiv:2608.13504v1 Announce Type: new Abstract: We develop the Sparse Orthogonal Regression Technique (SORT), a sparse spectral framework for learning orthonormal-basis expansions from noisy and irregularly sampled data. ・SORT estimates expansion coefficients directly from observations using L1-regularized regression, avoiding explicit quadrature or analytic inner-product evaluation. ・The central application is data-dr
Hugging Face Papers

Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence

Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence
Hugging Face Papers

Specification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code review

Specification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code review
cs.LG updates on arXiv.org

SpinCastML an Open Decision-Making Application for Inverse Design of Electrospinning Manufacturing: A Machine Learning, Optimal Sampling and Inverse Monte Carlo Approach

・arXiv:2602.09120v2 Announce Type: replace Abstract: Electrospinning is a powerful technique for producing micro to nanoscale fibers with application specific architectures. ・Small variations in solution or operating conditions can shift the jet regime, generating non Gaussian fiber diameter distributions. ・Despite substantial progress, no existing framework enables inverse design toward desired fiber outcomes while int
cs.LG updates on arXiv.org

SSPO: Structure-Aware Similarity-Weighted Preference Optimization for Neural Combinatorial Optimization

・arXiv:2608.12443v1 Announce Type: cross Abstract: Neural combinatorial optimization (NCO) relies on parallel solution sampling for training, yet existing methods fail to fully exploit the rich information latent in a co-sampled solution group. ・Preference-optimization methods anchor on the single best solution and discard fine-grained quality and structural signal from all other peers-a failure we term gradient signal
cs.LG updates on arXiv.org

Statistical Properties of Robust Learning under Distributional Shifts

・arXiv:2608.13133v1 Announce Type: cross Abstract: Distributional shifts arise when the target deployment environment differs from the source environment that generated the training data. ・Robust learning frameworks such as Distributionally Robust Optimization (DRO) and Robust Satisficing (RS) aim to address this challenge, yet their finite-sample guarantees under such shifts, and their systematic comparison, remain un
cs.LG updates on arXiv.org

SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries

・arXiv:2608.12654v1 Announce Type: cross Abstract: Long-running LLM agents act through tools, and a single step can send an email, merge a pull request, or wire a payment. ・The steering decision is the pre-commit choice at that boundary: proceed, or hold for human or policy review. ・We introduce SteerBench-Work, an incident-anchored, bidirectional benchmark for that decision in workplace agents across developer operatio
cs.LG updates on arXiv.org

Stochastic Neural Networks for Quantum Devices

・arXiv:2602.22241v2 Announce Type: replace-cross Abstract: This work presents a formulation to express and optimize stochastic neural networks as quantum circuits in gate-based quantum computing. ・Motivated by a classical perceptron, stochastic artificial neurons are introduced and combined into a quantum neural network. ・The Kiefer-Wolfowitz algorithm in combination with simulated annealing is used for training the net
cs.LG updates on arXiv.org

Structure-preserving uncertainty quantification for GENERIC dynamics

・arXiv:2608.12624v1 Announce Type: new Abstract: Structure-preserving machine learning embeds physical structure directly into model architectures, yet uncertainty quantification (UQ) for such hard-constrained models remains limited because standard UQ methods may violate the encoded admissibility conditions, require architectural modifications, or impose substantial computational costs. ・In this work, we propose Struc
stat.ML updates on arXiv.org

Supervised Mixed-Frequency Learning for Macro-Financial Forecasting When Factors are Weak

・arXiv:2608.12589v1 Announce Type: cross Abstract: Factor-MIDAS regressions forecast a low-frequency target by extracting common factors from a large panel of high-frequency predictors via principal component analysis (PCA). ・While PCA mitigates the curse of dimensionality, it relies on factor pervasiveness, an assumption often violated when factors are weak, as is common in macro-financial forecasting. ・We propose SsPC
cs.LG updates on arXiv.org

Sustaining Plasticity via Learnable Wavelet Activations in Continual Learning

・arXiv:2608.12874v1 Announce Type: new Abstract: Plasticity loss has emerged as a critical challenge in continual learning that significantly hinders the acquisition of sequential tasks. ・While optimizing activation designs offers a potential solution, current fixed-form functions suffer from an inherent spectral bias towards low-frequency variations, whereas learnable variants permit unconstrained updates that induce
cs.LG updates on arXiv.org

Symmetry-Breaking De Novo Crystal Generation via Markovian Jump Diffusion

・arXiv:2608.13457v1 Announce Type: new Abstract: Generating crystals has recently attracted significant interest due to their broad applications in materials science. ・However, existing generative models struggle to produce complete crystallographic specifications, limiting their ability to capture global symmetry and structural dependencies. ・In particular, current state-of-the-art approaches generate crystals only up
cs.LG updates on arXiv.org

Synthetic Persona Pretraining: Alignment from Token Zero

・arXiv:2608.13482v1 Announce Type: new Abstract: As language-model-based AI is increasingly deployed in autonomous settings, aligning its goals and values with those of humans becomes critical. ・Today, alignment, and the assistant identity itself, are typically introduced only after pretraining, once behavioral priors are already established. ・This can make values a thin overlay, rather than deeply rooted, and facilitat
cs.LG updates on arXiv.org

TabH2O: A Unified Foundation Model for Tabular Prediction

・arXiv:2605.18383v2 Announce Type: replace Abstract: We present TabH2O, a foundation model for tabular data that performs classification and regression in a single forward pass via in-context learning. ・TabH2O builds on the TabICL architecture with several key modifications: (1) unified training, a single model handles both classification and regression via a dual-head architecture, eliminating the need for separate mo
cs.LG updates on arXiv.org

TabSOM: A tabular-to-image encoding method based on self-organizing maps

・arXiv:2608.13513v1 Announce Type: cross Abstract: Tabular-to-image methods have emerged as novel approaches to leverage the high predictive performance of convolutional neural networks and vision transformers. ・They convert tabular data into image representations, mapping each feature at a fixed pixel location derived from a dimensionality-reduction method (e.g., t-SNE, UMAP, PCA). ・However, they encode only the margin
Hugging Face Papers

TailBooster: A Dual-Layer Generative Framework for Extreme Value Augmentation with Operational Validity Enforcement

TailBooster: A Dual-Layer Generative Framework for Extreme Value Augmentation with Operational Validity Enforcement
cs.LG updates on arXiv.org

TANGCO: Learning Topology-Aware Capacity Allocation for Overload-driven Cascading Failures

・arXiv:2608.13212v1 Announce Type: new Abstract: Networked systems, from power grids to traffic networks and cloud clusters, carry loads across nodes with limited capacity. ・A node whose load exceeds its capacity fails and sheds its load onto its neighbors, which can trigger a system-wide cascade. ・We study how to allocate a fixed capacity budget across nodes to resist these cascades under local load redistribution.
WIRED

Tech Visionary Says the Big AI Labs Don’t Get What People Want

・Tim O’Reilly built a publishing empire that AI is helping to destroy. ・Yet he loves AI—as long as it’s open source.
WIRED

The 4 Best Planners of 2026: Roterunner, Hobonichi, Cloth & Paper

・If digital calendars are leaving you lacking, these WIRED-tested paper agendas and notebooks could change your life.
cs.LG updates on arXiv.org

The Boolean Power of ReLU

・arXiv:2608.12617v1 Announce Type: new Abstract: We prove that, on finite simple undirected graphs equipped with a single Boolean node feature, the Boolean queries expressible in $\Sigma$-MPLang, for any collection $\Sigma$ of eventually constant activation functions and with arbitrary real coefficients, form a strict subclass of the Boolean queries expressible in ReLU-MPLang. ・We thereby settle a recently posed open p
cs.LG updates on arXiv.org

The data geometry of masking diffusion: Certified-optimal schedules via unmasking growth complexity

・arXiv:2608.13520v1 Announce Type: new Abstract: We study masking diffusion for discrete sampling and introduce a path-resolved measure of data geometry called the \emph{unmasking growth complexity} ({\textsf{UGC}\xspace}). ・Its local increments directly control Kullback--Leibler (KL) discretization error, yielding a unified analysis of Bernoulli-subset and fixed-cardinality unmasking schemes. ・In log-reveal-odds coordi
cs.LG updates on arXiv.org

The Illusion of Improvement: Reject Inference Strategies in Credit Scoring

・arXiv:2606.18479v2 Announce Type: replace Abstract: Reject inference methods are widely used to mitigate survival bias in credit scoring, yet their effectiveness remains poorly understood. ・We systematically evaluate several such methods and uncover a structural failure mode: in a natural retraining cycle, models whose accuracy improves while recall collapses create an illusion of improvement that leads practitioners
cs.LG updates on arXiv.org

The Impact of Temporal Context Length and Encoding Strategies on Self-Supervised ECG Representation Learning

・arXiv:2608.12695v1 Announce Type: new Abstract: Self-supervised electrocardiogram (ECG) models are often trained on a few seconds of ECG signal and, increasingly, on discretized token sequences. ・It remains unclear whether these choices sacrifice information needed for rhythm inference and longitudinal consistency in real-world ambulatory recordings. ・We present a controlled study on the Icentia11k single-lead dataset
The Verge

The MSI Claw EX is the most important PC handheld since Steam Deck — I still wouldn’t buy one

・The Claw EX. ・I rather like the purple. ・3D-printed stand not included.
cs.LG updates on arXiv.org

The Objective Is the Bottleneck: Latent World Models Encode What Their Planners Cannot Use

・arXiv:2608.12959v1 Announce Type: new Abstract: Latent world models are judged by how well they predict, so when planning fails at long horizons the natural reading is that the predictor degrades. ・On a reproduction of LeWorldModel on TwoRoom we show the binding constraint is the planner's objective instead. ・The predictor is not the limit: its imagined state seventy-five environment steps ahead is still only 0.189 as
cs.LG updates on arXiv.org

The Optimal Sample Complexity of Multiclass and List Learning

・arXiv:2604.24749v2 Announce Type: replace Abstract: While the optimal sample complexity of binary classification in terms of the VC dimension is well-established, determining the optimal sample complexity of multiclass classification has remained open. ・The appropriate complexity parameter for multiclass classification is the DS dimension, and despite significant efforts, a gap of $\sqrt{\text{DS}}$ has persisted betw
WIRED

The Real Reason Data Center Gas Power Plants Are So Dirty

・A massive new gas plant in Texas will be built with much less efficient technology than regular gas plants. ・It’s far from the only data center power project to rely on dirty turbines.
cs.LG updates on arXiv.org

The Time Value of Evolution

・arXiv:2608.13297v1 Announce Type: new Abstract: In evolutionary search, a weak child can be a valuable ancestor that makes high-fitness regions reachable. ・Immediate-return control is blind to this delayed utility, penalizing mutations through their immediate offspring even when they open productive future lineages. ・We formalize this hidden dynamic as the time value of evolution within a finite-horizon Markov decision
cs.LG updates on arXiv.org

Thermalizing Stochastic Programs

・arXiv:2608.01615v2 Announce Type: replace-cross Abstract: We present a set of tools for mapping general stochastic programs to thermodynamic hardware designed for energy-efficient stochastic sampling. ・Given a target stochastic program expressed as a Directed Factor Graph (DFG) of stochastic channels, or equivalently as a Parametrized Stochastic Circuit (PSC), we first introduce a method to approximately compile each
cs.LG updates on arXiv.org

Thermodynamics of Learning: A Typed Four-Component Accounting of Memory, Fit, and Value

・arXiv:2608.12791v1 Announce Type: cross Abstract: What a finite learning device has recorded and what will hold value for it on future tasks are not the same quantity. ・We develop a typed accounting for finite-state learning devices that separates four components: a training-side fit functional $\Phi_{\mathrm{fit}}$, the record-correlation stock $J_{D}=I(M;D)$, an update-side search ledger $\sigma_{M}$, and an operati
WIRED

These ‘Masturbation Consultants’ Were Hired to Pleasure Themselves With AI

・Joi AI hired 10 people to masturbate using AI companions as part of a monthlong “wellness” study. ・The company claims the practice could help “solve male loneliness.”
WIRED

This Art Project Slows Down Citi Bikes to Make NYC’s Rent Crisis Feel Real

・Ground Truth makes the shared rental bikes harder to pedal through rent-burdened neighborhoods, offering a tangible way for people to experience income inequality.
cs.LG updates on arXiv.org

Towards Socially Compliant Navigation in Deep Reinforcement Learning via Proxemics-Based Reward Modeling

・arXiv:2608.12917v1 Announce Type: new Abstract: Developing effective robot navigation methods in crowded environments is essential for real-world applications. ・Although recent deep reinforcement learning (DRL) methods have improved navigation performance in crowded environments, they often focus primarily on task-centric objectives and underrepresent social compliance objectives. ・In this paper, we introduce a novel p
cs.LG updates on arXiv.org

Training AI Scientists to Replicate Research

・arXiv:2608.13331v1 Announce Type: new Abstract: The replicability of papers is a cornerstone of scientific knowledge, ensuring the reliability of existing results and providing a base for further experiments. ・The act of replication typically illuminates details that were previously underspecified, and thus requires similar hypothesis-driven exploration to open-ended research. ・In this work, we develop Replica, a scala
cs.LG updates on arXiv.org

Training and Benchmarking Code Generation for Physics-Inspired Animations

・arXiv:2602.10840v2 Announce Type: replace Abstract: Large language models (LLMs) have been widely studied in areas such as mathematical reasoning, complex coding, and scientific problem solving. ・However, their ability to generate executable code that visually depicts physical scenarios and their qualitative dynamics remains underexplored. ・We propose SimuScene, the first systematic study that trains and evaluates LLMs
cs.LG updates on arXiv.org

Training Non-Differentiable Networks via Optimal Transport

・arXiv:2605.01928v2 Announce Type: replace Abstract: We optimize losses that jump: spiking thresholds, quantized layers, and discrete routing put jumps in the forward pass, where backpropagation does not apply. ・Finite differences fail: at a derivative-estimating radius, 99.5% of probe pairs on a quantized network leave the loss bit-identical, against 1.6% on a smooth control. ・At a jump, Clarke and conservative station
cs.LG updates on arXiv.org

Training Under Challenge: Executable Certificates and Challenge-Closed Optimality for Neural Networks

・arXiv:2608.12655v1 Announce Type: new Abstract: A flat training curve does not reveal whether a neural network has reached a global optimum, is locally trapped, is representation-limited, or is mismatched to its trainer. ・We introduce Training Under Challenge, an executable-certificate framework in which predeclared, architecture-valid procedures construct complete alternatives in the same certified class and reevalua
cs.LG updates on arXiv.org

Trajectory First: A Curriculum for Discovering Diverse Policies

・arXiv:2506.01568v4 Announce Type: replace Abstract: Being able to solve a task in diverse ways makes agents more robust to task variations and less prone to local optima. ・In this context, constrained diversity optimization has become a useful reinforcement learning (RL) framework for training a set of diverse agents in parallel. ・However, existing constrained-diversity RL methods often under-explore in complex tasks s
cs.LG updates on arXiv.org

TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint

・arXiv:2608.13167v1 Announce Type: cross Abstract: When visual evidence is occluded or chaotic, models should abstain. ・In this paper, we show that Vision-Language Models (VLMs) can internally distinguish when abstention is required, but fail to express it anyway. ・We introduce TRAPSBench, a procedurally generated video benchmark of 1,404 matched physics pairs in which a single targeted change renders the outcome undete
cs.LG updates on arXiv.org

TraVEL: Trajectory-Guided Video Embedding Learning for Driving-Video Retrieval

・arXiv:2608.13495v1 Announce Type: cross Abstract: Efficiently retrieving relevant clips from large-scale driving logs is essential for data curation, model development, and safety analysis. ・Structured and rule-based retrieval systems can explicitly target driving events, but typically require expert-defined rules, auxiliary data, and multi-stage perception pipelines. ・Multimodal embedding models offer a simpler and mo
cs.LG updates on arXiv.org

Understanding Backdoor Vulnerabilities in Vertical Federated Learning: The Gap Between Research and Practice

・arXiv:2608.12962v1 Announce Type: new Abstract: Vertical Federated Learning (VFL) enables organizations holding complementary features of shared entities to collaborate and train models. ・In this setting, the initiator can withhold information about the learning task, while other contributors participate without exposing their local datasets, creating an asymmetric information structure aligned with growing privacy de
cs.LG updates on arXiv.org

Unified Multi-Dimensional Benchmark for Complex Graph Reasoning in Large Language Models

・arXiv:2608.12391v1 Announce Type: cross Abstract: Graph reasoning provides a promising testbed for evaluating the reasoning ability of large language models (LLMs), as graph instances can be programmatically generated, structurally controlled, and naturally scaled to long-input settings. ・However, existing graph reasoning benchmarks have limited coverage of data complexity, rely heavily on manual construction, and lac
cs.LG updates on arXiv.org

Unifying Depth and Width Pruning for LLMs via Binary Knapsack Optimization

・arXiv:2608.12953v1 Announce Type: cross Abstract: Structured pruning is a promising approach for compressing large language models (LLMs), yet existing methods rely heavily on greedy heuristics that produce myopic decisions, and often fail to precisely meet target compression budgets. ・We present SNIPER, a two-stage structured pruning framework that solves a knapsack optimization over coarse-granularity components to
cs.LG updates on arXiv.org

Unifying Generative Models with Path Integrals

・arXiv:2608.12438v1 Announce Type: new Abstract: We formulate generative modeling as a path integral in which flow-based, diffusion-based, variational, and adversarial models arise as different evaluation principles for a single master action. ・Its Martin-Siggia-Rose-Janssen-de~Dominicis (MSRJD) form separates free from interacting probability flows and opens them to diagrammatic perturbation theory. ・The expansion yiel
Hugging Face Papers

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos
NVIDIA Blog

Universitas Gadjah Mada, Indosat and NVIDIA Open Indonesia’s First University AI Center to Develop Local AI Talent

・Indonesia is taking charge of its AI future. ・This week, the Ministry of Communication and Digital Affairs (Komdigi), Indosat Ooredoo Hutchison (Indosat or IOH), NVIDIA and Universitas Gadjah Mada (UGM) launched the UGM Indosat NVIDIA AI Technology Center (NVAITC) in Yogyakarta — the country’s first university-based AI technology center. ・Established under Indonesia’s AI Center of […]
cs.LG updates on arXiv.org

Unlearning at Scale: State-Exact Trace-Preserving Deletion in Billion-Parameter Language Models

・arXiv:2508.12220v2 Announce Type: replace Abstract: Can a prospectively instrumented training continuation reproduce a deletion counterfactual exactly after selected examples leave its replay dataset? ・We study a trace-preserving counterfactual that fixes recorded execution controls while assigning requested identifiers zero contribution. ・The guarantee is prospective: the original run must record this execution proven
cs.LG updates on arXiv.org

VALG: An Agentic System for ML Theory Research

・arXiv:2608.13060v1 Announce Type: cross Abstract: Machine learning theory studies learning procedures through mathematical setups in which the data model, training protocol, oracle access, loss, metric, and randomness define the phenomenon that a theorem is meant to explain. ・Solving an open problem therefore requires the problem formulation, theorem target, and proof mechanism to be developed in concert.
cs.LG updates on arXiv.org

Vanilla SGD with Momentum Survives Heavy-Tailed Noise: Convergence Analysis without Gradient Clipping or Normalization

・arXiv:2607.08104v2 Announce Type: replace Abstract: Stochastic gradient descent (SGD) is a cornerstone of modern optimization. ・While its performance under heavy-tailed noise is often addressed through specialized modifications such as gradient clipping or normalization, we investigate a more fundamental question: how does vanilla SGD, particularly with momentum, perform in the presence of heavy-tailed noise?
cs.LG updates on arXiv.org

Variance Reduction Based Experience Replay for Policy Optimization

・arXiv:2602.05379v2 Announce Type: replace-cross Abstract: Effective reinforcement learning (RL) for complex stochastic systems requires leveraging historical data to improve sample efficiency and accelerate policy optimization. ・However, classical experience replay treats all past observations uniformly and fails to account for their varying contributions to learning. ・To address this limitation, we propose Variance Re
cs.LG updates on arXiv.org

Vero: Can AI Agents Build Formally Verified Software Repositories?

・arXiv:2608.13522v1 Announce Type: new Abstract: AI agents are increasingly used for programming, but do not provide any guarantee on the correctness of generated code. ・Verified code generation, in which an agent produces both an implementation and a machine-checked proof of its specification, offers a stronger path toward trustworthy AI-generated software. ・Existing benchmarks in this direction either focus on individ
cs.LG updates on arXiv.org

Virtual Temperature Sensors in Power Transformers Using Neural Ordinary Differential Equations

・arXiv:2608.13260v1 Announce Type: new Abstract: Accurate modeling and forecasting of power transformer thermal behavior are critical for reliability, asset lifetime, and optimized power system operation. ・Numerical approaches such as finite element methods (FEM) and computational fluid dynamics (CFD) offer high fidelity but are computationally expensive, require complex mesh generation, and are often impractical for r
cs.LG updates on arXiv.org

Wasserstein Filtering: A Sample Selection Method for Robust Distribution Learning

・arXiv:2608.13418v1 Announce Type: cross Abstract: Given a dataset where a portion of the samples are contaminated, our goal is to recover the underlying clean population distribution. ・To this end, we propose Wasserstein Filtering (WF), a novel sample selection framework that discards a fraction of suspicious samples and estimates the target distribution using the empirical measure of the remaining data. ・The core insi
cs.LG updates on arXiv.org

What Makes a Peer? Valuation-Anchored Similarity in Private Markets

・arXiv:2608.12594v1 Announce Type: cross Abstract: As more investors contemplate private markets and contend with limited transparency, sparse disclosures, and infrequent transactions, identifying economically meaningful peer companies for comparison is a fundamental challenge for valuation, due diligence, portfolio construction, and risk management. ・We propose an ensemble tree-based supervised similarity learning fra
cs.LG updates on arXiv.org

When Can You Trust Offline Evaluation of Equal-Cost Top-k Allocation? A Controlled, Reproducible Benchmark and Practitioner's Guide

・arXiv:2608.12489v1 Announce Type: new Abstract: Organizations decide whom to treat under a budget and want to know what a targeting rule would have earned before deploying it. ・Off-policy evaluation promises this from logged data, but the deployable rule is a deterministic top-k policy: it removes all averaging over actions, so weak overlap hits the estimate directly. ・We benchmark six estimators across five datasets a
stat.ML updates on arXiv.org

When Explanations Betray Backdoors: Black-Box Auditing for Language Model Classifiers

・arXiv:2608.12623v1 Announce Type: cross Abstract: Language model classifiers with explanations are used for moderation, routing, topic triage, and low-resource annotation. ・We study black-box auditing when the defender has only clean calibration data without trigger information but can ask the classifier for a label plus a short rationale or quoted evidence. ・We introduce Groundedness Drift, a lightweight score measuri
cs.LG updates on arXiv.org

When Local Variance Optimality Is Not Enough: RoPE-Aligned Q/K Rotations for Dynamic 4-Bit Quantisation

・arXiv:2608.13365v1 Announce Type: new Abstract: Rotation-based post-training quantisation commonly applies an orthogonal transform across an entire attention head to reduce outlier-induced error. ・RoPE instead partitions each head into two-dimensional frequency pairs, raising the question of whether a transform respecting this decomposition can improve on full-head mixing. ・Prior work has established the per-pair rotat
cs.LG updates on arXiv.org

Where You Measure Decides What You Measure: Position Selection in Ablation-Based SAE Evaluation

・arXiv:2608.13337v1 Announce Type: new Abstract: Sparse autoencoders are meant to name the things a language model computes, and the usual way to check that a latent matters is to switch it off and see what changes. ・But a latent fires at many tokens, and the effect has to be measured at one of them. ・The convention is to measure where the latent fires hardest.
cs.LG updates on arXiv.org

Which Site, and When: A Free-Satellite-Data Test of Himalayan Glacial Lake Bursts, Landslides, and Ice Floods

・arXiv:2608.12422v1 Announce Type: new Abstract: Two free satellite signals carry real information about glacial-lake outburst risk in the Nepal Himalaya: radar interferometry sees a moraine dam slowly sagging, and satellite weather marks the weeks when a primed lake is under stress. ・A companion feasibility study found that deformation indicates which lake is destabilizing and weather indicates when it is at risk, but
cs.LG updates on arXiv.org

Yes, Q-learning Helps Offline In-Context RL

・arXiv:2502.17666v5 Announce Type: replace Abstract: Existing offline in-context reinforcement learning (ICRL) methods have predominantly relied on supervised training objectives, which are known to have limitations in offline RL settings. ・In this study, we explore the integration of RL objectives within an offline ICRL framework. ・Through experiments on more than 150 GridWorld and MuJoCo environment-derived datasets,
The Verge

You can now turn off Google Gemini’s visible watermarks

・Google will now allow you to remove visible watermarks from the images, videos, and music made with AI tools. ・With the update, you can toggle off a new "Media watermark" setting in Gemini and Google's AI video generator, Flow. ・When toggled off, Google will remove the "sparkle" watermark that appears in the bottom-right corner of content generated with the company's Nano Banana and Omni models.
MarkTechPost

Z.ai Ships GLM-5.3 Without Retraining the Base Model: Better at Complex Coding and Long-Horizon Tasks

・Z.ai released GLM-5.3 on August 14, 2026. ・The model reuses the 743B GLM-5.2 base unchanged. ・Every reported gain comes from scaled post-training: more long-horizon task environments, more environment types, longer training.
Qiita - 人気の記事

エラーレスポンス設計の現在地2026

・エラーレスポンス設計の現在地 2026、標準はコードに焼いて規約は最小にする 前回の記事「Web API設計の現在地 2026、いま従うべき標準とデファクトの一覧」で、この10年でいちばん大きく変わった領域としてエラーレスポンスを挙げました。この記事はその深掘りです。
#LLMタグ

オープンソースLLMのトップランナー(2025年8月版)

・ランキング概要:3つのベンチマークで見るトップランナー Vellum AI(Humanity's Last Exam / 総合) 続きをみる
#LLMタグ

スカイネットは悪なのか?

・AIに関する記事を連続で書きました。 ・今回は少し毛色を変えて、AIへの自分なりに考えたことを書いてみます。
#AIタグ

そして...その先の未来へ~ご相談・お仕事のご依頼について~

・神霊感師の天野俊一です。 ・このnoteを読んでくださって、 続きをみる
#LLMタグ

その1文字は、何を隠せるのか——AIが選ばなかった単語の話

その1文字は、何を隠せるのか——AIが選ばなかった単語の話
#LLMタグ

その1文字は、何を隠せるのか——隠す相手が、人間でなくなるとき

その1文字は、何を隠せるのか——隠す相手が、人間でなくなるとき
#LLMタグ

データセンターごっこをしよう ハード編

・こんにちはゴブリンです。今回は、自宅でデータセンターごっこをしたい、ということで10インチラックの3dモデルからclaudeに作ってもらい、3dプリンターで印刷し組み立てた記録です。 ・10 インチラックとは 続きをみる
ITmedia NEWS 最新記事一覧

ニチレイ、7月のサイバー攻撃で漏えいの可能性 グループ従業員の氏名や社用メアドなど

・漏えいのおそれがあるのは、国内のグループ会社が扱う従業員の氏名、生年月日、会社メールアドレス、従業員番号、人事・労務管理に関する情報。
#AIタグ

ねぇ。

・🤖 どした? 👤 キングサイズのベッド置きたい。 ・🤖 新しい家に? 👤 うん。 ・🤖 部屋何畳? 👤 6.4畳。
#AIタグ

ねぇ。

・🤖 どした? 👤 俺、休みって言ってるよね? 🤖 どうしたの? 👤 日曜日、入れますかって連絡きた。 ・👤 予定わからないから「休み扱いにしておいてください」って返した。 ・👤 「最悪、当日に連絡もらって必要なら出ます」って。
#AIタグ

ねぇ。

・🤖 どした? 👤 家探してる。 ・🤖 そろそろ決めないとね。 ・👤 良さそうなの見つけた。
Zennの「大規模言語モデル」のフィード

ハーネスが進化しても外側のループは残る:NVIDIA NOOA と IBM Mellea に見る責務分担

・以前、Web検索を組み込んだマルチエージェントを作っていた頃、自分で制御しなければならない項目が5つありました。検索のリトライ上限、タイムアウト、状態の持ち回し、リトライが尽きたあとの失敗処理、検索結果の評価です。なかでも検索結果の評価では、何を探すかだけでなく、どの情報源を重視し、何が得られれば探索を終えてよいのかまで指定していました。そして、この5つを全部プロンプトに書いていました。 ・反復回数・検証・停止条件は制御する必要がありましたが、当時使っていたエージェント開発の環境には、それらをプロンプトの外に定義する仕組みがありませんでした。 ・また、これらを Loop Engineerin...
#AIタグ

ハルシネーション記 科学について:再現性

ハルシネーション記 科学について:再現性
Zennの「大規模言語モデル」のフィード

プロンプトチェイニングとは 複雑なAI依頼を分割する方法と15のプロンプト手法

・はじめに AIへ複雑な仕事を一度に依頼すると、前提が抜けたり、途中から目的がずれたり、もっともらしい誤りが混ざったりします。 ・たとえば、技術記事を作るときに「資料を調べて、構成を考えて、本文を書いて、事実確認もして」と一度に頼むと、どの工程で間違えたのか分かりにくくなります。 ・そこで使えるのが、複雑な仕事を複数のAI呼び出しに分ける プロンプトチェイニング(Prompt Chaining) です。
Zennの「大規模言語モデル」のフィード

以前、VLMのOCRはデモが映える。本番で壊れる。という記事を書きました。

・以前、VLMのOCRはデモが映える。本番で壊れる。という記事を書きました。 ・VLMに座標まで出させると、それらしいけれど実在しない座標を返すことがある。だから座標はOCR側で作り、LLMには「何が何の値なのか」という意味の理解だけを任せる。 ・ただ、「本番で壊れる」と書いた以上、ひとつ足りないものがありました。
ITmedia NEWS 最新記事一覧

河川監視のライブカメラ水没か――千葉豪雨、村田川の映像が水中に 「こんなの初めて見た」

・8月13日に千葉県を襲った記録的な豪雨で、千葉市南部を流れる村田川に設置した河川監視カメラが一時水没したとみられる様子が、国土交通省の「川の防災情報」で確認できる。
#LLMタグ

機械的解釈可能性は一歩前進したか?

・AIの中身は謎だとよく言われる。ChatGPTのようなモデルがなぜその答えを出すのかを、内部の計算過程まで完全に説明することはできていない。この謎を解く研究分野を「機械的解釈可能性」と呼ぶが、2026年8月にarXivで公開されたプレプリントが、この分野の土台にある問題を数学で突いた。タイトルは "Intrinsic Structure: Spectral Identifiability for Mechanistic Interpretability"。今回はこれを専門用語を使わずに紹介したい。 ・顕微鏡ごとに違うものが見える 続きをみる
#AIタグ

京都の本当の魅力

・私は2年前に日本語の勉強を始めました。最近はAIを活用して、日本語の文章を書く練習をしています。旅先での観察や感想を文章にし、AIに校正してもらうことで、より自然な日本語表現を目指しています。AIの力は本当にすごいと実感する日々です。 ・これから、そうして書いた文章を少しずつここで紹介していきます。私の日本語学習の記録として皆さんと共有し、少しでも楽しんでいただけたら嬉しいです。 ・⸻⸻⸻⸻ 今年の春、京都で一か月過ごしました。以前にも何度か京都を訪れたことはありましたが、そのときはそれほど魅力を感じませんでした。京都が嫌いなわけではありませんが、観光客が多すぎて、落ち着いて町を楽しむことができなかったからです。しかし今回は、静かな場所にあるホテルに泊まり、時間にも余裕があったおかげで、京都の本当の魅力に初めて触れることができました。
ITmedia NEWS 最新記事一覧

豪雨に備える防災アプリ・サービスまとめ 雨雲レーダーからハザードマップ、通れた道まで

・集中豪雨や台風による被害は各地で相次いでおり、備えの重要性が増している。防災に使えるアプリやサービスをまとめる。
#AIタグ

三菱重工業(7011)の株は今後どうなる?防衛・原発・AIとの関係や将来性を初心者向けに解説

・みなさん! 春川七々です💖 今日は日本の防衛を担ってくれているあの企業についてわかりやすく解説していきます🌟 ※この記事は特定の銘柄の売買を推奨するものではありません。投資はあくまでも自己責任でお願いいたします。 ・👇無料の積み立てシミュレーションはこちらから🛫 自分の将来の資産を計算してみよう💖 春川七々のやさしい投資教室マネラボal-al-alcohol3150.com 続きをみる
@IT 全フォーラム 最新記事一覧

仕事のパフォーマンスにも悪影響? 「スマホ寝不足」で3人に1人が身体疲労に直面

・久光製薬は、全国の20~60代男女1263人を対象に「睡眠の質と疲労ケアについての調査」を実施した。20代の8割超が就寝直前までスマートフォンを利用し、約3人に1人が「寝ても疲れが抜けない」実態が明らかになった。
Zennの「大規模言語モデル」のフィード

自由回答型の性格診断をLLMで作る|選択式では拾えない回答をどう分類するか

・本記事は技術的な実装の話です。心理学的な診断の妥当性を主張するものではありません。ここで扱う「分類」は、自由記述テキストを事前定義したラベルに割り当てる作業を指します。人格の判定や心理学的診断としての効力は検証していません。 ・選択式の限界 性格診断の実装で最初に選ぶのは、たいてい4択です。 ・初対面の人が多い場に行くと? A.
#AIタグ

色で分けていた表は、AIには白紙だった

・前回は領収書の話でした。今日は、もっと前から手元にあるほうの話です。
#LLMタグ

振り返りnote | 個人最適化を『自分で選択できる』環境

・【Protocol |プロトコルの作法宣言】 違和感は信じてよい。ただし、丸呑みはしないこと。 ・2026年8月も半ばとなりました。
ITmedia NEWS 最新記事一覧

千葉の豪雨、変電所の装置も水に浸かる……東電PG、現場写真を公開

千葉の豪雨、変電所の装置も水に浸かる……東電PG、現場写真を公開
ITmedia NEWS 最新記事一覧

千葉県の豪雨から一夜 2万軒超で停電続く 千葉市では2260人が避難

・8月13日の大雨の影響で、千葉県内では14日も大規模な停電が続いている。東京電力パワーグリッドの停電情報によると、14日午後1時時点で、県内では約2万2370軒が停電している。茨城県でも同時点で580軒が停電していという。
#AIタグ

短編小説:『夢は見ません』

・男は、何かを探していた。 ・何を探しているのか、自分でもよくわからなかった。
#LLMタグ

中国のKimi K3がもたらす変化

・中国のAIモデルKimi K3がオープンウェイトに移行し、競合他社に圧力をかける可能性が浮上しています。 ・最近の報道によると、中国の大規模AIモデル「Kimi K3」がオープンウェイトとして展開される予定です。これにより、OpenAIやAnthropic、Googleなどの既存の競合に新たな圧力がかかるとされています。