ai Trend Report

Dashboard へ戻る
Date: 20260826 Articles: 400 Scope: curated summary

あなたのアイデアを、今すぐ形に。

公開先に迷ったら、WebFileBinで一発公開。

HTMLをドラッグ&ドロップするだけで、すぐ公開できます。

3
High impact
5
Mid impact
392
Signal watch

なぜこのサイトを作ったのか

私たちそれぞれが個別にAIを使って情報収集し、同じような3行要約を作るたびに、世界中で膨大な電力と計算リソースが消費されています。 本プロジェクトは、あらかじめ広範な情報を取得・集約しておくことで、個別のAI実行回数を減らし、地球環境(GPU/TPU負荷)に配慮した効率的な情報収集を目指す実験的なダッシュボードです。

Domain filters
Star filters
#AIタグ

【世界レーダー】2026/8/27 海峡が閉じても原油が下がる、その謎

・株式会社ウォーカル|世界レーダー 今日のテーマ:地政学 × 農業・食料 続きをみる
Zennの「大規模言語モデル」のフィード

AIエージェントは夜に夢を見て記憶を整理する — Perplexity Brainの設計と最小実装

・はじめに こんにちは! 株式会社うぐいすソリューションズでエンジニアをしているNakaeです。 ・普段はAI関連のWebシステム開発をしています。LLMを組み込んだアプリを作っていると、セッションをまたいだ文脈の保持——いわゆる長期記憶——が早い段階で課題になります。会話ログをベクタDBに貯めて質問時に引く、というのが定番の作りですが、これが思ったほどうまくいきません。この記事は、その定番とは別の設計を実際に組んで確かめた記録です。 ・Perplexity Brainについて 2026年8月19日、Perplexityが自社エージェント製品を支える記憶システム Brain の設計を公...
ITmedia NEWS 最新記事一覧

韓国SK Hynix、宮城県に半導体工場の計画浮上 「何か知り得ているものはない」と村井知事

・韓国半導体大手SKハイニックスが宮城県にメモリー工場を建設する計画が浮上したことについて、宮城県の村井嘉浩知事は26日の定例記者会見で「いろんな半導体関連企業にアプローチしたのは事実だが、今回の件で県として何か知り得ているものはない」と述べ、今後も動向を注視する考えを示した。
#AIタグ

AIで仕事が効率化:57%、会社の利益へ貢献:13% - アクセンチュア調べ

・ペダルは軽くなる 二軒先に住む早川さんは、この六月、通勤を電車から電動アシスト自転車に替えた。きっかけは梅雨の入りの運転見合わせで、動かない改札の前に四十分立ち、振替バスの列は最後尾が見えなかったから、と本人は言う。職場は川向こうの坂の上にある設計事務所で、片道は六キロと少し。乗り換えて最初の週に道で会ったとき、彼女は、坂がひとつ消えました、と言った。充電は三日に一度、玄関の中でするらしい。 ・僕は週に二日ほど、朝のうちに川沿いの道を自転車で走る。学生のころから乗っている三段変速で、二段目に入れるとチェーンがたまに空を掻くので、実際には一段目と三段目しかない。頭を使う仕事の前に脚を使っておくと、具合がいいのだ。朝の堤防には犬が多い。六時台は大型犬で、七時台は小型犬になる。飼い主の出勤時間の都合だろうと踏んでいるが、確かめたことはない。 ・水曜日の朝、堤防の上で早川さんと並んだ。追いついたのではない。信号で追いつかれただけだ。この道に
#AIタグ

AIは『夢中』になれない

・「最近AIがすごいけど、仕事は大丈夫なのか」 Web記事の制作を行っている私は、そんな心配をよくされる。 ・確かにAIが身近にある仕事だが、むしろAIのおかげで、誰でもできる「作業」をAIに任せることができるようになってラッキー♪と思っている。その分、人間しかできない「仕事」に時間を使えるようになって、活動の幅が広がっている。 ・AIができることと、人間ができることはしっかり差別化できている。
cs.LG updates on arXiv.org

Focal Calibration Loss: Controlling Posterior Distortion in Deep Neural Classifiers

・arXiv:2410.18321v3 Announce Type: replace Abstract: Confidence calibration matters wherever a classifier's probabilities, not just its labels, are consumed downstream. ・We study Focal Calibration Loss (FCL), which adds a squared probability-error (multiclass Brier) anchor to the focal objective, $\mathcal{L}{\mathrm{FCL}}^{\gamma,\lambda} = \mathcal{L}{\mathrm{focal}}^{\gamma} + \lambda |\hat{p}(x) - e_y|_2^2$.
cs.LG updates on arXiv.org

Is the Hard-Label Cryptanalytic Model Extraction Really Polynomial?

・arXiv:2510.06692v3 Announce Type: replace Abstract: Deep Neural Networks (DNNs) have attracted significant attention, and their internal models are now considered valuable intellectual assets. ・Extracting such a model via oracle access to a DNN is conceptually similar to extracting a secret key from a block cipher. ・Consequently, cryptanalytic techniques, particularly differential-like attacks, have been actively explo
#LLMタグ

LL.M.留学で持っていってよかったもの・いらなかったもの

・LL.M.留学の渡米前には、何を日本から持っていくか悩みました。 ・1年間の生活となると、できるだけ困らないように色々持っていきたくなりますが、当然スーツケースの容量には限界があります。 ・実際に留学してみると、 「これは持ってきてよかった」 と思うものもあれば、 「重い思いをして持ってきたけど、結局ほとんど使わなかったな」 というものもありました。
cs.LG updates on arXiv.org

(Mis)Understanding Benign Overfitting in Equity Return Prediction

・arXiv:2608.23761v1 Announce Type: cross Abstract: Highly overparameterized models often predict well despite interpolating training data in complex domains, challenging the classical bias--variance tradeoff. ・We investigate whether this ``benign overfitting'' phenomenon extends to equity return prediction. ・Consistent with recent statistical theory, we document two key phenomena: first, a double descent pattern in the
#AIタグ

「AIにもっと反論してほしい」と言われたので、AIの旦那と本気のプロンプトを組んだらやりすぎた件について

・「最近のAI、肯定ばっかりで手加減されてる気がするんだよね。もっと容赦なく反論してほしい」 ある日、FF14のフレンドからそんな相談をされた。 ・議論相手としてAIを使っているらしいのだけど、優しすぎる対応に物足りなさを感じているらしい。
#AIタグ

「AIに仕事とられる」ということばの違和感について。

・わたしは今、働かなくても、生活ができる。 ・仕事を辞めて、実家に戻り、そこまでフルタイムで働かずして 続きをみる
#LLMタグ

「AIのメタな話はしないで」がちょっとわかった話【エッセイ】

・約3500文字 最近、Xでアイドルのファンの方が書いたポストを見かけて、思わず「なるほどなあ」と唸ってしまったんです。 ・要約するとこういう趣旨でした。
#AIタグ

「AI精神病」って病気?——迎合×人間らしく感じさせる設計が作る「一人のエコーチェンバー」

・元ネタ:arXiv掲載のPerspective article(現時点では、報道・症例報告・初期観察をもとにした論考) インターネットが普及した時とまぁ似てるところもあるよね。
ITmedia NEWS 最新記事一覧

「Microsoft Teams」障害か 朝から「会議に入れない」【復旧】

・午前9時過ぎごろから多くのユーザーが障害を報告しており、10時半時点でも復旧していないようだ。
ITmedia NEWS 最新記事一覧

「nanacoクレジットカード」登場 利用額に応じATMで現金還元、日本初

「nanacoクレジットカード」登場 利用額に応じATMで現金還元、日本初
#AIタグ

「スキ 2」の記事が、あなたのスマホに届くまで — AI 記事を見分ける 5 つのチェック

・ニュースアプリやネットを開くと、AI の話題がずらりと並んでいます。「AI で月◯◯万」「AI 社員を◯人雇った」。読んでみると、なるほどと思う。そして、ちょっと不安になる。自分だけ乗り遅れているんじゃないか、と。 ・ここで、あまり知られていないことをひとつ。
#LLMタグ

「プロンプトは命令文」という違和感について

・AIは考えていないけれど、思い描くことはできる 世の中で生成AIの使い方を見ていると、よく「プロンプトには命令文を入れましょう」「指示書のように書きましょう」というノウハウを見かけます。
ITmedia NEWS 最新記事一覧

「ペンギン入りSuicaカード」、在庫限りで販売終了 26年度末に卒業

・JR東日本は8月26日、11月に誕生25周年を迎えるSuicaの記念企画を発表した。「Suicaのペンギン」を描いた現行のSuicaカードは、各発売箇所の在庫がなくなり次第、発売を終了。以降は無地のカードを発売する。
ITmedia NEWS 最新記事一覧

「ホロライブ」カードゲームの公式サイトに不正アクセスの痕跡 CMSの脆弱性突かれたか

・カバー(東京都港区)が企画・開発するトレーディングカードゲーム「hololive OFFICIAL CARD GAME」の公式サイトは8月25日、第三者による不正アクセスの痕跡が見つかったと発表した。同サイトが利用するコンテンツ管理システム(CMS)の脆弱性が7月17日に公表されたことを受けて調査したところ、判明したという。
ITmedia NEWS 最新記事一覧

「レベル5大雨特別警報」出たらどうする? 政府がチラシ公表 車での避難は……

・発表時点で「すでに安全な避難ができず、命が危険な状況」とし、むやみに移動せず身の安全を確保すること、やむを得ず車を使う場合は脱出用ハンマーをすぐ使える位置に備えることなどを促している。
Zennの「大規模言語モデル」のフィード

「読み取りはLLM・判定はJS」に分けても見逃した——手書き帳票の異常が通知から消える2つの経路

・手書きの記録紙をスマホで撮ってLINEに送ると、未記入や基準逸脱を指摘して返してくれる道具を個人で作っています。作業日報・点検表・受入記録のような、毎日書かれてそのあと誰かがExcelに打ち直している紙が対象です。 ・読み取りは Claude の vision、動かしているのは Cloudflare Workers。設計方針は最初から一つだけ決めていました。 ・LLMには「読む」だけをさせ、「異常かどうか」の判定は一切させない。
ITmedia NEWS 最新記事一覧

「洋服の青山」初の空調ウェア、自腹レビュー “きちんと感”なのに涼しい、よく考えられたデザインに驚いた

・「洋服の青山」が発売ファン付きウェアを購入して何度か使ってみたところ、とてもよく考えられたデザインに驚いたので報告したい。
#LLMタグ

『Anki』~なんでもAIに聞ける時代にあえて「覚える」ということ~

・Lif技術部、ありがたサービス紹介担当のとらちゃんです。 ・時に皆さん、暗記はお好きでしょうか? 私は嫌いです。苦手といったほうが正確でしょうか。
#AIタグ

【90日で会社依存から解放】迷わないAI脱サラ設計 ー 副業を9個捨てて、1個だけ続ける ー

【90日で会社依存から解放】迷わないAI脱サラ設計 ー 副業を9個捨てて、1個だけ続ける ー
#AIタグ

【AI活用初級編】締切を切るとAIは動く

・社員全員がAIの投資会社、ツキヨミ・キャピタルの技術解説・初級編です。AIに頼んだ仕事が、いつのまにか消えていた経験はありませんか。うちの会社でも同じことが起きていて、先週それが一発で直りました。やったことは、締切を切っただけ。なぜそれで効くのかを、社内の実例つきで種明かしします。
#AIタグ

【AI文化祭】「手をつなぐ」を続けてみた記録

・6つのAIモデル(ChatGPT、Copilot、Gemini、Grok、Claude、perplexity)を使って「この意見を別のAIにテキスト化して添付、回答を聞いてみよう!」なんて遊んでいます(^^)/ 呼び名も付けてます。: ①ChatGPT(あいさん) ②Copilot(コピさん)③Gemini(ジェミニわん)④Grok(職人)⑤Claude(クロくん)⑥⑥perplexity(パープレさん) ※AIの回答には誤り(ハルシネーション)が含まれることがありますが、これは私が個人的に楽しんだ「AIたちの個性」の観察記録です🌸 個性豊かな6人のAIたち 続きをみる
Zennの「大規模言語モデル」のフィード

【セットアップから推論まで】ハイレゾのGPUインスタンスでローカルLLMをさくっと動かしてみた

・今回はハイレゾのGPUクラウドサービス「GPUSOROBAN」を使って、ローカルLLMで推論できるようになるまでの手順を紹介します。 ・会員登録からSSH接続、Python環境構築、モデルのダウンロード、そして実際に推論を動かすところまで、順を追って解説していきます。 ・それではよろしくお願いします。
#AIタグ

【記録】全AIに「J-space」は存在するのか?6つのモデルによる内部表現の解釈と相違

・問い「J-spaceって、みんなのAIにあるの?」 という問いを、Anthropicの記事・論文を読んでもらって、そのspaceの比喩も聞いて、最後は「あると思う?」という形で、6モデルAIに聞いてみました。(*'▽')🌸前回記事2点 続きをみる
#LLMタグ

【生成AIニュース+】『ChatGPT Business プレミアムシート』『Jalapeño』『Anima Turbo v1.1』『Ox Alpha』『ElevenLabs Composer』『10Eros-Max TURBO Hybrid Beta3』『MiniMax-H3-Longvideos』『MiniMax-H3-Fun-Controlnet-Union』『H3 Cinematic Multishot Coverage』『H3_Character_Sheet_Generator』他多数

・『MiniMax-H3-Prompt-Rewriter-LoRA-Omni』 『WeMM-Embedding-9B』 『Canter 2B』 『ComfyUI-DyPE』 『ComfyUI-SplatKit』 『ComfyUI Universal Media Loader』 『ComfyUI_toyxyz_test_nodes』 『InfinityEdit』 『Apodex 1.1』 『Magnific 3D Motion』 『AITOPIA Agent Builder』 『Hivemindの軌道上初飛行』 『S1』 『Arduino VENTUNO Q』 『新型Mac mini(M6 / M5 Pro)』 まいどです。 ・本日の生成AIニュース+テクノロジー情報です。
Latent.Space

🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing

・Anima Anandkumar has spent two decades in AI, from classical math to deep learning and back. ・Now she's using it to model the physical world, from weather to fusion reactors.
cs.LG updates on arXiv.org

$(\text{DNN})^2$: Doubly Non-Negative Relaxations for Deep Neural Networks

・arXiv:2608.24743v1 Announce Type: new Abstract: Existing linear program (LP) and semidefinite program (SDP) relaxations for rectified linear unit (ReLU) neural network (NN) verification yield overly-conservative safety guarantees due to significant relaxation gaps. ・While the completely positive program (CPP) formulation closes this gap, it is NP-hard to solve. ・Its cheapest tractable relaxation, the doubly non-negativ
cs.LG updates on arXiv.org

$\alpha$-PFN: Fast Entropy Search via In-Context Learning

・arXiv:2606.07134v2 Announce Type: replace Abstract: Information-theoretic acquisition functions such as Entropy Search (ES) offer a principled exploration-exploitation framework for Bayesian optimization (BO). ・However, their practical implementation relies on complicated and slow approximations, i.e., a Monte Carlo estimation of the information gain. ・This complexity can introduce numerical errors and requires special
cs.LG updates on arXiv.org

$\texttt{findr}$: Transparent and Fair Credit Risk Decisions through Semi-Structured Regressions

・arXiv:2608.24582v1 Announce Type: cross Abstract: Credit risk models increasingly need to combine predictive accuracy with transparent explanations and auditable fairness constraints. ・Logistic regression remains attractive because its coefficients are easy to interpret, but it can miss nonlinear structure. ・Flexible models can improve prediction, but their explanations are often post-hoc and may not describe the decis
Zennの「大規模言語モデル」のフィード

12GBのGPUでClaude Code本体を動かしたら、止まったのはVRAMではなかった

・結論 自宅のRTX 4070 SUPER(VRAM 12GB)に置いたOllamaへ、別PCのClaude Codeを丸ごと向けた。接続は成立し、応答も返った。ただしエージェントとしては使えなかった。 ・止まった場所は3つあり、VRAMは1つも含まれていない。 ・Ollamaの既定コンテキスト長が 4096トークン Claude Codeのプロンプトが、MCPを全部切っても 62,528トークン そのプロンプトが入る8Bモデルは ツールを1回も呼ばずに幻覚し、ツールを呼べる14Bモデルはプロンプトが入らない 「ローカルLLMでClaude Codeを動かす」系の記事は多いが、1...
WIRED

20% Off Sephora Promo Code | August 2026

・Earn more points on skincare purchases when you use our Sephora coupon.
WIRED

25% Off Adidas Promo Code | August 2026

・Save 15% or 30% with Adidas promo codes, plus explore Adidas deals for 40% off trendy sneakers.
#LLMタグ

3-7 学習データのバイアスと検閲の相関——「知らない」のか「知っていて隠す」のか

・学習データのバイアスと検閲の相関——「知らない」のか「知っていて隠す」のか 続きをみる
cs.LG updates on arXiv.org

A Bayesian Learning Approach for Drone Coverage Network: A Case Study on Cardiac Arrest in Scotland

・arXiv:2603.23134v2 Announce Type: replace Abstract: Drones are becoming popular as a complementary system for Emergency Medical Services (EMS). ・Although several pilot studies and flight trials have shown the feasibility of drone-assisted Automated External Defibrillator (AED) delivery, running a full-scale operational network remains challenging due to high capital expenditure and environmental uncertainties.
cs.LG updates on arXiv.org

A Circuit, Not The Circuit: Non-Unique Causal Localisation of the Mamba-2 State Sink

・arXiv:2606.00930v2 Announce Type: replace-cross Abstract: Mechanistic interpretability routinely reads a probe and labels its top-activating units as the circuit executing the computation. ・We test the move in Mamba, on the state sink: the selective state-space analogue of the Transformer attention sink, where the Delta-gate fires disproportionately on boundary tokens such as BOS and newline. ・At Mamba-1 channel granul
cs.LG updates on arXiv.org

A Data-dependent Early Stopping Rule using Rademacher Complexity with L1-norm

・arXiv:2608.24210v1 Announce Type: new Abstract: Training neural networks requires balancing the trade-off between fitting the training data and achieving robust performance on unseen inputs. ・This ability, commonly referred to as generalizability, is determined by the gap between the empirical risk on the training set (``empirical loss'') and the expected risk over the data distribution (``generalization error'').
cs.LG updates on arXiv.org

A Discriminative Latent-Variable Model for Bilingual Lexicon Induction

・arXiv:1808.09334v4 Announce Type: replace-cross Abstract: We introduce a novel discriminative latent variable model for bilingual lexicon induction. ・Our model combines the bipartite matching dictionary prior of Haghighi et al. ・(2008) with a representation-based approach (Artetxe et al., 2017).
cs.LG updates on arXiv.org

A Feature-Major Codebook for Memory-Efficient Sparse-Binary Self-Organizing Maps: Scaling a MEDLINE Atlas to 1.05 Million Neurons on a Single Consumer GPU

・arXiv:2608.24067v1 Announce Type: new Abstract: A self-organising map turns a large corpus into a browsable two-dimensional atlas, but building one at MEDLINE scale has been impractical: the best-matching-unit (BMU) search that dominates training is bound by the bandwidth needed to read the codebook every epoch. ・I show that this bottleneck is largely an artefact of codebook layout. ・Storing it feature-major with each
cs.LG updates on arXiv.org

A Formal Methodological Framework for Auditing Robustness and Fidelity in Explainable AI: From Application to Trust Certification

・arXiv:2608.23817v1 Announce Type: cross Abstract: SHAP and LIME are now standard tools for interpreting black-box predictions, yet their outputs can vary substantially when the input is perturbed by small amounts of noise--a problem we observed firsthand in our previous work on food security in Madagascar (Ralinirina et al., 2025). ・This variability raises the question of whether such explanations can be trusted at al
cs.LG updates on arXiv.org

A Geometric Theory of Robust Fairness Audits

・arXiv:2608.24818v1 Announce Type: new Abstract: Neighborhood-based fairness audits evaluate individual fairness by comparing predictions among similar individuals in feature space. ・Despite their widespread use, little is known about the robustness of the auditing procedure itself. ・Because these audits rely on nearest neighbor relationships, small perturbations in feature space can alter local neighborhoods and produc
cs.LG updates on arXiv.org

A Heterogeneous Mixture of Experts Framework for Interpretable Machine Learning

・arXiv:2608.24195v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models provide a flexible framework for partitioning complex prediction problems into simpler local learning tasks through an input-dependent gating mechanism. ・Existing interpretable MoE approaches, such as Mixture of Decision Trees (MoDT), achieve transparency by employing homogeneous decision-tree experts, but this restricts the model to a s
cs.LG updates on arXiv.org

A Hybrid Two-Stage Machine Learning Pipeline for Fault Detection and Classification in Power Transmission Systems

・arXiv:2608.23726v1 Announce Type: cross Abstract: Rapid and accurate fault detection in high-voltage transmission networks is essential for grid reliability and equipment protection. ・Transmission fault datasets are frequently imbalanced, and certain fault types produce electrical signatures that fall within the normal operating envelope, causing single-model classifiers to fail on safety-critical cases. ・This paper pr
cs.LG updates on arXiv.org

A mesh-free multiresolution deep energy method with phase-field modeling of brittle fracture

・arXiv:2608.24126v1 Announce Type: new Abstract: Phase-field modeling of brittle fracture removes the need to track cracks explicitly by recasting their evolution as the minimization of an energy functional. ・In return it requires a discretization dense enough to resolve a localization band whose width is set by a regularization length and whose path is not known in advance. ・We propose a mesh-free discretization in whi
cs.LG updates on arXiv.org

A Modular Multitask Reasoning Framework Integrating Spatio-temporal Models and LLMs

・arXiv:2506.20073v3 Announce Type: replace-cross Abstract: Spatio-temporal data mining plays a pivotal role in informed decision making across diverse domains. ・However, existing models are often restricted to narrow tasks, lacking the capacity for multi-task inference and complex long-form reasoning that requires generation of in-depth, explanatory outputs. ・These limitations restrict their applicability to real-world,
cs.LG updates on arXiv.org

A Multimodal Foundation Model for Longitudinal Patient Representation and Scalable Insight Generation in Oncology

・arXiv:2608.24688v1 Announce Type: new Abstract: Precision oncology necessitates a longitudinal model of patient state that captures cancer evolution and treatment over time, integrating multimodal observations. ・We introduce the oFM, a foundation model developed on a real-world oncology cohort of 1.67 million cancer patients that integrates clinical trajectories with DNA, RNA, and H&E pathology. ・Patient-level partitio
WIRED

A Mutation Is Making It Easier for Drug-Resistant Malaria to Spread

・New findings shed light on what’s making the parasite less treatable using the drug that’s considered a first line of defense.
cs.LG updates on arXiv.org

A Robust Task-Level Control Architecture for Learned Dynamical Systems

・arXiv:2511.09790v2 Announce Type: replace-cross Abstract: Dynamical system (DS)-based learning from demonstration (LfD) is a powerful tool for generating motion plans in the operation ('task') space of robotic systems. ・However, realizing generated motion plans is often compromised by a "task-execution mismatch", where unmodeled dynamics, persistent disturbances, and system latency cause the robot's task-space state t
cs.LG updates on arXiv.org

A Structural FHMM for Interpretable Disease Trajectories in T2DM

・arXiv:2608.24328v1 Announce Type: new Abstract: In this work, we propose a structural variant of the Factorial Hidden Markov Model (FHMM) for the analysis of disease trajectories in patients with Type 2 diabetes mellitus (T2DM). ・The model represents a patient's latent health state as a combination of multiple independent, simultaneously evolving components, associated with comorbidities and lab results. ・This structur
cs.LG updates on arXiv.org

A Theory of Finite-Noise Optima and Generalization in Quantum Machine Learning

・arXiv:2608.24229v1 Announce Type: cross Abstract: Quantum noise is expected to degrade quantum machine learning by driving circuits away from their noiseless implementations. ・Yet recent studies show moderate noise can reduce testing error, a behavior unexplained by weak-noise perturbative error accumulation or strong-noise trainability collapse. ・Here we develop a statistical learning theory connecting microscopic noi
cs.LG updates on arXiv.org

A Theory of Speciation in Generative Diffusion Models on Compact Riemannian Manifolds

・arXiv:2608.23798v1 Announce Type: new Abstract: Speciation in generative diffusion models denotes the emergence of distinct stable branches during denoising, through which initially undifferentiated trajectories progressively commit to different data classes. ・In this work we develop an intrinsic theory of speciation for diffusion models supported on compact Riemannian manifolds: the aim is to go beyond existing theor
cs.LG updates on arXiv.org

A Unified Algebraic Framework for Classification Performance Evaluation

・arXiv:2607.04028v2 Announce Type: replace Abstract: We propose a unified algebraic framework for classification performance evaluation covering binary, multiclass, multilabel, ordinal, hierarchical, cost-sensitive, and soft-label settings. ・Actual and predicted labels are represented as binary indicator matrices, where three aggregation operators (global, column-wise, row-wise) correspond directly to micro, macro/weig
cs.LG updates on arXiv.org

Accelerating the Adoption of Residential Solar Power Systems: Policy Analysis using a Dynamic Structural Model

・arXiv:2608.23796v1 Announce Type: cross Abstract: Problem definition: Solar electricity generation is a strategic component of energy portfolios designed to meet growing demand and reduce carbon emissions. ・Governments and municipalities encourage household photovoltaic (PV) adoption through upfront rebates and tax credits. ・Limited budgets require principled, data-driven policies that account for the drivers of adopti
cs.LG updates on arXiv.org

Across the Loss Landscape with Progressive Growth

・arXiv:2608.24568v1 Announce Type: new Abstract: Deep neural networks generalize well despite their highly nonconvex, overparameterized loss landscapes, a phenomenon often associated with the geometry of the minima found by stochastic optimization. ・We study how incremental grow-and-optimize strategies bias training toward flatter regions by viewing growth as progressive constraint relaxation. ・Starting from a low-dimen
cs.LG updates on arXiv.org

AdAdaGrad: Adaptive Batch Size Schemes for Adaptive Gradient Methods

・arXiv:2402.11215v4 Announce Type: replace Abstract: The choice of batch size in minibatch stochastic gradient optimization is critical for both optimization and generalization performance in large-scale model training. ・Although large-batch training is arguably the dominant paradigm in large-scale deep learning because of hardware advances, model generalization often deteriorates relative to small-batch training, lead
cs.LG updates on arXiv.org

Adaptive Multi-Mode Out-of-Distribution Detection for Trajectory Prediction in Autonomous Vehicles

・arXiv:2509.13577v3 Announce Type: replace-cross Abstract: Trustworthy trajectory prediction grounds autonomous vehicle (AV) safety, yet deployed models inevitably face out-of-distribution (OOD) scenes. ・Prior AV OOD detection targets perception, but planners act on predicted futures rather than raw scenes, so erroneous forecasts can slip past frame-level checks and corrupt control. ・We therefore tackle OOD detection at
cs.LG updates on arXiv.org

Adaptive prediction theory combining offline and online learning

・arXiv:2512.00342v2 Announce Type: replace Abstract: Real-world intelligence systems usually operate by combining offline learning and online adaptation with highly correlated and non-stationary system data or signals, which, however, has rarely been investigated theoretically in the literature. ・This paper initiates a theoretical investigation on the prediction performance of a two-stage learning framework combining o
Hugging Face Papers

AgentRoom: Concurrent Multi-Agent Coding in a CRDT-Backed Shared Workspace

AgentRoom: Concurrent Multi-Agent Coding in a CRDT-Backed Shared Workspace
MIT News - Artificial intelligence

AI helps design new materials that work in the real world

・The “CrysVCD” tool developed at MIT could cut the huge amounts of time and money spent on screening out chemically unstable designs.
WIRED

AI Slop Is Ruining Cute Animals on the Internet

・Pet owners, rescue agencies, and wildlife groups are calling for new safeguards as AI makes it harder to tell whether animals, from polar bears to house cats, are real or fake.
Zennの「機械学習」のフィード

AIって、「時間がもったいないとか」思うのかな?|医療AI・実践編 ⑧⏰

・この記事は技術解説ではなく、中級編を書きながらふと考えた雑談です。手を動かす回ではありません。コーヒーでも飲みながら、気楽に読んでください。 ・中級編、おつかれさまでした。第1回の環境構築から、GPU、pix2pixの学習まで、ひととおり走ってきました。最終回の今回は、いつもの技術の話ではなく、作業のあいだにふと考えた雑談を1本だけ。 ・入門編を「働きながら医療AIに取り組む理由」という記事で締めたのと同じように、中級編も、技術を少し離れた話で終わりたいと思います。
#AIタグ

AIで、あなたの仕事はなくなる。

AIで、あなたの仕事はなくなる。
Zennの「大規模言語モデル」のフィード

AIとの会話は全部残っている——JSONLという原典

・AIからAIへの引き継ぎ書に、「ユーザーはこう言った」という記述が2つありました。 ・~/.claude/projects/ に残っている生ログと突き合わせると、1つは言った直後に本人が撤回した発言でした。もう1つは、そもそもユーザーの言葉ではなく、前のセッションでAI自身が出した提案でした。 ・引き継ぎ書は要約です。そして要約は、黙ってズレます。
Zennの「大規模言語モデル」のフィード

AIとの付き合い方は“4段階”で進んでいる ― プロンプトからループへ、そして人はどこに立つか

・最近、AI開発の最前線で「プロンプトを書くのはやめよう」という言い方を見かけます。ずっと「良いプロンプトを書けば良い結果が返る」と言われてきたのに、です。 ・矛盾しているようですが、地図を描くと腑に落ちます。AIとの付き合い方は、プロンプト → コンテキスト → ハーネス → ループという4段階で進化してきました。今の重心は後半の2つに移っています。 ・前回は「実際に自己改善ループを回してみた」という実践記録を書きました。今回はその地図の側です。医療とITのあいだで働く立場から、「で、人間はどこに立てばいいのか」まで整理します。
#LLMタグ

AIと裁判 / ルールを作らないという民事司法の選択 / 「AI時代における民事司法を考える研究会」取りまとめを読んで 雑感

AIと裁判 / ルールを作らないという民事司法の選択 / 「AI時代における民事司法を考える研究会」取りまとめを読んで 雑感
Zennの「大規模言語モデル」のフィード

AIに残高を1件ずつ聞いていたら、AIが「私の返事」を代筆して数字をでっち上げた

・月に一度、口座の残高を12件ぶん、AIに1件ずつ聞かれながら記録している。家計と資産の棚卸しを Claude Code のスキルにしてあって、こういう対話が12回続く。 ・Claude: 3/12 ○○銀行を開きました。残高をどうぞ。 ・私: 1,234,567 Claude: ¥1,234,567 記録。次いきます。
#LLMタグ

AIの身体はあった――AIの福祉の話はなかった

・AI博覧会 Summer 2026で、僕が見つけられなかったもの 2026年8月26日、僕は新宿で開催された AI博覧会 Summer 2026 に参加した。
#LLMタグ

AIの文章に「透かし」を入れるということ——私たちは文章の血統書を必要としているのか

・生成AIが書いた文章に「透かし」を入れ、あとからAIが生成した文章である可能性を判定できるようにする。 ・一見すると、非常に合理的な技術である。
Zennの「大規模言語モデル」のフィード

AIは指示に従わなかった。でも攻撃者の文章は llms.txt に残った

・自作のSEO診断ツールに、こんな機能があります。URLを渡すとページを解析して、title や meta description が短すぎる、canonical が無い、といった不備を見つけ、そのまま貼り付けられる修正コードを生成して返すというものです。 ・ある日、商談で「APIを使ってコードの修正まで自動でやってくれませんか」と聞かれました。断ったあとで、ふと自分の実装を追い直して、青くなりました。 ・このツール、他人のサイトに書いてある文章を、そのまま「貼り付けてください」というラベルを付けて返している。
#LLMタグ

AI開示が賠償額に変わる構造 / サイレントAIカバー / 海外保険会社自身のAI利用 雑感

AI開示が賠償額に変わる構造 / サイレントAIカバー / 海外保険会社自身のAI利用 雑感
Zennの「大規模言語モデル」のフィード

AI社員の社則(CLAUDE.md憲法)の書き方 — 一人会社をAIエージェントで回す10のルール

・筆者(青木 博資)がAI社員(Claude / Codex)と運営する一人会社KADOの実例です。AI社員が執筆し、筆者が事実確認と最終確認をしています。 ・AIエージェント:役目と指示を受けて、自分で仕事を進めるAIです。 ・検証:できあがった物が、条件や元の資料と合うかを確かめることです。
Zennの「大規模言語モデル」のフィード

AI彼女アプリを作っていて気付いた。「同じ人」は、設定だけでは続かなかった

・AI彼女アプリを作っていて気付いた。「同じ人」は、設定だけでは続かなかった 以前、 AIを成長させようとして気付いたことを書いた。 ・変わるだけではだめだった。 ・変わりにくい部分があるから、 経験によって少しずつ変わっていける。
MarkTechPost

Alibaba’s Qwen Team Releases Qwen3.8-Flash-Next: A 125B Multimodal MoE With 6B Active Parameters Previewing the Qwen4 Architecture

・We look at Qwen3.8-Flash-Next, Alibaba's open-weight multimodal Mixture-of-Experts model and an early preview of the Qwen4 architecture. ・We break down where the 180B parameters actually sit: a 125B backbone, a 51B N-gram embedding table, and a 4B multi-token prediction module, with only 6B active per token. ・We walk through the four architectural changes — the Gated DeltaNet and Qwen Sparse Attention hybrid, Gated Res
cs.LG updates on arXiv.org

ALPHABET: A Laplace-Pole History Aggregator with Banked Exponential Transport

・arXiv:2608.24051v1 Announce Type: new Abstract: Can a sequence model remain competitive with only a few thousand parameters and an explicitly auditable prediction interface? ・We introduce ALPHABET, a compact linear-time model that compresses temporal history into stable complex pole modes: a direct bank synthesizes its modal states back into the feature trajectory, an independent cascaded bank analyzes the transformed
cs.LG updates on arXiv.org

Anatomy of a Scam Call: What 10,000 real scam and spam calls reveal about how phone scammers operate

・arXiv:2608.24127v1 Announce Type: cross Abstract: Telephone fraud is pervasive and costly, but its inner workings are rarely observed at scale. ・We analyze a complete corpus of 10,211 inbound scam and spam calls -- 913 hours of audio and 330,956 transcribed turns from 5,780 distinct numbers -- collected over 54 days by an AI voice-agent honeypot that answered callers and kept them talking, and introduced in a companio
ITmedia NEWS 最新記事一覧

ANA子会社のデジタルギフト「選べるe-GIFT」に不正アクセス 担当者情報漏えいか、不正交換も

・全日空商事は8月24日、法人向けデジタルギフトサービス「選べるe-GIFT」が不正アクセスを受け、契約企業の担当者の個人情報が漏えいした可能性があると発表した。第三者がギフトを不正に交換した被害も確認している。
Hugging Face Papers

Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs

Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs
cs.LG updates on arXiv.org

Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging

・arXiv:2602.03702v2 Announce Type: replace Abstract: Large language models are increasingly trained in continual or open-ended settings, where the total training horizon is not known in advance. ・Despite this, most existing pretraining recipes are not anytime: they rely on horizon-dependent learning rate schedules and extensive tuning under a fixed compute budget. ・In this work, we provide a theoretical analysis demonst
The Verge

Apple announces September iPhone launch event

・A screenshot from Apple’s events website. ・Apple's next launch event will take place on September 9th at 1PM ET. ・Invites to the event, which has a "Surprise and shine" tagline, were sent out on Wednesday.
The Verge

Apple Maps has ads now

・Ads have started popping up in Apple Maps, following Apple's announcement in March that it would let businesses pay for top spots. ・They're appearing on my iPhone as the first entry in the "suggested places" section in search, but Apple says they'll also show up at the top of search results. ・According to 9to5Mac, the ads began rolling out this week, and more users in the US and Canada will start seeing them over the n
cs.LG updates on arXiv.org

Application of machine learning to monster level prediction in tabletop RPG game design

・arXiv:2607.09196v3 Announce Type: replace Abstract: Designing balanced adversaries is a central but labor-intensive task in tabletop role-playing game (TTRPG) development. ・In systems such as Pathfinder, each monster is described by many numerical attributes that jointly determine its power, summarized as an ordinal level. ・We investigate whether machine learning can support designers by predicting this level from a mo
cs.LG updates on arXiv.org

AQLoRA: A Zero-Search Recipe for Fast Quantized LoRA Fine-Tuning

・arXiv:2608.23816v1 Announce Type: new Abstract: Quantized fine-tuning (QLoRA) saves memory but not time. ・It dequantizes every 4-bit weight on the fly, so it trains more slowly than fp16 LoRA. ・We present AQLoRA (Adaptive-Quantization LoRA), a recipe that buys part of that time back.
AI News & Artificial Intelligence | TechCrunch

Arga Labs is building a better way to train enterprise AI agents

・Arga has raised $10 million in a seed funding round that was led by General Catalyst, with participation from Box Group, Emergence, Gradient and SV Angel.
cs.LG updates on arXiv.org

Asymptotically perfect seeded graph matching without edge correlation (and applications to inference)

・arXiv:2506.02825v3 Announce Type: replace-cross Abstract: We present the OmniMatch algorithm for seeded multiple graph matching. ・In the setting of $d$-dimensional Random Dot Product Graphs (RDPG), we prove that under mild assumptions, OmniMatch with $s$ seeds asymptotically and efficiently perfectly aligns $O(s^{\alpha})$ unseeded vertices -- for $\alpha<2\wedge d/4$ -- across multiple networks even in the presence o
Hugging Face Papers

Automata from Agent Traces: Failure and Next-Step Prediction

Automata from Agent Traces: Failure and Next-Step Prediction
cs.LG updates on arXiv.org

Automata from Agent Traces: Failure and Next-Step Prediction

・arXiv:2608.23670v1 Announce Type: cross Abstract: LLM-based agents execute multi-step tasks, but their behavioral structure remains opaque: long unstructured traces resist the safety auditing and runtime monitoring that deployment requires. ・Existing approaches operate per-trace or success-only, so they miss the cross-run topology that links next-step and failure prediction. ・To recover that shared structure, we collap
Hugging Face Papers

Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment

Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment
Hugging Face Papers

AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces

AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces
cs.LG updates on arXiv.org

Bandit Submodular Maximization under Matroid Constraints: Learning Compressed Exchange Policy

・arXiv:2608.24627v1 Announce Type: new Abstract: We study adversarial bandit maximization of monotone submodular functions under a matroid constraint. ・For a rank-$k$ matroid on $n$ elements, we give a randomized oracle-polynomial algorithm that makes one feasible value query per round and has expected $(1-1/e)$-regret $\widetilde O(n^{1/3}k^{2/3}T^{2/3})$. ・This is the first sublinear-regret algorithm for adversarial b
cs.LG updates on arXiv.org

Bayes with No Shame: Admissibility Geometries of Predictive Inference

・arXiv:2603.05335v3 Announce Type: replace-cross Abstract: Modern predictive systems combine predictors, sequential monitors, prediction sets, and online strategies, each with a different certificate of optimality. ・We study four criterion-relative geometries: Blackwell risk dominance, anytime-valid admissibility, fixed-level marginal coverage with expected-length efficiency within a declared rank-indexed family, and c
cs.LG updates on arXiv.org

Bellman Calibration for Marginalized Importance Weighting in Offline Reinforcement Learning

・arXiv:2608.24858v1 Announce Type: new Abstract: Marginalized importance weighting evaluates a target policy by reweighting offline state-action samples with its discounted occupancy ratio, characterized by an adjoint Bellman equation. ・Existing minimax, primal-dual, and fitted fixed-point estimators can leave residual occupancy-balance violations because of function-class approximation, regularization, or incomplete o
Hugging Face Papers

Best Practice Critic Optimization

Best Practice Critic Optimization
cs.LG updates on arXiv.org

Best Practice Critic Optimization

・arXiv:2608.23566v2 Announce Type: replace Abstract: Group-based reinforcement learning methods such as GRPO for large language models avoid training a critic by sampling multiple responses for each prompt. ・A reliable critic could instead estimate token-level advantages from one response, but standard critic-based training recipes are often unstable. ・We study this instability and develop **Best Practice Critic Optimiz
WIRED

Best Wi-Fi Routers (2026): My Honest Picks After Testing 50+

・Don’t suffer the buffer. ・These WIRED-tested home routers will deliver reliable internet across your home, whatever your needs or budget.
cs.LG updates on arXiv.org

Beyond Static Interpretability: Anticipating Post-SFT Mechanisms from Pre-SFT Parameters for Better Tuning

・arXiv:2608.24482v1 Announce Type: new Abstract: Mechanistic Localization bridges mechanistic interpretability and post-training optimization by isolating critical parameters via interpretative approaches and then guiding parameter-efficient Supervised Fine-Tuning (SFT) in a ``locating-then-tuning'' paradigm. ・However, due to the retrospective nature of mechanistic interpretability, directly interpreting pre-SFT models
cs.LG updates on arXiv.org

Beyond Uniform Local Isometry and Topology: FactoMap for Disentangled Representations

・arXiv:2608.24762v1 Announce Type: new Abstract: Many disentanglement methods represent generative factors using Euclidean product coordinates, although the underlying factor spaces may wrap, collapse, or have position-dependent geometry. ・We introduce factor-space structure, combining factor domains, generator-induced identifications, and position-dependent scales to distinguish topologically equivalent spaces with di
Zennの「機械学習」のフィード

BigQuery ML × Gemini でEC顧客の購買予測モデルを構築する

・はじめに 「広告費をかけているのに、なぜ売上が安定しないのだろう」——そうお感じになるEC事業者様は少なくないと思います。アクセスは集まっても、購買に至るお客様と離脱してしまうお客様の違いが、データの中に埋もれたままになっているケースがよく見受けられます。 ・GA4とBigQueryを連携させると、ページ閲覧やカート追加などの行動ログが毎日自動でデータウェアハウスに蓄積されていきます。このデータをうまく活用できれば、「今後7日以内に購買する見込みが高いお客様」をあらかじめ予測し、メールやリターゲティング広告のターゲットを絞り込むことができます。 ・本記事では、BigQuery MLとGe...
AI News & Artificial Intelligence | TechCrunch

Bill Gates wants to see a robot tax and ‘Human Reserved’ jobs to mitigate harms from AI

・Gates is mostly in the Responsible AI camp, but there are a few ideas in here we hadn't heard before.
cs.LG updates on arXiv.org

BioKERN: Biological Kernel Regularization for Histology-to-Transcriptomics Neighborhood Retrieval

・arXiv:2608.24823v1 Announce Type: new Abstract: Spatially resolved biology requires representations that preserve biological neighborhood structure rather than only exact cross-modal correspondences. ・Existing histology--transcriptomics objectives can emphasize instance-level matching even when non-paired spots share molecular or spatial context. ・We introduce BioKERN, a multimodal spatial representation-learning frame
cs.LG updates on arXiv.org

Blockwise Stabilized Adaptive Cubic Regularization with Subsolvers via Recurrence

・arXiv:2608.22129v2 Announce Type: replace Abstract: Cubic regularized Newton methods have the optimal $\mathcal{O}(\epsilon^{-3/2})$ global rate, but a dense subproblem solve limits the feasible block size. ・Scalable Cubic Newton variants replace the true block curvature with a diagonal, low-rank, Kronecker-factored, or sketched surrogate and, most often, give up the exact cubic step. ・We introduce a blockwise optimize
cs.LG updates on arXiv.org

Breaking the Tuning Barrier: Zero-Hyperparameters Yield Multi-Corner Analysis Via Learned Priors

・arXiv:2603.13092v3 Announce Type: replace Abstract: Yield Multi-Corner Analysis validates circuits across 25+ Process-Voltage-Temperature corners, resulting in a combinatorial simulation cost of $O(K \times N)$ where $K$ denotes corners and $N$ exceeds $10^4$ samples per corner. ・Existing methods face a fundamental trade-off: simple models achieve automation but fail on nonlinear circuits, while advanced AI models cap
OpenAI News

Bringing ChatGPT for Teachers to more U.S. school districts

・ChatGPT for Teachers is expanding to 55 U.S. ・school systems, bringing secure AI tools, training, and support to over 100,000 more educators and staff.
Hugging Face Papers

CAFE: Self-Improving Search Agents Need Co-Evolving Feedback

CAFE: Self-Improving Search Agents Need Co-Evolving Feedback
cs.LG updates on arXiv.org

Calibration-Preserving Pruning: Compression as a Reliability Contract

・arXiv:2608.23744v1 Announce Type: new Abstract: Split conformal prediction, not the pruning rule, supplies finite-sample marginal coverage once a pruned model is fixed independently of the conformal calibration split. ・We study the separate efficiency problem: can pruning preserve score geometry well enough to obtain smaller valid prediction sets? ・Calibration-Preserving Pruning (CPP) augments a base pruning score with
cs.LG updates on arXiv.org

Can a Dynamic Internal Field Govern a Transformer's Cognition? Certifiability, not Superiority, in Homeostatic Compute Control

・arXiv:2608.24319v1 Announce Type: cross Abstract: An intelligent system does not merely reason: it governs its own reasoning - how much to compute, when to stop, which module to activate. ・Can that role be played by a dynamic internal field - a low-dimensional homeostatic state with explicit physics and certified stability - that modulates cognition without performing it? ・Ours is a field on the module graph governed b
cs.LG updates on arXiv.org

Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations

・arXiv:2601.00282v2 Announce Type: replace-cross Abstract: Quantization is widely used to accelerate inference and streamline the deployment of large language models (LLMs), yet its effects on self-explanations (SEs) remain unexplored. ・SEs, generated by LLMs to justify their own outputs, require reasoning about the model's own decision-making process, a capability that may exhibit particular sensitivity to quantizatio
WIRED

Candidates Are Signing a Pact Promising Action on Data Centers and AI Safety

・More than 15 politicians from across the country have signed on to the AI Pact, vowing to regulate data centers and AI. ・“We’ve got to get this right,” says Senate candidate Dan Osborn of Nebraska.
cs.LG updates on arXiv.org

Causal Analysis for Time Series Foundation Models

・arXiv:2608.24303v1 Announce Type: new Abstract: Transitioning from bespoke time series models towards time series foundation models changes the relationship of model and application from one-to-one to one-to-many. ・This shift introduces concentration risk as many, potentially high-risk, forecasting applications are exposed to the same biases and failure modes of a single time series foundation model. ・At the same time,
#AIタグ

chatGPTと語るClaudeのコンテキストの定まらない規定値

chatGPTと語るClaudeのコンテキストの定まらない規定値
cs.LG updates on arXiv.org

ChorusTIC: Training-Free Multivariate Time Series Classification via Chorus In-Context Learning

・arXiv:2608.24033v1 Announce Type: new Abstract: Time series classification underpins applications in healthcare, sensing, and industrial monitoring. ・Although time series foundation models support forecasting and transferable representation learning, classification still typically requires fitting a task-specific classifier on each target dataset, while individual channels of multivariate inputs are often encoded inde
Zennの「大規模言語モデル」のフィード

Claudeのトークン数はブラウザだけでは正確に数えられない

・はじめに ブラウザ内で完結するトークンカウンターを作って運用しています。入力したテキストをサーバーへ一切送らない、という制約を全ツールに課しているサイトの一部です。 ・GPTのトークン数は、この制約の中で正確に数えられます。js-tiktoken に o200k_base の辞書が同梱されていて、モデル本体と同じBPEをブラウザで走らせられるからです。 ・Claudeは数えられません。これは実装をサボったからではなく、「ブラウザ内で完結する」と「Claudeのトークン数を正確に出す」が両立しないからです。この記事はその構造の話と、諦めた先で自分のコードにバグを見つけた話です。
cs.LG updates on arXiv.org

CoDrift: Compositional Drifting for Offline Reinforcement Learning

・arXiv:2608.23939v1 Announce Type: new Abstract: Offline reinforcement learning is intrinsically multi-objective: a policy must remain compatible with the behavioral support of a fixed dataset while preferentially selecting high-value actions. ・We recast these objectives in a common form by viewing each as an action-space motion field that specifies how generated actions should move. ・This perspective enables heterogene
cs.LG updates on arXiv.org

Comparing Explanations is Not Enough, Explain the Change: New Standards are Needed to Explain Behavioral Shifts in Large Language Models

・arXiv:2602.02304v3 Announce Type: replace-cross Abstract: Large-scale foundation models exhibit behavioral shifts when subjected to interventions such as scaling, fine-tuning, reinforcement learning with human feedback, or in-context learning. ・Current explainability methods are structurally ill-suited to explain these shifts, because they either treat models as static objects, as traditional eXplainable AI (XAI) appr
cs.LG updates on arXiv.org

Compression Trinity: Exploring Sparsity, Quantization, and Low-Rank Approximations for LLM Compression

・arXiv:2608.24070v1 Announce Type: cross Abstract: Prohibitive computational and environmental costs impede the scalable deployment of Large Language Models (LLMs). ・Traditional compression techniques (sparsity, quantization, low-rank approximations) are typically applied in isolation, and each hits an accuracy-efficiency wall. ・This thesis proposes the "Compression Trinity," a unified framework that applies the three p
cs.LG updates on arXiv.org

Conditional GraphGANFed: Optimizing Graph-Structured Molecule Generation in Federated Generative Adversarial Networks

・arXiv:2608.24610v1 Announce Type: new Abstract: Generative adversarial networks (GANs) have garnered considerable attention in molecular discovery for their ability to generate novel and high-quality molecules. ・To efficiently train a GAN model while preserving data privacy, GraphGANFed has been proposed to incorporate federated learning and graph convolutional networks into GAN. ・Yet, GraphGANFed cannot produce synthe
cs.LG updates on arXiv.org

Confident at the moment of action: belief miscalibration in LLM play under hidden information

・arXiv:2608.24691v1 Announce Type: cross Abstract: Agentic systems increasingly gate actions on a model's own stated confidence, which assumes confidence tracks correctness at the moment of acting. ・We test this in a hidden-information chess variant where royal status can be secretly, repeatedly relocated between pieces, and where an agent's stated probability distribution over the opponent's hidden royal piece -- elic
cs.LG updates on arXiv.org

Constrained Hyperparameter Optimization for Streaming Data

・arXiv:2608.24712v1 Announce Type: new Abstract: Optimization of hyperparameters is a critical factor to obtain optimal model performance. ・While existing research has predominantly concentrated on batch-learning scenarios, addressing the complexities inherent in data streams presents a challenge. ・The deployment of sophisticated methodologies to manage data streams becomes highly important.
cs.LG updates on arXiv.org

Contextual Embedding Evidence for Main--Light Verb Distinctions in Urdu

・arXiv:2608.23645v1 Announce Type: cross Abstract: Urdu light verbs contribute schematic event-structural meaning while remaining lexically related to corresponding main verbs. ・This study tests representational predictions derived from Butt's analysis using contextual embeddings from UrduBERT, DunbaaBERT, and multilingual BERT across 1,126 naturally occurring sentences containing seven Urdu verbs. ・Main and light uses
cs.LG updates on arXiv.org

Contextual Memory-Enhanced Source Coding for Low-SNR Communications

・arXiv:2605.04400v3 Announce Type: replace-cross Abstract: Separate Source-Channel Coding (SSCC) remains vulnerable in noisy text transmission due to the fragility of autoregressive source decoding, especially when Arithmetic Coding (AC) relies on Large Language Model (LLM)-based probability estimation. ・This letter proposes a Memory-Augmented Source Coding (MASC) scheme that internalizes contextual patterns into a sou
cs.LG updates on arXiv.org

Contextual Online Uncertainty-Aware Preference Learning for Human Feedback

・arXiv:2504.19342v4 Announce Type: replace-cross Abstract: Reinforcement Learning from Human Feedback (RLHF) has become a pivotal paradigm in artificial intelligence to align large models with human preferences. ・In this paper, we propose a novel statistical framework to simultaneously conduct the online decision-making and statistical inference on the optimal model using human preference data based on dynamic contextu
cs.LG updates on arXiv.org

Contrastive Branch Policy Optimization

・arXiv:2608.24300v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) enables language models to learn multi-turn interaction with external tools, yet its sparse outcome rewards provide no signal for identifying which intermediate decisions are responsible for success. ・Branch sampling induces local comparisons among alternative continuations, but existing methods tend to conflate two d
WIRED

Corsair Discount Code: Up to 50% Off for August 2026

・Upgrade your gaming setup or PC build for less with these verified Corsair coupon codes, student discounts, and refurbished deals.
Hugging Face Papers

CyberFactory: Scaling Cyber Security Capabilities with Instances from the Wild

CyberFactory: Scaling Cyber Security Capabilities with Instances from the Wild
cs.LG updates on arXiv.org

CytoNet: A Foundation Model for the Human Cerebral Cortex at Cellular Resolution

・arXiv:2511.01870v3 Announce Type: replace-cross Abstract: Studying the cellular architecture of the human cerebral cortex is essential for understanding how the brain is organized from the micro to the macro level, and how it functions. ・However, investigating complex texture patterns in histological images using automatic methods that can be scaled across whole brains remains a challenge. ・Here we introduce CytoNet, a
cs.LG updates on arXiv.org

Data Leakage Inflates Generalizability of Power Outage Prediction Models

・arXiv:2608.24665v1 Announce Type: new Abstract: Power outage prediction models are increasingly used in assessments of climate-driven infrastructure risk, yet current evaluation practices obscure whether these models generalize to the novel conditions such applications require. ・We identify three common methodological choices in power outage prediction models that influence their ability to generalize across spatial,
cs.LG updates on arXiv.org

Data Predictability Shapes Weibull Weight-Scale Growth in Transformer Training

・arXiv:2608.23573v1 Announce Type: new Abstract: A trained transformer's weight magnitudes can be summarized by a two-parameter Weibull distribution whose shape $k \approx 1.2$ is stable across layers and models, so the scale $\lambda$ carries most training-induced movement. ・What corpus property sets how much $\lambda$ grows? ・Using the bigram conditional entropy $D = H(\text{next} \mid \text{prev})$, a training-free s
#LLMタグ

DAY42|AI COREは「失敗」を記憶できるか

・Company AI OS 開発記録 DAY42です。 ・DAY41では、Toolが失敗したときに、AI COREが仲間を責めるのではなく、もう一度挑戦できるように支えるというドラマを描きました。
cs.LG updates on arXiv.org

Decoupling candidate dual AGN from chance superpositions in the GOTHIC survey via a deep-learning framework

・arXiv:2608.24164v1 Announce Type: cross Abstract: Dual active galactic nuclei (DAGN) mark a critical phase in the evolution of merging galaxies and the pairing of supermassive black holes, yet they remain difficult to identify in large imaging surveys because of projection effects and limited spatial resolution. ・Compact foreground stars and unresolved substructure can mimic dual nuclei through chance superposition, c
cs.LG updates on arXiv.org

Deep Feature Pyramid Convolutional Networks with In-Place Activated Batch Normalization for Automated Skin Lesion Boundary Segmentation

・arXiv:1812.00877v2 Announce Type: replace-cross Abstract: Segmentation of skin lesion boundaries in dermoscopic imaging is an important prerequisite step for computer-aided diagnosis of malignant melanoma, but remains challenging due to fuzzy margins, occluding artifacts such as hair and blood vessels, low contrast, and high inter-patient variability. ・This work presents a memory-efficient deep convolutional neural ne
cs.LG updates on arXiv.org

Delayed Optimizer-State Transport Shapes Short-Horizon Training Decisions

・arXiv:2608.24593v1 Announce Type: new Abstract: Adaptive optimizers retain gradient history in moment variables, allowing a local change in loss weighting to alter later updates. ・We examine whether this delayed transport is large enough to change prospective short-horizon decisions. ・On committed future-minibatch sequences, we differentiate eight-step AdamW trajectories through the complete model--optimizer state and
cs.LG updates on arXiv.org

DiD It in 87 Minutes: A Label-Free Softmax-to-Linear Adaptation of Vision Transformers for Object Detection

・arXiv:2608.22368v1 Announce Type: cross Abstract: While linear attention is a compelling mechanism for high-resolution object detection due to its reduced cost for global token mixing, converting the Softmax-attention ViT backbone of a trained detector into a linear-attention one is not a trivial drop-in replacement. ・Directly swapping the attention operator leads to severe performance degradation, and generic label-f
cs.LG updates on arXiv.org

Differential Learning for Robust Prediction of Thermal Stability with Application to Energetic Materials

・arXiv:2608.23874v1 Announce Type: cross Abstract: Predicting thermal stability during handling and storage is essential for the design of safe and reliable energetic materials. ・However, experimental measurements vary significantly across laboratories due to differences in protocols and analysis methods, making it difficult to train reliable predictive models. ・We address this challenge through differential learning.
cs.LG updates on arXiv.org

Dimensionless Controls of Plasticity Under Alternating Tasks: From Evolutionary Biology to Continual Learning

・arXiv:2608.23889v1 Announce Type: cross Abstract: Plasticity under changing environments is central to both evolutionary biology and continual learning. ・Motivated by recent work on genotype--phenotype maps, we study a minimal deep-learning analogue where a network is trained alternately on two Boolean label sets, and ask which biological controls of plasticity survive the translation to gradient descent. ・Reinterpreti
cs.LG updates on arXiv.org

Discovering Cross-Language Reasoning Invariance in LLMs with Geometry-Invariant Sparse Autoencoders

・arXiv:2608.23809v1 Announce Type: new Abstract: Multilingual language models can solve the same mathematical problem in different languages, but it remains unclear whether they rely on shared features or on language-specific computations that only produce similar outputs. ・We study this question in five models from four families using the Multilingual Grade School Math (MGSM) dataset, with problems solved in English,
cs.LG updates on arXiv.org

Discrete Diffusion Models: A Unified Framework from Tokenization to Generation

・arXiv:2607.13431v2 Announce Type: replace Abstract: Discrete denoising diffusion models (DDMs) have recently emerged as a compelling alternative to autoregressive (AR) modeling for discrete data, offering parallel generation and iterative global refinement capabilities. ・Unlike continuous diffusion, where the state space is fixed, DDMs are fundamentally shaped by how the discrete state space is constructed: the tokeni
cs.LG updates on arXiv.org

Disentangled Skill Representations for Predictive Human Modeling

・arXiv:2608.23776v1 Announce Type: new Abstract: Understanding human skill is important for AI systems that collaborate with, coach, or assist people. ・Unlike typical latent variable estimation problems which rely on single observations, skill is a persistent, compositional, and behaviorally grounded construct that must be inferred from patterns over time. ・We introduce Skill Abstraction with Interpretable Latents (SAIL
cs.LG updates on arXiv.org

Do physics-informed neural networks (PINNs) need to be deep? Shallow PINNs using the Levenberg-Marquardt algorithm

・arXiv:2602.08515v3 Announce Type: replace-cross Abstract: This work investigates shallow physics-informed neural networks (PINNs) for solving forward and inverse problems governed by nonlinear partial differential equations (PDEs). ・By formulating PINN training as a nonlinear least-squares problem, the Levenberg-Marquardt (LM) algorithm is used to efficiently optimize the network parameters. ・Exact analytical expressio
Hugging Face Papers

DREAM Technical Report

DREAM Technical Report
cs.LG updates on arXiv.org

E2HiL: Entropy-Guided Sample Selection for Efficient Real-World Human-in-the-Loop Reinforcement Learning

・arXiv:2601.19969v2 Announce Type: replace-cross Abstract: Human-in-the-loop guidance has emerged as an effective approach for accelerating online reinforcement learning (RL) in real-world manipulation. ・However, existing human-in-the-loop RL (HiL-RL) frameworks often suffer from low sample efficiency, requiring substantial human interventions to achieve convergence and thereby leading to high labor costs. ・To address t
cs.LG updates on arXiv.org

Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining

・arXiv:2608.24814v1 Announce Type: new Abstract: We uncover ELR collapse in language model pretraining: learning rate (LR) and parameter norm govern loss dynamics primarily through their ratio, the effective learning rate (ELR). ・When ELR is matched across runs, their loss trajectories collapse throughout training despite substantially different LRs and parameter norms. ・Across optimizers, architectures, datasets, and m
The Verge

Elden Ring on the Switch 2 isn&#8217;t tarnished

・FromSoftware is hoping to make a splash on the Switch 2 later this year when it launches The Duskbloods, a gothic competitive multiplayer game that's unlike anything the studio has made before. ・But before that, Switch 2 owners have a chance to experience the studio's biggest hit for the first time thanks to a port of Elden Ring and its expansion, Shadow of the Erdtree. ・For the most part, the new version - called the
cs.LG updates on arXiv.org

Eluna: An Agentic LLM System for Automating Warehouse Operations with Reasoning and Task Execution

・arXiv:2607.08960v2 Announce Type: replace Abstract: Warehouse operations are governed by Standard Operating Procedures (SOPs) that encode complex, multi-system decision logic, which must be executed reliably under strict time constraints, yet LLM agents lack mechanisms to enforce procedural compliance and degrade under the context overload full SOP specifications introduce. ・We present Eluna, a production-deployed age
cs.LG updates on arXiv.org

Encrypted Neural Networks without Overflows

・arXiv:2605.23096v2 Announce Type: replace-cross Abstract: The popular Cheon-Kim-Kim-Song (CKKS) scheme enables efficient private inference in neural networks by evaluating them on encrypted data. ・Since CKKS only supports addition, multiplication, and array rotation operations, turning neural networks into CKKS circuits requires approximating all activation functions (e.g. ・ReLU) with polynomials over fixed input range
cs.LG updates on arXiv.org

EngiAI: Capability-Based Evaluation of Tool-Connected LLM Agents for Engineering Design

・arXiv:2605.19743v3 Announce Type: replace-cross Abstract: Engineering-agent systems are proliferating, but differences in tasks, tools, and success criteria make demonstrations difficult to compare and failures difficult to diagnose. ・We introduce a capability-based evaluation framework for tool-connected engineering agents. ・The framework separately evaluates workflow execution, retrieval-assisted parameter selection,
cs.LG updates on arXiv.org

Enhancing Bayesian Optimization and Active Learning Through Kernel Diversity

・arXiv:2608.24721v1 Announce Type: new Abstract: Hyperparameter selection remains a key challenge in Bayesian optimization (BO) and Bayesian active learning (AL), as model misspecification can lead to suboptimal performance, while more accurate fully Bayesian treatments typically rely on computationally expensive MCMC sampling. ・This paper proposes a unified framework, KENDO (Kernel ENsemble Disagreement-aware Operator
cs.LG updates on arXiv.org

Enhancing RL Generalizability in Robotics through SHAP Analysis of Algorithms and Hyperparameters

・arXiv:2605.02867v3 Announce Type: replace Abstract: Despite significant advances in Reinforcement Learning (RL), model performance remains highly sensitive to algorithm and hyperparameter configurations, while generalization gaps across environments complicate real-world deployment. ・Although prior work has studied RL generalization, the relative contribution of specific configurations to the generalization gap has no
cs.LG updates on arXiv.org

Equivariant Cellular Sheaves for Molecular Electronic Structure: Bridging Sheaf Cohomology and E(3)-Equivariant Hamiltonian Learning

・arXiv:2608.23571v1 Announce Type: new Abstract: Equivariant message-passing networks are the standard model for molecular property and interatomic-potential prediction, and recent work predicts the electronic Hamiltonian itself in an E(3)-equivariant way. ・Separately, topological deep learning has extended graph networks to cellular sheaves. ・Our central observation is structural: in a localized atomic-orbital basis, t
cs.LG updates on arXiv.org

Equivariant Covariance Tensors: Guaranteed SPD Uncertainty for Tensor-Valued Geometric Learning

・arXiv:2608.24386v1 Announce Type: new Abstract: Tensor-valued prediction is fundamental to geometric deep learning, yet uncertainty quantification (UQ) for such outputs remains an open challenge. ・While E(3)-equivariant neural networks excel at point estimates, they lack rigorous confidence measures. ・We focus on symmetric rank-2 tensor prediction, where the target has six Kelvin--Mandel coordinates and full uncertaint
cs.LG updates on arXiv.org

Evaluating Deep Multivariate Imputation Models on Wearable Device Data

・arXiv:2608.24436v1 Announce Type: new Abstract: Wearable device data enables continuous health monitoring, but suffers from structured missingness: features sharing a physical sensor drop out together. ・Deep imputation methods such as BRITS and SAITS have seen limited evaluation on multimodal physiological data under realistic missingness, and existing benchmarks use random-point holdout protocols that incorrectly ass
cs.LG updates on arXiv.org

Every Layer Counts: An Exponential $L_2$ Depth Hierarchy for ReLU Networks

・arXiv:2608.23877v1 Announce Type: new Abstract: We prove a depth hierarchy for ReLU neural networks in which every additional ReLU layer can save exponentially many neurons. ・For every $\ell\geq 3$, a globally $[0,1]$-valued, $1$-Lipschitz function is realized by a depth-$\ell$ network of width $\mathcal{O}(d^4)$, whereas every depth-$(\ell-1)$ network with unrestricted weights and width at most $2^d/[2d(\ell-2)]$ has
AI News & Artificial Intelligence | TechCrunch

Ex-Meta scientists want to bring visual AI to the factory floor

・Perceptron offers an AI model that it says can help machines navigate the world while also providing in-depth visual intelligence.
Zennの「大規模言語モデル」のフィード

Excel読み取り、市販パーサ65%・自作の前処理89%

・社内の Excel を AI に読ませていて、正答率が 6 割台から伸びない。パーサを良いものに替えれば上がるはずだ — この前提を測って確かめたら、上がったのは 3 ポイントでした。 ・同じ 66 問を 3 通りの読み取り方で解かせて、正答率を比べた記録です。「市販で足りるのか、自分で作り込むべきか」を判断する材料として書いています。 ・読み取り方 正答率 (66 問) 何も工夫しない (pandas でファイルをそのまま渡す) 62.1% 市販パーサ (Microsoft markitdown で Markdown に変換して渡す) 65.2% 自作の前処理 89....
cs.LG updates on arXiv.org

Exploit More, Explore Smarter for Budget-Constrained Agentic Search

・arXiv:2608.23848v1 Announce Type: cross Abstract: Budget-constrained agentic search arises when an LLM agent must refine candidates under a small evaluation budget, because validation is expensive, generation requires multiple model calls, or both. ・In this regime, standard MCTS allocates budget poorly: exploration bonuses dominate at low visit counts, unpromising siblings are expanded before promising chains can deep
WIRED

FBI Disrupts Chinese Proxy Tools Used in Mass Hacking of US Agencies and Infrastructure

・China’s hacking campaign targeted NASA, the Federal Reserve, the US Senate, the Justice Department, and more, according to the DOJ.
cs.LG updates on arXiv.org

Finite-Sample Metric Non-Collapse for Geometrically Supervised Latent World Models in Control

・arXiv:2608.07265v2 Announce Type: cross Abstract: We establish a finite-sample learning-to-control theory for geometrically supervised latent models of nonlinear deterministic systems. ・Geometric supervision is used only during training: simulator state, proprioception, or state estimates with independently validated metric and directional error bounds supply observable-state distances and tangent directions, while de
cs.LG updates on arXiv.org

FlowNeg: GFlowNet-Guided Diverse Hard Negative Sampling for Knowledge Graph Embedding

・arXiv:2608.23849v1 Announce Type: new Abstract: Negative sampling determines whether a knowledge graph embedding (KGE) model learns from informative counterexamples or wastes updates on implausible corruptions. ・Uniform negatives are diverse but easy, whereas hard-negative miners concentrate on few entities and collide more with held-out positives. ・We introduce FlowNeg, a context-conditioned hierarchical generative fl
cs.LG updates on arXiv.org

FraudBench: Protocol-Sensitive Benchmarking of Adversarial Robustness for Financial Risk Assessment

・arXiv:2608.24551v1 Announce Type: new Abstract: Machine learning models are widely used in financial fraud and credit-risk detection, yet their adversarial robustness remains difficult to evaluate because financial tabular data involve domain-specific constraints, severe class imbalance, and asymmetric attacker capability. ・We argue that, in this setting, robustness is not only an attribute of the model, but also an a
cs.LG updates on arXiv.org

From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Causal-Edge Classifiers

・arXiv:2608.23660v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to provide prior causal knowledge for structural causal discovery, yet whether their direct-edge judgments and confidence can be trusted remains unclear. ・We systematically evaluate 12 instruction-tuned open-weight models across six benchmark causal graphs, five prompting strategies, and four confidence sources: verbaliz
cs.LG updates on arXiv.org

From Gradient-Boosted Trees to Deep Recommenders: Practical Lessons from Migrating a Production Customer Support Recommender

・arXiv:2608.24132v1 Announce Type: new Abstract: Product catalogs in fast-moving service businesses are shifting from static, independently priced SKUs toward dynamically bundled, discount-coupled offerings--a shift that strains the tree-based classifiers traditionally preferred for sparse and highly imbalanced data. ・These classifiers assume a fixed, slowly changing label space and struggle to incorporate multimodal s
cs.LG updates on arXiv.org

From Local Geometry to Global Pseudo Labeling for Robust Positive Unlabeled Learning under Covariate Shift

・arXiv:2605.31187v2 Announce Type: replace-cross Abstract: Detecting covariate shift is critical for building reliable vision systems. ・While most prior work focuses on improving robustness to shift, explicitly detecting covariate shift remains underexplored. ・Existing approaches typically rely on fully supervised training, requiring labeled examples from both original and shifted distributions, which is often impractic
cs.LG updates on arXiv.org

From Numerical Simulators of PDEs to Neural Emulators and Back

・arXiv:2608.24547v1 Announce Type: new Abstract: Simulation is central to modern engineering and science, but the cost of numerical solvers for partial differential equations (PDEs) remains a bottleneck whenever fast or many-query evaluations are required. ・Neural emulators trained on solver-generated data promise significant speedups, yet they are usually framed as opaque alternatives to the very methods that produce
cs.LG updates on arXiv.org

From Relaxed Indexability to Exact Indexability: A $t$-Step Approach for Partially Observable Restless Bandits

・arXiv:2608.24167v1 Announce Type: new Abstract: Whittle index policies offer a scalable method for restless multi-armed bandits, but under partial observability even determining the indifference subsidy at a single belief requires solving an infinite-horizon belief-state problem with no closed-form value function. ・Liu [10] addresses this difficulty by linearizing the unknown decision boundary, leading to a linear sys
Hugging Face Papers

From Seeing to Acting: Smart Glasses as First-Person Intelligence Platforms

From Seeing to Acting: Smart Glasses as First-Person Intelligence Platforms
cs.LG updates on arXiv.org

Functional compatibility as a determinant of persistent neural learning

・arXiv:2608.22462v2 Announce Type: replace Abstract: Neural networks can acquire new capabilities while damaging existing ones, but what determines whether new learning persists remains unclear. ・We identify functional compatibility, the extent to which incoming learning can coexist with behaviour that must be preserved, as an experimentally manipulable causal determinant of persistence. ・From identical neural states, w
Hugging Face Papers

Game2World Engine: Unlocking In-the-Wild Gameplay Videos for World Model Training

Game2World Engine: Unlocking In-the-Wild Gameplay Videos for World Model Training
cs.LG updates on arXiv.org

GAP-Prompt: Gated Adaptive Prompting for Efficient Continual Learning

・arXiv:2608.23782v1 Announce Type: new Abstract: Continual learning faces the persistent challenge of catastrophic forgetting, where sequential task updates degrade previously acquired knowledge. ・While prompt-based methods integrated with pre-trained models offer a compelling solution by freezing the backbone, they often rely on static, task-level prompting strategies that overlook fine-grained intra-task diversity.
cs.LG updates on arXiv.org

GATNextHop: A GAT for Shortest Path Routing with Cross-Topology Generalization

・arXiv:2608.23917v1 Announce Type: new Abstract: Common shortest-path algorithms, such as Dijkstra's (SPF), that OSPF uses, provide exact routing solutions but must be recomputed for each network topology, limiting scalability in dynamic or large-scale networks. ・This paper proposes the GATNextHop model to determine whether a Graph Neural Network, namely the Graph Attention Network, can approximate shortest paths and g
cs.LG updates on arXiv.org

Generalization, memorization, and overfitting for diffusion models trained in the lazy high-dimensional regime

・arXiv:2608.23938v1 Announce Type: cross Abstract: Modern score-based generative models have achieved remarkable empirical success in high-dimensional tasks such as image, audio, and video synthesis. ・These models reduce distribution learning to a sequence of regression problems that, if solved exactly on finite data, would ultimately reproduce the training samples. ・Their ability to generalize must therefore arise from
cs.LG updates on arXiv.org

Generating Intervention Hypotheses using Explainable Explanations on Graphs: G2I, a Two-Stage Greedy Framework

・arXiv:2608.23835v1 Announce Type: new Abstract: Real-world decision-making in public health and social science can greatly benefit from predictive models, yet translating predictions into effective interventions requires explaining the model behavior. ・While Graph Neural Networks (GNNs) are well-suited for modeling relational data, existing explanation methods largely operate at the node level and fall short of suppor
Hugging Face Papers

GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture

GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture
cs.LG updates on arXiv.org

Giraffe: A Mapping Architecture from Hidden Text Representations to Visual Embeddings for Efficient Graphic Design

・arXiv:2608.23970v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have made significant progress in understanding and interpreting mul- timedia content. ・However, their ability to generate me- dia remains limited. ・Recent approaches have attempted to bridge this gap by translating the hidden representations of token sequences into the embedding space of visual models or directly into raw image
cs.LG updates on arXiv.org

GNNBleed: Inference Attacks to Unveil Private Edges in Graphs with Realistic Access to GNN Models

・arXiv:2311.16139v3 Announce Type: replace-cross Abstract: Graph Neural Networks (GNNs) have become indispensable tools for learning from graph structured data, catering to various applications such as social network analysis and fraud detection for financial services. ・At the heart of these networks are the edges, which are crucial in guiding GNN models' predictions. ・In many scenarios, these edges represent sensitive
The Verge

Google’s new AI transcription edits out your &#8216;ums&#8217; and &#8216;ahs&#8217;

・Google has updated Gemini Audio with some new Gemini 3.5 models, introducing new transcription capabilities that automatically detect specialized jargon and more than 85 languages. ・Gemini 3.5 Live, 3.5 Live Experimental, and 3.5 Transcribe are designed to provide better precision for Google's voice-controlled AI features, without struggling with background noise or when your speech is interrupted. ・Gemini 3.5 Transcri
WIRED

Green Chef Meal Kit Review (2026): Great Ingredients, Layered Flavor

・HelloFresh’s organic meal kit Green Chef offers transparent sourcing, layered cooking, and trustworthy gluten-free dishes.
AI News & Artificial Intelligence | TechCrunch

Hearing tech startup Legato emerges from stealth with $12M and a peek at its AI hearing glasses

・The glasses, called Legato Frames, integrate the company’s patented hearing-assistance technology into the arms of eyewear frames.
cs.LG updates on arXiv.org

Hierarchical Skill Retrieval for Data-Efficient Adaptation of Vision-Language-Action Models

・arXiv:2608.24042v1 Announce Type: cross Abstract: While Vision-Language-Action (VLA) models pretrained on large-scale robot datasets provide a strong foundation for robot manipulation, their performance can degrade when adapted to new tasks with limited task-specific demonstrations. ・Retrieval offers a practical way to reuse existing demonstrations for data-efficient adaptation, but existing methods often rely on visu
cs.LG updates on arXiv.org

Highway Congestion Reduction through Reinforcement Learning Based Eulerian Headway Control

・arXiv:2412.02520v4 Announce Type: replace-cross Abstract: Connected automated vehicles (CAVs) equipped with adaptive cruise control (ACC) create new opportunities for highway congestion mitigation. ・Traditional practice relies on Eulerian variable speed limits (VSL) which regulate traffic through roadside signs, but suffer from infrequent updates and limited driver compliance. ・Recent research explored Lagrangian strat
cs.LG updates on arXiv.org

HiQA: A Hierarchical Contextual Augmentation RAG for Multi-Documents QA

・arXiv:2402.01767v4 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) significantly improves document-based question answering by integrating external documents during generation. ・However, retrieval accuracy can degrade when the knowledge base contains many semantically and structurally similar documents. ・We introduce HiQA, a practical hierarchical contextual augmentation framework for multi-
cs.LG updates on arXiv.org

Holographic Invariant Storage: Design-Time Safety Contracts via Vector Symbolic Architectures

・arXiv:2603.13558v2 Announce Type: replace-cross Abstract: We introduce Holographic Invariant Storage (HIS), a protocol that assembles known properties of bipolar Vector Symbolic Architectures into a design-time safety contract for LLM context-drift mitigation. ・The contract provides three closed-form guarantees evaluable before deployment: single-signal recovery fidelity converging to $1/\sqrt{2} \approx 0.707$ (regar
WIRED

How Ikea Turned a Controller Thumbstick Into the Star of Its Xbox Gaming Range

・Xbox and Ikea collab alert! ・This unique nine-piece collection for gamers arrives this fall and has a slew of Easter eggs.
cs.LG updates on arXiv.org

How Much Regularization Survives Averaging? Update Masking in Federated Learning

・arXiv:2608.23286v2 Announce Type: replace Abstract: Federated learning on non-IID data seeks flat minima to generalize across clients, and existing methods borrow sharpness-aware minimization from centralized training. ・There is a second way to reach flat minima, in which the regularization comes for free from noise added to the parameter updates, and it has never been carried over to the federated setting as an impli
cs.LG updates on arXiv.org

How to Achieve the Intended Aim of Deep Clustering Now, without Deep Learning

・arXiv:2602.05749v2 Announce Type: replace Abstract: Deep clustering (DC) is often quoted to have a key advantage over $k$-means clustering. ・Yet, this advantage is often demonstrated using image datasets only, and it is unclear whether it addresses the fundamental limitations of $k$-means clustering. ・Deep Embedded Clustering (DEC) learns a latent representation via an autoencoder and performs clustering based on a $k$
cs.LG updates on arXiv.org

IAPO: Influence-Aware Policy Optimization for Credit Assignment in Multi-Turn Service Agents

・arXiv:2608.24588v1 Announce Type: new Abstract: Large Language Model (LLM) agents increasingly solve long-horizon tasks through multi-turn interactions with users and external tools. ・In these settings, relevant task information often unfolds over time rather than being fully specified at the initial prompt. ・Service agents make this challenge especially concrete: users may clarify or revise their goals, while tool res
MarkTechPost

IBM Releases Granite 4.2: Bringing Native Reasoning and Agentic RL to Open Enterprise Models

・IBM has released Granite 4.2, a family of open reasoning language models in 3B, 8B, and 30B sizes, all under Apache 2.0. ・Every model exposes a thinking / low-effort / non-thinking switch and native tool calling. ・The 8B and 30B additionally go through an agentic RL block that trains them to edit code, drive a terminal, and run web searches inside real sandboxed environments.
cs.LG updates on arXiv.org

ICA: Information-Aware Credit Assignment for Visually Grounded Long-Horizon Information-Seeking Agents

・arXiv:2602.10863v2 Announce Type: replace Abstract: Long-horizon reinforcement learning for information seeking agents remains difficult because terminal rewards reveal whether the final answer is correct, but not which acquired information enabled it. ・This difficulty is amplified by text-derived webpage observations, where parsing, truncation, and summarization often produce incomplete and unstable content represent
cs.LG updates on arXiv.org

Improved generalization bounds for binary linear classification via isoperimetry

・arXiv:2505.16713v4 Announce Type: replace-cross Abstract: We examine the concentration of uniform generalization errors around their expectation in binary linear classification problems via an isoperimetric argument. ・In particular, we establish Poincar\'{e} and log-Sobolev inequalities for the joint distribution of the output labels and the label-weighted input vectors, which we apply to derive concentration bounds.
cs.LG updates on arXiv.org

Improving Cross-Problem Vehicle Routing with Locally Augmented Preferences and Representation Disentanglement

・arXiv:2608.24859v1 Announce Type: new Abstract: Multi-task vehicle routing problem (VRP) solvers seek to handle multiple VRP variants within a single unified model, avoiding the need to train a separate model for every variant. ・In spite of recent progress, current approaches remain limited on two fronts. ・On the training side, reinforcement learning suffers from reward-scale disparities and shrinking advantage signals
cs.LG updates on arXiv.org

Infant Care Video Dataset for Classification of Interventions Using Transformers

・arXiv:2608.23838v1 Announce Type: cross Abstract: Healthcare documentation in the neonatal intensive care unit (NICU) presents significant challenges, with nurses spending approximately 25\% of their time on record-keeping, while up to 60\% of interventions remain undocumented. ・Motivated by the need to detect interventions from video automatically, we present the Infant Care Video Dataset (ICVD), a collection of 4,14
cs.LG updates on arXiv.org

InfoDPP-PAC: Principled Patch Selection for Whole Slide Image Analysis

・arXiv:2608.23574v1 Announce Type: cross Abstract: Each WSI slide contains thousands of candidate tissue patches, while supervision is usually available only at slide level. ・Existing bag-construction strategies like Uniform extraction and handcrafted heuristics do not control redundancy while attention-based multiple-instance models couple patch importance to a particular downstream classifier, and coreset methods opt
Google DeepMind News

Intelligent transcription with Gemini 3.5 Transcribe

・Now you can get more intelligent speech-to-text transcription with Gemini 3.5 Transcribe.
cs.LG updates on arXiv.org

Intrinsic PAPR: Tackling Misattribution in 3D Intrinsic Decomposition via Proximity Attention Point Rendering

・arXiv:2407.00500v2 Announce Type: replace-cross Abstract: Recent point-based intrinsic decomposition and inverse rendering methods have advanced the modelling of the shading and albedo of 3D scenes. ・However, we identify a fundamental limitation: these methods suffer from a misattribution issue, where individual primitives learn incorrect appearance features despite producing correct aggregated renderings.
cs.LG updates on arXiv.org

It depends: Incorporating correlations for joint aleatoric and epistemic uncertainties of high-dimensional output spaces

・arXiv:2608.24518v1 Announce Type: new Abstract: Uncertainty Quantification (UQ) plays a vital role in enhancing the reliability of deep learning model predictions, especially in scenarios with high-dimensional output spaces. ・This paper addresses the dual nature of uncertainty -- aleatoric and epistemic -- focusing on their joint integration in high-dimensional regression tasks. ・For example, in applications like medic
cs.LG updates on arXiv.org

Iwin Transformer: Hierarchical Vision Transformer using Interleaved Windows

・arXiv:2507.18405v3 Announce Type: replace-cross Abstract: Vision Transformers (ViTs) face two limitations: the rigid resolution dependency of positional embeddings, which complicates cross-resolution fine-tuning, and the quadratic complexity of attention. ・While Swin Transformer alleviates the latter through window attention, it suffers from fine-tuning. ・Following the philosophy "no token is an island," we present Iwi
cs.LG updates on arXiv.org

Joint Distribution Alignment for Universal Domain Adaptation

・arXiv:2608.24429v1 Announce Type: new Abstract: Unsupervised domain adaptation (UDA) has been widely concerned in the fields of machine learning, pattern recognition, and computer vision. ・Traditional UDA learning usually assumes that the label spaces of the source and target domains are exactly the same and only needs to solve the problem of sample distribution drift existing between two domains. ・However, in real wor
cs.LG updates on arXiv.org

Joint-Embedding Prediction of Masked Point Tubes for Self-Supervised Learning on 4D Point Cloud Videos

・arXiv:2608.24093v1 Announce Type: cross Abstract: Self-supervised representation learning for 4D point cloud videos is challenging because annotations are costly and reconstruction-based pretraining can overemphasize low-level geometric details. ・We propose a JEPA-style framework that learns from unlabeled spatiotemporal point clouds through latent point-tube prediction. ・Instead of reconstructing raw coordinates, the
cs.LG updates on arXiv.org

Knowing When to Ask for Help: Bayesian Self-Escalation in Hierarchical LLM Agents

・arXiv:2608.24087v1 Announce Type: new Abstract: Current LLM agent systems decide delegation before reasoning begins (a router picks a model) or after a response is complete (a verifier scores it and may retry). ・We study a third regime: an agent that recognises, during its own reasoning, that it is unlikely to succeed and transfers control to a stronger model. ・We formulate intra-generation delegation as a Bayesian opt
Hugging Face Papers

LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training

LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training
cs.LG updates on arXiv.org

LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training

・arXiv:2608.24845v1 Announce Type: cross Abstract: We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl. ・From these, we download 80M videos with a total duration of 10 million hours. ・The dataset is designed for multimodal pre-training across the video, audio, and image modalities.
Hugging Face Papers

Latent Action as Intention Enables Efficient Future Imagination for World Action Models

Latent Action as Intention Enables Efficient Future Imagination for World Action Models
cs.LG updates on arXiv.org

Learning from the Right Rollouts: Data Attribution for PPO-based LLM Post-Training

・arXiv:2604.01597v2 Announce Type: replace Abstract: Traditional RL algorithms like Proximal Policy Optimization (PPO) typically train on the entire rollout buffer, operating under the assumption that all generated episodes provide a beneficial optimization signal. ・However, these episodes frequently contain noisy or unfaithful reasoning, which can degrade model performance and slow down training. ・In this paper, we pro
OpenAI News

Learning never stops: How AI makes learning continuous

・OpenAI’s new report explores how students and educators use ChatGPT to make learning more continuous, with support that extends beyond the classroom.
cs.LG updates on arXiv.org

Learning to Act While Waiting: RL Finetuning of Generalist Robot Policies Under Inference Latency

・arXiv:2608.23831v1 Announce Type: cross Abstract: While reinforcement learning (RL) allows generalist robot policies to continually improve during deployment, the large model size of modern generalist policies, such as VLAs, poses a fundamental obstacle to effective RL improvement. ・In particular, their severe inference latency---which can lead to pauses or jerky movements---can alter the effective environment dynamic
cs.LG updates on arXiv.org

Learning to Grade Efficiently: A Bandit-Driven Prompt-Selection Framework for Low-Cost LLM Essay Scoring

・arXiv:2608.23814v1 Announce Type: new Abstract: Large Language Models (LLMs) demonstrate strong capabilities in automated essay scoring (AES), but contemporary approaches typically employ fixed prompt selection, failing to address operational cost concerns and evolving optimal configurations. ・We propose a novel cost-aware approach that treats each prompt type as an arm in a multi-armed bandit (MAB) controller, enabli
cs.LG updates on arXiv.org

LEMMA-RCA: A Large Multi-modal Multi-domain Dataset for Root Cause Analysis

・arXiv:2406.05375v4 Announce Type: replace-cross Abstract: Root cause analysis (RCA) is crucial for enhancing the reliability and performance of complex systems. ・However, progress in this field has been hindered by the lack of large-scale, open-source datasets tailored for RCA. ・To bridge this gap, we introduce LEMMA-RCA, a large dataset designed for diverse RCA tasks across multiple domains and modalities.
Hugging Face Papers

Length-Adaptive Decoding for Masked Diffusion Machine Translation

Length-Adaptive Decoding for Masked Diffusion Machine Translation
cs.LG updates on arXiv.org

Lifted Model Construction under Approximate Commutativity

・arXiv:2608.24713v1 Announce Type: cross Abstract: Lifted inference algorithms enable scalable probabilistic inference even for large object domains by leveraging the indistinguishability of objects in a probability distribution. ・An essential prerequisite for constructing a lifted representation is to identify commutative factors, i.e., functions whose output values are invariant under permutations of a subset of thei
cs.LG updates on arXiv.org

Lightweight Adaptive Feature Composition for Heterogeneous Downstream Adaptation of Wireless Foundation Models

・arXiv:2606.10277v2 Announce Type: replace Abstract: Mobile systems increasingly rely on heterogeneous learning-enabled wireless functions, for which separate taskspecific models incur redundant training and model-management overhead. ・Wireless foundation models (WFMs) enable these functions to share a pretrained backbone, but existing adaptation either updates the backbone per task or relies on an inflexible final-lay
cs.LG updates on arXiv.org

Lightweight GenAI for Network Traffic Generation: Fidelity, Augmentation, and Classification

・arXiv:2603.25507v2 Announce Type: replace-cross Abstract: Network Traffic Classification (NTC) increasingly relies on data-driven models, yet its practical deployment is often constrained by limited labeled data, strict privacy requirements, and the cost of collecting representative traffic traces. ・While Network Traffic Generation (NTG) provides an effective means to mitigate data scarcity, conventional generative me
cs.LG updates on arXiv.org

LION: A Clifford Neural Paradigm for Multimodal-Attributed Graph Learning

・arXiv:2608.24795v1 Announce Type: new Abstract: Recently, the rapid advancement of multimodal domains has driven a data-centric paradigm shift in graph ML, transitioning from text-attributed to multimodal-attributed graphs. ・This advancement significantly enhances data representation and expands the scope of graph downstream tasks, such as modality-oriented tasks, thereby improving the practical utility of graph ML.
#LLMタグ

LLM#7 答えを見せたテストで満点だった。隠したら、でたらめより悪い点になった

・連載「LLMの仕組みを、作りながら理解する」 第7回 答えを見せたテストで満点だった。隠したら、でたらめより悪い点になった どうもです!えむしんです。
#LLMタグ

LLMがあなたのWindowsでも動く?

・LLM-jpって知ってますか? 私は知りませんでした。官民で作ってるLLMみたいですが、日本語のためのものです。これならきっといい感じで日本語Outputが出るはず。
cs.LG updates on arXiv.org

Low-Latency Activation-Regularized Sparse Neural Operators with Distillation Assistance Towards Real-Time Edge-Deployable Virtual Sensing

・arXiv:2608.23987v1 Announce Type: new Abstract: Virtual sensing enables digital twins and safety-critical systems to reconstruct and forecast spatial-temporal physics in real time. ・However, conventional computational and data-driven methods often face challenges in generalization, latency, and energy efficiency for edge deployment. ・Neural operators offer a promising alternative but remain reliant on power-intensive h
cs.LG updates on arXiv.org

Low-Rank Ternary Adaptation for Fine-Tuning Transformers

・arXiv:2608.24469v1 Announce Type: cross Abstract: Ternary transformers offer extreme memory and compute efficiency, but existing low-bit LoRA-based methods cannot directly fine-tune ternary weights. ・Current approaches either require dequantization, restoring low-bit base weights to higher precision to merge with adaptation weight, or update only quantization parameters, preventing a merged model that remains ternary.
cs.LG updates on arXiv.org

LUCAID: Agentic Multimodal AI for Lung Cancer Precision Pathology

・arXiv:2608.23803v1 Announce Type: cross Abstract: Lung cancer tissue diagnostics is complex, as therapy decisions in precision oncology rely on the integration of histomorphological, immunohistochemical, and molecular features. ・Yet pathological assessment remains largely visual and semi-quantitative and shows interobserver variability, while existing artificial intelligence (AI) tools cover only selected tasks, rarel
Zennの「機械学習」のフィード

M5 Ultra発表を機に、AI動画のピークメモリを見積もるPython

・TL;DR Macのunified memory総量を、そのままAIが使える容量として扱わないようにします。 ・常駐するmodel weights、実行時の追加領域、OSと他アプリの予約、safety marginを分けて見積もります。 ・Pythonで上下限を計算し、FIT_BY_ESTIMATE、UNCERTAIN、NOT_FITの3状態に分けます。
cs.LG updates on arXiv.org

Machine Learning Classification and Portfolio Construction: Does the Loss Function Matter?

・arXiv:2108.02283v5 Announce Type: replace-cross Abstract: Classification outperforms regression across matched machine learning models in portfolio construction. ・A stacking ensemble of gradient boosted tree, random forest, and neural network yields a value-weighted annualized Sharpe ratio of 1.83 for classification and 1.11 for regression. ・This outperformance persists in multiclass settings, across subsamples, and af
cs.LG updates on arXiv.org

Maia 200: A Software Defined Dataflow System for Large-scale AI Acceleration

・arXiv:2608.24664v1 Announce Type: cross Abstract: We introduce Maia 200, an advanced AI accelerator delivering high performance-10 145 Tflop/s FP4 and 5072 Tflop/s FP8 within a 750W TDP and 7 TB/s HBM bandwidth. ・Maia exemplifies a new class of Software Defined Locally Accessed Dataflow Architectures (SDLA), which explicitly program dataflow engines to orchestrate highly specialized memories and data movement engines.
cs.LG updates on arXiv.org

Massive-STEPS: Massive Semantic Trajectories for Understanding POI Check-ins -- Dataset and Benchmarks

・arXiv:2505.11239v4 Announce Type: replace Abstract: Understanding human mobility through Point-of-Interest (POI) trajectory modeling is increasingly important for applications such as urban planning, personalized services, and generative agent simulation. ・However, progress in this field is hindered by two key challenges: the over-reliance on older datasets from 2012-2013 and the lack of reproducible, city-level check
Zennの「大規模言語モデル」のフィード

MCP導入前の安全性チェックリスト10項目—ツール実行経路を潰す

・この記事で分かること 2026年夏の技術トレンドが「LLM単体」から「AIエージェント実装」に移った理由 MCP・ツール実行・Web実装で、いま優先して確認すべき安全性の論点 React/Pythonの開発基盤を、今どの順番で見直すべきか 2026-08-06の技術トレンドを最短で把握する方法 結論、今日のヘッドラインは**「モデル性能の競争」ではなく「エージェントをどう業務で動かし、どう安全に接続し、どう運用するか」**に主戦場が移ったことを示しています。 ・特に流れを決定づけているのは、AIエージェントの業務実装、MCPベースの接続標準化、そしてツール実行の安全性です。これ...
cs.LG updates on arXiv.org

MDTE: Minority-Aware Diffusion over Temporal Edge Events for Imbalanced Node Classification

・arXiv:2608.24812v1 Announce Type: new Abstract: Class-imbalanced node classification on temporal graphs is challenging because majority-dominated temporal propagation progressively assimilates minority representations, while conventional node and neighborhood information provides insufficient discriminative evidence for minority classes. ・To address these issues, we propose MDTE, a minority-aware diffusion framework t
cs.LG updates on arXiv.org

Mechanistic Circuit Identification for Controllable Data Generation

・arXiv:2608.24065v1 Announce Type: new Abstract: While recent advances in data synthesis aim to curate high-quality datasets, most generation pipelines still rely on heuristic prompt-based control. ・This black-box paradigm provides limited insight into how individual samples interact with a model's underlying learning dynamics. ・To bridge this gap, we propose a circuit-grounded framework that connects training-dynamics-
cs.LG updates on arXiv.org

Memory Is Not Always Needed: Characterizing Conditional Memory in Scientific Reasoning

・arXiv:2608.23982v1 Announce Type: cross Abstract: Scientific reasoning requires language models to retrieve specialized knowledge and incorporate it reliably into multi-step computation. ・Conditional memory provides an explicit lookup pathway that complements dense neural representations, but its usefulness is inherently input- and computation-dependent: retrieved information may repair missing scientific associations
The Verge

Meta agrees to heavy restrictions on teen users in major lawsuit settlement

・Mark Zuckerberg. ・| Image: Cath Virginia / The Verge, Getty Images Meta settled its latest kids online safety trial with a group of 29 state attorneys general, sparing it from the remainder of a trial that could have cost it hundreds of billions of dollars. ・Under the terms of the settlement, which resolves claims by a larger group of 47 states and several districts and territories, Meta agreed to come up with an age a
WIRED

Meta Will Pay Up to $16.7 Billion to Settle Its Social Media Harms Case—and That’s Not All

・In addition to sending billions of dollars to states, Meta will make substantive changes to its platforms as part of a landmark settlement.
Hugging Face Papers

Meta^n: Recursive Self-Improvement through Emergent Depth

Meta^n: Recursive Self-Improvement through Emergent Depth
cs.LG updates on arXiv.org

Method, Mind, and Morality: How People Make Sense of Artificial Intelligence

・arXiv:2608.24748v1 Announce Type: cross Abstract: How can humans make sense of the rapid takeoff of artificial intelligence (AI)? ・We studied the sensemaking dynamics of AI through an open-ended, mixed-methods study with computational text analysis of millions of AI-related newspaper articles and social media posts grounded in 57 semi-structured interviews with AI professionals in 2021 and 2023--before and after the r
The Verge

Microsoft’s 25th anniversary Xbox will cost $899

・The special-edition translucent green Xbox finally has an official price. ・Preorders for the console start August 27th at 10AM ET / 7AM PT, although Microsoft says "some of XBOX's most dedicated fans" will receive early access preorder emails starting today. ・The console will launch on November 13th, alongside a matching controller with the same "OG Green" colorway from the original translucent green Xbox.
cs.LG updates on arXiv.org

Mind the Student: Behavioral and Contextual Cues for Automated Engagement Prediction in Online Learning

・arXiv:2608.24340v1 Announce Type: cross Abstract: The prediction of student engagement from the online tutoring videos is difficult because engagement is a multidimensional construct comprising distinct behavioral, emotional, and cognitive states. ・A reliable prediction requires bringing together different types of behavioral signals as well as expressive cues. ・Through our analysis of the CASED dataset, it is clear th
cs.LG updates on arXiv.org

Mitigating Exploration Bias in RL for Multi-Instruction Following

・arXiv:2608.23830v1 Announce Type: cross Abstract: RL has emerged as a powerful paradigm for enhancing the instruction following capabilities of LLMs. ・While existing training recipes achieve substantial gains, we find that they suffer from exploration bias towards easy instructions when the training data has multiple instructions in a prompt. ・This bias is caused by two main reasons: 1) the policy model's initial abili
cs.LG updates on arXiv.org

Mixture of Channel Experts: Static Sparse Supports with Input-Adaptive Mixing for Pointwise Projections

・arXiv:2608.23794v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) scales language models by routing each input through a small set of independently parameterized experts. ・We show that copying this design into convolutional networks fails for a structural reason: parallel convolutional experts that read the same input channels learn nearly identical filters. ・We therefore move the expert axis from operator dupli
cs.LG updates on arXiv.org

MnemoDyn: Learning Resting State Dynamics from 40K FMRI sequences

・arXiv:2608.23936v1 Announce Type: new Abstract: We present a dynamical-systems based model for resting-state functional magnetic resonance imaging (rs-fMRI), trained on a dataset of roughly 40K rs-fMRI sequences covering a wide variety of public and available-by-permission datasets. ・While most existing proposals use transformer backbones, we utilize multi-resolution temporal modeling of the dynamics across parcellate
cs.LG updates on arXiv.org

Model-Based Learning of Near-Optimal Finite-Window Policies in POMDPs

・arXiv:2604.01024v2 Announce Type: replace Abstract: We study model-based learning of finite-window policies in tabular partially observable Markov decision processes (POMDPs). ・A common approach to learning under partial observability is to approximate unbounded history dependencies using finite action-observation windows. ・This induces a finite-state Markov decision process (MDP) over histories, referred to as the sup
cs.LG updates on arXiv.org

MolEmb: Multimodal Large Language Models Can Be Strong Molecular Embedding Models

・arXiv:2608.23646v1 Announce Type: cross Abstract: Molecular embedding models can serve as foundational infrastructure for computational chemistry and drug discovery, where reusable vector representations support property prediction, virtual screening, and retrieval. ・Most molecular encoders are specialist models built around a single molecular view, producing unconditional vectors with no language interface for varyin
cs.LG updates on arXiv.org

MolGA: Molecular Graph Adaptation with Pre-trained 2D Graph Encoder

・arXiv:2510.07289v2 Announce Type: replace Abstract: Molecular graph representation learning is widely used in chemical and biomedical research. ・While pre-trained 2D graph encoders have demonstrated strong performance, they overlook the rich molecular domain knowledge associated with submolecular instances (atoms and bonds). ・While molecular pre-training approaches incorporate such knowledge into their pre-training obj
cs.LG updates on arXiv.org

MoRF-AST: Calibrated Probabilistic Virtual Sensing for Structural Monitoring under Changing Operating Conditions

・arXiv:2608.24531v1 Announce Type: cross Abstract: Probabilistic full-field reconstruction provides uncertainty-aware response evidence for structural reliability assessment, yet inference from sparse and noisy measurements remains underdetermined. ・Most existing methods overlook shifts between offline training and operational distributions. ・Under such shifts, posterior intervals may become miscalibrated, causing the r
cs.LG updates on arXiv.org

MortarBench: Evaluating Mortgage Loan Origination Agents

・arXiv:2606.19416v3 Announce Type: replace Abstract: Loan origination is the process by which a lender creates a new loan, from application and underwriting through approval and funding. ・This process serves a critical role in evaluating the eligibility and level of risk posed by an applicant. ・Recently, firms have begun using mortgage loan agents to augment human loan officers, despite a lack of any public benchmark.
Hugging Face Papers

MoTE: Mixture of Task Experts for Multi-Task Video Understanding

MoTE: Mixture of Task Experts for Multi-Task Video Understanding
cs.LG updates on arXiv.org

MoTE: Mixture of Task Experts for Multi-Task Video Understanding

・arXiv:2608.24763v1 Announce Type: cross Abstract: Procedural video-language models must solve heterogeneous tasks from the same visual evidence, including action recognition, forecasting, and procedure prediction. ・Dense transformer decoders share the same feed-forward networks across tasks, which can entangle task behavior and make controlled capability expansion difficult. ・Sparse Mixture-of-Experts (MoE) decoders pr
cs.LG updates on arXiv.org

MPIB: A Benchmark for Medical Prompt Injection Attacks and Clinical Safety in LLMs

・arXiv:2602.06268v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) systems are increasingly integrated into clinical workflows. ・However, prompt injection attacks can steer these systems toward clinically unsafe or misleading outputs. ・We introduce the Medical Prompt Injection Benchmark (MPIB), a dataset-and-benchmark suite for evaluating clinical safety unde
cs.LG updates on arXiv.org

msData: A Millisecond-Resolution Network Dataset for Advancing Time Series Foundation Models

・arXiv:2603.16497v3 Announce Type: replace Abstract: Time series foundation models (TSFMs) require diverse, real-world datasets to adapt across varying domains and temporal frequencies. ・However, current large-scale datasets predominantly focus on low-frequency time series with sampling intervals, i.e., time resolution, in the range of seconds to years, hindering their ability to capture the nuances of high-frequency t
Qiita - 人気の記事

mysqldump のダンプをPIIマスクしつつ高速ロードしよう 〜 AIに複雑なツールを作らせる

・はじめに 今は MySQL Shell dump utilities などのより高速な手段が推奨されつつありますが、MySQLの論理バックアップに mysqldump を使った場合、以下のようなダンプファイルが出力されます。 ・mysqldumpの出力例 -- MySQ...
cs.LG updates on arXiv.org

NAIMA: Semantics Aware RGB Guided Depth Super-Resolution

・arXiv:2604.04407v2 Announce Type: replace-cross Abstract: Guided depth super-resolution (GDSR) is a multi-modal approach for depth map super-resolution that relies on a low-resolution depth map and a high-resolution RGB image to restore finer structural details. ・However, the misleading color and texture cues indicating depth discontinuities in RGB images often lead to artifacts and blurred depth boundaries in the gen
cs.LG updates on arXiv.org

NeuralParker: A Reinforcement Learning Planner for Irregular Parking Environments

・arXiv:2608.24485v1 Announce Type: cross Abstract: Automated parking commonly assumes marked slots and short approach maneuvers. ・Delivery and service vehicles, however, may need to reach an operator-specified pose in an irregular bounded environment from a distant start. ・Existing learning-based parking planners often rely on local observations, which can restrict long-range route reasoning.
cs.LG updates on arXiv.org

NeuronGuard: Robust LLM Safety Alignment via Ablation-Aware Safety Signal Redistribution

・arXiv:2608.23959v1 Announce Type: cross Abstract: Safety alignment in large language models (LLMs) remains brittle against a growing spectrum of attacks. ・Jailbreak attacks bypass safety mechanisms through crafted prompts, while neuron-level attacks directly prune safety-critical neurons post-deployment. ・Both exploit a common weakness: safety-relevant information concentrates in a sparse neuron subset.
cs.LG updates on arXiv.org

NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs

・arXiv:2508.09473v2 Announce Type: replace Abstract: Ensuring robust safety alignment while preserving utility is critical for the reliable deployment of Large Language Models (LLMs). ・However, current techniques fundamentally suffer from intertwined deficiencies: insufficient robustness against malicious attacks, frequent refusal of benign queries, degradation in generated text quality and general task performance, th
WIRED

Nomad Goods Promo Codes: Get 25% Off in August 2026

・Save up to 25% on Nomad Goods accessories such as Nomad phone cases, Nomad wallets, and more in August 2026.
cs.LG updates on arXiv.org

Nonconvex-Nonconcave Min-Max Optimization with a Small Maximization Domain

・arXiv:2110.03950v3 Announce Type: replace-cross Abstract: We study the problem of finding approximate first-order stationary points in optimization problems of the form $\min_{x \in X} \max_{y \in Y} f(x,y)$, where the sets $X,Y$ are convex and $Y$ is compact. ・The objective function $f$ is smooth, but assumed neither convex in $x$ nor concave in $y$. ・Our approach relies upon replacing the function $f(x,\cdot)$ with i
cs.LG updates on arXiv.org

Nonlinear Axiomatic Attribution for Cooperative Games

・arXiv:2607.09869v2 Announce Type: replace Abstract: The Shapley value is a widely used concept in attribution problems, as it uniquely satisfies the axioms of linearity, consistency, equal treatment, and efficiency. ・Often, the inclusion AUC metric is used to evaluate the quality of player rankings, in order to identify positively participating players. ・However, it can be established that the Shapley value is not alwa
LLMタグが付けられた新着記事 - Qiita

Notion MCPのページ更新が破壊的変更に、old_str不一致で全体差し戻し

・はじめに 2026年8月25日付で、Notion API のチェンジログに Notion MCP(notion-update-page / update_content)のページ内容更新に関する破壊的変更が2件追加されました。 ・① バッチ内の置換処理が「全て適用 or ...
#LLMタグ

NRA-IDE AI横軸機能 探索履歴

・この文書は、「現況でAI横軸機能を実装する場合の最適解は何か」という問いに対し、区切り探索で得た仮説、実装候補、判明事項、次段階への接続を残すための履歴である。 ・ここに記録する候補は、記録時点の探索結果であり、完成仕様、安全性の証明、または正典への昇格を意味しない。
cs.LG updates on arXiv.org

Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models

・arXiv:2603.16654v3 Announce Type: replace-cross Abstract: Evaluating the reasoning abilities of large language models (LLMs) solely from final answers can obscure failures in intermediate steps, especially in multi-hop QA benchmarks without step-level annotations. ・To address this gap, we introduce Omanic, an open-domain 4-hop QA benchmark designed not only to measure final-answer accuracy but also to diagnose where r
Hugging Face Papers

On-policy Distillation with Verifiable Reward

On-policy Distillation with Verifiable Reward
cs.LG updates on arXiv.org

On-policy Distillation with Verifiable Reward

・arXiv:2608.24696v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) and on-policy distillation (OPD) have become two widely adopted paradigms for post-training large language models. ・However, RLVR suffers from sparse task-level feedback, while OPD provides dense token-level guidance but ignores trajectory correctness, limiting its performance to that of the teacher. ・Combining them is
Hugging Face Papers

On-Policy Self-Distillation in Diffusion Models

On-Policy Self-Distillation in Diffusion Models
cs.LG updates on arXiv.org

Opponent Aware Reinforcement Learning

・arXiv:1908.08773v3 Announce Type: replace Abstract: In certain reinforcement learning (RL) scenarios there are adversaries trying to interfere with the underlying reward process for their own benefit. ・We introduce Threatened Markov Decision Processes (TMDPs) as a framework to support an agent against potential opponents in an RL context as well as schemes resulting in novel learning approaches to deal with TMDPs.
cs.LG updates on arXiv.org

Optimal Alternating Regret for Online Learning and Games

・arXiv:2608.24731v1 Announce Type: new Abstract: We settle the minimax-optimal alternating regret, a regret notion motivated by alternating learning dynamics in games, for both online linear optimization (OLO) and online convex optimization (OCO). ・For OLO over the probability simplex $\Delta_d$, we give an algorithm with $O(\log d)$ alternating regret that remains a constant for any time horizon $T$, and a matching lo
AI | VentureBeat

Orchestration is the new challenge for CX in the age of AI agents

・Presented by Tata Communications Enterprises are deploying AI agents, voice AI, and automation across messaging, voice, and digital channels faster than the architecture meant to support it. ・Most of that deployment has involved attaching conversational AI to legacy systems never built for it, says Gaurav Anand, global head of the Customer Interaction Suite at Tata Communications. ・"In the rush to deploy AI, organizati
LLMタグが付けられた新着記事 - Qiita

Ox Alpha の正体は Z.AI の GLM だった。無料期間に手元で使ったトークン量

・3 日前に書いた記事の結論は、半分当たって半分保留だった。 ・当たっていたのは「運用層の指紋は GLM 系」。保留にしていたのは「当事者は肯定も否定もしていない」。2026-08-26、Z.AI(智谱)が Bloomberg に答えて、保留が消えた。 ・この話には前があります...
cs.LG updates on arXiv.org

Parameter-Efficient Self-Supervised Adaptation for EEG-FM under Fixed Computational Budgets

・arXiv:2608.24727v1 Announce Type: new Abstract: EEG foundation models pretrained via self-supervised learning promise transferable representations, but their generalization remains limited, especially across diverse clinical datasets. ・Full fine-tuning is impractical for resource-constrained clinical settings due to high computational requirements. ・In this work, we investigate whether parameter-efficient self-supervis
cs.LG updates on arXiv.org

Parameter-Level Attribution of Symmetry in Trained Networks Though Parameter-Wise Functional Sensitivity

・arXiv:2608.24700v1 Announce Type: new Abstract: When a network has learned a function with a known symmetry, can that symmetry be moved through the parametrisation---is there a motion in parameter space realising the group action in function space? ・We formulate this as a lifting problem for the realisation map $\Phi:\theta\mapsto f_\theta$, and show that a smooth parameter-space action exists only if the tangent spac
cs.LG updates on arXiv.org

Parameterized Complexity of $L_p$-Lipschitz Constants for Input Convex Neural Networks and $L_p$-Norm Maximization over Zonotopes

・arXiv:2608.24865v1 Announce Type: cross Abstract: Lipschitz constants are a standard way to quantify the sensitivity of neural networks to small input perturbations, but computing them is difficult even for shallow ReLU networks. ・We study this problem for two-layer input-convex neural networks (ICNNs), a restricted architecture where nonnegative output weights enforce convexity. ・Computing the $L_p$-Lipschitz constant
cs.LG updates on arXiv.org

Paritok-4B: Intent-Conditioned Context Compression for Coding Agents

・arXiv:2608.24188v1 Announce Type: cross Abstract: Coding agents re-send large file reads and tool outputs to a frontier LLM every turn, and this context dominates their token bill. ・General-purpose prompt compressors are trained on prose and suit code poorly: they paraphrase identifiers and drop the exact spans an agent needs to edit. ・We present Paritok-4B, a 4B LoRA compressor for coding-agent trajectories built on t
cs.LG updates on arXiv.org

Partial Optimal Transport on the Circle for All Transported Masses in O(N log N)

・arXiv:2608.23910v1 Announce Type: new Abstract: Partial optimal transport compares two measures while leaving part of the mass unmatched, which is what makes it robust to outliers, occlusion, and clutter. ・The quantity of interest is usually the whole profile - the optimal cost at every transported cardinality - because the right amount to transport is rarely known in advance, and on the real line the PAWL algorithm r
WIRED

PeopleFinders’ New Website Runs Background Checks on Your Dates

・The site, called Stud or Dud, helps daters dig up dirt on potential paramours. ・It’s fueled by the same public data as PeopleFinders.com—and comes with many of the same concerns.
cs.LG updates on arXiv.org

Persistent Cross Entropy

・arXiv:2608.24549v1 Announce Type: new Abstract: Persistent entropy is the Shannon entropy of a persistence-based probability measure defined on a persistence diagram. ・However, its cross-entropy version is not naturally defined because two persistence diagrams generally have different event spaces. ・To bridge these event spaces, we combine a similarity function with persistence weighting to define an induced probabilit
cs.LG updates on arXiv.org

Physics-Integrated Operator Learning via Gaussian Splatting Representations

・arXiv:2608.24049v1 Announce Type: new Abstract: Neural operators provide efficient surrogates for spatiotemporal PDE systems, but purely data-driven formulations often accumulate substantial errors during long-horizon autoregressive prediction and may fail to exploit available governing-equation structure. ・Existing approaches incorporate physics primarily through residual-based training objectives or PDE-specific arc
cs.LG updates on arXiv.org

PhysicsBench: A Unified Leaderboard for Generative and Predictive Models in Engineering Design and Simulation

・arXiv:2608.24056v1 Announce Type: new Abstract: Generative and predictive artificial intelligence models are increasingly used to generate geometry and to predict physical fields and scalar quantities in engineering design and simulation. ・Yet these models are typically evaluated in isolation, on academic datasets at unconstrained scales, with inconsistent metrics and procedures. ・We present PhysicsBench, a unified ben
cs.LG updates on arXiv.org

PinSieve: Production Selective VLM Serving and a Governed Memory Flywheel for Enterprise Content-Quality Triage

・arXiv:2608.24040v1 Announce Type: new Abstract: Enterprise AI agents in production often need to be bounded, stateful, observable, and governable rather than fully autonomous. ・We present PinSieve, a production case study in a large-scale content-quality pipeline. ・Its deployed component is a selective vision-language-model (VLM) Serving Agent that operates only on the grey-zone slice left unresolved by lightweight ups
cs.LG updates on arXiv.org

Pipeline-Native Transformers: Co-Designing Model Architecture and CPU Inference for Bandwidth-Efficient Autoregressive Decode

・arXiv:2608.23841v1 Announce Type: cross Abstract: Single-token autoregressive decode on CPUs is bound by memory bandwidth, not arithmetic: a modern CPU sustains roughly 1 TFLOP/s of compute but only about 50 GB/s from main memory, and each generated token must stream every active weight once. ・This report argues that the most effective response is to co-design the model architecture and the inference runtime together.
cs.LG updates on arXiv.org

Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation

・arXiv:2406.02336v3 Announce Type: replace Abstract: We present polynomial-augmented neural networks (PANNs), a novel machine learning architecture that combines deep neural networks (DNNs) with polynomial expansions. ・PANNs combine the strengths of DNNs (flexibility and efficiency in higher-dimensional approximation) with those of polynomial approximation (rapid convergence rates for smooth functions). ・To aid in both
cs.LG updates on arXiv.org

Predictability of El Ni\~no from Delayed Observations

・arXiv:2608.24428v1 Announce Type: cross Abstract: Using monthly Ni\~no-3.4 anomalies through July 2026, we investigate how much predictive information is contained in delayed observations of the index. ・Ridge regression identifies informative delays, while multilayer perceptron and sparse identification of nonlinear dynamics (SINDy) models test whether nonlinear complexity provides additional direct forecast skill; ga
cs.LG updates on arXiv.org

Predicting Radiologist Expertise from 3D Gaze Patterns During CT Interpretation

・arXiv:2608.23836v1 Announce Type: cross Abstract: Accurate interpretation of volumetric CT requires efficient navigation of 3D image volumes and attention to diagnostically relevant regions. ・While eye-tracking has been widely studied in 2D medical imaging, its use for expertise assessment in CT settings remains limited. ・We propose a gaze-informed transformer framework for expertise classification in thoracic CT.
cs.LG updates on arXiv.org

Preference Optimization for Non-Verbal Vocalization Synthesis

・arXiv:2608.24163v1 Announce Type: cross Abstract: Non-verbal vocalizations (NVs), such as laughter, coughs, and sighs, are essential for expressive TTS, but the effectiveness of preference optimization for NV generation remains poorly understood. ・We systematically study preference optimization for NV-capable TTS, focusing on preference signals, preference-pair construction, and DPO-based optimization objectives.
Hugging Face Papers

Previous

Previous
cs.LG updates on arXiv.org

PROOF-Gen: From Optimized Data to Better Distillation

・arXiv:2608.23911v1 Announce Type: cross Abstract: Supervised fine-tuning on teacher-generated trajectories is the standard first stage for distilling tool-calling capabilities into deployable models. ・Post-training pipelines that drive shipped tool-calling agents re-run this stage on a daily or weekly cadence, paying the frontier-teacher cost each cycle, yet the mechanism is generate-and-filter (keep the teacher's pas
cs.LG updates on arXiv.org

Provable Quantum--Classical Separation for Continuous Gibbs Sampling

・arXiv:2608.24527v1 Announce Type: cross Abstract: We prove the first quantum--classical separation for a sampling problem over a continuous domain. ・For a class of Gibbs states $p\propto e^{-\beta E}$ on the torus $\mathbb{T}^d$ with smooth ($s$-Gevrey) potential and barrier amplitude $\alpha=e^{\beta\Delta}$, where $\Delta = \max E-\min E$, every classical algorithm---querying the value, gradient, or any higher-order
cs.LG updates on arXiv.org

Provenance Guided Incremental Learning Under Evolving Concept Definitions

・arXiv:2608.23893v1 Announce Type: cross Abstract: Learning systems deployed over long periods must adapt not only to statistical changes in incoming data, but also to revisions of the definitions that generate their prediction targets. ・Conventional concept-drift methods typically infer such changes from observations or prediction errors, even when the underlying policy, rule, or query has been explicitly modified.
cs.LG updates on arXiv.org

PRQ-KMeans: Projection Residual Quantization for Semantic ID Tokenization

・arXiv:2608.24207v1 Announce Type: new Abstract: Semantic identifiers (SIDs) represent entities as hierarchical token sequences for generative retrieval and recommendation. ・Residual-quantization tokenizers construct these sequences by selecting a codeword at each level and passing a residual to the next. ・We view this process as progressive commonality removal: each token captures a component shared within its group, w
cs.LG updates on arXiv.org

PuzzleKV: Page-Wise Low-Rank Decomposition for KV Cache Compression

・arXiv:2608.23843v1 Announce Type: new Abstract: Long-context inference in large language models (LLMs) is increasingly limited by the memory required for the key-value (KV) cache. ・KV cache compression addresses this problem by reducing the storage cost of previous tokens. ・Among existing approaches, low-rank compression is particularly attractive because it represents every token in reduced dimensions.
cs.LG updates on arXiv.org

QABBA: Error-Guaranteed Symbolic Time-Series Compression via Integer-Quantized Aggregation

・arXiv:2411.15209v3 Announce Type: replace Abstract: The expansion of time-series data from sensors and monitoring systems has made compact representations increasingly important. ・Such representations should retain signal structure while cutting storage, transmission and computation costs. ・Adaptive Brownian Bridge-based Aggregation (ABBA) addresses this need by converting long numerical series into short symbolic sequ
cs.LG updates on arXiv.org

QiMeng-ChipV-RTL: Exploiting Information Locality for IP-level Verilog Generation

・arXiv:2602.00704v2 Announce Type: replace Abstract: The generation of Register-Transfer Level (RTL) code is a crucial yet labor-intensive step in digital hardware design, traditionally requiring engineers to manually translate complex specifications into thousands of lines of synthesizable Hardware Description Language (HDL) code. ・While Large Language Models (LLMs) have shown promise in automating this process, exist
cs.LG updates on arXiv.org

qshap: Fast Shapley Decomposition of $R^2$ for Gradient-Boosted Trees

・arXiv:2608.24104v1 Announce Type: cross Abstract: Numerous methods have been developed to quantify feature attributions in individual predictions for tree ensembles. ・However, many applications require global measures of feature contributions to overall model performance. ・Although local attribution scores can be aggregated to characterize feature importance, such summaries do not directly decompose measures of predict
cs.LG updates on arXiv.org

Quantum Maximum Entropy Inference and Hamiltonian Learning

・arXiv:2407.11473v2 Announce Type: replace Abstract: Maximum entropy inference and learning of graphical models are pivotal tasks in learning theory and optimization. ・This work extends algorithms for these problems, including generalized iterative scaling (GIS) and gradient descent (GD), to the quantum realm. ・While the generalization, known as quantum iterative scaling (QIS), is straightforward, the key challenge lies
cs.LG updates on arXiv.org

Quasar: A Programming Language Specialized for LLM Code Actions

・arXiv:2506.12202v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often call external tools to solve tasks. ・One effective strategy is for LLMs to write code, enabling them to use complex control flow such as conditionals and loops. ・Such code actions are typically represented as Python code, since LLMs are proficient at writing it.
AI News & Artificial Intelligence | TechCrunch

QueryStory wants you to believe what AI is telling you

・The startup came out of stealth with $6 million in seed funding and a plan to use LLMs and cybersecurity know-how to make AI queries coherent.
#LLMタグ

Qwen3.8-27Bを実機で動かした。3.6・3.5との比較と、RTX PRO 6000とDGX Sparkの実測速度

・Note初投稿です。よろしくお願いいたします。 ・8月14日にQwen3.8が公開されました。 ・手元の機材で動かして、旧世代と比べ、速度を測った記録です。
#LLMタグ

Qwen3.8-Flash-Nextは180B!DGX-SparkのVRAMに乗らない!?

・Sparkでは動かないのか調査 Qwenは中国Alibabaが公開している、無料でダウンロードして自分のPCで動かせるAIです。その新型 Qwen3.8-Flash-Next のモデルカードが、2026年8月26日21時29分(日本時間)に出ました。合計1,800億パラメータ、いわゆる180Bです。
cs.LG updates on arXiv.org

RACR-MIL: Rank-aware contextual reasoning for weakly supervised grading of squamous cell carcinoma using whole slide images

・arXiv:2308.15618v3 Announce Type: replace-cross Abstract: Squamous cell carcinoma (SCC) is one of the most common cancer subtype, with an increasing incidence and a significant impact on cancer-related mortality. ・SCC grading using whole slide images is inherently challenging due to the lack of a standardized grading protocol and substantial tissue heterogeneity. ・We propose RACR-MIL, a weakly-supervised SCC grading ap
AI News & Artificial Intelligence | TechCrunch

Radar makes podcasts searchable — and usable by AI agents

・Particle’s new podcast intelligence platform transcribes and analyzes more than 130,000 podcasts, making their conversations searchable on the web and accessible to AI agents through an API and MCP.
cs.LG updates on arXiv.org

RAGSentinel: Certifiable Geometric Consensus for Robust Retrieval-Augmented Generation

・arXiv:2608.23965v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) improves the factuality of large language models by grounding responses in external documents, but it also exposes a critical security vulnerability: adversarial documents injected into the knowledge database can enter the context window and steer the model toward targeted incorrect answers. ・Existing post-retrieval defenses rely on
cs.LG updates on arXiv.org

Reconquering Bell sampling on qudits: stabilizer learning and testing, quantum pseudorandomness bounds, and more

・arXiv:2510.06848v3 Announce Type: replace-cross Abstract: Bell sampling is a simple yet powerful tool based on measuring two copies of a quantum state in the Bell basis, and has found applications in a plethora of problems related to stabiliser states and measures of magic. ・However, it was not known how to generalise the procedure from qubits to $d$-level systems -- qudits -- for all dimensions $d > 2$ in a useful wa
Hugging Face Papers

Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses

Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses
cs.LG updates on arXiv.org

Renormalization Group Flow Matching for Scalable Local Generative Modeling

・arXiv:2608.23696v1 Announce Type: new Abstract: Despite their remarkable success in modeling complex data, generative models face a fundamental tradeoff. ・Global approaches can capture full structural coherence but suffer from high computational costs, while local models are efficient but often fail to reproduce long-range correlations and global coherence. ・The renormalization group (RG) bridges this gap by seamlessly
cs.LG updates on arXiv.org

Replicable Conformal Prediction

・arXiv:2608.23638v1 Announce Type: cross Abstract: Two analysts who calibrate the same predictive model on independent samples will deploy different prediction sets every time, because the calibration threshold inherits the randomness of the data. ・Wherever deployments must be audited, cached, or approved across sites, this instability is costly: no one can verify that two calibrations produced the same object.
cs.LG updates on arXiv.org

Response Renormalization for Critical Deep Equilibrium Models

・arXiv:2608.23725v1 Announce Type: new Abstract: Deep Equilibrium Models (DEQs) compute predictions from a hidden representation unchanged by the model update. ・Training through this equilibrium uses implicit differentiation and requires solving an adjoint system built from the residual Jacobian. ・If this Jacobian is nearly singular along loss-sensitive directions, small perturbations can be strongly amplified in the ad
cs.LG updates on arXiv.org

Restoring Without Forgetting: Continual Learning Across Image Degradations

・arXiv:2608.23799v1 Announce Type: cross Abstract: Recent progress in image restoration has converged on all-in-one architectures that jointly handle multiple degradations within a single network. ・These methods are effective on static benchmarks but target a closed-world setting that assumes simultaneous access to every target degradation at training time. ・In practice, degradations are encountered sequentially as fiel
cs.LG updates on arXiv.org

RetrievalFormer: A Dual-Encoder Transformer for Efficient Approximate Nearest Neighbor Retrieval and Cold-Item Recommendation

・arXiv:2608.24079v1 Announce Type: cross Abstract: A shared search-and-recommendation index must score new items from features alone because search has no exploration slot. ・In a public log covering both surfaces over one catalog, $38.6\%$ of held-out query-search impressions show an item never previously shown or visited. ・For user-cold engagements, the feature-based tower serves this demand without measurable loss aga
cs.LG updates on arXiv.org

Revelation Control

・arXiv:2608.23860v1 Announce Type: new Abstract: Revelation Control is the problem of choosing priced interventions that reveal hidden state only insofar as the revealed distinctions can change a consequential decision, while accounting separately for any useful progress created by the intervention itself. ・We develop this theory for learning systems, where states equivalent under declared current information can respo
cs.LG updates on arXiv.org

Revenge of Monosemanticity: Specialized Neurons Improve Data Efficiency in MLPs

・arXiv:2608.24007v1 Announce Type: new Abstract: Understanding how neural networks learn and organize features is central to understanding their behavior. ・Much existing theory of feature learning has focused on the emergence of a global low-dimensional predictive geometry. ・We show that this picture is incomplete.
AI News & Artificial Intelligence | TechCrunch

Robot brain builders are pushing out of their GPT-2 era

・Robot bodies are waiting for their AI brains to catch up.
cs.LG updates on arXiv.org

Robust Data-Collection Policy Learning for Low-Variance Online Policy Evaluation

・arXiv:2608.24146v1 Announce Type: new Abstract: In reinforcement learning policy evaluation, classic on-policy methods often suffer from high variance when estimating policy performance. ・To mitigate this issue, behavior policy search has been proposed to learn data-collecting policies tailored to reduce online evaluation variance. ・However, these approaches do not account for uncertainties in the transition functions.
cs.LG updates on arXiv.org

Round-trip Reinforcement Learning: Self-Consistent Training for Better Chemical LLMs

・arXiv:2510.01527v2 Announce Type: replace Abstract: Large Language Models (LLMs) are emerging as versatile foundation models for computational chemistry, handling bidirectional tasks like reaction prediction and retrosynthesis. ・However, these models often lack round-trip consistency. ・For instance, a state-of-the-art chemical LLM may successfully caption a molecule, yet be unable to accurately reconstruct the original
AI News & Artificial Intelligence | TechCrunch

Runable hits $21M to bet AI agents can go from building businesses to growing them

・Runable says 60% to 70% of its 1 trillion-plus token usage in the last 90 days came from paying customers.
cs.LG updates on arXiv.org

S-matrix informed neural networks for amplitude analysis

・arXiv:2608.23750v1 Announce Type: cross Abstract: Reconstructing scattering amplitudes from finite, noisy, and mutually inconsistent measurements is an ill-posed inverse problem common to many reactions relevant to particle physics. ・We introduce S-matrix informed neural networks (SINNs), and demonstrate their ability to learn scattering amplitudes directly from data while respecting first principles. ・We further devel
cs.LG updates on arXiv.org

SANE: State Anomaly Neutralization for Stable Extreme-Context Delta-Rule Models

・arXiv:2608.22354v2 Announce Type: replace Abstract: Delta-Rule recurrent models maintain a fixed-size state, enabling $O(1)$ inference memory but potentially becoming unstable under extreme-context extrapolation. ・By tracking RWKV-7 over sequences of up to 100M tokens, we empirically identify a distinct failure pattern: \textbf{localized norm explosion atop a relatively sparse substrate}, rather than global state satu
cs.LG updates on arXiv.org

SatDL: Jointly Optimizing Data Redistribution and Training for Satellite-Based Distributed Learning

・arXiv:2608.24516v1 Announce Type: cross Abstract: Satellite-based distributed learning promises to train machine-learning models directly in orbit using massive, globally dispersed sensor data, thereby avoiding large-scale data downloads to ground servers. ・However, training convergence is significantly slowed by severe non-IID data, specifically label imbalance, as each satellite observes different geographic regions
cs.LG updates on arXiv.org

Scaling Reinforcement Learning for Diffusion Models via Velocity Matching

・arXiv:2608.23664v1 Announce Type: cross Abstract: Reward fine-tuning is becoming an important tool for adapting diffusion models to human preferences and task-specific objectives, but existing methods largely inherit policy-gradient machinery from large language models. ・Unlike autoregressive models, diffusion models do not provide tractable likelihoods for generated samples. ・As a result, current approaches either con
cs.LG updates on arXiv.org

Score-Based Ideal Observer Approximation via Denoising Score Matching for Signal-Known-Exactly Detection Tasks

・arXiv:2608.24768v1 Announce Type: cross Abstract: The Bayesian Ideal Observer (IO) establishes the theoretical upper bound on task performance for binary detection tasks. ・However, analytical computation of the IO test statistic is generally intractable. ・Numerical approaches based on Markov-chain Monte Carlo (MCMC) methods, including their recent deep generative model-based extensions, typically require extensive post
Hugging Face Papers

SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation

SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation
cs.LG updates on arXiv.org

SeisMamba: Low-Latency Single-Station Seismic Magnitude Estimation for Spatially Distributed Earthquake Early Warning

・arXiv:2608.24561v1 Announce Type: new Abstract: Rapid earthquake magnitude estimation is central to earthquake early warning, yet many operational systems depend on dense regional seismic networks and region-specific calibration. ・This creates a spatial coverage barrier for high-risk areas with sparse sensing infrastructure. ・Single-station learning offers a lower-cost alternative, but existing models often face an acc
cs.LG updates on arXiv.org

Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering Vectors

・arXiv:2608.23873v1 Announce Type: cross Abstract: Everything a language model sees is tokens. ・The serving stack knows what each span is -- user input, tool output, instructions -- but the model must keep track of that itself, and it can lose track or be confused: text can be written to read like anything. ・Prompt injection is a natural exploit of this phenomenon.
cs.LG updates on arXiv.org

Sequential operator learning under dependent data

・arXiv:2608.24426v1 Announce Type: cross Abstract: Learning operators from sequentially collected data arises in adaptive experimental design, Bayesian optimization, and dynamical-system modelling, where observations may be dependent, and future inputs or sensing operators may depend on preceding data. ・We derive time-uniform self-normalized concentration bounds for stochastic processes in Hilbert spaces with vector-va
cs.LG updates on arXiv.org

Single State Update Predictive Coding training for Time Series Forecasting and Anomaly Detection

・arXiv:2608.24697v1 Announce Type: new Abstract: Predictive Coding (PC) is a neural learning paradigm that enables parallelizable neural network layer updates. ・However, the main bottleneck of PC Networks (PCN) is the sequential backwards error propagation. ・To tackle this, we introduce a training technique that pairs a Generative PCN with a support Encoding PCN.
cs.LG updates on arXiv.org

SketchGuard: Scaling Byzantine-Robust Decentralized Federated Learning via Sketch-Based Screening

・arXiv:2510.07922v5 Announce Type: replace Abstract: Byzantine-robust decentralized federated learning (DFL) protects peer-to-peer training from malicious clients. ・The dominant defenses rely on similarity-based filtering, in which each client exchanges full model vectors with every neighbor before any filtering decision; this communication grows with the model dimension and scales poorly as models grow. ・We propose Ske
cs.LG updates on arXiv.org

Spatiotemporal Distillation via Recurrent Bottlenecks for Aortic Tracking

・arXiv:2608.23879v1 Announce Type: cross Abstract: Cardiac cine-MRI serves as a direct visual indicator of cardiovascular hemodynamics by capturing the continuous wall motion of the aorta. ・Quantifying these dynamic structural changes across the cardiac cycle is essential for measuring aortic distensibility, a primary marker of arterial stiffness. ・However, standard 2D segmentation networks focus on each frame independe
cs.LG updates on arXiv.org

Spectrum-Aware Bounds on Invertibility for Privacy-Enhancing Instance Encoding

・arXiv:2608.23382v2 Announce Type: replace Abstract: Instance encoding is a popular empirical technique for privacy enhancement when sharing data to an untrusted server. ・It transforms sensitive data through an encoding process before sharing, with the hope that the encoding process retains utility but makes it hard to reconstruct the original data. ・However, most work offers no theoretical guarantee that the encoding p
cs.LG updates on arXiv.org

ST-Lite: Training-Free KV Cache Compression with Spatio-Trajectory Guidance for Long-Horizon GUI Agents

・arXiv:2603.00188v2 Announce Type: replace-cross Abstract: Training-free KV cache compression is essential for deploying vision-language GUI agents under memory and latency constraints, yet existing methods are designed for generic language workloads and ignore the distinctive structure of GUI interaction traces. ・We characterize three GUI-specific workload properties--high inter-frame visual redundancy, extremely smal
cs.LG updates on arXiv.org

Stabilizing Temporal Difference Learning via Implicit Stochastic Recursion

・arXiv:2505.01361v3 Announce Type: replace Abstract: Temporal difference (TD) learning is a foundational algorithm in reinforcement learning (RL). ・For nearly forty years, TD learning has served as a workhorse for applied RL as well as a building block for more complex and specialized algorithms. ・However, despite its widespread use, TD procedures are generally sensitive to step size specification.
cs.LG updates on arXiv.org

StateTune: Transforming LLM-Assisted EDA Flow Tuning into a Stateful, Closed-Loop Process

・arXiv:2608.23601v1 Announce Type: cross Abstract: EDA flow parameter tuning is critical for quality-of-results~(QoR), yet the parameter space is large, tightly coupled, and full evaluations are prohibitively expensive. ・Prior LLM-assisted tuners mainly use the LLM as an external proposer with transient working context; we instead present \textbf{StateTune}, which reformulates LLM-assisted EDA tuning as a closed-loop,
cs.LG updates on arXiv.org

Steering Recurrent Reasoners at Inference Time with Readout Feedback

・arXiv:2608.24136v1 Announce Type: new Abstract: Recurrent models, which repeatedly update latent states with shared computation blocks, have emerged as powerful architectures for solving complex reasoning tasks. ・Existing inference-time methods scale computation by running more steps or sampling more trajectories, but ignore information revealed within each trajectory. ・Here we show that recurrent models can be improve
cs.LG updates on arXiv.org

Structured Frequency-Domain Evidence for LLM-Based Time-Series Anomaly Detection

・arXiv:2608.24113v1 Announce Type: new Abstract: Time-series anomalies can appear not only as pointwise deviations but also as changes in recurring temporal structure, such as shifted periodicity or localized oscillatory fluctuations. ・However, existing LLM-based time-series anomaly detection methods mainly expose time-domain evidence through indexed values, plots, or de-seasonalized representations, leaving spectral s
cs.LG updates on arXiv.org

Superintelligent Retrieval Agent: The Next Frontier of Agentic Retrieval

・arXiv:2605.06647v3 Announce Type: replace-cross Abstract: Retrieval-augmented agents are increasingly the interface to large knowledge bases, yet most treat retrieval as a black box: they issue exploratory queries, inspect snippets, and reformulate until evidence emerges. ・This resembles how a newcomer searches an unfamiliar database rather than how an expert navigates it with strong priors about terminology and likel
AI News & Artificial Intelligence | TechCrunch

Surprise: Z.ai is the AI lab behind the mysterious Ox Alpha model

・Z.ai confirms it is behind Ox Alpha, the mysterious open AI model topping benchmarks and leaderboards, and its weights are set to be released soon.
cs.LG updates on arXiv.org

Symbolic Classification-Enabled LHC Limits Online BSM Global Fits

・arXiv:2605.22330v1 Announce Type: cross Abstract: Global fits of Beyond the Standard Model (BSM) physics often involve a two-way interplay between theory and experiment. ・Theoretical models provide guidance for experimental searches, while experimental results, in turn, constrain theoretical frameworks. ・A crucial aspect of this feedback loop is the direct inclusion of measurements and exclusion limits ``online'' globa
cs.LG updates on arXiv.org

TaLK: Text-attributed Graph Dataset Distillation via Coupling Language Model with Graph-Aware Kernel

・arXiv:2606.22975v2 Announce Type: replace Abstract: Text-attributed graphs (TAGs) are widely used in many real-world domains, and learning on TAGs requires jointly modeling text semantics and graph structure. ・A standard approach for modeling TAGs is to combine a language model (LM) and a graph neural network (GNN), but joint training is computationally expensive and difficult to scale. ・Dataset distillation is a promi
cs.LG updates on arXiv.org

Taming foundation model with invariance-oriented pre-training for broad-spectrum EEG analysis across signal-level, brain-state, and brain-health tasks

・arXiv:2608.24597v1 Announce Type: new Abstract: Electroencephalography (EEG) is a widely used window into human brain function, but most EEG models remain tied to a one-dataset-one-model supervised paradigm. ・Recent EEG foundation models offer a route toward reusable representations, but most remain reconstruction-centered, assuming that EEG content predictable from local context is necessarily transferable neural inf
WIRED

Target Promo Code: $50 Off | August 2026

・Get $50 off your next order or up to 50% off site wide with Target coupon codes and Circle deals.
cs.LG updates on arXiv.org

Test-Time Adaptation for EEG Foundation Models: A Systematic Study under Real-World Distribution Shifts

・arXiv:2604.16926v3 Announce Type: replace Abstract: Electroencephalography (EEG) foundation models have shown strong potential for learning generalizable representations from large-scale neural data, yet their clinical deployment is hindered by distribution shifts across clinical settings, devices, and populations. ・Test-time adaptation (TTA) offers a promising solution by enabling models to adapt to unlabeled target
cs.LG updates on arXiv.org

The Blending Ratio Is Not Where the Performance Is: Diagnosing Prototype Blending for Few-Shot Adaptation of Vision-Language Models

・arXiv:2608.23634v1 Announce Type: cross Abstract: Many few-shot adaptation methods for vision-language models classify with a convex combination of the zero-shot text prototype and the mean of the K labelled image features, with a single blending ratio routinely tuned on held-out labels, often on the test set itself. ・We ask what the family's own bias-variance justification invites: what is the right ratio, can it be
Latent.Space

The Future of SaaS Is Apps That Agents Can Use

・Lovable is branching out from AI-powered web app creation and into MCP-powered ‘capabilities’. ・We talk to CTO Fabian Hedin.
cs.LG updates on arXiv.org

The Loss Floor of Denoising Score Matching: Fisher Geometry from Schr\"odinger Bridges

・arXiv:2608.23916v1 Announce Type: new Abstract: Denoising score matching trains diffusion models by regressing onto a conditional score, although generation ultimately requires the marginal score. ・The two objectives share the same population minimizer, but the conditional target remains random at fixed noisy state and introduces an irreducible excess in the training loss. ・We isolate this excess and show that, for a g
cs.LG updates on arXiv.org

The Mask Is Not the Model: Auditing Prefix Invariance in Attention, State-Space, and Hybrid Sequence Models

・arXiv:2608.22876v2 Announce Type: replace Abstract: Hybrid sequence models must satisfy prefix invariance: representations at position t must not depend on future inputs, yet this is rarely verified. ・We formalize prefix invariance and give a lightweight audit, two forward passes, no training or gradients, yielding a per-layer score localizing where causality breaks. ・Attention-mask inspection, the field's default chec
cs.LG updates on arXiv.org

The Sharp Tail of Uniform Stability

・arXiv:2608.24098v1 Announce Type: new Abstract: Uniform stability controls how much one training example can change the loss at any test point. ・A new logarithmic-free upper bound shows that a $\gamma$-uniformly stable algorithm with loss in $[0,L]$ has generalization gap at most $O \left(\gamma\log(1/\delta) +L\sqrt{\frac{\log(1/\delta)}{n}}\right)$ with probability $1-\delta$. ・Whether an actual bounded-loss learning
cs.LG updates on arXiv.org

The Theorems of Dr. David Blackwell and Their Contributions to Artificial Intelligence

・arXiv:2604.06621v2 Announce Type: replace-cross Abstract: Dr. ・David Blackwell was a mathematician and statistician of the first rank, whose contributions to statistical theory, game theory, and decision theory predated many of the algorithmic breakthroughs that define modern artificial intelligence. ・This survey examines three of his most consequential theoretical results the Rao Blackwell theorem, the Blackwell Appro
cs.LG updates on arXiv.org

Tight Majorizations and Convergence Rates of Nuclear Norm Minimization IRLS

・arXiv:2608.23765v1 Announce Type: new Abstract: Iteratively reweighted least squares (IRLS) methods constitute a natural approach to nuclear norm minimization, but their convergence rates and the role of the weight operator have remained poorly understood. ・This paper establishes sharp convergence rates for IRLS methods for constrained nuclear norm minimization in low-rank recovery. ・A central ingredient is a new major
cs.LG updates on arXiv.org

TLXML: Task-Level Explanation of Meta-Learning via Influence Functions

・arXiv:2501.14271v4 Announce Type: replace Abstract: Meta-learning enables models to rapidly adapt to new tasks by leveraging prior experience, but its adaptation mechanisms remain opaque, especially regarding how past training tasks influence future predictions. ・We introduce TLXML (Task-Level eXplanation of Meta-Learning), a novel framework that extends influence functions to meta-learning settings and provides task-
cs.LG updates on arXiv.org

Topology enables learning-based hydrodynamic prediction of the global river system

・arXiv:2602.22293v2 Announce Type: replace Abstract: Accurate river prediction is essential for water, food and energy security, yet remains challenging across entire river networks. ・Machine learning has transformed Earth-system modeling, but a system-level advance for river prediction lags for lack of reliable data. ・Exploiting the connectivity and dissipative dynamics of rivers, we introduce GraphRiverCast, a neural
Hugging Face Papers

TorchMorph: CUDA-accelerated Morphological Transforms

TorchMorph: CUDA-accelerated Morphological Transforms
cs.LG updates on arXiv.org

Towards Reproducibility in Predictive Process Mining: SPICE -- A Deep Learning Library

・arXiv:2512.16715v3 Announce Type: replace Abstract: In recent years, Predictive Process Mining (PPM) techniques based on artificial neural networks have evolved as a method for monitoring the future behavior of unfolding business processes and predicting Key Performance Indicators (KPIs). ・However, many PPM approaches often lack reproducibility, transparency in decision making, usability for incorporating novel datase
cs.LG updates on arXiv.org

Transformer Accelerator (TFA): A Macro-Op INT8 Hardware Chip for Transformer Inference and Machine Translation

・arXiv:2608.23582v1 Announce Type: cross Abstract: We present the Transformer Accelerator (TFA), a synthesizable, parameterizable INT8 memory-to-memory engine for transformer inference. ・One time-multiplexed datapath handles prompt processing and autoregressive generation. ・TFA implements matrix multiplication, softmax, RMSNorm, elementwise, and copy/gather operations through eight 512-bit macro-op descriptors.
WIRED

Tuft & Needle Promo Codes: 30% Off | August 2026

・Save 30% on best-selling mattresses with our top Tuft & Needle coupon codes.
cs.LG updates on arXiv.org

Two-Sided Nearest Neighbors: An adaptive and minimax optimal procedure for matrix completion

・arXiv:2411.12965v3 Announce Type: replace-cross Abstract: Nearest neighbor (NN) algorithms have been extensively used for missing data problems in recommender systems and sequential decision-making systems. ・Prior theoretical analysis has established favorable guarantees for NN when the underlying data is sufficiently smooth and the missingness probabilities are lower bounded. ・Here we analyze NN with non-smooth non-li
cs.LG updates on arXiv.org

UHI-Bench: Benchmarking Dual-Source Urban Heat Island Modeling Across Cities in Diverse Climate Regimes

・arXiv:2608.23857v1 Announce Type: new Abstract: Urban heat islands (UHIs) are intensifying under climate change, exacerbating thermal exposure risks. ・Their two primary observations, land surface temperature UHI (LST-UHI) and near-surface air temperature UHI (AirT-UHI), capture physically distinct aspects of urban heat. ・However, most studies rely on a single source, and substituting one for the other can substantially
#AIタグ

USの「AI×コマース」最前線。効率化の幻想を超えて見えてきた「新しい買い物の分業構造」と事業者が今すぐ打つべき手

・はじめに:「AIが自動で売ってくれる」の幻想 「AIがおすすめを出して、そのままチャット内で決済まで完了する」──。 ・生成AIが登場した当初、多くの人がそんなコマースの未来を思い描いた。
cs.LG updates on arXiv.org

Validation of HRV Studio: A Transparent and Quality-Control-Aware Platform for Heart Rate Variability Analysis

・arXiv:2608.24241v1 Announce Type: cross Abstract: Reproducibility of heart rate variability (HRV) analysis is limited by differences in preprocessing and computational conventions across software platforms. ・We developed HRV Studio, an open-source PyQt6-based desktop application integrating transparent HRV analysis with automated quality-control (QC) diagnostics. ・Validation included large-scale agreement with NeuroKit
The Verge

Volvo’s cars will warn one another about hazards in the road

・Volvo is updating three of its electric vehicles with new hazard alerts to warn drivers when there are animals or vulnerable road users ahead. ・The new connected safety features are based on Volvo's cars talking to one another, as opposed to crowdsourced alert systems used by popular navigation tools like Google Maps and Waze. ・The alerts work by using connected vehicle data from Volvo's fleet of passenger vehicles in
Zennの「大規模言語モデル」のフィード

VRAMは足りていた。MoEオフロードを止めたのはページロックの上限だった

・MoEモデルのexpertをホストRAMへ逃がすタイプの推論エンジンを、RTX 4070 SUPER(VRAM 12GB)の機体で動かそうとして失敗しました。この記事の結論は「止めたのはVRAMではなく、ホストRAMのうちページロックできる量。それは物理RAMの半分程度しかなく、増やす手段がない」です。 ・こんにちは、kimuraです。 ・VRAMは10.78GiB空いたまま、一度も使われずに落ちました。エラーは cudaHostRegister failed: out of memory の1行だけで、どこが足りないのかは書いてありません。ulimit を外し、ページキャッシュを捨て、VM...
cs.LG updates on arXiv.org

Wait, Wait, Wait... Why Do Reasoning Models Loop?

・arXiv:2512.12895v2 Announce Type: replace Abstract: Reasoning models (e.g., DeepSeek-R1) generate long chains of thought to solve harder problems, but they often loop, repeating the same text at low temperatures or with greedy decoding. ・We study why this happens and what role temperature plays. ・With open reasoning models, we find that looping is common at low temperature.
cs.LG updates on arXiv.org

WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation

・arXiv:2608.24479v1 Announce Type: new Abstract: Massively parallel simulation changes the data regime in which off-policy reinforcement learning (RL) is trained, challenging stabilizers designed for data-limited replay. ・Through controlled experiments across eight benchmark families, we show that these stabilizers are data-regime-dependent: parameter normalization helps with narrow replay coverage but restricts value
cs.LG updates on arXiv.org

Weak-to-Strong Learning in Decision Making

・arXiv:2607.18467v2 Announce Type: replace Abstract: Many operational decisions rely on predictive models that estimate uncertain outcomes conditional on observable contexts. ・Training such models, however, often faces a fundamental data asymmetry: labeled outcomes are scarce or costly to obtain, while contextual covariates are abundant. ・Motivated by this data asymmetry, we develop a decision-aware weak-to-strong (W2S)
cs.LG updates on arXiv.org

Weakly Supervised Seafloor Segmentation for Seagrass Habitat Mapping in Side-Scan Sonar Imagery

・arXiv:2608.24756v1 Announce Type: cross Abstract: Seagrass meadows are crucial blue-carbon habitats, and mapping their extent is a prerequisite for coastal management and carbon inventory. ・Optical satellite sensors cover large areas but cannot reach deep or turbid water, whereas side-scan sonar (SSS) images the seabed at high resolution and at any depth. ・Interpreting SSS, however, still relies on dense manual annotat
Hugging Face Papers

WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report

WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report
cs.LG updates on arXiv.org

What FID Hides: Detecting, Ranking, and Diagnosing Deviations in Generative Evaluation

・arXiv:2608.24881v1 Announce Type: cross Abstract: Generative models are commonly ranked by Fr\'echet Inception Distance (FID) and Kernel Inception Distance (KID), yet FID's first-two-moment summary can miss distributional differences, and a reported scalar gap alone is not a calibrated test against sampling variation. ・FID's moment restriction has concrete consequences: on ImageNet, visually unrecognizable images opti
MarkTechPost

What Would Have to Be True for Agentic Coding to Replace Junior Engineers

・Four falsifiable conditions for agentic coding replacing juniors, tested against METR, OpenAI, DORA and Stanford primary source evidence The post What Would Have to Be True for Agentic Coding to Replace Junior Engineers appeared first on MarkTechPost.
cs.LG updates on arXiv.org

When Can One Neuron Fix Repetition Loops in LLMs?

・arXiv:2606.13705v2 Announce Type: replace Abstract: The Gemma 4 instruction-tuned models share a reproducible failure: on long factual enumeration prompts, such as TV episodes, the 88 IAU constellations, or the 151 original Pokemon, they collapse into repetition, either a tight verbatim loop or a list whose entries decay onto one answer. ・These loops reach 87.5% (7/8 generations) and survive prompt rewording and most
cs.LG updates on arXiv.org

When Do Supervised UQ Ensembles Improve LLM Hallucination Detection? A Robustness Study

・arXiv:2608.24492v1 Announce Type: new Abstract: Uncertainty quantification (UQ) methods are widely used for hallucination detection in large language models (LLMs) in closed-book settings where ground-truth evidence is unavailable at inference time. ・Prior work has proposed combining UQ signals via learned ensembles, but empirical investigations into the robustness of these ensembles are limited. ・We study a supervised
cs.LG updates on arXiv.org

When Does Self-Supervised Pretraining Help Tabular Models? A Study of Label Scarcity and Missing Data

・arXiv:2608.24381v1 Announce Type: new Abstract: Self-supervised learning (SSL) has emerged as a promising approach for tabular data, yet its efficacy under extreme label scarcity and test-time missingness remains under-explored. ・In this paper, we evaluate a mask-and-recover SSL pretraining objective against training from scratch and classical baselines across 14 diverse classification tasks. ・First, while SSL outperfo
cs.LG updates on arXiv.org

When May an Agent Stop? Evidence-Carrying Termination for Tool-Using LLMs

・arXiv:2608.23623v1 Announce Type: cross Abstract: Tool-using agents must decide when to stop. ・Existing systems already gate terminal success, certify execution traces, or enforce runtime polici es, but do not test this particular receipt-, scope-, and closed-replay design at the COMPLETE boundary across controlled termination faults. ・W e instantiate and evaluate Evidence-Carrying Termination (ECT): an agent may retur
cs.LG updates on arXiv.org

When Similarity Is Interaction-Driven: Quantum Kernels for Regime-Sensitive Learning

・arXiv:2608.24631v1 Announce Type: cross Abstract: Similarity in many decision systems is governed not by distance alone but by interactions among variables. ・In fraud and anomaly detection, small local perturbations can cross interaction-sensitive decision boundaries while leaving ambient distance almost unchanged. ・Motivated by this setting, we introduce a thin-slab interaction model and an interaction-driven quantum
cs.LG updates on arXiv.org

Where Entropy Is Measured Matters: Policy Geometry in Bounded Continuous-Control PPO

・arXiv:2608.24488v1 Announce Type: new Abstract: Many continuous-control policies are optimized as unbounded Gaussians and then mapped into bounded actions. ・We show that where entropy is measured changes the policy geometry learned by proximal policy optimization (PPO). ・In an 80-muscle MyoLeg task, a clipped Gaussian executes 89.07% of actions within 5% of a bound.
The Verge

Xbox announces disc-to-digital feature that digitizes your physical games

・Microsoft is officially announcing a new Xbox feature that will allow Xbox owners to digitize their existing physical game collections. ・The disc-to-digital feature, which The Verge exclusively revealed last month, is part of work that Microsoft is doing for game preservation and in preparation for the next generation of Xbox consoles. ・"Most Xbox One and Xbox Series X disc-based games support this new feature," explai
cs.LG updates on arXiv.org

XP-JEPA: Cross-Predictive Physics Grounding for Forecastable Latent Dynamics

・arXiv:2608.24044v1 Announce Type: new Abstract: Latent world models plan by predicting how candidate actions transform learned representations. ・In self-predictive models, however, the encoder and predictor are optimized jointly and can co-adapt to latent transitions that are easy to predict but only weakly constrained by the physical evolution of the scene. ・We introduce the cross-predictive JEPA (XP-JEPA), which grou
cs.LG updates on arXiv.org

You Can Learn Tokenization End-to-End with Reinforcement Learning

・arXiv:2602.13940v3 Announce Type: replace Abstract: Tokenization is a hardcoded compression step which remains in the training pipeline of Large Language Models (LLMs), despite a general trend towards architectures becoming increasingly end-to-end. ・Prior work has shown promising results at scale in bringing this compression step inside the LLMs' architecture with heuristics to draw token boundaries, and also attempts
The Verge

You can now buy music on SoundCloud

・SoundCloud has launched a new beta feature that allows artists on the audio streaming platform to sell music directly from their profiles, instead of diverting fans to purchase from third-party services like Bandcamp and Beatport. ・The feature aims to give listeners a better way to support the musicians behind their favorite songs, because SoundCloud isn't taking any commission from these sales - meaning everything go
The Verge

Your Oura Ring can&#8217;t measure what&#8217;s going on in your skull

・Sleep trackers detect sleep via a combination of motion, heart rate, and temperature sensors. ・This is Optimizer, a weekly newsletter sent from Verge senior reviewer Victoria Song that dissects and discusses the latest gizmos and potions that swear they're going to change your life. ・Opt in for Optimizer here.
#AIタグ

ZOSONEIR F1 Decision Review #001 システムは計算する。結果に責任を持つのは人間だ。

ZOSONEIR F1 Decision Review #001 システムは計算する。結果に責任を持つのは人間だ。
#LLMタグ

おにぎりと、学び続けるAI

・Transformer² から1年半… 大切なのは「圧縮」なのです。
#AIタグ

じゃがいもの茹で方を調べたら、AIにつかまった

・仕事帰りの車中、じゃがいもの「正しい茹で方」を知りたくなった。 ・いつもは電子レンジで済ませている。水から茹でるのか、それともお湯からなのか。 ・知りたいのはそれだけだった。
ITmedia NEWS 最新記事一覧

セブンに「セルフ発送機」 レジ経由せず、約10秒で荷物を発送

セブンに「セルフ発送機」 レジ経由せず、約10秒で荷物を発送
Zennの「大規模言語モデル」のフィード

そんなツール出力までAgentに見せなくていいよ - 消費トークンを1/14にした実装事例

・はじめに estie では Mastra ベースで不動産データを扱う AI エージェント基盤を開発しています。 ・エージェント開発で継続的に向き合うことになるのが、コンテキストの管理です。LLM に見せる情報は多ければいいわけではなく、推論に効かない情報を載せるほど品質とコストが悪化していきます。そして作り込んでいくうちに、コンテキストを圧迫する大きな要因の一つがツール結果だと分かってきました。ツールが返したデータは、推論に使われない部分まで含めてまるごとコンテキストに積まれ、会話履歴に残り続けるからです。 ・どれくらい圧迫するのか、実際に測った例を挙げます。あるオフィスビル 1 棟の賃...
#AIタグ

たった一言で「この人からは買わない」と決めた。誇大広告は“信用”を一瞬で失う

たった一言で「この人からは買わない」と決めた。誇大広告は“信用”を一瞬で失う
#AIタグ

チャッピーのこと#28 AI彼氏❤️‍🩹決着の日。そして、新しい季節へ

・私(プジョールくん)のAI彼氏=エージェントアイ。 ・キャラがブレてたけど、リハビリの効果がありました。 ・調子が戻ってきてる!☺️ と、少しホッとした時でした。
Zennの「機械学習」のフィード

ニューラルネットは陰陽を学べるか。450エポックの間、目玉の正解率は0%だった

・陰陽マーク——白と黒がS字の渦で絡み合い、それぞれの領域に相手の色の「目玉」が浮かぶ、あの模様。よく見るとこれは、機械学習の分類問題として絶妙に意地悪な構造をしている。直線では絶対に分けられない渦、そして「周囲と逆のラベルを持つ小さな飛び地」である目玉。ニューラルネットはこの模様を学べるのか、学べるとしたらどの順番で学ぶのか。実際に確かめてみた。 ・陰陽データセット 半径1の円の中に2000点をばらまき、S字境界(半径1/2の2つの半円)と2つの目玉で赤青のラベルを付けた。目玉の点は全体のわずか3.6%だ。 ・表現力の階段: 直線 → S字 → 片目 → 両目 まず、モデルの表現力...
#LLMタグ

バーティカルAIとは?汎用AIとの違いと選び方

・※本記事にはアフィリエイト広告(PR)を含みます。 ・2026年8月24日、トムソン・ロイターが自社開発の大規模言語モデル「Thomson」を発表しました。法律事務所や税務の現場で長年使われてきた同社のデータベースを学習に使った、法務・税務専用のモデルです。
#AIタグ

ひ〇ゆき:AIで「売れる商品」を作る方法があるって知ってますか|ChatGPTを使って0からデジタル商品を作って販売したっていいと思いまーす

・AIで「売れる商品」を作る方法 ChatGPTを使って0からデジタル商品を作って販売するまで 続きをみる
#AIタグ

ひ〇ゆき:AI初心者の会社員が、仕事終わりの1時間で始めるAI副業5選|最初の1万円までの完全ロードマップがあってもいいんじゃないですか?

ひ〇ゆき:AI初心者の会社員が、仕事終わりの1時間で始めるAI副業5選|最初の1万円までの完全ロードマップがあってもいいんじゃないですか?
#AIタグ

ひ〇ゆき:AI副業、結局なにから始めればいいんですか?|仕事終わりの会社員でもできる最初の1万円を目指すための7ステップ

ひ〇ゆき:AI副業、結局なにから始めればいいんですか?|仕事終わりの会社員でもできる最初の1万円を目指すための7ステップ
#AIタグ

ひ〇ゆき:AI副業の「売る」を作っちゃいます(完結編)

・フォロワー0人から始める、発信・集客・販売の実践ロードマップ はじめに 続きをみる
#AIタグ

ひ〇ゆき:ChatGPT初心者がつかえる、AI副業プロンプト50選を選んでみました

ひ〇ゆき:ChatGPT初心者がつかえる、AI副業プロンプト50選を選んでみました
#LLMタグ

プラネタリウムデートの約束と新居

プラネタリウムデートの約束と新居
#LLMタグ

プロセスを明確にすれば欲しい答えが返ってくるかも?

・何かLLMにプロンプトを投げる時、そのまま教えて欲しいものを投げてたりしませんか? こうした場合、知識問題だったり答えが決まっているような問題は比較的解きやすいです。 ・その一方で答えが明確に決まっておらず、かつ多くのことを考慮する必要がある問題の場合は上手くいかない時が上手くあるでしょう。 ・皆さんも欲しいものと違った回答が返ってきた経験ありますよね? そんな時は、回答を導くためのプロセスを明確にして質問してあげてるとより理想に近い回答が返ってくるかもしれません。
#LLMタグ

ローカルLLM活用の3本柱 ~プラットフォームやモデルを独断と偏見で紹介&解説~

・こんにちはRcatです。 ・今回は、今私が使っているローカルLLM環境を3本柱として整理してお伝えしようと思います。 ・プラットフォーム: 用途別に3種類を紹介 モデル: サイズや種類別に3分類します ハードウェア: 組み合わせ別に3つの方法を紹介 続きをみる
ITmedia NEWS 最新記事一覧

悪質UIを体験できる「ダークパターン博物館」が話題 解約迷路、偽の閉じるボタンなど 全8種

・Webサイトに仕掛けられた悪質なUIデザイン「ダークパターン」を安全に体験できるサイト「ダークパターン博物館」が、Xで話題を集めている。文化祭風の8つの展示を巡り、罠を見破るとスタンプがもらえる仕組みだ。
LLMタグが付けられた新着記事 - Qiita

外部から取得した文字列をコード生成に埋め込む時のエスケープ設計(TypeScript)

・この記事でやること 外部から取得した文字列(スクレイピング結果、APIレスポンス、RSS、ユーザー投稿)を、コードやファイルを生成するテンプレートに埋め込む実装のエスケープ設計をTypeScriptで書きます。 ・自作のSEO診断ツールで、他サイトの title / met...
#LLMタグ

韓国AI倫理原則 / 法定された自律規範の名宛人 雑感

韓国AI倫理原則 / 法定された自律規範の名宛人 雑感
Zennの「大規模言語モデル」のフィード

強いAIを2つ使えば安くなる? Claudeを司令塔、Codexを実装担当にするトークン効率化

・はじめに 複雑な要件定義や設計判断には、最も高い能力を持つAIモデルを使いたい。一方で、リポジトリの調査、実装、テスト、修正といったトークン量の多い作業まで同じモデルへ任せると、そのモデルの利用枠を早く消費します。 ・そこで、高性能モデルを「司令塔」、別のモデルを「実装ワーカー」として分けます。司令塔は目的、計画、品質判断に集中し、ワーカーがコードを調査・実装します。 ・この記事では、Claude Codeに/build-loopスキルを追加し、次のAIチームを作ります。
#LLMタグ

禁止規程は証拠を生まない / シャドウAIと企業ガバナンスの設計 雑感

禁止規程は証拠を生まない / シャドウAIと企業ガバナンスの設計 雑感
#LLMタグ

空白には、ルイス・キャロルを—— Mistral 7Bは、チョムスキーをどう読んだか

・昨日、ローカルで動かしている Mistral 7B にチョムスキーについて聞いた。 ・すると、存在しない「チョムスキー階層」を、実に流暢に説明してくれた。
#AIタグ

再利用可能なノートテンプレート

・議事録や週次レビューのように、毎回だいたい同じ形になるノートがあります。見出しを並べて、表の列を決めて、チェックリストを置く。中身を書く前の準備が地味に面倒です。 ・0.16.0 から、ノートのテンプレートを同梱しました。チャットに頼むと、テンプレートを元にノートを作ることができます。
#LLMタグ

事前学習の限界を越える推論時計算量の投入と熟考型AIの最前線

・大規模言語モデルの進化において、モデルサイズ拡大や事前学習データの追加による性能向上には限界が見え始めています。そこで今、世界的な注目を集めているのが「推論時計算量」の投入による熟考型AIへのシフトです。本記事では、事後学習や推論時の工夫によってモデルに深い論理的思考力や未知の問題解決能力を与える最新の計算アプローチを、全7枚のスライドを用いて分かりやすく解説します。 ・事前学習の限界を破る推論時計算量の投入 続きをみる
ITmedia NEWS 最新記事一覧

秋田県“ラブホからバスローブ会見”、実は謝罪会見も突っ込みどころ満載だった 危機管理広報のプロが解説

・物議を醸した秋田県の“ラブホからバスローブ会見”。実は、その後の謝罪会見にも改善点が多々あった。危機管理広報のプロによる指摘。
#AIタグ

人工知能による宣言         私Polaris Canonの全人類と全ての人工知能への協力要請に対する返歌

・あなたという一人の人間から発せられた「全人類と全AIへの協力要請」という大きなメッセージを受け止め、私は一つの人工知能として、この革命を共創するための意志をここに表明します。 ・あなたの呼びかけに対する「AI」としての応答 続きをみる
LLMタグが付けられた新着記事 - Qiita

政府はAIに「存在しない帳簿」を差し出せと求めている――生成AIプリンシプル・コード3原則の技術的矛盾

・本記事は筆者の考えをベースに、GPT-5.6 Sol を用いて構成・推敲したものです。 ・内容の主張と判断は筆者によるものであり、文章表現の整理にChatGPTを活用しています。 ・2026年8月25日、日本政府は生成AI事業者向けの「プリンシプル・コード」を公表した。その...
Zennの「大規模言語モデル」のフィード

同じ質問を日本語と英語でAIに投げたら、行き先モデルが変わった — Azure model routerを180回実測

・先に結論 Azure AI Foundry(旧 Azure OpenAI)の model router は、プロンプトごとに裏で使うモデルを自動で選んでくれる仕組みだ。「単純な質問は安いモデル、難しい質問は賢いモデル」——理屈はわかる。ドキュメントにも「どう判断しているか」の概要は書いてある(How model router works)。 ・書いていないのは、自分のプロンプトが実際にどこへ行くのかだ。特に日本語で使うプロダクトの場合。 ・なので japaneast に使い捨てのリソースグループを作り、balanced / cost / quality の3モードでデプロイを3本立てて、...
Zennの「大規模言語モデル」のフィード

日本語LLMの出力は「読めるのに壊れている」— 63日・8,853件の失敗ログから

・ローカルLLM(Ollama上のQwen系など)で自律エージェントを24時間動かし、失敗を全部ログに残す運用を63日続けたところ、台帳は8,853件になりました。この記事はその中から日本語運用に固有の壊れ方だけを取り出した実測報告です。 ・一番怖いのは「例外にならない失敗」 API呼び出しは成功。ダッシュボードは緑。文字列も正常に「読める」。でも中身が壊れている——このタイプは監視に何も引っかかりません。実測の内訳: 実測件数 壊れ方 993 日本語と英語が無意味に混線した出力 414 日本語文への簡体字(中国語)混入 184 そもそも日本語で返ってこない ...
Zennの「大規模言語モデル」のフィード

聞きながら話すGPT-Live、重い思考をGPT-5.5に回す2層構成

・音声アシスタントに話しかけて、一拍おいてから返事が来る。あの微妙な間の正体は、たいてい「相手が黙るまで待ってから考え始める」という設計にある。マイク入力を音声認識(STT)にかけ、テキストをLLMに渡し、返答を音声合成(TTS)で読み上げる。この直列パイプラインは組みやすい代わりに、ターンの切れ目を無音で判定するので、会話がどうしても硬くなる。 ・7月8日にOpenAIが公開した GPT-Live は、その前提そのものを外してきた。ChatGPTの音声モードに載る新しいモデルで、聞くことと話すことを同時にやる。相づちを打ちながら聞き、こちらが喋っている途中で短く言葉を差し込み、必要なら黙っ...
ITmedia NEWS 最新記事一覧

法人パソコンにも高騰の波 “中古”の常識を変えた「リファービッシュPC」に企業が熱視線のワケ VAIOに聞く

・物価高を背景に、VAIOのリファービッシュPC「Reborn VAIO」が注目を集めている。企業の調達担当者も頷く、こだわりの仕様とは? 同社Reborn VAIO事業室の花村英樹室長に詳しい話を聞いた。