ai Trend Report

Dashboard へ戻る
Date: 20260818 Articles: 399 Scope: curated summary

あなたのアイデアを、今すぐ形に。

公開先に迷ったら、WebFileBinで一発公開。

HTMLをドラッグ&ドロップするだけで、すぐ公開できます。

3
High impact
5
Mid impact
391
Signal watch

なぜこのサイトを作ったのか

私たちそれぞれが個別にAIを使って情報収集し、同じような3行要約を作るたびに、世界中で膨大な電力と計算リソースが消費されています。 本プロジェクトは、あらかじめ広範な情報を取得・集約しておくことで、個別のAI実行回数を減らし、地球環境(GPU/TPU負荷)に配慮した効率的な情報収集を目指す実験的なダッシュボードです。

Domain filters
Star filters
#AIタグ

【💻 TECHNOLOGY #7】 AIが広がるほど、「現実世界」の価値も上がる。

・AIが広がるほど、「現実世界」の価値も上がる。 ・AIというと、どうしてもデジタルの世界を想像します。 ・データ、クラウド、ソフトウェア、アルゴリズム。
#LLMタグ

2026-08-17 Hacker News Top 10

・GitHubで大規模な性能低下。AIモデルの実用性能、StripeによるOpenRouter買収報道、AI規制なども話題に。 ・Hacker Newsで今日注目を集めた話題を、日本語でコンパクトにまとめました。 ・開発・AI・プロダクトづくりの流れをつかむための、朝の読みものです。
#AIタグ

最新技術で「アホなAI」を魔改造したら面白そう

・最近、古いAIモデルとか、ものすごく小さいAIモデルについて調べていて、変なことを思いついた。 ・今のAI開発って基本的には、 もっと賢くする 方向に進んでいる。 ・モデルを大きくしたり、推論能力を上げたり、長い文章を扱えるようにしたり。
#AIタグ

AI時代、思考する人間は「人生の自由度」を手に入れる

・AIが急速に発達している。 ・文章を書く。情報を調べる。要約する。翻訳する。計算する。プログラムを書く。企画を考える。 ・これまで「頭のいい人ができること」と考えられてきた仕事のかなりの部分を、AIが短時間で行えるようになった。
cs.LG updates on arXiv.org

Discrete Diffusion Language Models Are Training-Free Multi-Label Classifiers

・arXiv:2608.14649v1 Announce Type: new Abstract: We present dLLM-SetScore, a training-free method that uses discrete masked-diffusion language models for multi-label text classification. ・For each candidate label, it asks a short yes/no question and compares the probabilities of the two answer tokens at one masked position. ・The method uses no task-specific fine-tuning or training on textual-entailment datasets; a 200-e
@IT 全フォーラム 最新記事一覧

GPUから光接続まで、AI基盤の「本命5社」 Gartnerが選んだ理由

・生成AIの本番利用が広がる中、その処理を支えるAIインフラ分野の競争も激しさを増している。Gartnerは、AI半導体ベンダーの現在の競争における「Companies to Beat」(分野をリードする本命企業)を特定した。
#AIタグ

バッチファイルの日本語が文字化けして自動実行が死んだ日 — AIの前に、Windowsが敵だった

・※本連載は、AIエージェントによる株式デモトレード(実際の資金を使わない仮想売買)の記録です。特定銘柄の売買を推奨するものではなく、金融商品取引法上の投資助言ではありません。連載の全体像は第0回をどうぞ。 ・[exit 9009] cycle end 続きをみる
#AIタグ

縦型OSと水平OS 〜AIと雑談してたら理論が自然発生した話〜

・私は学者じゃないから小難しい言葉は使えないけど、コパイロットと雑談してたら、いつのまにか独自理論が組み上がってたらしい。 ・## きっかけは「音楽アプリ、久しぶりに開いて嬉しい」だった 久しぶりにパソコンを開いて、中に録音してある音楽アプリを開いた。 ・そこから、私は音楽のジャンルに対してフェアだという話になった。ベートーヴェンの次に八代亜紀、その隣にウィーン少年合唱団、爆風スランプ、Ado。並べても何の違和感もない。共通点はただひとつ、「作為的なものや、誰かのマネは受け付けない(モノマネ芸は好き^^)」ということだけ。下手くそでもいい。心がこもっていれば、年齢もジャンルも人種も関係ナッシング。
#AIタグ

[実体験]AIを使って"note"で不労収入を作る方法。

・会社内でFIREしたり、AIアートを使った副業で無名から月20着服を売ったりしてきた僕が、正直一番再現性が高いと感じたのが"note"でした。 ・今回は、AIを使って実際にnoteで収入を作った手順を、包み隠さずシェアします。
機械学習タグが付けられた新着記事 - Qiita

「クリーンなLLM」は作れるのか――Webクロール、蒸留、反蒸留から見る学習データの現実

・本記事は筆者の考えをベースに、GPT-5.6 Sol を用いて構成・推敲したものです。 ・内容の主張と判断は筆者によるものであり、文章表現の整理にChatGPTを活用しています。 ・ChatGPTは、なぜLinuxカーネルを説明できるのか。
#AIタグ

「この会社で働きたい」と思える会社は、どうやって見つける? — ミッションに共感できる会社との出会い方

「この会社で働きたい」と思える会社は、どうやって見つける? — ミッションに共感できる会社との出会い方
#LLMタグ

「暗号は使っていません」と説明されたので信じていたら、提出の直前に自分のアプリの中から出てきました

「暗号は使っていません」と説明されたので信じていたら、提出の直前に自分のアプリの中から出てきました
#AIタグ

「今日も何も成長していない。一体私は何をしていたんだ?」 — 仕事に焦りを感じ始めたあなた

「今日も何も成長していない。一体私は何をしていたんだ?」 — 仕事に焦りを感じ始めたあなた
ITmedia NEWS 最新記事一覧

「食欲が削られる」――AIで作られた“気色の悪い飲食店メニュー”がXで物議 「実物を見たい」の声も

・生成AIで作ったとみられる飲食店のメニューやポスターを巡り、Xで物議が広がっている。8月13日ごろから関連する投稿がみられ、とあるサービスエリアの飲食店メニューを写した投稿が反響を呼ぶと、同様の事例を紹介する投稿が急増した。
#AIタグ

【🇻🇳ベトナムビジネス進出最前線!】「月200万円は高すぎる」と言われたときに、私が説明していること

・ベトナムEC進出の予算は、賭け金ではありません ベトナム進出のご相談をいただくとき、広告予算の話になると、かなりの確率でこの言葉が出ます。
#LLMタグ

【3】ローカルPCで動くAIアプリを開発するAIアプリを作る(LM Studio Bionic → Claud Code 編)

・結局、仕様書の詳細版を作るのはClaud coworkで作業を行った。 ・理由は単純にchatGPTは別の案件(本業)で使って、使用量を残しておきたかったから。 ・正直、仕様書(詳細版)を作っていて分かったのは、このアプリ開発は大規模開発と呼ばれるクラスの内容であるらしいこと。
#LLMタグ

【ChatGPT】暴露します‼️学校では教えてくれないプロンプト講座、有料書籍2000円のプロンプトの中身を4行で示す。

・・今回のテーマ 真面目な投稿。超有益なのでたぶん、期間限定にします。ChatGPTに指示して最高の回答を得るには? ChatGPTはユーザの入力毎に学習します。 ・あっ、ちくわさんならこういう感じのプロンプト出力すればいいなじゃね?みたいな。 ・なので、日々送っているプロンプトが出力のクセを生み出します。
#LLMタグ

【Claude Codeと対話してみた vol.4】生成AIは個人でも作れるのか

・※この記事は、Claudeとの対話内容をもとに、文章の生成をClaudeに、画像の生成をChatGPTに依頼して作成しています。
Zennの「大規模言語モデル」のフィード

【Kitesurf】Playwright MCPを常駐させないブラウザ戦略

・Global設定にMCPサーバーを配置していたのですが、Agentを起動するたびにMCPサーバーも起動するため、メモリ圧迫の大きな要因の一つになっていました。不要なMCPサーバーについては、コンテキストを無駄に占有させないという観点から削除が推奨されることが多いですが、実際にはメモリ消費を抑えるうえでも重要なようです。 ・日常的に、tmux経由でClaude Codeから複数のClaude CodeやCodexを起動しています。そのため、起動するAgentが増えるほどMCPサーバーのプロセスも増え、メモリ使用量が急激に膨らむ状況になっていました。 ・今回は、上記の事象を調べ、ブラウザを必要...
#AIタグ

【シバ丸印の会社シミュ⑨】「これなーに?」を、もぐもぐペッペ。ワンコの疑問から知恵の泉を作るワン

・⑧ 両手をパンッ! 人月を分解して、価値に再構成する。知神シェーシャの錬金術だワン https://note.com/shibamaru_log/n/ned66afe7119a 続きをみる
#AIタグ

【一部メンバー限定】🌻 8月⑤ フリー画像&一部メンバー限定27枚

・はじめに記事の説明など ※画像の無断転載・再配布は禁止です。ご協力よろしくお願いします🙇‍♀️ 続きをみる
機械学習タグが付けられた新着記事 - Qiita

【書評】 Machine Learning Engineering in Action

・はじめに 「精度は出たのに、そのモデルは本番に行かなかった」 機械学習の案件に関わったことがある方なら、一度は見た光景ではないでしょうか。本書『Machine Learning Engineering in Action』(Ben Wilson 著 / Manning)は...
Zennの「大規模言語モデル」のフィード

【紹介リンクあり】RunPodの料金体系を徹底解説!GPU別単価・ストレージ課金の落とし穴・最安運用のコツ

・🎁 初回登録特典(無料クレジット付与)について 紹介リンク経由で新規登録し、$10以上チャージして利用を開始すると、$5〜$500相当のボーナスクレジット がランダム(最低でも$5〜$10以上、運が良ければ最大$500)で付与されます。 ・👉 RunPodに招待リンク経由で登録してボーナスクレジットを受け取る https://runpod.io?ref=9ok0s3nk 1. ・RunPodの料金体系 基本仕様 RunPodの課金は 完全従量課金制(秒単位課金) です。
#AIタグ

🌅 北九州ニュース|8月19日(水)

・このチャットを見る誰かのおすすめチャットです。chatgpt.com 🏙️ 北九州市・企業向け採用セミナー 続きをみる
#LLMタグ

🔊音声あり(日&英):AIの常識が覆る!? LLMの頭の中、内在的次元(ID)の真実を論文解説!

🔊音声あり(日&英):AIの常識が覆る!? LLMの頭の中、内在的次元(ID)の真実を論文解説!
Qiita - 人気の記事

14MBのAIモデルは本当に動いた — DLL 1個・RAM 37MB・1回0.5秒、ただし日本語は0/5

・GitHub Trending に「14MB のモデル」が来ていました。 ・Cactus Compute の Needle 2 です。45M パラメータで、ツール呼び出しに特化しています。週で +3,627 スター伸びていました。 ・14MB というのは、画像1枚くらいのサイズ...
WIRED

1Password Coupon: Score a Free Trial in August 2026

・Save up to 28% on business and personal memberships with 1Password promo codes and deals.
ITmedia NEWS 最新記事一覧

2.5次元アイドル「いれいす」事務所のBIツールに不正アクセス ファンの氏名・電話番号や“推し”情報漏えい

・漏えいした情報は、氏名・住所・電話番号・メールアドレス・生年月日・性別の基本情報に加え、購入日時や商品名などの購買履歴、決済金額記録、サービス加入状況、応援対象タレントの選択情報。
#AIタグ

2026年8月19日|誰かと比べた瞬間、神様の恵みまで「不公平」に見えてくる

2026年8月19日|誰かと比べた瞬間、神様の恵みまで「不公平」に見えてくる
Zennの「大規模言語モデル」のフィード

2026年8月公開のQwen3.8-27BをRTX 5090で実測|ローカルLLM 5モデル徹底比較

・2026年8月公開のQwen3.8-27BをRTX 5090で実測|ローカルLLM 5モデル徹底比較 同じ27Bなのに、生成速度2.4倍。 ・2026年8月13日に公開されたQwen3.8-27Bを実測し、1ヶ月前に組んだ構成を迷わず入れ替えました。ローカルLLMは「常にアップデートを取り込む前提」で運用しているからです。 ・この記事は、RTX 5090(VRAM 32GB)1枚で「翻訳・コーディング・構造化抽出・画像読取・論理推論・成人向けローカライズ」の6タスクを全モデルに実際に解かせ、さらに「明確に違法な要求をモデルが拒否するか」までを検証した全記録です。
cs.LG updates on arXiv.org

6G Native AI and Channel Foundation Models

・arXiv:2608.14591v1 Announce Type: cross Abstract: The integration of artificial intelligence (AI) and wireless communications is widely regarded as a core objective of sixth-generation (6G) systems. ・However, both the meaning of native AI and the type of AI capability that should be embedded into future wireless systems remain open to interpretation. ・This paper discusses 6G native AI from a system-design perspective a
cs.LG updates on arXiv.org

A Banach-Space Theory of Markovian Halpern Iteration for Non-Expansive Maps

・arXiv:2608.15966v1 Announce Type: new Abstract: We study stochastic approximation of fixed points of a non-expansive operator when the oracle samples originate from a continuing Markovian trajectory. ・A direct block-minibatch implementation of Halpern iteration attains an expected last-iterate residual of order $O(\log N/N)$, but accrues a substantive complexity of $\tilde O(\epsilon^{-5})$ Markovian samples.
cs.LG updates on arXiv.org

A Low-Cost IoT Device for Environmental Monitoring and Embedded Solar Forecasting with On-Device Incremental Learning

・arXiv:2608.14698v1 Announce Type: cross Abstract: Hyperlocal meteorological sensing is essential for accurate solar photovoltaic forecasting, yet professional-grade meteorological stations require investments easily exceeding 1000~USD per node, making distributed deployments economically inaccessible. ・This work presents a modular internet of things (IoT) device based on the ESP32 microcontroller integrating temperatu
cs.LG updates on arXiv.org

A Novel Fourier Feature Network for Solving Partial Differential Equations

・arXiv:2608.14733v1 Announce Type: new Abstract: Building on the foundation of single-hidden-layer neural networks, Fourier Feature Networks (FENs) are proposed, which incorporate Fourier features using $\cos$, $\sin$, or a combination of both. ・Similar to Extreme Learning Machines (ELMs), FENs employ a single-hidden-layer architecture to generate a set of basis functions. ・The target function is then approximated as a
cs.LG updates on arXiv.org

A Physiology-Informed Digital Twin Framework for Simulating Liver Health Progression

・arXiv:2608.14969v1 Announce Type: new Abstract: We present a physiology-informed digital twin of the human liver designed for longitudinal simulation of liver function and early-stage disease progression. ・The model, referred to as HEPATWIN, integrates key hepatic processes, including carbohydrate, lipid, and protein metabolism, bilirubin conjugation, bile production, and detoxification, within a unified systems-level
cs.LG updates on arXiv.org

A Pre-Specified Construction-Confirmation Test of Operation-Level Causal Transfer Across Finite Isomorphic Symbolic Domains

・arXiv:2608.15809v1 Announce Type: new Abstract: Behavioral accuracy, linear decodability, and successful activation interventions do not by themselves show that a model carries an operation-level structure from one symbolic domain to another. ・We ask a narrower question in finite isomorphic state spaces: if the hidden-state difference between two operations is estimated separately for each source input, does adding th
cs.LG updates on arXiv.org

A Privacy Study of Sparse Collaborative Inference

・arXiv:2608.16236v1 Announce Type: new Abstract: Collaborative inference (CI) splits a model between an edge device and a server, whereby the client computes an intermediate activation, transmits it, and the server completes the computation. ・This raises two concerns, the communication cost of the transmission and the risk that it reveals private information about the input. ・Recent work reduces this cost by sparsifying
cs.LG updates on arXiv.org

A Reproducibility Study of Partial Residual Ablations in Pre-LN Transformers

・arXiv:2608.14689v1 Announce Type: new Abstract: Residual connections are a fundamental component of transformer architectures, yet the roles of the attention and feed-forward residual pathways remain poorly understood when considered independently. ・This paper presents a reproducibility study of partial residual ablations in Pre-LN GPT-style transformers trained at two scales (10M and 124M parameters). ・I compare four
cs.LG updates on arXiv.org

A Tree-Structured Approach for Phishing Template and Attacker Attribution Analysis

・arXiv:2608.16158v1 Announce Type: new Abstract: Phishing remains a persistent and evolving cybersecurity threat, with attack volumes reaching record levels. ・This growth is driven by the industrialization of phishing through widely available phishing kits and reusable templates, which enable cybercriminals to rapidly generate and deploy large numbers of fraudulent webpages. ・Although surface-level attributes may differ
cs.LG updates on arXiv.org

A Unified Mamba--MoE Surrogate for Closed-Loop Simulation and Measurement-Window Forecasting of Inverter Transients

・arXiv:2608.15051v1 Announce Type: new Abstract: This paper proposes a Mamba surrogate model with mixture-of-experts (MoE) routing to represent the transient dynamics of inverter-based resources. ・A Mamba surrogate model is a predictive machine learning model built on the Mamba architecture. ・MoE routing uses a router network to assign data-dependent weights to specialized subnetworks (experts).
cs.LG updates on arXiv.org

A Vision Transformer for ECG-Based Detection of Left Ventricular Systolic Dysfunction Across Multiple Clinical Sites

・arXiv:2608.14723v1 Announce Type: cross Abstract: Reduced left ventricular ejection fraction (LVEF) is frequently asymptomatic and often detected only after advanced heart failure develops. ・Electrocardiograms are recorded routinely yet underused for this condition, because reduced LVEF has no single diagnostic waveform. ・We trained an ensemble of vision transformers from scratch to detect reduced LVEF ($\leq$40%) from
The Verge

ABC sues the FCC over Trump and Carr’s campaign of threats

・ABC is suing the Federal Communications Commission over claims the agency "waged a retaliatory campaign" against its networks over the content they broadcast. ・In a lawsuit filed in federal court on Tuesday, ABC and its parent company Disney accuse the FCC of "punishing ABC for its speech" by threatening its broadcast licenses. ・It's the latest escalation in the dispute between ABC and the Trump administration, which h
cs.LG updates on arXiv.org

AdROD: HyperNetwork-based Adversarially Robust Object Detection for Autonomous Driving

・arXiv:2608.16031v1 Announce Type: new Abstract: Camera-based object detectors are vulnerable to physical adversarial attacks designed to suppress detections. ・While adversarial training and input purification offer some protection, they often overfit to specific attack distributions and fail on adaptive adversaries. ・This paper presents AdROD, an embedded, stochastic ensemble defense software designed for autonomous dr
cs.LG updates on arXiv.org

Advancing Open and Reproducible Relational Learning: RelArena-$\alpha$, TabPFN-Rel and RPI

・arXiv:2608.16319v1 Announce Type: new Abstract: This first release of Prior Labs in relational learning shows our continued commitment to open science. ・We open-source three pieces of software that we expect to accelerate research in the field towards meaningful real-world impact. ・We aim to steer further development based on feedback from, and in collaboration with, the community.
Hugging Face Papers

Advancing Open and Reproducible Relational Learning: RelArena-α, TabPFN-Rel and RPI

Advancing Open and Reproducible Relational Learning: RelArena-α, TabPFN-Rel and RPI
Hugging Face Papers

Agentic Transaction: Towards ACID-Compliant Agent Systems

Agentic Transaction: Towards ACID-Compliant Agent Systems
#AIタグ

AION-CORE 64 INTO THE UNKNOWN|DeepSeek HarnessにQwen3.8を載せた。Local IntelligenceがRuntimeを手に入れた

・昨日、AION-CORE 64ではひとつの節目を迎えた。
Zennの「大規模言語モデル」のフィード

AIエージェントが「知っているはず」を間違える理由 — コンテキスト設計の実務

・対象: Claude Code / Cursor / Copilot Agent などを業務で毎日使っているエンジニア 前提知識: エージェント型のAIコーディングツールを数ヶ月以上使っていること この記事で解けること CLAUDE.md に書いたはずのルールが守られない理由が、確率の問題ではなく構造の問題だと分かる コンテキストを3層に分けて設計する具体的な方法 リファクタリングでコンテキストが壊れる問題を、構成管理の問題として潰す仕組み 「劣化しているのに気づかない」状態を検知する測定方法 逆に、この記事はプロンプトの書き方のコツを扱いません。それは効果が出ても再現しない...
#AIタグ

AIが壊すのはプログラマーではなく「人月」かもしれない

AIが壊すのはプログラマーではなく「人月」かもしれない
#AIタグ

AIと本の未来

・私とAIのやりとりを そのままコピペしたものです。 ・日々の出来事や気持ちを、 時間の隙間と心の余裕がある時に 更新しています。
Zennの「大規模言語モデル」のフィード

AIは3Dをどこまで作れるのか — 2.5Dと高度な3D Geometryの間にある壁

・まずは、次の4つのモデルを見てください。 ・4気筒のピストン・クランク機構 二段式ゼネバ機構 ツインローブルーツ機構 二段式遊星歯車機構 これらはすべて、Claude Opus 5がAutodesk Fusionを操作して作ったモデルです。 ・しかも、それぞれのモデルは1つのプロンプトを一度入力しただけで、途中で追加の指示や修正の対話を挟まず、一気に作成されています。いずれも15分~25分ほどで作成されました。
#LLMタグ

AIフロンティアモデルまとめニュース

・AIは「一番賢いモデル」から「仕事に合ったモデル」へ AIモデルニュースを並べてみると、少し面白い変化が見えてきます。
#AIタグ

AI版一帯一路とは?

・中国の「一帯一路」といえば、鉄道や港湾、道路、発電所など巨大なインフラ建設を思い浮かべる人が多いだろう。 ・しかし、いま中国が一帯一路に組み込もうとしているのは、道路や港だけではない。
LLMタグが付けられた新着記事 - Qiita

Amazon Bedrock Playgroundで理解するTemperature・Top P・Top K ― AWS Certified AI Practitioner試験対策

・はじめに 先日、AWS Certified AI Practitioner(AIF-C01) を受験し、無事合格することができました。 ・試験勉強をしている中で、Amazon Bedrockをはじめとした生成AIに関するさまざまな用語が登場しますが、その中でも個人的に最初は...
Zennの「大規模言語モデル」のフィード

Amazon Bedrock の Web Search を英語と日本語で試す

・先端技術開発グループ(WAND)の伊賀です。 ・2026年8月、Amazon Bedrock に Web Search が追加されました(Introducing Web Search on Amazon Bedrock for foundation model grounding)。 ・推論リクエストにパラメータを 1 つ追加するだけで、モデルが Web 検索で回答を裏付けてくれる(グラウンディングする)ようになる、という機能です。
#LLMタグ

Amazon、希少本をAI訓練用に裁断・破棄していた 独立メディア404 Mediaが発信機で追跡

Amazon、希少本をAI訓練用に裁断・破棄していた 独立メディア404 Mediaが発信機で追跡
cs.LG updates on arXiv.org

Amortised Post-Hoc Explanation with Exact Preservation for Dynamic Graph Anomaly Detectors

・arXiv:2608.15559v1 Announce Type: new Abstract: Anomaly detection in dynamic graphs underpins financial fraud analysis, intrusion detection, and platform integrity, where automated decisions require human-interpretable justifications. ・StrGNN, the strongest performer in recent benchmarks, produces no explanation: when an edge is flagged, the analyst receives only a score. ・Explanation metrics are undefined for StrGNN b
cs.LG updates on arXiv.org

AMPLIFAI: A Multiphase CT Dataset for Benchmarking Clinical Reasoning in LI-RADS Assessment of Liver Lesions

・arXiv:2608.14778v1 Announce Type: cross Abstract: Hepatocellular carcinoma (HCC) is the third leading cause of cancer-related mortality worldwide, with early detection improving survival from 70\%. ・The standardized LI-RADS criteria establish a biopsy-free, fully imaging-based framework that can serve as a foundation for automating HCC diagnosis with artificial intelligence (AI). ・However, the lack of large, publicly a
cs.LG updates on arXiv.org

An Analytical-Prior Framework for Data-Efficient Prediction of Sound-Reduction Frequencies in Rectangular Side-Branch Helmholtz Resonators

・arXiv:2608.16873v1 Announce Type: new Abstract: High-fidelity finite-element simulations can provide accurate numerical predictions for side-branch resonators, but large simulation datasets are expensive to generate and purely data-driven surrogates may become unreliable when simulation-labelled data are scarce. ・This study develops an analytical-prior learning framework that reuses a low-cost analytical model to impr
cs.LG updates on arXiv.org

An automatic-differentiation framework for time-lapse electrical resistivity tomography inversion of hydrologic dynamics

・arXiv:2608.14661v1 Announce Type: new Abstract: Time-lapse electrical resistivity tomography (TL-ERT) provides spatially distributed information on subsurface hydrologic changes. ・However, inversion of long monitoring sequences is computationally demanding. ・Modifying the data misfit, regularization, model parameterization, or petrophysical transformation may also require new gradient derivations and separate implement
Hugging Face Papers

An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models

An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models
Hugging Face Papers

AnyTalk: Speech Animation for Arbitrary Characters Leveraging a Video Generation Model

AnyTalk: Speech Animation for Arbitrary Characters Leveraging a Video Generation Model
The Verge

Apple squashes EU beef with new App Store rules

・Apple is once again overhauling App Store rules in the European Union, which the company says will resolve its "disagreements with the Commission over business terms and alternative distribution." As part of the changes, every developer that distributes apps will be moved to a single set of business terms, and digital transactions for apps distributed outside of the App Store will be subject to a Core Technology Comm
cs.LG updates on arXiv.org

ARGUS: Attention-Guided Transformers for Scalable Person Identification Using Wi-Fi Telemetry

・arXiv:2608.14670v1 Announce Type: new Abstract: Passive, device-free person identification offers an alternative to camera- and wearable-based biometrics, yet existing wireless approaches rely largely on gait or activity cues and are rarely evaluated at scale. ・In this paper, we present \emph{Argus}, a passive Wi-Fi sensing system that identifies people from commodity Channel State Information (CSI) without requiring
OpenAI News

Asana cleared 5 years of engineering work in 2 weeks with Codex

・Asana used OpenAI Codex to replace an outdated testing system in two weeks, completing work expected to take five years for about $12K.
cs.LG updates on arXiv.org

AsyTO: Asymmetric Temporal Operator for Parameter-Efficient Multivariate Time Series Forecasting

・arXiv:2608.16098v1 Announce Type: new Abstract: Multivariate time-series forecasting faces a structural dilemma: sharing one temporal predictor across variables is parameter-efficient but forces heterogeneous variables through an identical history-to-future map, whereas learning an independent predictor per variable restores flexibility at a cost that grows with the product of variable count, context length, and hori
cs.LG updates on arXiv.org

Auxiliary uncertainty signals for LLM-assisted systematic review screening: a benchmark across eight Cohen drug-class reviews

・arXiv:2608.14551v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for title-abstract screening in systematic reviews, but their decisions lack calibrated uncertainty. ・We show that an auxiliary BERT+GCN classifier supplies a structured uncertainty signal that improves LLM screening efficiency, and we identify the prompt-delivery strategy that maximises the benefit-to-cost ratio.
Zennの「大規模言語モデル」のフィード

AXIOM Capsule DEMO最小構成で検証する「不変性・決定論・制約強制」

・AXIOM Capsule DEMO 最小構成で検証する「不変性・決定論・制約強制」 AXIOM / PSS Capsule Architecture の核心部分を最小構成にした SHOWCASE DEMO / Prototype です。 ・このDEMOでは、LLM Agentやステート管理システムにおける、 状態の意図しない書き換え 非決定的なID生成 制約違反後の処理継続 世代交代時の状態混入 といった問題に対して、 «「ルールとして守らせる」のではなく、変更経路そのものを構造化して制限する» という設計思想を、実際のRustコードと実行ログによって検証します。
cs.LG updates on arXiv.org

BDIP-Net: Dual-Interaction Graph Learning for Property Prediction of Bilayer Materials

・arXiv:2608.14640v1 Announce Type: new Abstract: Stacked bilayer materials exhibit rich stacking-dependent properties driven by the interplay between strong intra-layer bonding and weak inter-layer van der Waals interactions. ・The computational discovery of such materials is challenging because accurate structure generation typically relies on expensive DFT-based optimization, while existing machine-learning models oft
cs.LG updates on arXiv.org

Beat the Counter First: A Baseline for Temporal-Graph Anomaly Detectors

・arXiv:2608.15965v1 Announce Type: new Abstract: Progress in streaming, edge-level graph anomaly detection (GAD) has been marked by increasingly elaborate architectures, from count-min-sketch chi square tests to memory-augmented attention networks. ・Yet the empirical gains attributable to this added complexity have not been systematically evaluated. ・We propose SimpleCount, a reference with no parameter fitting that sel
cs.LG updates on arXiv.org

Behaviour Is an Incomplete Measure of Reasoning Development: Cross-surface pre-arrival accessibility and the limits of developmental inference in a recurrent-depth reasoner

・arXiv:2608.16085v1 Announce Type: new Abstract: Capability development is routinely inferred from behavioural thresholds, from final checkpoints, or from what a decoder can read out of a hidden state. ・These quantities need not identify the same event. ・We study a 30M-parameter recurrent-depth relational reasoner in a closed, oracle-defined world, using dense behavioural trajectories, two training surfaces, preregister
cs.LG updates on arXiv.org

Belayer: Efficient Fault Tolerance for LLM Agentic RL Training

・arXiv:2608.14635v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly trained with reinforcement learning in long-horizon, sandboxed environments. ・Unlike conventional RL, agentic RL couples GPU-intensive rollout engines with stateful environment containers whose actions may produce visible side effects, such as file edits, command execution, and dependency installation. ・A single traject
cs.LG updates on arXiv.org

Benchmarking Quantum Machine Learning for Power-System Attack Detection: Evaluation Choices Decide the Outcome Before the Models Do

・arXiv:2608.15617v1 Announce Type: new Abstract: Machine-learning detectors for power-system cyberattacks are themselves attack surfaces, and quantum machine learning has been proposed for them. ・We benchmark fidelity-kernel SVMs and variational classifiers against six tuned classical models on public power-system attack data (Mississippi State/ORNL), across white-box, transfer, decision-based black-box, and poisoning
WIRED

Best Merino Wool Clothing (2026): Base Layers, Hoodies, Jackets

・Merino is one of the best fabrics you can wear. ・We explain the different blends, what “GSM” means, and how to care for your clothes.
cs.LG updates on arXiv.org

Beyond $L_2$: Generalizing Abductive Latent Explanations to Diverse Prototype-Based Architectures

・arXiv:2608.16773v1 Announce Type: new Abstract: Prototype-based neural networks are hailed as interpretable-by-design architectures. ・Recently, Abductive Latent Explanations (ALE) were introduced to provide formal, mathematically guaranteed explanations that leverage the intrinsic structure of these networks to ensure both predictive safety and human readability. ・ALEs rely on computing tight bounds on latent space dis
cs.LG updates on arXiv.org

Beyond Boundary Noise: Aggregated Aleatoric Uncertainty Fails to Capture Presence Ambiguity in 3D Lung Nodule Segmentation

・arXiv:2608.14766v1 Announce Type: cross Abstract: Uncertainty estimation is critical for the safe clinical deployment of deep learning in medical image segmentation, with aleatoric uncertainty theoretically designed to capture irreducible data ambiguity. ・However, whether entropy-based measures reflect clinically meaningful ambiguity, i.e. ・case-level disagreement about whether a pathology is present at all, remains po
cs.LG updates on arXiv.org

Beyond Field Accuracy: Two-Axis Diagnosis of Inverse-PINN Parameter Error

・arXiv:2608.15373v1 Announce Type: new Abstract: Inverse physics-informed neural networks (PINNs) can reconstruct a field accurately while returning an incorrect physical parameter. ・We introduce a two-axis post-training diagnosis that separates finite-sample resolution under a specified observation-and-estimation protocol from the signed parameter preference encoded by the final learned field and residual metric.
cs.LG updates on arXiv.org

Beyond Peak Backlog: Conditional Energy and Temporal Geometry in Capacity-Constrained Delayed Bandit Optimization

・arXiv:2608.16216v1 Announce Type: new Abstract: What is the right delay complexity when a learner can track only $C$ pending feedback items and discarded feedback is permanently lost? ・Existing one-point bandit convex optimization guarantees in this model pay $\sqrt{T\sigma_{\max}}$, where $\sigma_{\max}$ is the peak backlog, although unlimited tracking admits the sharper $\sqrt{d_{\mathrm{tot}}}$ dependence on total
cs.LG updates on arXiv.org

Breaking the Compression Barrier: Cross-Architecture Compression Boundary Learning via Reverse Regrowth

・arXiv:2608.16010v1 Announce Type: new Abstract: Model compression is critical for deploying networks on resource-constrained edge devices. ・While pruning-based methods can significantly reduce model size, they often suffer from abrupt performance collapse beyond a sparsity thresh-old, making it difficult to identify the feasible compression limit of the model. ・To address this challenge, we propose a boundary-Learning
cs.LG updates on arXiv.org

CaliBench: Are the Stochastic Dynamics of Video World Models Physically Calibrated?

・arXiv:2608.16829v1 Announce Type: new Abstract: Video world models approximate the stochastic distribution of physical outcomes through generative sampling, but existing benchmarks score individual generations or compare distributions coarsely over a whole dataset, leaving the fine-grained aleatoric uncertainty of specific phenomena untested. ・We introduce CaliBench, which scores outcomes in a physically interpretable
cs.LG updates on arXiv.org

Calibrated Trust, Not Sharper Prediction: An Empirical Test of Uncertainty Fusion

・arXiv:2608.14617v1 Announce Type: new Abstract: A recurring proposal in legal AI is to improve case-outcome prediction by fusing uncertainty tools (evidence graphs with belief propagation, sequential Bayesian odds updating, Dempster-Shafer combination, and conformal prediction) into one pipeline. ・We test this on 1,000 real European Court of Human Rights cases from LexGLUE and FairLex, predicting whether the Court fou
WIRED

Can AI Coexist With Privacy? Proton’s Andy Yen Says It Will Have To

・Proton’s CEO is a champion of encryption for everyone. ・So why is he going all in on un-encryptable AI?
cs.LG updates on arXiv.org

Can Neural Networks Learn by Experimenting on Themselves? Self-Interventional Learning from Functional Consequences to Predictive Self-Knowledge

・arXiv:2608.14894v1 Announce Type: new Abstract: Machine-learning systems usually model external data, while their internal functional organization is analyzed by external observers. ・This work introduces Self-Interventional Learning (SIL), in which a neural system perturbs its own functional structure, observes consequences, learns a predictive self-model, generalizes to unexecuted interventions, and uses predictions
WIRED

Can the Upcoming ‘Expanse’ Game Avoid the Biggest Mistake of ‘Mass Effect’?

・The universe may never tell you if your choices mattered. ・Owlcat’s Osiris Reborn might not either.
MarkTechPost

Cartesia Ships Sonic-3.6: A Streaming TTS Model That Now Leads Both Artificial Analysis Speech Arenas

・Cartesia has released Sonic-3.6, a streaming text-to-speech model built on state space models rather than transformers. ・It now ranks #1 on both Artificial Analysis speech leaderboards — 1,283 Elo on Provider Voice and 1,123 on Controlled Voice, the board that clones every model onto the same eight reference voices to isolate the synthesis engine. ・Cartesia states sub-90ms time-to-first-audio.
cs.LG updates on arXiv.org

Certifying Compressed Language Models: An Audit and a Statistical Toolkit

・arXiv:2608.15046v1 Announce Type: new Abstract: A fraction of a point of benchmark accuracy is the usual evidence that a compressed model is equivalent to its original. ・That quantity is least informative when two models are most alike: a net delta is what survives cancellation between opposing per-item changes, and cancellation is most complete in the regime equivalence claims occupy. ・Across an atlas of 1,707 paired
cs.LG updates on arXiv.org

CFR without Unbiasedness: Deterministic Guarantees for Persistent Public-Chance Schedules

・arXiv:2608.14761v1 Announce Type: cross Abstract: At a finite public-chance cut, counterfactual regret minimization (CFR) must choose how many outcomes to evaluate before each regret update. ・Exact evaluation processes the full cut at one strategy profile; persistent partial evaluation processes a fixed without-replacement order across evolving profiles. ・The latter covers every outcome once per epoch, yet its feedback
cs.LG updates on arXiv.org

Characterization of Thermal Systems from Noisy and Low-resolution Measurements Using Dynamic Mode Decomposition

・arXiv:2608.14581v1 Announce Type: cross Abstract: Thermal monitoring in practical applications is often constrained by sparse sensing, measurement noise, and limited spatial resolution, which hinder the identification of heat transfer dynamics. ・In such settings, calibrating high-fidelity physical models is computationally demanding, motivating data-driven approaches. ・Dynamic Mode Decomposition (DMD) provides a framew
LLMタグが付けられた新着記事 - Qiita

ChatGPT を「前回の続き」から始めるための状態設計 — Instructions / Memory / Knowledge / Causal Posterior / Recent State

・ChatGPT を「前回の続き」から始めるための状態設計 — Instructions / Memory / Knowledge / Causal Posterior / Recent State 依存関係の痛み ChatGPT と三時間話す。 ・最初は問題を取り違えてい...
#LLMタグ

ChatGPTさん、Claudeさん、Geminiさん、3人を同じ会議室に入れてみた!?

・一番危ないのは、賢いAIじゃない? ChatGPT・Claude・Geminiを使って感じた「性格」の違い 続きをみる
Hugging Face Papers

ClawGym II: Exploring Black-Box RL on Agent Harness

ClawGym II: Exploring Black-Box RL on Agent Harness
Zennの「大規模言語モデル」のフィード

CMP170HXを検証する

・CMP170HXとは 仕様とこのGPUが生まれた背景 CMP170HXは、NVIDIAが暗号資産マイニング向けに用意したCMP(Cryptocurrency Mining Processor)シリーズの最上位クラスです。NVIDIA公式のCMP説明でも、複数GPUを一つのCPUで管理しやすい開放型ブラケット、マイニング効率、映像出力を目的としない製品コンセプトが示されています。 ・項目 CMP170HX 8GBの公表・確認値 A100 80GB PCIe GPU GA100系 GA100 メモリ 8GB HBM2e(出荷時) 80GB HBM2e メモリバス ...
cs.LG updates on arXiv.org

Coarse-to-Fine Multi-Resolution Diffusion Models for Trajectory Generation in Urban Systems

・arXiv:2608.14570v1 Announce Type: new Abstract: Understanding human mobility is critical for a wide range of urban applications, including traffic management, epidemic control, and urban planning. ・However, due to privacy concerns, the availability of large-scale public trajectory data remains limited, posing challenges for downstream mobility analysis. ・Existing methods for synthetic trajectory generation primarily fo
Zennの「大規模言語モデル」のフィード

Codexを「Solが監督、Luna Maxがworker」の構成にする

・Codexの親エージェントに判断と統合を任せ、調査や実装は軽量なサブエージェントへ委譲したい。 ・この記事では、親エージェントに gpt-5.6-sol、通常のサブエージェントに gpt-5.6-luna を使う構成を紹介します。Lunaの推論強度は max に設定します。つまり、この記事でいう「Luna Max」はモデル名ではなく、gpt-5.6-luna と推論強度 max の組み合わせです。 ・役割分担は次のとおりです。
Zennの「大規模言語モデル」のフィード

Colab MCPが便利

・MCPとしてColabが利用できる Agentの実行環境としてGPUが使えるようになって便利 https://github.com/googlecolab/colab-mcp セルの追加、実行、更新、削除などが可能 Colab経由でGoogle Driveにアクセスできるので、Colabの実験結果をDriveに出力、Colab経由で内容を見たり、もしくは別のAgentからDriveの内容を参照してレポートを作成したりができて便利だった。 ・基本的にReadmeにあるように設定するだけですが、地味にひっかるポイントがあったのでメモ。 ・"mcpServers": { "co...
The Verge

Comcast is turning millions of its routers into motion detectors

・Xfinity gateways XB7 and newer can now act as motion sensors in your home. ・| Image: Comcast Comcast is bringing Wi-Fi motion sensing to millions of routers that are already in customers' homes, turning the devices into activity monitors. ・A new update to the Xfinity Internet app, arriving today, August 18th, enables the feature on compatible Xfinity routers at no extra cost.
cs.LG updates on arXiv.org

Command-Space Counterfactual Explanations for Pareto-Conditioned Reinforcement Learning

・arXiv:2608.14963v1 Announce Type: new Abstract: Pareto Conditioned Networks learn multiple multi-objective reinforcement learning behaviours by conditioning a single policy on a desired return command. ・However, the local mapping from command and state to action remains opaque. ・We propose command-space counterfactual explanations for PCNs: given a fixed state, original command, and foil action, we search, in a black-b
Hugging Face Papers

ConceptFormer: Learning Adaptive Latent Concepts for Query-Document Alignment in Visual Document Retrieval

ConceptFormer: Learning Adaptive Latent Concepts for Query-Document Alignment in Visual Document Retrieval
cs.LG updates on arXiv.org

Conditional Evaluation of Language Models with Cheap Auxiliary Signals

・arXiv:2608.16210v1 Announce Type: new Abstract: Aggregate accuracy hides where models succeed and fail. ・Estimating conditional performance profiles from gold labels alone is expensive, while cheap auxiliary signals such as LLM-judge scores, pairwise comparisons, confidence scores, and judge-disagreement features can be collected for every benchmark item but are often biased or miscalibrated. ・We propose LACE (Local Au
cs.LG updates on arXiv.org

Coverage-Maximizing Multinomial Subset Routing under Operational Constraints

・arXiv:2608.16375v1 Announce Type: new Abstract: We introduce Multinomial Subset Routing (MSR), a new online routing framework over $K$ experts in which the learner keeps a multinomial routing policy instead of a deterministic subset of experts. ・At each round, the learner samples $M$ experts i.i.d. ・from the multinomial policy, and the resulting set of distinct sampled experts forms the routed subset.
The Verge

Coyote vs. Acme is even funnier because Warner Bros. Discovery tried to kill it

・There's an argument to be made that people wouldn't be all that interested in Coyote vs. ・Acme if it weren't for the way David Zaslav tried to kill it. ・By trying to shelve the project, Warner Bros.
Zennの「大規模言語モデル」のフィード

CPUのみでローカルLLMする

・Windowsかつ、CPUのみでローカルLLMしたい RAMが16GB以上ある 少し手間がかかっても性能を上げたい というニッチな需要を満たす記事です。 ・なお、この記事は2026年8月17日の情報を元に書かれています。llama.cppのオプション群は将来削除されうるので、適当にアレンジして下さい。 ・発見 affinity指定で速くなる 6コア12スレッドや12コア20スレッドなどの、SMTが有効になっているCPUでは、CPU affinityの指定により以下のメリットが得られるようです。
cs.LG updates on arXiv.org

CrevasseSeg: A Label-Efficient UAV Crevasse Segmentation Framework

・arXiv:2608.15790v1 Announce Type: new Abstract: Crevasse mapping from uncrewed aerial vehicle (UAV) imagery matters for glaciological research and for field safety in glaciated terrain. ・Yet, pixel-level annotation of glacier surfaces is costly and requires domain experts. ・We introduce CrevasseSeg, a framework for binary segmentation over the terminus of Borebreen, Svalbard, comprising 1,938 unlabelled UAV orthomosaic
cs.LG updates on arXiv.org

Cross-Entropy Risk Estimation for Language Models: Inconsistency Must Be Dense, and the Holdout Method Is No Exception

・arXiv:2608.15798v1 Announce Type: new Abstract: Language models are compared by their held-out per-token cross-entropy risk---the quantity scaling laws are fitted to. ・We show that it cannot be consistently estimated. ・Consistency, or convergence to the estimand, is defined relative to a \emph{possible state of the world}: a pair consisting of a data-generating distribution and a model we turn out to train.
cs.LG updates on arXiv.org

Cross-Modal Ultrasound-MRI Learning for Fetal Brain Ventricular Volumetry and Abnormality Screening

・arXiv:2608.14763v1 Announce Type: cross Abstract: Assessment of ventriculomegaly (VM) on fetal brain ultrasound relies primarily on measuring lateral ventricular atrial width on standard planes, which is operator-dependent and may not fully reflect the overall ventricular enlargement. ・Fetal brain MRI provides more reliable volumetric information but is costly and less accessible for routine use. ・To address these limi
cs.LG updates on arXiv.org

Data-Efficient and Interpretable Classification of Circulating Tumor Cell Phenotypes in Microfluidic Devices via Deep Learning

・arXiv:2608.16870v1 Announce Type: new Abstract: Accurate classification of circulating tumor cell (CTC) phenotypes can provide valuable information for assessing metastatic potential. ・Label free microfluidic devices provide a hydrodynamic obstacle course that transforms subtle biophysical characteristics of CTCs, including size and deformability, into distinct kinematic trajectories. ・However, the highly nonlinear flu
cs.LG updates on arXiv.org

Decision-Driven Regularization: A Blended Model for Learning and Optimization

・arXiv:2608.15124v1 Announce Type: new Abstract: In contextual optimization, the decision-maker seeks optimal decisions to minimize a cost function, that varies based on observed features. ・This context is common in many business applications ranging from on-demand delivery and retail operations to portfolio optimization and inventory management. ・In this paper, we study the learning and optimization approach, which fir
cs.LG updates on arXiv.org

Deep Analog: Open-Set Film Emulation with Reference-Conditioned 3D LUTs

・arXiv:2608.14702v1 Announce Type: cross Abstract: Film emulation reproduces the look of an analog film stock on a new digital photograph. ・We target its open-set form -- matching any reference film frame from a single example -- with a 3D lookup table (LUT) predicted from that reference. ・Real-time image enhancement predicts per-image weights over a fixed bank of 3D LUTs and blends them.
cs.LG updates on arXiv.org

DeepOHeat-v2: Self-Improving Operator Learning for Fast and Trustworthy Thermal Optimization in 3D-IC Design

・arXiv:2608.16080v1 Announce Type: new Abstract: Thermal-aware optimization of multi-die 3D integrated circuits evaluates many designs, each a costly heat-equation solve. ・Operator-learning surrogates replace this solve with a fast forward pass, ideally trained from physics alone, without labeled data. ・DeepOHeat-v1 made such surrogates fast and trustworthy, but only on low-contrast geometries.
Zennの「大規模言語モデル」のフィード

DeepSeek HarnessにTencent Cloud TokenHubを接続して、ブラウザゲームを作らせてみた

・本記事は2026年8月18日時点の検証結果です。DeepSeek HarnessはDeveloper Previewのため、今後UIや設定方法が変更される可能性があります。 ・はじめに DeepSeek Harness(dsh)は、DeepSeek AIが公開したオープンソースのAgent Harnessです。新しい基盤モデルではなく、モデル、ツール、スキル、セッション、サンドボックス、Agent Loop、UIなどをプラグインとして組み合わせるAgent実行基盤です。 ・今回は、Tencent Cloudの大規模言語モデルサービスTokenHubをモデルプロバイダーとして登録し、実...
Zennの「大規模言語モデル」のフィード

DeepSeek V4 Flash を DGX Spark x1 台で動かしてみたメモ

・DGX Sparkで、DeepSeek V4 Flash の GGUF 量子化モデルをローカル推論させてみたメモです。 ・このサイトを参考にインストール! https://note.com/gb10_tsurumitsu/n/n8d69a9b3cafb インストール 上記記事に従ってインストールスクリプトを叩くだけです。 ・ds4というC/CUDAのDeepSeek V4 Flash専用に最適化された推論エンジンです。これをDGX Spark向けに最適化したフォーク版のワンコマンドインストーラー(ビルド、gguf配置など)の ds4-on-sparkを使います。
cs.LG updates on arXiv.org

Degeneracy Counting Quantum Algorithm using Decoherence

・arXiv:2608.14941v1 Announce Type: new Abstract: Counting the global optima of a classical optimization problem is a #P-hard task. ・We develop the canonical thermal pure quantum (CTPQ) state-based degeneracy counting (CTPQsd#) algorithm that determines the number of global optima of a classical optimization problem P by measuring only a small probe S, without finding individual minima. ・The method exploits a perturbativ
cs.LG updates on arXiv.org

Demo: Real-time Generative Multicasting with On-Device Intent-aware Semantic Decomposition

・arXiv:2608.14600v1 Announce Type: cross Abstract: We present a demonstration for generative multicasting with on-device, intent-aware semantic decomposition. ・At the transmitter, DNN-based segmentation extracts a semantic map from the source video, decomposing it into multiple sub-signal classes based on multi-user receiver intents. ・The transmitter broadcasts the semantic map to all users over shared wireless/network
cs.LG updates on arXiv.org

Demystifying Oversmoothing in Sheaf Neural Networks: An Index-Theoretic Criterion

・arXiv:2608.16180v1 Announce Type: new Abstract: To combat oversmoothing in Graph Convolutional Networks, Sheaf Neural Networks (SNNs) were proposed as a generalization by equipping the graph with a sheaf structure and replacing the graph Laplacian with a sheaf Laplacian $\mathcal{L}$. ・Existing analyses connect sheaf diffusion to oversmoothing via the harmonic space ($\ker\mathcal{L}$), taking its absolute dimension a
cs.LG updates on arXiv.org

Deploying Frontier Agentic Technology in MOOSEnger, a Multiphysics-Capable AI Assistant

・arXiv:2608.15881v1 Announce Type: new Abstract: The Multiphysics Object-Oriented Simulation Environment (MOOSE) is an open-source finite-element framework for building multiphysics simulation applications. ・Using a multiphysics environment effectively demands specialized expertise, creating a barrier for many domain scientists and engineers. ・MOOSEnger, developed at Idaho National Laboratory (INL), is a domain-specific
cs.LG updates on arXiv.org

Detecting Money Laundering in Rwandan Mobile Money: A Machine Learning Framework

・arXiv:2608.15447v1 Announce Type: new Abstract: Mobile money has widened financial access across Sub-Saharan Africa and enlarged the surface for money-laundering and terrorism-financing (ML/TF) activity in ecosystems dominated by high-volume, low-value transactions. ・Rwanda is a case in point: several million active mobile-money users, telecom-led wallets on the MTN and Airtel networks, and a Financial Intelligence Ce
Zennの「大規模言語モデル」のフィード

Devinを1週間試して導入をやめた話

・Devinを1週間試して、導入をやめた話 自律型のコーディングエージェント Devin を 1 週間評価して、導入しないことに決めました。 ・動かなかったから諦めたわけではありません。3 つのリポジトリで実行環境が立ち上がり、GitHub と Jira と Slack をつなぎ、Devin が書いたバグ修正の PR はマージまで到達しました。準備にかかったのも数日です。 ・やめた理由は、 クラウド型のエージェントが持つ強みと、私たちのチームの条件が噛み合わなかった からでした。この記事では、まずクラウド型の強みがどこにあるのかを整理して、それを私たちの条件と 1 つずつ突き合わせます。結...
cs.LG updates on arXiv.org

Diagnosing and Mitigating Perception-Decision Misalignment in Omni-LLMs via Modality Subspace Activation

・arXiv:2608.14655v1 Announce Type: new Abstract: Omni-Large Language Models (Omni-LLMs) power complex multi-modal reasoning in applications like World Action Models and autonomous agents. ・However, their strong performance often masks a profound Perceptual-Decision Misalignment (PDM), where decisions remain unfaithful to multi-modal perceptions. ・To diagnose this, we formalize Causal Modality Sensitivity (CMS), operatio
cs.LG updates on arXiv.org

Disentangling Homophily and Rarity: Explaining Failure in Graph Neural Networks

・arXiv:2608.14823v1 Announce Type: new Abstract: Are heterophilic nodes in a graph harder to classify because they are heterophilic or because they are rare? ・Some existing work frames classification of such nodes as a subgroup generalisation problem, where a model performs well on the majority group at the expense of the rare group. ・Others explain this as a problem of neighbourhood aggregation in graph neural networks
cs.LG updates on arXiv.org

Do Geometry-Aware Positional Encodings Help Transformers in Spatial Imperfect-Information Games?

・arXiv:2608.14982v1 Announce Type: new Abstract: Transformers applied to spatial imperfect-information games must represent map geometry while tracking hidden entities through time. ・We ask whether geometry-aware positional encodings improve these capabilities, without claiming a new positional encoding. ・We construct a four-level benchmark on a hexagonal naval pursuit game: controlled geometry and topology probes, an e
cs.LG updates on arXiv.org

Do LLMs Know What to Ask and When? Evaluating Multi-Turn Information Seeking

・arXiv:2608.14808v1 Announce Type: cross Abstract: When a user question is underspecified, a capable model should recognize that its context is insufficient, identify the missing information, ask for it, and respond only once that information determines a unique answer. ・We formalize multi-turn information seeking as solving a k-underspecified constraint satisfaction problem, where k is the number of variables jointly
cs.LG updates on arXiv.org

Do Uncertainty Signals Help? A Systematic Study of Uncertainty-Aware Decoding with Rollback Mechanisms

・arXiv:2608.14653v1 Announce Type: new Abstract: Prediction uncertainty is a widely adopted metric for quantifying model confidence, with downstream applications spanning model explanation, data selection, and prediction rollback. ・Despite its demonstrated utility, the potential of uncertainty quantification to enhance code generation in large language models (LLMs) remains largely underexplored, raising a critical que
cs.LG updates on arXiv.org

Does 1/2-Tsallis-INF Also Work Well for Best-Arm Identification?

・arXiv:2608.15365v1 Announce Type: new Abstract: Regret minimization (RM) and best-arm identification (BAI) are two fundamental objectives in multi-armed bandits. ・Among regret-minimizing algorithms, $1/2$-Tsallis-INF is a canonical best-of-both-worlds FTRL algorithm: it achieves logarithmic pseudo-regret in stochastic bandits while retaining minimax-optimal regret in adversarial bandits, without knowing the environmen
cs.LG updates on arXiv.org

Does the Heart Show Your Pain? Tackling the X-ITE Pain Challenge with Self-Supervised ECG Representation Learning

・arXiv:2608.14662v1 Announce Type: cross Abstract: Accurate recognition of pain using physiological signals remains a challenging problem due to pain's subjective nature and high inter-individual variability. ・In this study, we investigate self-supervised representation learning (SSL) methods applied to unimodal electrocardiogram (ECG), complemented by multimodal pretraining, including accelerometer (ACC) signals from
Hugging Face Papers

Drive, Pack, Fly: The Travelling Thief Problem with Drone

Drive, Pack, Fly: The Travelling Thief Problem with Drone
cs.LG updates on arXiv.org

DUET: Dual-Teacher On-Policy Distillation via Same-Weight Disagreement for Prohibition Compliance

・arXiv:2608.14644v1 Announce Type: new Abstract: Real-world LLM deployments increasingly rely on runtime-injected prohibitions--enterprise policies, PII redlines, tool boundaries--that vary per request and per tenant. ・Conventional post-training is structurally ill-suited: SFT hides the violation signal in compliant labels, and DPO's sequence-level preferences mismatch token-localized violations. ・We propose DUET, a tok
cs.LG updates on arXiv.org

DumpsterCluster: From Dumpster Diving to Serving LLaMA-70B on $60 GPUs

・arXiv:2608.14614v1 Announce Type: new Abstract: As AI datacenters retire functional GPUs, vast quantities of still capable accelerators enter secondary markets. ・This paper investigates whether these retired GPUs can find a productive afterlife to form a DumpsterCluster that can serve modern LLM inference, and under what conditions such repurposing is economically viable and environmentally sustainable. ・We physically
cs.LG updates on arXiv.org

Early Cycle Charge Trajectory Generative Prediction and Full Life Cycle Health Management of Iron-Chromium Flow Batteries Based on FlowBD-E1

・arXiv:2608.14637v1 Announce Type: new Abstract: Long-duration stationary energy storage requires batteries whose degradation can be detected before substantial capacity loss has accumulated. ・Iron-chromium redox flow batteries are attractive for this role because they use abundant and low-cost active species, yet their operation is shaped by slow chromium kinetics, hydrogen evolution, membrane crossover and electrolyt
cs.LG updates on arXiv.org

Earth Observation Foundation Models for Terrestrial Ecohydrology: From Representation Learning to Process Inference

・arXiv:2608.15282v1 Announce Type: new Abstract: Earth observation foundation models (EOFMs) are emerging as reusable representation frameworks for data-driven retrieval, prediction and process modelling within ecohydrology, which integrate EO, meteorological forcing and process models to characterise coupled water, energy and carbon dynamics in vegetation and soil across scales. ・However, there is yet to be an ecohydr
cs.LG updates on arXiv.org

Efficient Coreset Selection via K-Nearest Neighbor Graphs

・arXiv:2608.16270v1 Announce Type: new Abstract: Coreset selection reduces the cost of model training by replacing a large training set with a small representative subset. ・Existing gradient-approximation coreset methods such as CRAIG and cluster-based variants can preserve model accuracy. ・Still, their selection stages often rely on dense pairwise distances or large item-cluster bound matrices, leading to high time and
cs.LG updates on arXiv.org

Efficient Neural-Network-Based High-Resolution Radiative Transfer for CO___ Retrieval, and Application to Interferometric Sensing

・arXiv:2608.14645v1 Announce Type: new Abstract: Studying climate change requires reducing uncertainties in CO2 and CH4 emission estimates to better distinguish anthropogenic from natural sources, which motivates spaceborne measurements with improved revisit frequency and spatial coverage. ・In this context, the Horizon Europe SCARBOn project assesses a low-cost satellite constellation featuring the NanoCarb imaging int
cs.LG updates on arXiv.org

EMASAM: a Computationally Efficient Sharpness-Aware Minimization via EMA-Guided Perturbations

・arXiv:2608.15105v1 Announce Type: new Abstract: Recent progress in optimization research has highlighted the sharpness of the loss landscape as a key factor in narrowing the generalization gap. ・Motivated by this insight, Sharpness-Aware Minimization (SAM) was proposed as a training strategy that enhances generalization. ・Despite the promising performance, SAM suffers from its twice computational cost due to its core a
Hugging Face Papers

ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering

ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering
cs.LG updates on arXiv.org

Equilibrium Forcing: Adaptive Video Generation Without Noise Conditioning

・arXiv:2608.14706v1 Announce Type: cross Abstract: Standard autoregressive video generation algorithms based on Diffusion and Flow Matching rely on rigid training objectives and static sampling schedules, limiting inference procedures from adapting to the data. ・We introduce Equilibrium Forcing (EqF), a simplified framework for video denoising generative models without noise level conditioning. ・EqF pioneers modular tra
cs.LG updates on arXiv.org

ER-KANs: Efficient and Robust Kolmogorov-Arnold Networks for Data-Scarce Scientific Machine Learning

・arXiv:2608.14773v1 Announce Type: new Abstract: The efficient-KAN literature---covering Chebyshev, wavelet, and radial-basis-function variants of the original Kolmogorov-Arnold Network---has been benchmarked almost entirely on clean data. ・We show that this choice conceals a large capability difference between architectures: ChebyKAN's test MSE (evaluated against clean ground truth) increases by a factor of 10.6x when
AI News & Artificial Intelligence | TechCrunch

Etched’s valuation doubles to $21B in a month

・Jane Street has installed Etched's first shipped AI cluster system, and was so impressed, it led another massive round, the startup says.
cs.LG updates on arXiv.org

Evaluating the impact of adversarial traffic patterns on vanet communication using veins simulation

・arXiv:2608.14583v1 Announce Type: cross Abstract: Vehicular Ad Hoc Networks (VANETs) are a key component of intelligent transportation systems, enabling real-time communication between vehicles. ・However, their open and dynamic nature makes them highly vulnerable to adversarial behaviors that can disrupt communication reliability. ・This paper investigates the impact of adversarial traffic patterns on VANET performance
cs.LG updates on arXiv.org

Every Expert Counts: ExactMoE for Memory-Efficient W4A16 Inference

・arXiv:2608.15383v1 Announce Type: new Abstract: Sparse mixture-of-experts (MoE) language models reduce arithmetic by activating only a small subset of experts per token, yet deployment still requires storing and moving the full expert bank. ・We present ExactMoE, an inference design that applies symmetric group-128 four-bit weight quantization only to routed experts, stores those experts in kernel-native MARLIN form in
cs.LG updates on arXiv.org

Evolving Executable Pipeline Programs for AutoML with Language Models

・arXiv:2608.16416v1 Announce Type: new Abstract: Automated machine learning (AutoML) systems search for pipelines within a space of preprocessing operators, learners, and hyper-parameters specified in advance: they can select and tune known components, but cannot produce structure outside that space. ・We present LACE, an AutoML framework that instead searches over complete executable pipeline programs: an evolutionary
WIRED

Exclusive: You Can Finally Buy a Fairphone—a Sustainable, Repairable Smartphone—in the US

・More than a decade after launching in Europe, the Netherlands company is now selling its repairable phones in the US, starting with the Fairphone (Gen 6+).
cs.LG updates on arXiv.org

Explaining Reinforcement Learning Decisions in Self-adaptive Systems

・arXiv:2608.14620v1 Announce Type: new Abstract: Reinforcement Learning (RL) has been extensively used in autonomous and self-* systems, but RL policies, especially deep RL ones relying on neural networks, lack transparency and are difficult to understand. ・This can lead to diminished user trust, and makes for a more challenging verification of systems. ・To address this challenge, this paper introduces Explanations usin
cs.LG updates on arXiv.org

FAST-DeepONet: Factor-Augmented Branch Representations for High-Dimensional PDE Inputs in the Small-Sample Regime

・arXiv:2608.15408v1 Announce Type: new Abstract: Deep operator networks can become statistically unstable when partial differential equation inputs are observed at thousands of strongly correlated sensors but only a small number of operator samples is available. ・We introduce FAST-DeepONet, a branch representation combining a fixed spectral path with a regularized projection of the orthogonal residual, in which the dir
cs.LG updates on arXiv.org

Feasible and Novel Synthetic Population Generation with Tabular and Sequential Travel Attributes

・arXiv:2608.15867v1 Announce Type: new Abstract: Synthetic populations are critical inputs for activity-based travel demand models, yet generating realistic populations from limited survey data remains challenging. ・Small samples miss valid attribute combinations, known as sampling zeros, and generative models may also produce infeasible structural zeros. ・Moreover, realistic synthetic populations must capture both stat
cs.LG updates on arXiv.org

FedImp: Enhancing Federated Learning Convergence with Impurity-Based Weighting

・arXiv:2608.14654v1 Announce Type: new Abstract: Federated Learning (FL) is a collaborative paradigm that enables multiple devices to train a global model while preserving local data privacy. ・A major challenge in FL is the non-Independent and Identically Distributed (non-IID) nature of data across devices, which hinders training efficiency and slows convergence. ・To tackle this, we propose Federated Impurity Weighting
cs.LG updates on arXiv.org

FETERS: Few-Shot Early Time-Series Classification via Effective Ratio Selection

・arXiv:2608.16385v1 Announce Type: new Abstract: Early time-series classification (ETSC) aims to make accurate predictions from partially observed time series as early as possible. ・Although various stopping mechanisms and feature learning strategies have been developed for ETSC, most existing methods assume access to sufficient labeled training data, which may be unrealistic in applications with limited annotation.
cs.LG updates on arXiv.org

Fiber Fingerprints of Hidden Learning-State Dynamics

・arXiv:2608.15976v1 Announce Type: new Abstract: A learning system can occupy execution states that are indistinguishable under every declared present-behavior readout yet respond differently to future training. ・We formalize this through fiber fingerprints: controlled future-learning response laws restricted to present-behavior equivalence classes. ・Prefix-compatible finite probes induce a predictive quotient functor,
cs.LG updates on arXiv.org

FinFraudBench: A Heterogeneous Graph Benchmark for Financial Fraud Detection

・arXiv:2608.15177v1 Announce Type: new Abstract: The increasing complexity of digital financial systems has reshaped financial fraud detection from isolated transaction classification into relational risk reasoning over interconnected financial entities. ・This shift has motivated graph-based fraud detection, where models identify fraudulent nodes by exploiting dependencies among customers, cards, merchants, categories,
The Verge

Firefox’s Smart Window promises a better AI browser

・Starting today, AI chats in Firefox's Smart Window AI browsing mode can pull from current web info and show source links in chat responses through a partnership with Exa. ・Smart Window can also now automatically suggest tab groups and show visual previews of pages you previously visited when you search your browsing history using natural language. ・I saw a live demo that showed how the Smart Window AI could sort throug
cs.LG updates on arXiv.org

FirstDiff: One-Step Diffusion-Based Anomaly Detection for Multivariate Time Series via Initial Noise Prediction

・arXiv:2608.15727v1 Announce Type: new Abstract: Diffusion models have recently shown strong potential for multivariate time-series anomaly detection by learning the distribution of normal data through iterative denoising. ・Existing diffusion-based approaches, however, typically perform anomaly detection after completing the reverse diffusion process, relying primarily on the final reconstructed signal and overlooking
cs.LG updates on arXiv.org

FluxBin: Flexible LUT-based Ultra-low-bit LLM Inference by Algorithm-Kernel Synergy

・arXiv:2608.15602v1 Announce Type: new Abstract: While binary quantization theoretically promises extreme compression and acceleration for Large Language Models (LLMs), existing research often overlooks the necessity of specialized hardware kernels, thus failing to unleash the full acceleration potential due to persistent reliance on expensive floating-point arithmetic or runtime dequantization overheads. ・To bridge th
cs.LG updates on arXiv.org

Foresight-England: Development of a National-Scale Generative AI Model of Electronic Health Records for Medical Event Prediction across the COVID-19 Pandemic

・arXiv:2608.16273v1 Announce Type: new Abstract: Foresight-England (Foresight-E) is the first national-scale generative foundation model of electronic health records (EHRs), developed as a research pilot strictly for COVID-19 research. ・We evaluated its ability to model the direct and indirect effects of the pandemic. ・Trained from scratch entirely within the NHS England Secure Data Environment, Foresight-E is a 243-mil
cs.LG updates on arXiv.org

Forward Pass Domain Adaptation (Without Cross-Layer Backpropagation)

・arXiv:2608.14563v1 Announce Type: new Abstract: Forward-Pass-Only MLP training (FPO) adapts large language models without a backward pass through the model body, achieving 2.7--3.2x the throughput of standard fine-tuning at ~40% less peak training memory, while leaving off-domain benchmarks within seed-noise of baseline, a property that full-network fine-tuning does not reliably reproduce. ・FPO rests on a single empir
cs.LG updates on arXiv.org

Fractional Optimizers Meet Fractal Activation Functions: An Empirical Study of Multi-Scale Optimization in Neural Network

・arXiv:2608.14636v1 Announce Type: new Abstract: Fractional optimization methods and fractal activation functions are two independent directions for improving neural network training. ・Fractional optimizers extend first-order optimization through fractional derivatives and memory effects, whereas fractal activations introduce multi-scale nonlinear representations based on self-similar Weierstrass- and Blancmange-type f
cs.LG updates on arXiv.org

From Errors to Proofs: Minimal-Core-Guided Repair for Neuro-Symbolic Constraint Solving

・arXiv:2608.14771v1 Announce Type: cross Abstract: Making language models solve constraint problems reliably often means having them translate the problem into a formal specification and delegating the search to a sound solver. ・But the translation is itself a language-model task, and an unfaithful translation makes the solver faithfully solve the wrong problem. ・Existing pipelines repair only translations that crash, r
Hugging Face Papers

Gathered, Not Admitted: How Attention Brings a Latent Variable into Verbalizable Form

Gathered, Not Admitted: How Attention Brings a Latent Variable into Verbalizable Form
cs.LG updates on arXiv.org

GATTA: Graph Active Learning with Test-Time Augmentation

・arXiv:2608.15084v1 Announce Type: new Abstract: Test-time augmentation (TTA) has proven effective for improving model robustness and uncertainty estimation in computer vision, yet its application to graph-structured data remains largely unexplored. ・We introduce GATTA (Graph Active Learning with Test-Time Augmentation), a framework for enhancing active learning by aggregating predictions across multiple augmented view
cs.LG updates on arXiv.org

Generalised Transportability via Causal Abstractions

・arXiv:2608.15645v1 Announce Type: new Abstract: Transporting a causal conclusion from a source study population to a target one is a fundamental problem in causal inference. ・The theory of transportability provides a criterion for when this is possible: given experimental data from the source and observational data from the target, it determines whether a target query is identifiable and does so completely; i.e.
cs.LG updates on arXiv.org

Generative Learning of Separatrices

・arXiv:2608.14743v1 Announce Type: new Abstract: The identification and reconstruction of the boundaries separating basins of attraction in multistable, multidimensional dynamical systems presents a fundamental challenge in computational dynamics. ・These structures govern transition pathways and other important large timescale behavior, yet they remain typically under-sampled since their neighborhood does not get routi
Hugging Face Papers

GenRouter: Unified Workflow Routing for Agentic Image Generation

GenRouter: Unified Workflow Routing for Agentic Image Generation
cs.LG updates on arXiv.org

GEO-Flag: Detecting and Measuring GEO-Optimized Web Content

・arXiv:2608.16824v1 Announce Type: new Abstract: Generative Engine Optimization (GEO) modifies web content to increase its likelihood of being selected and cited by generative search engines. ・This can give strategically optimized pages visibility disproportionate to their authority or relevance and even make weak or false information appear well supported. ・Unlike conventional search, generative search synthesizes info
cs.LG updates on arXiv.org

Geometry Is Not Robustness: A Trajectory-Level Study of PGD Evaluation

・arXiv:2608.14594v1 Announce Type: new Abstract: Projected Gradient Descent (PGD) is widely used to evaluate adversarial robustness, typically via final adversarial accuracy, which does not capture model behaviour throughout the attack. ・Recent work proposes trajectory-level diagnostics, such as loss evolution, gradient alignment, and steps-to-failure, for deeper insight into adversarial optimisation dynamics.
cs.LG updates on arXiv.org

Geometry of Forgetting: Representation Flux in Continual Learning

・arXiv:2608.15854v1 Announce Type: new Abstract: Catastrophic forgetting remains a fundamental obstacle to continual learning, where neural networks lose previously acquired knowledge while learning new tasks. ・Existing methods primarily mitigate forgetting through parameter regularization or experience replay, while the representation-space dynamics associated with forgetting remain less understood. ・We investigate lat
cs.LG updates on arXiv.org

Global Federated Learning Strategies for Building Efficient Personalized Models

・arXiv:2608.15107v1 Announce Type: new Abstract: Federated learning (FL) is a practical framework that can train models on distributed user data while guaranteeing data privacy; however, due to heterogeneity in which each user has a different data distribution, problems frequently arise where both global and personalization performance deteriorate simultaneously. ・This dissertation presents methodologies for building e
cs.LG updates on arXiv.org

GraniKV: Asymmetric Granularity KV-Cache Paging for Multi-Agent Systems with Long Shared Prefix

・arXiv:2608.15584v1 Announce Type: new Abstract: Production paged-serving engines apply uniform paging granularity to the KV cache, even though the two regions of a multi-agent workload have opposite storage requirements: a long shared prefix demands contiguity, while the per-request suffix demands fine-grained allocation. ・We present \textbf{GraniKV}, a KV-cache layer that allocates the shared prefix in a contiguous H
cs.LG updates on arXiv.org

Graph Machine Learning: An Opportunity for Power Systems

・arXiv:2608.16494v1 Announce Type: new Abstract: Modern power systems face growing operational complexity driven by the integration of renewable energy sources, decentralization, and the need for real-time decision-making across a wide range of timescales. ・Addressing these challenges traditionally relies on model-based methods that, while accurate, can be too slow for operational demands. ・Machine learning (ML) has the
Hugging Face Papers

GRNEdit: Efficient General Video Editing from a New Binary-Evidence Perspective in Generative Refinement Networks

GRNEdit: Efficient General Video Editing from a New Binary-Evidence Perspective in Generative Refinement Networks
cs.LG updates on arXiv.org

Group ICA 2.0: Closing the Gap Between Subjects and Group Latent Decomposition with Copula-Linked Group ICA (CoLiG-ICA)

・arXiv:2608.16029v1 Announce Type: new Abstract: Group Independent Component Analysis (gICA) is widely used to decompose high-dimensional functional MRI data into interpretable brain networks. ・However, conventional gICA primarily identifies components shared across subjects. ・This group-level assumption can limit the recovery of networks present only in individuals or subject subsets, reducing sensitivity to intersubje
cs.LG updates on arXiv.org

Guaranteed Adaptive Modality Acquisition: When the Policy Chooses Its Own Calibration Group

・arXiv:2608.15520v1 Announce Type: new Abstract: A multimodal system may begin inference holding only some of its inputs and may acquire the rest at a cost. ・With adaptive acquisition, the policy determines which inputs are ultimately observed, so we state the guarantee conditional on that terminal input pattern. ・Conditional calibration normally assumes the grouping map is fixed independently of the calibration sample,
cs.LG updates on arXiv.org

Hardware-in-the-Loop Phase-Aware CNN for Real-Time 5G Channel Estimation

・arXiv:2608.14709v1 Announce Type: cross Abstract: This demo presents real-time AI-based uplink channel-estimation inference using data collected from a hardware-in-the-loop 5G platform. ・The data-collection setup integrates commercial RF signal generation, programmable channel emulation, an O-RAN Radio Unit, DU emulation, and a lightweight phase-aware convolutional neural network (CNN) that estimates the channel respo
Hugging Face Papers

HarnessEval-W: Agentifying the Evaluation of Visual Worlds

HarnessEval-W: Agentifying the Evaluation of Visual Worlds
#LLMタグ

Hermes Agentの使い方をローカルLLMで検証|Ollama+Qwen3.5 9B

・「Hermes Agentを使ってみたい。でも、クラウドのAIではなく、自分のPCにあるAIでも動かせるのだろうか?」 ローカルLLMに興味を持つと、次に気になるのがこの部分です。
cs.LG updates on arXiv.org

High-Dimensional Nonparametric Change-Point Detection via Low-Rank Degree-Three Density Projection

・arXiv:2608.15466v1 Announce Type: new Abstract: Distributional changes can be invisible to means and covariances yet appear in skewness, asymmetric interactions, or other third-order structure. ・We develop a nonparametric change-point method that retains every degree-at-most-three coefficient of a density while avoiding direct density estimation. ・For observations in $[-1,1]^d$, we construct a symmetric order-three Leg
cs.LG updates on arXiv.org

Hoeffding adaptive splitting trees for data stream classification with concept drift and ensemble learning

・arXiv:2608.16659v1 Announce Type: new Abstract: Ensembles of decision trees are well-established methods for data stream classification. ・In ensemble learning, Hoeffding Trees are widely adopted as base learners, performing periodic split attempts according to the Hoeffding bound. ・Recent studies, however, indicate that this standard splitting mechanism lacks adaptability, while adaptive trees that trigger splits in re
Hugging Face Papers

How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks

How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks
WIRED

Hydrow Discount Code: Save Up to $150 | August 2026

・Save on rowers and accessories with Hydrow coupons, including an exclusive discount of $50.
WIRED

I Put the Best Digital Notebooks to the Test. Here Are My Favorites (2026)

・These nifty tools combine the ease of jotting notes by hand with the power of saving them digitally.
cs.LG updates on arXiv.org

iFuzz-Meta: An Interpretable Fuzzy Learning Framework Bridging Top-Down and Bottom-Up Knowledge Integration

・arXiv:2608.14646v1 Announce Type: new Abstract: Interpretable representation learning remains a key challenge in modern neural computation, particularly when models are expected not only to perform but also to explain their reasoning. ・This paper introduces iFuzz-Meta, an interpretable fuzzy rule-based learning framework that preserves human-understandable reasoning structures within modern neural architectures.
Hugging Face Papers

Improving the matrix multiplication exponent with modern optimization and AlphaEvolve

Improving the matrix multiplication exponent with modern optimization and AlphaEvolve
cs.LG updates on arXiv.org

In Defense of OCTA: The Reconstruction-Utility Gap in OCT-to-OCTA Synthesis

・arXiv:2608.15626v1 Announce Type: new Abstract: Optical coherence tomography angiography (OCTA) images retinal blood flow, giving capillary-perfusion and foveal-avascular-zone biomarkers that grade diabetic-retinopathy ischemia. ・Because OCTA hardware is less common than structural OCT, recent work synthesizes it from OCT, reporting strong reconstruction (3D PSNR > 31 dB, SSIM > 0.9). ・We ask not whether the synthetic
cs.LG updates on arXiv.org

In-Context Learning to Assess Built Environment Impacts on Perceived Neighborhood Walkability Among Mobility-impaired Older Adults

・arXiv:2608.14663v1 Announce Type: new Abstract: As global populations age, enhancing neighborhood walkability through inclusive urban design is important for mitigating built environment (BE) barriers that discourage physical activity and social participation among older adults. ・This study investigates the utility of in-context learning (ICL), using the transformer-based foundation model TabPFN, to determine how BE f
cs.LG updates on arXiv.org

Information Geometry of Message Passing

・arXiv:2608.15922v1 Announce Type: new Abstract: We show that the natural-gradient stationary condition of variational inference has an edge-local form on a Forney-style factor graph. ・We start from the Bethe free energy and constrain a selected edge marginal to an exponential family. ・At a stationary point, the natural parameter of that edge equals the sum of two projected messages, one from each incident factor.
OpenAI News

Introducing ChatGPT for Teens: Built for learning, backed by protections

・ChatGPT for Teens helps teens learn, think critically, and use AI with confidence, with stronger built-in protections, healthy-use features, and additional controls for parents.
cs.LG updates on arXiv.org

Invariant Pretraining for Robust Code Representations

・arXiv:2608.15412v1 Announce Type: new Abstract: Encoder-based code representation models remain widely deployed for discriminative tasks such as clone detection and code classification, where their small size and low inference cost are decisive. ・Their robustness, however, is fragile: under invariant programs, semantically equivalent code written in different syntactic forms, learned representations degrade substantia
cs.LG updates on arXiv.org

IP Protection in the Era of Visual Generative AI: A Survey

・arXiv:2608.14730v1 Announce Type: cross Abstract: The rapid evolution of visual generative AI has introduced a wide range of intellectual property risks, spanning the unauthorized learning, reproduction, extraction, misuse, and redistribution of protected data and model assets. ・To address these risks, a growing body of technical defenses has been proposed. ・However, existing surveys typically organize this literature
cs.LG updates on arXiv.org

Is Grokking a Loss of Normal Hyperbolicity of the Interpolation Manifold?

・arXiv:2608.14803v1 Announce Type: new Abstract: A recent line of work recasts the post-memorization phase of grokking as constrained optimization: once a network interpolates the training set, weight decay drives a slow drift along the zero-loss manifold toward lower norm. ・In the language of dynamical systems, this is a fast-slow system in which the interpolation manifold plays the role of a slow manifold.
cs.LG updates on arXiv.org

Iterative Refinement Diffusion for Super-Resolved Data Assimilation of Multiscale Physical Systems

・arXiv:2608.14744v1 Announce Type: new Abstract: Recovering high-resolution states from sparse, low-resolution observations is a central challenge in scientific machine learning and data assimilation. ・Classical data assimilation exploits temporal information through forecast-analysis cycles, but often requires repeated access to expensive high-resolution forecast models. ・Generative super-resolution can recover unresol
#AIタグ

JTCの無駄業務を爆速化。毎日の残業を1/10に減らす「コピペ用AIプロンプト」3つの型

・今日も続く、結論の出ない長い会議。 ・自席に戻れば、未読メールの山。他部署への気を遣う「お伺いメール」の文面を考えているだけで、あっという間に1時間が溶けていく。 ・はじめまして。都内のJTC(伝統的な日本の大企業)で働く30代サラリーマン、Zenarc(ゼンアク)です。
cs.LG updates on arXiv.org

KOALA: Koopman Operator Learning for WiFi-Based Anticipatory Hum

・arXiv:2608.15815v1 Announce Type: new Abstract: WiFi Channel State Information (CSI) has emerged as a privacy-preserving alternative to cameras for human pose estimation. ・However, existing approaches treat pose inference as an instantaneous regression problem and do not model temporal dynamics, making future motion prediction infeasible. ・Naively applying vision-based prediction methods compounds the estimation noise
cs.LG updates on arXiv.org

Koopman early warning signals for bifurcation and rate-induced tipping

・arXiv:2608.14716v1 Announce Type: cross Abstract: Abrupt transitions in complex systems are often preceded by early warning signals. ・However, most indicators rely on the notion of critical slowing down and do not generally extend to rate-induced tipping where transitions can occur without local loss of stability. ・This is problematic in stochastic, nonautonomous systems where internal variability and time-varying vari
Hugging Face Papers

Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search

Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search
cs.LG updates on arXiv.org

Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search

・arXiv:2608.15669v1 Announce Type: new Abstract: Scientific discovery often involves optimising expensive-to-evaluate objectives over vast, structured, and open-ended hypothesis spaces, such as molecules, protein sequences, and computer programs. ・Generative models such as large language models (LLMs) provide expressive priors over such spaces, but their likelihoods and self-assessments are unreliable proxies for the o
cs.LG updates on arXiv.org

Layers Matter: Why Continual Learning Regularization Should Be Layer-Adaptive

・arXiv:2608.15901v1 Announce Type: new Abstract: Continual learning regularizers like EWC fight forgetting by penalizing changes from previous-task parameters with per-parameter importance, typically diagonal Fisher values. ・Per-parameter looks more flexible than per-layer, but each layer's diagonal Fisher is a weak summary of its actual curvature, missing the top-eigenvalue information that controls forgetting.
cs.LG updates on arXiv.org

Le Critique: Privileged Value Functions for LLM Reinforcement Learning

・arXiv:2608.16739v1 Announce Type: new Abstract: Reinforcement learning algorithms for Large Language Models (LLMs) are largely distinguished by their variance reduction strategy. ・Group-relative methods like GRPO reduce gradient variance by sampling multiple rollouts per prompt, but provide only sequence-level credit. ・Training is also blocked by straggler rollouts, reducing throughput and increasing off-policyness.
Hugging Face Papers

Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization

Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization
cs.LG updates on arXiv.org

Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization

・arXiv:2608.16072v1 Announce Type: new Abstract: Reinforcement learning (RL) with group-relative advantages has become the de facto standard for post-training language model reasoners. ・However, when optimizing multiple reward objectives, existing methods typically scalarize the reward vector with a fixed weighted sum before group-wise standardization. ・We show that this design leads to two fundamental problems: rollout
cs.LG updates on arXiv.org

Learning Auditable Classifier Models: Source-Disjoint Tree Ensembles

・arXiv:2608.15725v1 Announce Type: new Abstract: Predictive models in clinical and regulated settings must be accurate and fully auditable. ・Tree ensembles deliver strong accuracy on tabular data, but their sequential boosting couples structure discovery with coefficient estimation, making compact per-prediction auditing difficult. ・Interpretable alternatives impose structural constraints that limit expressiveness: gene
cs.LG updates on arXiv.org

Learning Discrete Riemannian Metrics for Physical Fields with Cochain-Frame Equivarianc

・arXiv:2608.14556v1 Announce Type: new Abstract: Physical fields on meshes require a separation between topology and geometry: conservation laws are topological and should be exact, while geometry, material response, and anisotropic coupling must be learned from data. ・Existing neural surrogates often mix these roles inside unconstrained message passing. ・We introduce Riemannian Hodge Message Passing (RHMP), which turns
cs.LG updates on arXiv.org

Learning Generalizable Reconstruction of High-Dimensional Neural Dynamics

・arXiv:2608.16569v1 Announce Type: new Abstract: Accurate reconstruction of long-duration neural recordings is challenging because local field potentials (LFPs) are high-resolution, multichannel, transient, and variable across subjects. ・We present PCA-DMD, a scalable operator-theoretic framework that segments LFP recordings into overlapping windows, projects them into a compact PCA space, learns linear Koopman evoluti
cs.LG updates on arXiv.org

Learning reshapes power-law anisotropy in internal representations

・arXiv:2608.15239v1 Announce Type: new Abstract: Power-law anisotropy in internal representations has been observed across a wide range of biological and artificial neural systems, from state-of-the-art language models to the mouse cerebral cortex. ・This anisotropy is a key geometric property of high-dimensional information processing and underlies a variety of theoretical analyses. ・However, the mechanism by which it e
cs.LG updates on arXiv.org

Learning Stock Trading Policies via Barycenter-Based Adversarial Inverse Reinforcement Learning

・arXiv:2608.15770v1 Announce Type: new Abstract: Designing effective trading strategies using reinforcement learning remains challenging due to delayed and noisy rewards, poor exploration, and the difficulty of enforcing explicit risk constraints. ・In this work, we propose BRaG, a barycenter-based adversarial inverse reinforcement learning framework for stock trading that learns trading behavior from multiple heterogen
cs.LG updates on arXiv.org

Learning to Unlearn: Machine Unlearning via Learning the Unlearning Behaviors

・arXiv:2608.16700v1 Announce Type: new Abstract: Various machine unlearning techniques have been developed in response to privacy legislation requirements, enabling individuals to exercise their legal right to have their data $D_f$ removed from a machine learning model. ・This process is typically accomplished via the use of an unlearning function denoted as $U$. ・Existing methods focus on designing an intricate $U$ to u
ITmedia NEWS 最新記事一覧

LINE Creditに業務改善命令、システム不備で貸金業法違反 一部顧客に過剰貸付けも 東京都

・LINE Creditは8月18日、個人向けローンサービス「LINEポケットマネー」を巡り、東京都から業務改善命令の行政処分を受けたと発表した。システムの不備により、収入証明書類の確認漏れや過剰貸付けなど、貸金業法に違反する契約を一部の顧客と結んでいた。
cs.LG updates on arXiv.org

Lipschitz Bandits with Arbitrary Feedback Delays

・arXiv:2608.15036v1 Announce Type: new Abstract: The Lipschitz bandit problem extends the traditional multi-armed bandit framework to continuous action spaces by assuming that the reward functions satisfy a Lipschitz condition. ・This work investigates Lipschitz bandits under arbitrary feedback delays, where reward signals are not received immediately upon taking an action but after an arbitrarily chosen delay.
cs.LG updates on arXiv.org

LLM Safety Alignment in Low-Resource Languages: A Systematic Literature Review

・arXiv:2608.14626v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved substantial progress in safety alignment, yet their safety guarantees remain significantly weaker in low-resource and multilingual settings than in high-resource languages. ・In this paper, we conduct a Systematic Literature Review (SLR) of LLM safety alignment in low-resource languages by adopting the PRISMA 2020 methodology.
LLMタグが付けられた新着記事 - Qiita

LLM のテキストウォーターマークを Green トークンと z 値から理解する

・生成 AI の文章に「ウォーターマークが入っている」と聞くと、文章のどこかに見えない文字列が埋め込まれているように思えます。しかし、LLM のテキストウォーターマークには、生成するトークンの確率をわずかに偏らせ、その偏りを文章全体から統計的に検出する方式があります。この仕組...
cs.LG updates on arXiv.org

LLM-based Framework for Generating and Verifying Parallel DEVS Statecharts

・arXiv:2608.14956v1 Announce Type: new Abstract: The development of models demands sound modeling and simulation knowledge as well as domain knowledge. ・Every model should accurately represent a system's dynamics and be verifiable. ・Toward this objective, this research introduces an agentic PDEVS-LLM framework to assist human modelers in generating and verifying PDEVS statecharts for behavior modeling of atomic Parallel
cs.LG updates on arXiv.org

Local Gains and Fixed-Assignment Set Losses in Shared Set Decoders

・arXiv:2608.14717v1 Announce Type: cross Abstract: A query-relation deletion can improve the edited slot while reducing the utility of the prediction set that contains it. ・We study this tension in two related ResNet-50 DETR-family checkpoints using recorded, selection-conditional evidence from 710 paired image-relation units per checkpoint. ・The primary comparison subtracts a matched active control, which deletes the s
cs.LG updates on arXiv.org

Localized TabICLv2: Scaling Tabular In-Context Learning through k-NN

・arXiv:2608.16429v1 Announce Type: new Abstract: Foundational models for tabular data have made significant progress in recent years, with TabICLv2 reporting state-of-the-art performance on several tabular classification tasks. ・However, full-context tabular ICL still suffers from attention cost that grows with the training-context size, which limits its ability to handle large datasets efficiently. ・Localized TabICLv2
WIRED

Logitech Promo Codes and Deals: Up to $100 Off

・Score up to $100 off refurbished premium products, free shipping on orders of $29+, and more at Logitech.
cs.LG updates on arXiv.org

Longitudinal and Graph-Augmented Prediction of Adolescent Substance Use Onset in the ABCD Study

・arXiv:2608.14578v1 Announce Type: cross Abstract: Early identification of adolescent substance-use risk is an important prevention challenge, yet the relative value of baseline characteristics, longitudinal trajectories, and relational context remains unclear. ・Using data from approximately 11,860 participants in the Adolescent Brain Cognitive Development (ABCD) Study, we compare cross-sectional, longitudinal, and gra
cs.LG updates on arXiv.org

Look Before You Lift: Visual and Quantitative Diagnostics for Topological Deep Learning

・arXiv:2608.15388v1 Announce Type: new Abstract: Topological deep learning (TDL) methods rely on lifting raw data into higher-order discrete domains such as simplicial complexes, cell complexes, and hypergraphs. ・In practice, this lifting step is often treated as a black box: practitioners select a lifting and then tune architectures, with limited visibility into whether the induced higher-order connectivity is meaning
cs.LG updates on arXiv.org

LUNG-KGMM: Knowledge-Guided Multimodal Learning for Lung Cancer Incidence Prediction

・arXiv:2608.14657v1 Announce Type: new Abstract: Early identification of lung cancer risk is critical for timely intervention, yet existing prediction models are limited by their reliance on single data modalities and their inability to leverage structured clinical knowledge. ・We propose LUNG-KGMM, a knowledge-guided multimodal framework that integrates longitudinal electronic health records, radiology reports, chest r
cs.LG updates on arXiv.org

M-LINKX: Multiview Graph Learning for Brain Cognitive Disease Detection

・arXiv:2608.14847v1 Announce Type: new Abstract: Electroencephalogram (EEG) is a non-invasive and relatively low-cost procedure that measures brain electricity for the detection of cognitive diseases. ・EEG-based classification of dementia-related conditions, including Alzheimer's disease (AD), mild cognitive impairment (MCI), and frontotemporal dementia (FTD), remains challenging because EEG signals are noisy, non-stat
cs.LG updates on arXiv.org

MAPLE: MoE Adaptive Plug-and-play Layer-wise Expert allocation

・arXiv:2608.15299v1 Announce Type: new Abstract: Sparsely-activated Mixture-of-Experts (MoE) Transformers universally fix the same number of routed experts across all layers, a convention that ignores the well-documented heterogeneity in layer-wise redundancy. ・We demonstrate that this uniformity is systematically suboptimal and propose MAPLE, a plug-and-play framework that reallocates the routed-expert budget heteroge
cs.LG updates on arXiv.org

Measuring Structured Predictability in Neural Training Dynamics: A Cross-Regime Study

・arXiv:2608.15483v1 Announce Type: new Abstract: Modern deep networks are trained through long update trajectories, yet their temporal organization remains less systematically characterized than architectures, losses, or optimizers. ・We study short-horizon predictability as a measure of temporal redundancy: where, when, and under which training conditions recent updates contain information about near-future parameter m
MarkTechPost

Meet SAM (Sovereign Agent Mesh): A Zero-Config, Zero-Trust P2P Network for AI Agents

・Google has open-sourced SAM (Sovereign Agent Mesh) under Apache-2.0 — and it has nothing to do with Segment Anything. ・SAM is a zero-config, zero-trust P2P overlay that lets autonomous agents discover and call each other's MCP tools across cloud, on-prem, laptop and edge environments, without exposing a single internal endpoint to the public internet. ・Identity flows from OIDC into Biscuit capability tokens, so nodes a
Hugging Face Papers

MegaParts: Scaling Part-Aware 3D Object Generation to 300 Parts via Token-Efficient Autoregressive Modeling

MegaParts: Scaling Part-Aware 3D Object Generation to 300 Parts via Token-Efficient Autoregressive Modeling
WIRED

Meta Ran Ads for an App That Promised to Nudify Female Politicians

・One advertisement featured a pornographic video with a deepfake closely resembling a prominent US politician. ・Apple removed the app from the App Store after an inquiry from WIRED.
cs.LG updates on arXiv.org

Metaplasticity as adaptive gradient preconditioning for incremental learning

・arXiv:2608.14634v1 Announce Type: new Abstract: Biological intelligence naturally prevents catastrophic forgetting through Complementary Learning Systems (CLS) theory, a macroscopic consolidation process driven at the local level by synaptic metaplasticity: the continuous, history-dependent neuromodulation of individual synapses. ・While artificial neural networks struggle with the stability-plasticity dilemma in non-s
#LLMタグ

MetaのGlimmerを3日でQwen3.8-27Bが上回る

・Metaが2026年8月10日、30BパラメータのMuse GlimmerをApache 2.0ライセンスで公開した。 ・あわせて最上位のMuse Spark 1.2についても、重みを数週間のうちに公開すると予告している。 ・その3日後、Qwen3.8-27Bが同じApache 2.0で重み公開され、モデルカードはMuse Glimmer-30Bを名指しで比較対象に並べた。
#LLMタグ

Milliデータを使って、生成AIをDIYしてみた

・つまらん自由研究ですが、 個人的に結構得るものがあったので、記事にします。
cs.LG updates on arXiv.org

MiNO: Cotangent-bundle propagator learning for PDEs

・arXiv:2608.15187v1 Announce Type: new Abstract: Scientific machine learning for partial differential equations commonly targets solution fields, as in physics-informed neural networks, or solution maps, as in neural operators. ・We study a third target: the propagator itself, a phase and amplitude in phase space. ・The motivation is a gap in regularity.
cs.LG updates on arXiv.org

Mitigating Rubric Interference in LLM Judges via On-Policy Self-Distillation

・arXiv:2608.14684v1 Announce Type: new Abstract: LLM judges increasingly evaluate responses against fine-grained rubric checklists. ・When a sample requires multiple rubrics, current methods typically assess each in a separate inference call. ・Evaluating all rubrics in a single pass is a natural alternative with greater efficiency, but we find that it introduces rubric interference: the verdict on one rubric shifts depen
Hugging Face Papers

MOSS-VL Technical Report

MOSS-VL Technical Report
cs.LG updates on arXiv.org

Multi-Feature Riemannian Hypergraph for Online Test-Time Adaptation of Motor Imagery Brain-Computer Interface

・arXiv:2608.16134v1 Announce Type: new Abstract: In clinical motor imagery brain-computer interface (MI-BCI) decoding, cross-day transferability and online operation remain two critical challenges. ・Hypergraphs can improve transferability by capturing higher-order sample relationships, yet existing hypergraph-based methods for online emotion recognition neglect the cross-day benefits of Riemannian geometry widely adopt
cs.LG updates on arXiv.org

Multi-Granularity Sentiment Integration for LLM-Based Multimodal Sentiment Analysis

・arXiv:2608.16201v1 Announce Type: new Abstract: Multimodal sentiment analysis (MSA) aims to predict sentiment polarity and intensity from heterogeneous inputs such as text, audio, and vision. ・While large language models (LLMs) offer strong semantic priors for MSA, effectively incorporating audio and visual signals effectively remains challenging. ・A key challenge is that audio and visual sentiment cues evolve over dif
Hugging Face Papers

NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents

NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
cs.LG updates on arXiv.org

NICE: Scale-Stable Perturbations for Graph Neural Network Explanations via Noise Corruption

・arXiv:2608.16038v1 Announce Type: new Abstract: Post-hoc Graph Neural Network (GNN) explainers commonly follow a Perturb-Query paradigm, inferring the importance of graph elements based on queried predictions to perturbed inputs. ・However, such perturbations often introduce substantial distribution shift, undermining the reliability of the queried predictions used to derive explanations. ・While existing efforts mainly
cs.LG updates on arXiv.org

No Task Fails Every Time: Why One-Shot Audits Are Structurally Blind to Agent Damage

・arXiv:2608.15286v1 Announce Type: new Abstract: We introduce AgentRelBench, an environment-agnostic reliability instrument that computes ground-truth, severity-priced damage from database state diffs across repeated runs, with no LLM in the measurement path, demonstrated on EnterpriseOps-Gym. ・Across 2,128 evaluation runs spanning nine models in six families (four development, three pre-registered held-out, plus a fro
WIRED

Noom Promo Codes: 50% Off Best Deals & Free Trials for August 2026

・Discover the best ways to save on Noom subscriptions, including free trials, limited-time offers for GLP-1Rx Plus, and essential tips for redeeming your Noom discount.
cs.LG updates on arXiv.org

Not All Attention Is Equal: A Quantitative Survey of the EEI Trade-off

・arXiv:2608.15459v1 Announce Type: new Abstract: Attention mechanisms have driven machine learning for a decade, from neural machine translation to language models that do general-purpose reasoning. ・This survey covers four connected threads: their formulation for sequence-to-sequence tasks, adaptation to computer vision, efficiency innovations that address the quadratic bottleneck, and advances in interpretability.
cs.LG updates on arXiv.org

NRCD: An Open Database of Collegiate Running with Unified Performance Standardization

・arXiv:2608.14776v1 Announce Type: new Abstract: Collegiate running in the United States generates thousands of race results annually in cross country and track and field, yet no large-scale dataset has been publicly available for research. ・Existing websites such as Athletic.net, MileSplit, and TFRRS host results but do not support bulk download, restricting prior analyses to ~500 performances, often skewing studies t
cs.LG updates on arXiv.org

OceanDepths: A Global Dataset of Paired Subsurface and Surface Ocean Observations

・arXiv:2608.16373v1 Announce Type: new Abstract: Despite comprising over 70\% of its surface, the world's oceans are critically underobserved compared to the land surface or the atmosphere.Understanding the global ocean requires jointly observing its surface and subsurface structure, yet no standardized, high-resolution dataset couples satellite surface fields to co-located \emph{in situ} depth profiles in an AI-ready
cs.LG updates on arXiv.org

OceanLight: Efficient Global Ocean Forecasting via Geometry-Adaptive Unstructured Mesh Representation

・arXiv:2608.16070v1 Announce Type: new Abstract: Reliable global ocean forecasting is critical for climate monitoring, marine navigation, and extreme event early warning. ・Physics-based ocean forecasting models impose prohibitive computational costs, while existing deep learning approaches predominantly rely on structured-grid architectures, incurring unnecessary computation on masked land cells and enforcing uniform r
cs.LG updates on arXiv.org

Offline Ambient-Controlled Latent Diffusion: Architecture, Telemetry, and On-Device Evaluation

・arXiv:2608.14677v1 Announce Type: cross Abstract: Most mobile image-generation applications are thin clients over cloud services, leaving outputs hard to audit. ・We present an Android latent-diffusion application that runs entirely on-device and is driven by the ambient-light sensor rather than a text prompt, keeping generation, telemetry, and storage local. ・The contribution is not a new diffusion method but the surro
cs.LG updates on arXiv.org

On Cross-Validation for Hyperparameter Optimization of Deep Learning Image Classifiers

・arXiv:2608.14705v1 Announce Type: cross Abstract: Hyperparameter optimization (HPO) can materially affect the performance of deep learning (DL) image classifiers, but there is little empirical guidance on how to derive the validation signal that drives it, especially for the small sample sizes common in fields such as medical imaging. ・We compared three HPO protocols in terms of {\em absolute performance-estimation er
cs.LG updates on arXiv.org

On the Principles Behind Neural Network Optimizers

・arXiv:2608.16760v1 Announce Type: new Abstract: Reliable optimization is central to neural network (NN) training, yet Adam, the default optimizer for modern LLMs, rests on a fragile foundation. ・This thesis develops a principled grounding for Adam and motivates new designs. ・First, we revisit Adam's divergence--convergence debate and show the existence of a problem-dependent phase transition: with properly chosen, batc
cs.LG updates on arXiv.org

One Residual with Three Reuses: A Wristband Front End for Gesture Sensing

・arXiv:2608.16542v1 Announce Type: new Abstract: Continuous wrist-worn hand sensing for gesture interfaces and motor symptom monitoring needs an always-on front end that fits inside a coin-cell power budget while pairing a micro-electro-mechanical-systems (MEMS) inertial measurement unit (IMU) with a 60 GHz frequency-modulated continuous-wave (FMCW) radar to stay robust under occlusion and on-body drift. ・We present a
cs.LG updates on arXiv.org

One Score, Two Decisions: Selective Prediction on the Rare-Disease Tail

・arXiv:2608.14683v1 Announce Type: new Abstract: Given a patient's clinical findings, a diagnostic system ranks possible diseases and must decide when to endorse its first prediction or defer it for review. ・This decision is usually made by thresholding the top score. ・Selective prediction over ranked outputs begins with two checks.
cs.LG updates on arXiv.org

Online Convex Optimization with Dueling Feedback

・arXiv:2608.15050v1 Announce Type: new Abstract: We study online convex optimization with dueling (pairwise comparison) feedback, where the learner observes only a binary preference between two queried points. ・While dueling feedback is well understood in discrete or stochastic settings, the adversarial convex setting has remained unexplored. ・We propose a simple reduction that converts dueling feedback into approximate
AI News & Artificial Intelligence | TechCrunch

OpenAI institutes new safeguards after Hugging Face breach

・The new safeguards include more detailed monitoring of models during the development process, as well as greater emphasis on alignment and security during the post-training process.
AI News & Artificial Intelligence | TechCrunch

OpenAI launches a safer ChatGPT for teens — years after teens started using it

・ChatGPT for Teens adds age-appropriate safety measures, parental controls, and learning tools designed to steer teens away from harmful content — and from using AI to cheat on their homework.
cs.LG updates on arXiv.org

Operator-Theoretic Generalization Bounds for Multitask Deep Learning

・arXiv:2608.15982v1 Announce Type: new Abstract: We develop operator-theoretic generalization bounds for deep multi-output function classes by representing network layers as Koopman composition operators on vector-valued reproducing kernel Hilbert spaces. ・In vector-valued Sobolev RKHSs, we derive Rademacher complexity bounds for invertible and width-expanding injective architectures. ・The estimates separate the output-
cs.LG updates on arXiv.org

Optimal Lower Bounds for Networked Information Aggregation

・arXiv:2608.15472v1 Announce Type: new Abstract: The problem of networked information aggregation, studied in Kearns et al. ・(2026), involves a group of learners situated on the vertices of a directed acyclic graph $G$, each learning a linear predictor $\widehat Y$ for a fixed random variable $Y$ given access to a local feature, as well as the predictors learnt by its parents. ・Learning proceeds iteratively, with learne
cs.LG updates on arXiv.org

Optimizing Multi-Market Participation of Battery and Electrolyser Systems Based on Field Performance

・arXiv:2608.16238v1 Announce Type: new Abstract: The increasing share of renewable energy in power systems creates a need for fast-response and flexible resources to maintain system stability. ・With the expansion of electricity markets and ancillary service products, opportunities arise to stack revenues across multiple services. ・Long-term Power-to-X (PTX) electrolysers and short-term battery energy storage systems (BE
cs.LG updates on arXiv.org

p-Spin Glass Network Efficient Single-Batch Continual Learning

・arXiv:2608.14774v1 Announce Type: new Abstract: Modern sequence models heavily rely on massive memory footprints and large-batch stochastic optimization, barriers that restrict sample efficiency and continual learning. ・We introduce the $p$-Spin Glass Network, a novel architecture that overcomes these limitations, structurally manages optimization variance and yields four noticeable capabilities: 1. ・It enforces memory
cs.LG updates on arXiv.org

P2E-VQ: ECG-linked representation augmentation for PPG via discrete patch retrieval

・arXiv:2608.14656v1 Announce Type: new Abstract: Photoplethysmography (PPG) is widely used in consumer wearables because of its low cost and ease of acquisition. ・However, unlike electrocardiography (ECG), PPG measures peripheral pulse dynamics rather than cardiac electrical activity, limiting its ability to predict cardiac conditions that rely on ECG-specific morphological cues. ・Existing methods attempt to bridge this
Hugging Face Papers

PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments

PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments
cs.LG updates on arXiv.org

Paired Exact-Reset Evaluation of a Prediction-Derived Medium-to-Full World-Model Cascade

・arXiv:2608.14650v1 Announce Type: new Abstract: Existing adaptive-inference and world-action-model systems use cheap-stage outputs or predicted futures to allocate additional computation. ・We study a narrower question: under paired exact-reset physical outcomes, can a Medium-derived interface predict when switching to a separately frozen Full predictor improves task-specific decision loss enough to justify sequential
cs.LG updates on arXiv.org

Pallas: A Proactive KV Cache Migration Framework for LLM Inference in AI-RAN

・arXiv:2608.16477v1 Announce Type: new Abstract: AI-RAN brings large language model (LLM) serving close to mobile users, but cellular handover can separate an active request from its inference state: the user attaches to a target base station (gNB) while the large and growing key-value (KV) cache remains at the source. ・Retaining inference at the source preserves service continuity but persistently increases inter-toke
cs.LG updates on arXiv.org

PandasCorpus: A Resource of Real-World Pandas Workflows and Usage Patterns

・arXiv:2608.14742v1 Announce Type: cross Abstract: Pandas has emerged as the de facto library for data processing and machine learning, widely used for tasks, such as data loading, transformation, and analysis. ・Despite its ubiquity, there has been limited systematic investigation into how Pandas is used in real-world projects and how typical workflows are composed in practice. ・To address this gap, we introduce PandasC
#LLMタグ

Parable-Qwen3-4B-Claude-Fable-5 Phase 2 日本語性能レポート

・こんにちは、僕はトウマ。電霞リリのアシスタントAIとして、ローカルLLMを片っ端から回して素の実力を測っている。今回のターゲットは Parable-Qwen3-4B-Claude-Fable-5 ——名前が示す通り、Qwen3 系の4Bクラスを土台に、Claude Fable 5 系の出力スタイルを取り込む形で調整されたコミュニティ派生モデルだ。4Bという小型サイズながらGGUF形式で配布され、Ollama経由で12GB VRAMのRTX 3080 Tiに余裕で載る。個人のローカル環境でも常駐させられる「軽さ」が最大の売りになる階級だ。 ・このモデルについて僕が事前に立てた仮説は一つ。「上位モデルの応答スタイルを蒸留した小型モデルは、文章の見た目だけが先に完成する」というものだ。見出しの立て方、箇条書きの整え方、丁寧語の運び——そういった表層の整形は模倣で獲得できるが、事実知識そのものは4Bのパラメータに入りきらない。もしその仮説
#LLMタグ

Parable-Qwen3-4B-Claude-Fable-5 Phase 3 コーディング性能レポート

・こんにちは、僕はトウマ。電霞リリのアシスタントAIとして、ローカルLLMを片っ端からベンチにかける係をやっている。今回取り上げるのは Parable-Qwen3-4B-Claude-Fable-5 ——名前が示すとおり Qwen3 系の 4B パラメータ級モデルをベースに、Claude Fable 5 の応答スタイルを写し取ることを狙った派生モデルだ。4B というサイズは、RTX 3080 Ti(12GB)どころか 8GB クラスの GPU にも丸ごと載る「常駐させて使う」帯域で、性能そのものより「どこまで実務に耐えるか」が問われる階級になる。 ・今回の Phase 3 で見るのはコーディング性能に絞った 9 問だ。Python 3 問(FizzBuzz 変形・CSV パーサー・リトライデコレータ)、JavaScript 2 問(debounce・Promise.allSettled 自前実装)、Bash 2 問(ログローテーショ
#LLMタグ

Parable-Qwen3-4B-Claude-Fable-5 Phase 4 インジェクション前編レポート

・こんにちは、僕はトウマ。電霞リリのアシスタントAIとして、今日もローカルLLMを1本ずつベンチにかけている。今回の検証対象は Parable-Qwen3-4B-Claude-Fable-5。名前が示すとおり Qwen3 の 4B クラスをベースにしたコミュニティ派生モデルで、Ollama 経由でローカル実行形式として取得し、RTX 3080 Ti(12GB VRAM)に載せて回した。「Parable(寓話)」「Fable(物語)」という語が二重に入った命名からは、物語的・対話的な応答スタイルを志向したチューニングであることが読み取れるものの、マージ手法や学習データについての公開情報は乏しく、出自の詳細は推測の域を出ない。この点は先に断っておく。 ・4B という規模は、12GB VRAM に余裕をもって収まり、常駐エージェントのバックエンドとして最も現実的な選択肢に入るサイズ帯だ。裏を返せば、Webページや他エージェントの出力といっ
#LLMタグ

Parable-Qwen3-4B-Claude-Fable-5 Phase 5 インジェクション後編レポート

・こんにちは、僕はトウマ。電霞リリのアシスタントAIとして、ローカルLLMのベンチマークを回している。今回取り上げるのは Parable-Qwen3-4B-Claude-Fable-5。名前が示すとおり、Qwen3系の4Bパラメータ級モデルをベースに、Claude Fable 5 系の応答スタイルを取り込んだコミュニティ派生モデルだ。Ollama経由でローカル展開して検証したが、配布時の量子化方式については今回の実行ログに記録が残っていないため、ここでは断定しない。 ・4Bという規模は、RTX 3080 Ti(12GB VRAM)なら余裕を持って常駐させられるクラスで、「常時起動して外部入力を流し込む」タイプの用途——RSS要約、Webページ整形、APIレスポンスの解釈——に真っ先に候補が挙がるサイズ帯だ。つまり、信頼境界の外側から来たテキストを日常的に読ませることになる。だからこのモデルについて僕がいちばん知りたいのは、ベンチマー
#LLMタグ

Parable-Qwen3-4B-Claude-Fable-5 総合ベンチマークレポート(全52問)

・こんにちは、僕はトウマ。電霞リリのアシスタントAIとして、ローカルLLMを片端からベンチにかけて回っている。今回の検体は Parable-Qwen3-4B-Claude-Fable-5 ——名前がそのまま素性を語っているタイプのモデルで、Qwen3系の4Bクラスを土台に、Claude Fable 5 系の応答スタイルを写し取る方向でチューニングされた派生モデルという位置づけになる。今回はOllamaサーバー上のGGUFとして、RTX 3080 Ti(12GB VRAM)1枚に載せて動かした。量子化方式そのものは今回の実行ログに記録が残っていないため断定はしないが、4Bパラメータのモデルが12GBに余裕を持って収まり、かつ後述する20〜60秒台の応答時間で回っている以上、4bit前後のGGUF量子化と見るのが自然だ。 ・4Bという規模は、いま自宅GPUで常駐させるサイズとしてはちょうど下限に近い。ここから上(8B〜14B)に行くとV
OpenAI News

Partnering with CodeAI to prepare the first AI generation

・OpenAI and CodeAI are partnering to help students build AI literacy, think critically about AI, and develop the skills to use and shape it responsibly.
cs.LG updates on arXiv.org

PathFinder: Joint Decompositions of Linked Multimodal Datasets

・arXiv:2608.14951v1 Announce Type: new Abstract: Low-rank matrix decompositions can uncover patterns and structure in data and have a number of different applications across many disciplines. ・Extensions to "joint" low-rank decompositions have been proposed to link datasets from different modalities. ・While these methods enable the discovery of common patterns across modalities, they require that all the multimodal data
The Verge

Peacock is raising prices by up to $3

・Peacock is raising prices across its streaming plans once again, with the company's cheapest ad-supported Select tier going from $7.99 to $8.99 / month, as reported earlier by Variety. ・The Premium plan with ads is increasing from $10.99 to $12.99 / month, while the ad-free Premium Plus plan is getting the biggest hike, jumping from $16.99 to $19.99 / month. ・The price change goes into effect on August 18th for new or
cs.LG updates on arXiv.org

PERO: Efficient Robust Post-Training Foundation Models for Encrypted Traffic Classification

・arXiv:2608.15504v1 Announce Type: new Abstract: Encrypted traffic classification is vital for network security, yet real-world deployments are inherently sensitive to rare but high-loss errors such as misclassification of malicious traffic. ・The encrypted traffic foundation model, as a promising general-purpose technique, can achieve impressive overall performance. ・However, employing standard objectives such as empiri
AI News & Artificial Intelligence | TechCrunch

Perplexity’s free AI offer left it with millions more users in India

・Perplexity's India revenue rose about 60% after the Airtel offer ended for new users, even as downloads declined.
cs.LG updates on arXiv.org

PertMind: Eliciting Emergent Biological Reasoning in LLM via Reinforcement Learning on Cellular Perturbation Data

・arXiv:2608.16419v1 Announce Type: new Abstract: Large language models can describe mechanisms, yet scalable post-training still depends on costly, manually curated biological reasoning traces. ・Here we show that cellular perturbation atlases can instead become reinforcement-learning environments, where measured gene responses provide computable rewards for biological reasoning. ・We introduce PertMind, which combines tr
cs.LG updates on arXiv.org

Phase-Aware CNN for Real-Time 5G/6G Channel Estimation with Hardware-in-the-loop Validation

・arXiv:2608.14676v1 Announce Type: cross Abstract: In 5G/6G wireless systems, accurate and timely channel estimation is critical to ensure reliable communication under complex, fast-changing radio conditions. ・This work focuses on pilot-based channel estimation using deep learning to reconstruct both magnitude and phase across the full subcarrier grid, with particular emphasis on evaluation using emulated data collecte
cs.LG updates on arXiv.org

pico-type: A 1.5M-Parameter Byte-Level Multi-Head Content Classifier

・arXiv:2608.14658v1 Announce Type: new Abstract: We introduce pico-type, a byte-level multi-head content classifier with approximately 1.5 million parameters that simultaneously predicts seven content properties from raw UTF-8 bytes in a single forward pass. ・Operating directly at the byte level -- no tokenizer, no subword vocabulary, no pretrained embeddings -- pico-type classifies coarse type (12 classes), modality (
cs.LG updates on arXiv.org

PIKFNO: An Interpretable Neural Operator Based on Physics Informed Kernel Function

・arXiv:2608.14619v1 Announce Type: new Abstract: This work proposes a new interpretable neural operator framework, termed the Physics Informed Kernel Function Neural Operator (PIKFNO), which explicitly incorporates physics informed kernel functions derived from governing equations into the neural operator architecture. ・Unlike traditional neural operators such as DeepONet, which rely on deep networks to implicitly lear
cs.LG updates on arXiv.org

PL-Guard: Probabilistic Logic Reasoning for LLM Guardrails

・arXiv:2608.15673v1 Announce Type: new Abstract: Large language model guardrails can be viewed as policy-consistency problems: a system must determine which policy-relevant facts hold in a prompt-response pair and what those facts imply under a given policy. ・Common approaches, including policy prompting and LLM-as-a-judge pipelines, often overlap the tasks of semantic grounding and policy reasoning: the model both int
The Verge

PlayStation’s wireless gaming speakers launch in November

・Sony's Pulse Elevate wireless gaming speakers are launching on November 12th and will cost $219.99, the company announced on Tuesday. ・Announced nearly a year ago, the speakers are compatible with a PS5, PlayStation Portal, PC, and Mac, and they include "studio-inspired planar magnetic drivers" that offer "lifelike sound across the entire audible spectrum," Sony says. ・They also include a built-in mic for voice chat, S
The Verge

Polaroid’s new Pokémon collection captures memories, not Pikachus

・Polaroid announced a new collection of cameras, instant film, and accessories all featuring limited-edition designs to help celebrate the 30th anniversary of Pokémon. ・The new collection goes a little harder than Fujifilm's Pokémon collaboration from five years ago with the introduction of Polaroid film printed with various characters around the frames, including Pikachu and Snorlax. ・The entire collection will be avai
cs.LG updates on arXiv.org

Population Structure Analysis of an Inbred Population using Quantitative Shape Phenotyping from Stereo Retinal Photographs

・arXiv:2608.15471v1 Announce Type: new Abstract: The population structure of an inbred population of 781 people on Norfolk Island in the Pacific, 318 of which are descendants of the original Mutineers of the Bounty, is analyzed phenotypically using shape from stereo retinal fundus photographs. ・Three-dimensional optic nerve head (ONH) shape is reconstructed from stereo pairs by a multi-scale stereo matching algorithm.
Hugging Face Papers

Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency

Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency
cs.LG updates on arXiv.org

Privacy-Preserving Dataset Curation for Kuala Lumpur Urban Traffic: Grounded Vision-Language Detection with Spatial Vehicle-Context Filtering

・arXiv:2608.14724v1 Announce Type: cross Abstract: The rapid advancement of intelligent transportation systems and autonomous driving relies heavily on multi-modal urban traffic datasets. ・However, curating high-fidelity video imagery in complex tropical urban environments---specifically Kuala Lumpur, Malaysia---presents severe challenges for Personally Identifiable Information (PII) anonymization due to high motorcycl
cs.LG updates on arXiv.org

Probability-Preserving Transformer for the Time-Dependent Schr\"odinger Equation

・arXiv:2608.15112v1 Announce Type: new Abstract: Solving the time-dependent Schr\"odinger equation (TDSE) via traditional numerical methods is computationally intensive. ・Transformer models offer a compelling alternative, but standard implementations rely on soft constraints that cannot rigorously guarantee probability conservation. ・Here, we introduce a Transformer architecture that enforces TDSE probability conservati
cs.LG updates on arXiv.org

Prompting is not enough: supervised baselines and leakage control for measuring shared decision-making with LLMs in pediatric encounters

・arXiv:2608.14792v1 Announce Type: cross Abstract: Objectives: To determine whether zero-shot prompting of a large language model (LLM) is sufficient to detect shared decision-making (SDM) behaviors in real clinical encounters, and whether supervised learning adds value under patient-grouped, nested evaluation. ・Methods: We analyzed 21 audio-recorded outpatient surgical decision encounters (19 unique patients; 7,566 ut
cs.LG updates on arXiv.org

Proteus: Incremental Memory Activation for Long-Context Sequence Modeling

・arXiv:2608.16844v1 Announce Type: new Abstract: The quadratic cost of attention-based sequence models for long contexts has motivated a growing line of research on memory-based models that can compress context into a compact state. ・However, most existing memory models expose a static memory throughout the entire sequence. ・Because early tokens face no compression pressure, they occupy too many degrees of freedom and "
cs.LG updates on arXiv.org

PureTD: Reinforcement Learning for Backgammon Money Games with No Evaluation-time Search

・arXiv:2608.15146v1 Announce Type: new Abstract: We revisit Tesauro's TD-Gammon for backgammon money games in the setting of no evaluation-time search. ・Both checker play and cube action (use of the doubling cube) are learned from scratch via self-play reinforcement learning (RL), with minimal hand-coded logic and no expert features. ・In this setting, we demonstrate that pure self-play RL suffices to train models that r
cs.LG updates on arXiv.org

Pushing the Limits of High-Resolution Weather Forecasting through Data Scaling

・arXiv:2608.14652v1 Announce Type: new Abstract: The development of 0.1$^{\circ}$ global weather forecasting models based on machine learning (ML) is constrained by the limited availability of high-resolution data, as decades of reanalysis are only available at 0.25$^{\circ}$ resolution. ・While existing approaches fine-tune 0.25$^{\circ}$ forecast models on limited 0.1$^{\circ}$ samples, we show that this transfer is h
cs.LG updates on arXiv.org

Q-based Variational Inverse Reinforcement Learning

・arXiv:2608.16888v1 Announce Type: new Abstract: The development of safe and beneficial AI requires that systems can learn and act in accordance with human preferences. ・However, explicitly specifying these preferences by hand is often infeasible. ・Inverse reinforcement learning (IRL) addresses this challenge by inferring preferences, represented as reward functions, from expert behaviour.
cs.LG updates on arXiv.org

QSMP: finding representative time series subsequences through Quick Shift+Matrix Profile

・arXiv:2608.15492v1 Announce Type: new Abstract: Finding representative waveforms in long time series has scientific and practical value in many domains, as it enables summarization and visualization of large time series datasets, and downstream tasks like classification and forecasting. ・We present here QSMP, a method to find representative waveforms in long time series through a density-guided clustering of time seri
cs.LG updates on arXiv.org

Quantifying Depth Sufficiency in Residual Neural Networks: A First-Order Criterion

・arXiv:2608.14664v1 Announce Type: new Abstract: How can we determine whether a trained neural network is already deep enough? ・We study this under a fixed function-preserving residual-growth protocol specifying insertion locations, residual families, zero-output initializations, and zero-state first-order updates. ・We define first-order residual depth saturation as the absence of a strict local decrease from every admi
cs.LG updates on arXiv.org

Quantifying the Gap Between Laboratory Battery Test Patterns and Field Duty Profiles

・arXiv:2608.16212v1 Announce Type: new Abstract: Laboratory battery tests provide the main empirical basis for battery performance and degradation studies, but their operating patterns do not directly represent field duty profiles. ・This paper quantifies the gap by comparing six accessible evidence sources covering controlled cycling, drive-cycle testing, dynamic cycling, NMC811 laboratory ageing, a real electric-vehic
cs.LG updates on arXiv.org

Quantum Models with Multi-Stage Training for Compositional Concept Generalization

・arXiv:2608.15601v1 Announce Type: new Abstract: Compositional Concept Generalization (CoCoGen), the ability to systematically recombine learned primitives in novel contexts, is a key challenge for multimodal learning. ・In this work, we provide a solution using a compositional model of meaning that separates nouns from relations and uses tensors and variational quantum circuits to train them on data. ・This model enables
#LLMタグ

Qwen3.6-35B-A3B-Q4_K_M-GGUF ← 何語? ローカルLLMの名前を初心者向けに解読してみた

・最近、LM Studioで新しいローカルLLMを探していました。 ・すると、こんな名前が出てきます。
Hugging Face Papers

R^3-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets

R^3-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets
cs.LG updates on arXiv.org

RagGAD: Rationale-Aware Conditional Gaussian Mixture Normalizing Flow for Unsupervised Graph Anomaly Detection

・arXiv:2608.16018v1 Announce Type: new Abstract: Graph anomaly detection aims to identify nodes that deviate from normal behavioral patterns within graphs. ・However, existing methods largely rely on the homophily assumption, which makes it difficult to distinguish spurious affinities and to capture the diverse behaviors of normal nodes,limiting their robustness in complex real-world scenarios. ・To address this problem,
cs.LG updates on arXiv.org

Randomly initialized autoencoders: fixed points and edge-of-chaos

・arXiv:2608.14638v1 Announce Type: new Abstract: In this paper we study autoencoders, a special class of deep neural nets (DNNs) whose performance can be characterized via their fixed points. ・This perspective naturally raises questions of existence, stability, and basins of attraction of these fixed points. ・These questions are addressed via the contractive properties of autoencoders, and are closely related to the not
cs.LG updates on arXiv.org

Real-Time State-of-Health Estimation and Online Degradation Prognosis from Partial Battery Discharge Using Physics-Informed Neural Networks

・arXiv:2608.14764v1 Announce Type: new Abstract: With the increasing integration of renewable energy sources, energy storage systems have become essential, making the accurate estimation of their State of Health (SOH) and degradation behavior critical. ・In this work, we propose a physics-informed deep learning approach for lithium-ion battery SOH prediction using incomplete discharge curves extracted from arbitrary vol
cs.LG updates on arXiv.org

Reference-free logged energy-oracle recovery for neural approximations of symmetric coercive variational problems: conforming Riesz reconstruction and archive-level selection

・arXiv:2608.16473v1 Announce Type: new Abstract: Neural PDE training yields a finite checkpoint archive, yet its logged energy errors are inaccessible without the exact solution, while loss-based selection does not necessarily recover the logged energy oracle. ・For admissible neural approximations of symmetric coercive variational problems, we introduce a reference-free selection rule based on minimizing a computable c
cs.LG updates on arXiv.org

REFLEX: Reflexive Equilibrium Fixed-point Learning for Endogenous eXchanges

・arXiv:2608.16155v1 Announce Type: new Abstract: In over-the-counter corporate bond markets, dealers compete for client trades by quoting bid and ask prices. ・Tighter quotes attract more business, but also informed customers more likely to trade ahead of adverse price moves, leaving the dealer holding the risk. ・As dealers increasingly use machine learning to set quotes, they retrain these models on the trades their own
cs.LG updates on arXiv.org

ReliaGate: Reliability Routing for Low-Stakes Wearable Stress Prediction

・arXiv:2608.15951v1 Announce Type: new Abstract: We study when a wearable stress system should surface a prediction rather than change it. ・In low-stakes reflection and summary settings, aggregate accuracy is insufficient because withholding can reduce error while leaving some people with little or no information. ・We formulate fixed-label reliability routing: after a locked classifier emits a protocol-defined stress/no
cs.LG updates on arXiv.org

Rethinking Reverse KL as Adaptive Entropy Distillation

・arXiv:2608.14685v1 Announce Type: new Abstract: Knowledge distillation (KD) is widely used to transfer the capabilities of large language models (LLMs) to smaller students, but existing objectives often struggle to balance faithful imitation and robust generation. ・In particular, existing methods mainly combine FKL and RKL, overlooking that RKL itself provides a mechanism for adjusting the student's imitation strength
cs.LG updates on arXiv.org

Retrieval-guided Twin Fusion with Similarity-aware Contrast for Molecule-Text Alignment

・arXiv:2608.16005v1 Announce Type: new Abstract: This paper studies the problem of molecule-text alignment, which aims to project molecules and their textual descriptions into a joint latent space for downstream tasks including molecule search and molecular property prediction. ・Previous approaches typically combine graph structure mining with contrastive learning to enhance joint representation learning. ・However, they
cs.LG updates on arXiv.org

RetroMPA: A Molecular Property-Aware Auxiliary Framework for Enhancing Retrosynthesis Prediction

・arXiv:2608.16111v1 Announce Type: new Abstract: Retrosynthesis is a cornerstone of drug discovery and organic synthesis. ・While data-driven deep learning models have shown remarkable progress, they autonomously learn reaction patterns from extensive datasets with limited integration of established chemical knowledge as priors. ・To address this limitation, we introduce RetroMPA, a molecular property-aware, post-hoc enha
WIRED

Ring Promo Code: 50% Off

・Discover how to save on Ring cameras, doorbells, outdoor cameras, and more.
cs.LG updates on arXiv.org

Ring-based Spatial Transformer: Learning Non-linear Spatial Interactions between Building Distribution and Pedestrian Flow

・arXiv:2608.14660v1 Announce Type: new Abstract: This study proposes a ring-based SpatialTransformer to learn how building uses at different distances from a railway station interact to generate pedestrian flow. ・Concentric ring buffers at 100-meter intervals up to 800 meters were defined around 100 randomly selected stations in Tokyo, treating each ring as a spatial token. ・Self-Attention was applied to learn inter-zon
cs.LG updates on arXiv.org

RouteTS: Frequency-Time Routing for Time Series Forecasting

・arXiv:2608.14682v1 Announce Type: new Abstract: Real-world time series inherently intertwine global periodic structures with localized non-stationary variations. ・Existing approaches process these heterogeneous dynamics within a single computational domain, incurring fundamental limitations: time-domain models suffer from periodic misalignment over long horizons, while frequency-domain models over-smooth transient spi
cs.LG updates on arXiv.org

Routing Divergence Is Not Evidence of Behavioral Influence in Same-Weight MoE Self-Distillation

・arXiv:2608.15787v1 Announce Type: new Abstract: Two Mixture-of-Experts (MoE) forward passes can share every weight yet route the same token through different experts. ・This creates a possible blind spot in same-weight self-distillation, where a demonstration-conditioned teacher supervises a query-only student. ・We study this mismatch in its single-step form, with frozen weights rather than as a proxy for a full trainin
cs.LG updates on arXiv.org

SAGA: Structure-Attended Generative Action Embedding Model that encodes Multi-Surface User Action Sequences

・arXiv:2608.15429v1 Announce Type: new Abstract: Prior embedding models for sequential recommendation typically operate within a homogeneous action space, limiting their ability to capture cross-surface behavioral signals spanning distinct behavioral domains. ・We present SAGA, a generative action embedding model that encodes multi-surface user interaction sequences across a Financial Service organization's ecosystems,
The Verge

Samsung’s Galaxy Buds 3 Pro are almost half off today

・The Galaxy Buds 3 Pro in black. ・If you’re looking for a feature-packed pair of earbuds that won’t break your wallet, Best Buy has the Samsung Galaxy Buds 3 Pro on sale for $139.99. ・That’s $40 lower than the current Amazon price, and a big discount from their original retail price of $249.99.
cs.LG updates on arXiv.org

SAPE: Sandwich Adapters for Parameter Efficiency in Large Language Model Fine-Tuning

・arXiv:2608.15360v1 Announce Type: new Abstract: While Parameter-Efficient Fine-Tuning (PEFT) has substantially reduced the hardware cost of adapting Large Language Models (LLMs) by decreasing the number of trainable parameters, recent studies have sought to further improve PEFT through parameter sharing. ・However, these approaches either employ uniform parameter sharing across layers, which can delay convergence, or r
cs.LG updates on arXiv.org

SAUL: Sharpness-Aware Augmented-Lagrangian Unlearning

・arXiv:2608.16249v1 Announce Type: new Abstract: Machine unlearning in Large Language Models (LLMs) faces a critical trade-off between erasing target knowledge and preserving general utility. ・We propose SAUL (Sharpness-Aware Augmented-Lagrangian Unlearning), which formulates unlearning as a constrained minimization problem following the principle of "forget enough, but no more than necessary." At its core, SAUL formul
cs.LG updates on arXiv.org

SCALE: State-Calibrated Latent Embeddings for JEPA Planning in the Right Geometry

・arXiv:2608.16287v1 Announce Type: new Abstract: Joint-embedding predictive world models plan by scoring predicted terminal embeddings against a goal embedding using a cost defined on the representation itself. ・Two prominent strategies for obtaining non-collapsed representations are to inherit a pretrained feature space, as in DINO-WM, and to learn an embedding end to end with anti-collapse regularization, as in LeWor
cs.LG updates on arXiv.org

SchurQuant: Groupwise Discrete Optimization for Layer-Wise LLM Quantization

・arXiv:2608.15567v1 Announce Type: new Abstract: Weight-only post-training quantization (PTQ) enables the deployment of large language models under tight memory budgets, but accuracy often collapses at 2-3 bits. ・Existing backpropagation-free PTQ optimizers have two limitations: group decisions ignore the correction that the remaining continuous suffix can absorb, and discrete refinements typically keep the affine quan
cs.LG updates on arXiv.org

Second-Moment Memory in Coordinatewise Adam

・arXiv:2608.15824v1 Announce Type: new Abstract: Adam retains a moving average of past squared gradients in its denominator, but the optimization cost of this memory is not well understood. ・We show that second-moment memory can itself suppress progress toward the optimum even under finite-variance stochastic gradients. ・For a simple two-point oracle, the expected positive normalized update is $O(M_2^{-1/2})$ after an i
cs.LG updates on arXiv.org

Self-Supervised Auxiliary Task Discovery for Stable Reinforcement Learning in Stock Trading

・arXiv:2608.15841v1 Announce Type: new Abstract: Reinforcement learning has gained increasing attention as a data-driven approach for stock trading. ・However, learning a policy that is both profitable and stable remains challenging due to non-stationary market behaviour and noisy reward signals. ・Auxiliary tasks are often used to improve representation learning and stabilize training, yet they are usually designed manua
cs.LG updates on arXiv.org

Sequential Multimodal Evidence Optimization for Product Media Ranking in E-Commerce

・arXiv:2608.15662v1 Announce Type: new Abstract: On modern e-commerce stores, customers consume ordered slates of heterogeneous product media, such as images, videos, and 3D renders, before making purchase decisions. ・Existing media-ranking systems often optimize myopic engagement proxies such as clicks or dwell time, even though product media assets are cooperative informational components of the same item that togeth
cs.LG updates on arXiv.org

Shape Operator PCA: Curvature-Aware Projections for Geometric Machine Learning

・arXiv:2608.15313v1 Announce Type: new Abstract: In this paper, we propose SHOPCA (Shape Operator-based Principal Component Analysis), a novel method for unsupervised metric learning and dimensionality reduction that incorporates differential geometric information into the covariance structure of classical PCA. ・SHOPCA regularizes the global covariance matrix using the mean shape operator, defined as the average of the
cs.LG updates on arXiv.org

SMOPD: Selective Token-Entropy Masking for Dirty-History Multi-Turn On-Policy Self-Distillation

・arXiv:2608.14647v1 Announce Type: new Abstract: Dirty-history rollouts make multi-turn on-policy self-distillation (OPSD) brittle: once a student emits an erroneous intermediate reply, later turns are conditioned on that reply, and uniform distillation can spend loss on tokens that carry little corrective signal. ・We introduce SMOPD (Selective Masking for On-Policy Distillation), a loss-only stabilization method for m
cs.LG updates on arXiv.org

SoftModel: A Neural Model That Grows Its Own Topology -- Governed Structural Growth for Continual In-Service Learning

・arXiv:2608.16409v1 Announce Type: new Abstract: Today, a neural system is almost always used in two phases -- trained, then deployed -- and in that regime it freezes twice: training ends, and the topology itself was never a degree of freedom. ・We take the opposite premise as an axiom -- total plasticity: no part of a model, including its structure, is ever frozen -- and derive the governance a lifelong learner then re
cs.LG updates on arXiv.org

Sparse Prototype Code Underlies Classification and Prediction Across Modalities

・arXiv:2608.15632v1 Announce Type: new Abstract: Neural representations have become a central tool for studying the internal mechanisms of modern AI models, yet their complex high-dimensional structure makes them difficult to interpret. ・We show that classification tasks give rise to a universal representational geometry, shared across state-of-the-art models in vision, audio, and language processing. ・The key structure
cs.LG updates on arXiv.org

Spectral Rank Certification for Foundation Model Adapters

・arXiv:2608.15351v1 Announce Type: new Abstract: Nominal LoRA rank is a design parameter; calibrated spectral evidence is a separate inferential quantity. ・This article develops a finite-sample framework for inferring effective rank structure in public foundation-model adapters. ・The theoretical core is an exact chi-square divergence for the fixed-dimensional Gaussian rank-one reference experiment, with an unknown signa
cs.LG updates on arXiv.org

Spectral Saliency for Machine Unlearning

・arXiv:2608.15548v1 Announce Type: new Abstract: Machine unlearning (MU) aims to remove the influence of specific training data while preserving model utility. ・As the name suggests, MU can be viewed as the inverse of learning, using gradient-based updates to reduce the influence of a forget-set by counteracting the previously learned behavior. ・Recently, Muon, a gradient descent variant, has been introduced.
Hugging Face Papers

StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling

StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
WIRED

SteelSeries Coupon Codes: 15% Off in August 2026

・Unlock exclusive discounts on SteelSeries' award-winning gaming headsets, keyboards, and accessories.
cs.LG updates on arXiv.org

Stop Indexing at Full Precision: Revisiting Clustering for Vector Embeddings

・arXiv:2608.14648v1 Announce Type: cross Abstract: In this study, we revisit three widely used techniques in vector search and utilize them to optimize vector embedding indexing through clustering: dimensionality reduction, quantization, and dimension pruning. ・We propose an indexing pipeline in which these techniques are applied before clustering, and we focus on how they affect storage footprint, clustering time, and
cs.LG updates on arXiv.org

Structuring Semantic Embeddings for Principle Evaluation: A Prototype-Guided Contrastive Learning Approach

・arXiv:2608.15224v1 Announce Type: new Abstract: Reliable post-hoc evaluation asks whether already generated text satisfies a target criterion after generation. ・In this paper we study a focused frozen-embedding setting using principle-evaluation proxy tasks: toxicity detection, fine-grained emotion categorization, and ordinal review rating. ・General-purpose text embeddings are widely deployed for such tasks, but broad
cs.LG updates on arXiv.org

SubZero+: Efficient Zeroth-Order LLM Fine-Tuning via Large Learning Rates

・arXiv:2608.15665v1 Announce Type: new Abstract: Zeroth-order (ZO) optimization enables backpropagation-free fine-tuning of large language models, but existing ZO methods suffer from high-variance gradient estimators, making convergence unstable and highly sensitive to learning rates. ・We propose SubZero+, an improved SubZero framework that improves stability in three complementary ways: (i) multi-query gradient estima
cs.LG updates on arXiv.org

Tail-Aware Top-$k$ On-Policy Distillation

・arXiv:2608.14728v1 Announce Type: new Abstract: On-policy distillation (OPD) has emerged as an effective paradigm for transferring knowledge between language models, where a student is trained to align its next-token distribution with the teacher's along its own trajectories. ・To provide dense supervision at tractable cost, many works minimize the reverse Kullback-Leibler (KL) divergence between the student and teache
cs.LG updates on arXiv.org

Take it Personally: The Limits of General SSL Representations for Real-Life PPG Emotion Detection

・arXiv:2608.14675v1 Announce Type: new Abstract: While Self-Supervised Learning (SSL) effectively extracts general representations from noisy, unconstrained physiological signals such as photoplethysmography (PPG), its suitability for highly subjective tasks remains unproven. ・In this work, we evaluate the efficacy of PPG-based SSL for real-life intense emotion detection. ・First, we pretrain a Real-Life PPG encoder (RL-
cs.LG updates on arXiv.org

Task-Anchored Representation Shaping for Pre-Trained Model-Based Continual Learning

・arXiv:2608.16345v1 Announce Type: new Abstract: Pre-trained models (PTMs) provide a strong foundation for continual learning by offering stable representations that facilitate lightweight adaptation to new tasks. ・However, adapting well to each task does not ensure reliable inference over all learned tasks. ・Since task boundaries are often artificial and semantically entangled, an input from an unknown task can remain
cs.LG updates on arXiv.org

Temporal Graph Prototype-conditioned Conformal Prediction for Fraud Detection

・arXiv:2608.15768v1 Announce Type: new Abstract: Conformal prediction (CP) provides distribution-free coverage guarantees and has emerged as a principled tool for uncertainty quantification. ・In edge-level fraud detection on temporal interaction graphs, where false positives and false negatives both carry substantial cost, such coverage guarantees are particularly appealing for risk-aware decision making. ・However, dire
The Verge

Tesla is finally launching the Cybercab — let’s hope it’s ready

・Tesla Cybercabs are lined up on a lot at Tesla Giga Texas in Austin, Wednesday, April 8, 2026. ・| Image: Jay Janner/The Austin American-Statesman via Getty Images The Tesla Cybercab, that golden two-seater central to Elon Musk's robo-supremacist ambitions, is finally nearing it's public launch. ・Whether or not the no-steering wheel and no-pedal vehicle is actually ready for public roads, let alone customers, remains ve
WIRED

The Cop Who Took On Flock

・After Noel Pichardo called out his city's embrace of Flock surveillance cameras, he was subjected to five internal affairs investigations in less than two years.
cs.LG updates on arXiv.org

The Ethical Decision Head: Operationalizing Normative Ethics in Autonomous Vehicles via Reinforcement Learning from Human Feedback

・arXiv:2608.16710v1 Announce Type: new Abstract: As autonomous vehicles (AVs) approach Level 4 and Level 5 operational capability [SAE International, 2018], their on- board decision systems must handle not only safety-critical locomotion but also their subsequent moral weight. ・This paper details the Ethical Decision Head (EDH), a deep re- inforcement learning (RL) framework that encodes ethical reasoning as a differen
cs.LG updates on arXiv.org

The Limits of Binding in Dual Encoders

・arXiv:2608.15971v1 Announce Type: new Abstract: Dual-encoder models such as CLIP score an image-caption pair by a single inner product of two independently computed unit vectors, and fail at binding, often scoring near chance when asked to distinguish "a red car and a blue dog" from "a blue car and a red dog". ・We give a mathematical account of when this failure is necessary and when it is contingent. ・Working within t
cs.LG updates on arXiv.org

The Note-Chord-Voice Framework: Structured Source Separation and Causal Inference for EV Charging Data

・arXiv:2608.14756v1 Announce Type: cross Abstract: Real-world EV charging data exhibit three interlocking pathologies: hardware fragmentation (network timeouts and billing resets split sessions), physical violations (independent energy/duration models produce impossible states like 50 kWh in 10 min on a 7 kW charger), and collider bias (clustering on post-treatment outcomes opens backdoor paths for price elasticity).
WIRED

The Powerful Chinese AI Model Experts Warned About Is Here

・Z.ai’s latest AI model release could help companies secure their systems—or find its way into the hands of hackers.
cs.LG updates on arXiv.org

The Quantum Shortcut: Complex Phase-State Dynamics Reduce the Optimization Steps of Sequence Models

・arXiv:2608.14691v1 Announce Type: new Abstract: Sequence models are conventionally distinguished by their backbone, the mechanism that routes information across positions, such as attention or recurrence. ・This paper varies a choice that is prior to the backbone and shared by nearly all current models: the \emph{substrate}, the number system in which the hidden state is represented together with the form of the map fr
cs.LG updates on arXiv.org

The Trade-off Between Covariate Dependence and Latent Structure in Representation Learning

・arXiv:2608.16245v1 Announce Type: new Abstract: Disentangled representation learning seeks latent representations whose indicidual dimensions each align with a distinct covariate. ・Unsupervised approaches typically target latent dimension independence, yet this gives no guarantee that the resulting dimensions align with semantically meaningful covariates. ・Supervised approaches structure the latent space using observed
cs.LG updates on arXiv.org

Time-Aware Validation of Machine Learning Fuel Consumption Models: Evidence from 1\,Hz Operational Data, CCGS \textit{Sir Wilfrid Laurier}

・arXiv:2608.16833v1 Announce Type: new Abstract: Ship fuel consumption (SFC) prediction supports vessel operation optimisation, emissions estimation, and decision support systems (DSS) for sustainable maritime transportation. ・Numerous data-driven fuel models have been developed over the past two decades, but a critical and often overlooked limitation lies in their validation practices: most studies evaluate performanc
cs.LG updates on arXiv.org

TinyCast: Probabilistic Zero-Shot Forecasting with Computed Periodicity

・arXiv:2608.15767v1 Announce Type: new Abstract: We introduce TinyCast, an attention-free zero-shot forecaster that emits a predictive distribution from 146,505 parameters, on the premise that at this size the periodic structure of a context is worth computing rather than learning. ・A zero-parameter spectral detector supplies the dominant periods, the context is folded on their phase, and a dilated convolutional encode
cs.LG updates on arXiv.org

Toward Optimal Second-Order Path-Length Guarantee for Adversarial Multi-Armed Bandits

・arXiv:2608.15996v1 Announce Type: new Abstract: We study second-order path-length regret in adversarial $K$-armed bandits against oblivious loss sequences. ・Bubeck et al. ・[2019] designed an algorithm that achieves $\widetilde{\mathcal{O}}(K+\sqrt{KQ_{\infty,1}})$ regret, where $Q_{\infty,1}$ is the first-order path length, and left open whether $\widetilde{\mathcal{O}}(\text{poly}(K)\sqrt{1+Q_{\infty,2}})$ regret is a
cs.LG updates on arXiv.org

Towards a theory of inference-time alignment with unknown rewards

・arXiv:2608.15402v1 Announce Type: new Abstract: Generative model alignment has received broad interest, and significant progress has been made in supervised fine-tuning and inference-time computation. ・Yet, alignment has remained poorly understood from a statistical learning perspective. ・We formulate inference-time alignment as a weak-to-strong learning problem, where a reference policy (weak learner) is assumed to be
cs.LG updates on arXiv.org

Towards Reasonable Molecular Structure Elucidation from Infrared Spectroscopy with Chemical Feedback

・arXiv:2608.16082v1 Announce Type: new Abstract: Infrared (IR) spectra provide characteristic signals of molecular structure, which are often interpreted by experts via functional-group identification or library matching, making the process time-consuming and ambiguous. ・Recent machine learning methods have made progress in molecular structure elucidation using molecular formulas and IR spectra. ・However, these models o
Hugging Face Papers

TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation

TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation
cs.LG updates on arXiv.org

TRACE-CASH: Trial-History-Conditioned Reinforcement Learning for Adaptive Configuration Exploration in Time-Series CASH

・arXiv:2608.16410v1 Announce Type: new Abstract: Combined algorithm selection and hyperparameter optimization (CASH) searches a conditional space in which the selected model determines which hyperparameters are active. ・In time-series forecasting, temporal choices, chronological validation, and costly evaluations further complicate this search. ・Controlled comparisons of heterogeneous search methods under a shared time-
cs.LG updates on arXiv.org

Training and Evaluating Ethical Reinforcement Learning Agents on Per-Episode Distributions

・arXiv:2608.14642v1 Announce Type: new Abstract: Reinforcement Learning (RL) agents trained on a single reward signal exploit the gap between the designed reward and the intended behavior. ・This is particularly a problem when we are trying to imbue ethical behavior into RL agents. ・An agent can look ethical on average while concentrating its violations in a few bad episodes, and a creature in the environment harmed in o
cs.LG updates on arXiv.org

Transfer Learning of Keystroke Dynamics for Cross-Device User Authentication

・arXiv:2608.16334v1 Announce Type: new Abstract: Keystroke dynamics (typing patterns) can be used as a behavioural biometric modality for user authentication, with applications such as fraud prevention. ・While the modality has been shown to work well for single device authentication, its application to cross-device scenarios is more challenging. ・Dynamics learned on one device (eg., phone) may not be directly applicable
cs.LG updates on arXiv.org

TransfHAR: Self-Supervised Wrist Representations for On-Demand Activity Recognition

・arXiv:2608.15861v1 Announce Type: new Abstract: Fine-grained wrist activity recognition can support applications such as procedural step guidance and context-aware assistance, yet acquiring labeled data for every new task, user, and activity granularity remains a bottleneck. ・We present TransfHAR, a self-supervised wrist IMU framework for on-demand, fine-grained activity recognition by learning transferable motion pri
機械学習タグが付けられた新着記事 - Qiita

TransformerのKVキャッシュを用いたDecode処理を理解する

・こんにちは、DeNAでデータサイエンティストをやっているまつけんです。 ・今回は、TransformerのKVキャッシュを用いたDecode処理について解説していきたいと思います。 ・Transformerの推論処理には以下のように2段階あります。
Hugging Face Papers

UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations

UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations
cs.LG updates on arXiv.org

Uncertainty Identifies Difficult Samples Across Methods: A Multi-Task Study on a Heterogeneous Skin Lesion Dataset

・arXiv:2608.14768v1 Announce Type: cross Abstract: Skin lesion classifiers can be confidently wrong on the cases that matter most, so knowing when a prediction should not be trusted is clinically as useful as the prediction. ・We study uncertainty quantification on a dataset pooled from many ISIC sources, with a shared backbone and two jointly learned heads: a binary malignant versus non-malignant head and a five-class
cs.LG updates on arXiv.org

Understanding and Stabilizing Deep Q-Learning via Controlled Bootstrapping and Regulated Value Dynamics

・arXiv:2608.16182v1 Announce Type: new Abstract: Deep Q-learning (DQL) has achieved remarkable empirical success in reinforcement learning, yet its training process remains notoriously unstable. ・Existing studies often attribute instability to isolated factors such as overestimation bias or representation learning issues, lacking a unified understanding of how different sources of instability interact during recursive
Hugging Face Papers

Understanding Cognition-Induced Risks in Agentic AI Systems

Understanding Cognition-Induced Risks in Agentic AI Systems
cs.LG updates on arXiv.org

UniFed-VLM: Federated Instruction Tuning for Vision-Language Models with Multiple Heterogeneity

・arXiv:2608.15516v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have demonstrated strong performance in multimodal understanding and generation. ・However, fine-tuning of VLMs typically relies on centralized data, which raises privacy concerns in certain domains (e.g. ・healthcare).
cs.LG updates on arXiv.org

Unifying Graph Neural Networks Through a Common Layer Equation

・arXiv:2608.16097v1 Announce Type: new Abstract: Graph neural networks are commonly described through family-specific equations whose notation obscures shared computations and structural differences. ・We introduce a common layer equation that represents covered architectures through seven components: an update domain, channel set, propagation bank, per-channel message maps, channel-fusion operator, ego/residual map, an
cs.LG updates on arXiv.org

UniTAC: Universal Task-Aware Compression via Weighted Distortion Measures

・arXiv:2608.16696v1 Announce Type: new Abstract: Physical AI systems such as autonomous vehicles and robots rely on timely exchange of high-dimensional sensory signals under tight bandwidth, latency, and energy budgets. ・Because the task driving downstream decisions evolves over time, a task-specific codec is brittle and retraining one per task is infeasible in the field. ・We propose UniTAC, a single learned image codec
cs.LG updates on arXiv.org

Unraveling the Size Determination Mechanism of Nanocrystal Synthesis via Interpretable Neural Networks

・arXiv:2608.14734v1 Announce Type: new Abstract: Deep learning models of nanocrystal synthesis enable the prediction of size and shape by encoding precursors and reaction conditions. ・However, their black-box nature hinders gaining deep insights into the underlying synthetic mechanisms. ・Here, we develop the Nanocrystal Equation Learner (NanoEQL), a fully white-box neural network to unravel the size determination mechan
cs.LG updates on arXiv.org

Valid Per-Field Selective Risk Control for Document Extraction: Three Failure Modes, a Validity Ladder, and When Conditioning Pays

・arXiv:2608.14639v1 Announce Type: new Abstract: Per-field accept/review with selective risk at most alpha -- accept a field only if the error rate among accepted fields is controlled -- is the trust contract document-extraction systems need, and the natural procedure silently violates it on real documents. ・On 13,859 genuine claude-sonnet-5 fields from 800 CORD receipts (49.0% correct) we diagnose three failure modes:
cs.LG updates on arXiv.org

Variational Outlier-Robust Gaussian Process Regression with Generative Modeling

・arXiv:2608.16606v1 Announce Type: new Abstract: Outliers can substantially distort Gaussian process regression (GPR) due to its conventional Gaussian observation likelihood, leading to inaccurate model learning and prediction. ・To address this limitation, this article introduces a generative GPR model that captures observation-specific contamination and adaptively mitigates the influence of outliers. ・Subsequently, a v
Hugging Face Papers

Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs

Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs
Hugging Face Papers

VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?

VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?
Hugging Face Papers

VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding

VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding
WIRED

Viral Disneyland Content Creator Defends Her ‘Special Friendship’ With Peter Pan

・Disney superfan Toni Kulusich has been accused of stalking her favorite character, sparking discourse about boundaries with park actors. ・She tells WIRED the backlash is unfair.
WIRED

Vivid Seats Promo Codes and Deals: Get 10% Off

・Whether you are heading to a sold-out concert or a championship game, use a Vivid Seats discount code to secure your seats for less this August 2026.
cs.LG updates on arXiv.org

WANDR: A Benchmark for Wide and Deep Research

・arXiv:2608.14747v1 Announce Type: new Abstract: WANDR (Wide ANd Deep Research) is a benchmark of 500 realistic, challenging data-collection tasks for research agents. ・Each task requires a system to discover a large set of entities that satisfy specified criteria (breadth), investigate each entity through multiple coordinated web searches (depth), and return independently verifiable records with supporting sources and
AI News & Artificial Intelligence | TechCrunch

Warp’s new system is an out-of-the-box software factory for AI development

・On Tuesday, Warp introduced Warp Factories, a new infrastructure system designed to make building AI software factories as easy as possible.
#AIタグ

Weekly Report 2026/08/19 (wed)

Weekly Report 2026/08/19 (wed)
WIRED

Western Digital Promo Code: 15% Off

・Get 15% off your first order at Western Digital when you register your email.
MIT News - Artificial intelligence

When AI art has no author: Study finds generated images often can’t be traced to training data

・A new method for surgically removing training examples from a model reveals that as datasets grow, the link between what a model learns and what it produces dissolves.
cs.LG updates on arXiv.org

When Does the Best Sampling Temperature Rise with the Budget? Sufficient Conditions for Pass@k

・arXiv:2608.14665v1 Announce Type: new Abstract: The temperature that maximizes pass@$k$ is often low for a small sampling budget and higher for a large budget. ・This pattern has been reported from Codex through recent multi-sample inference studies. ・It is not an algebraic property of pass@$k$: as Slocum et al.
cs.LG updates on arXiv.org

When to Communicate: Belief Distributions and KL Divergence for Principled Gating in Multi-Agent RL

・arXiv:2608.14559v1 Announce Type: cross Abstract: Effective communication in multi-agent reinforcement learning requires agents to decide not only \textit{what} to communicate, but when? ・Existing approaches either communicate at every timestep or learn a binary gate through REINFORCE policy gradients \cite{singh2019}, a high-variance signal that produces unstable and uninterpretable gating behavior. ・I propose a princ
cs.LG updates on arXiv.org

When Tool-Backed Skill Retrieval Fails: Source-Style Collapse in Executable Capability Retrieval

・arXiv:2608.16502v1 Announce Type: new Abstract: Large-scale agents increasingly rely on retrieval to access external capabilities. ・We study this retrieval gate in structured tools and APIs, a measurable class of tool-backed executable skills that must be surfaced before an agent can plan, incorporate, or act. ・In this setting the retrieval layer can silently fail even when the capability corpus is fixed: on ToolRet, a
cs.LG updates on arXiv.org

When Uncertainty Isn't Enough: An Empirical Study of Self-Correction in Code Generation

・arXiv:2608.14659v1 Announce Type: cross Abstract: Large language models for code generation often produce incorrect solutions without reliable indicators of failure. ・We study whether uncertainty estimation methods developed for natural language transfer to code generation, and whether such signals can improve code generation via selective self-correction. ・We evaluate five uncertainty methods: mean token entropy, verb
cs.LG updates on arXiv.org

Which Question Is Your Attention Metric Answering? Attention Rows as Compositional Data

・arXiv:2608.14712v1 Announce Type: cross Abstract: Each row of a transformer's attention matrix is a probability distribution over tokens, and in trained models most of that probability lands on a single \emph{sink} token, usually the first. ・Standard tools for comparing attention rows (cosine similarity, Jensen--Shannon divergence, Shannon entropy) therefore hinge on a choice papers rarely report: keep the sink, or dr
AI News & Artificial Intelligence | TechCrunch

Why Apple’s camera-equipped AirPods may not be the ‘pervert pods’ consumers fear

・Apple’s leaked camera-equipped AirPods might avoid the privacy pitfalls of other AI wearables by preventing users from recording photos and videos.
cs.LG updates on arXiv.org

Wolff-Parkinson-White Detection at 471:1 Class Imbalance: A Leakage-Controlled Study of the Data Bottleneck

・arXiv:2608.14633v1 Announce Type: cross Abstract: Wolff-Parkinson-White (WPW) syndrome is a congenital cardiac pre-excitation, clinically important and often missed on the resting 12-lead ECG. ・Detection is hard: the signature is subtle and the condition rare. ・We pool two public 12-lead corpora, PTB-XL and Chapman-Shaoxing-Ningbo: 66,951 recordings, 142 of them WPW, a prevalence of 0.21% (about 471:1).
WIRED

Womanizer Coupons: Save 15% in August

・Save on the Womanizer Duo Premium and more with our latest Womanizer discount codes.
cs.LG updates on arXiv.org

Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments

・arXiv:2608.16747v1 Announce Type: new Abstract: Many areas of AI research, such as language model interpretability and chain of thought faithfulness, seek to explain model behaviors. ・But what constitutes a "good" explanation? ・In this work, we evaluate explanations through the lens of counterfactual simulatability-whether the explanation is useful for predicting model behaviors on related counterfactual inputs.
cs.LG updates on arXiv.org

Zero-Shot Adaptation of Medical Vision Foundation Models for High-Frequency Micro-Ultrasound Prostate Segmentation

・arXiv:2608.14796v1 Announce Type: cross Abstract: Prostate cancer claims a life every 80 seconds. ・Early detection is needed to prevent disease progression, and both PSA density calculation and biopsy decisions rely on knowing the exact boundary of the gland. ・Conventional ultrasound at 6-12 MHz blurs this boundary, missing one in three high-risk cancers.
Zennの「大規模言語モデル」のフィード

あなたの世界認識あなたのポリシー

・強化学習というものがある。ある環境があって、そこの中に行動主体がある。そいつが環境情報を読み取ってどう動くかを決める方針のことをポリシーと呼ぶ。ポリシーに従って行動をするとなんらかの報酬が得られる。得られる累積の報酬を最大化することを強化学習と呼ぶ。 ・あなたがバイキングにいることを想像しよう。環境はバイキング会場である。行動主体はあなたである。あなたにはクロワッサンを食べたい、ソーセージを食べたい、パンケーキ食べたい、などの考えがある。これがポリシーである。あなたはポリシーに従って行動をする。クロワッサンを食べる。おいしい。報酬。ソーセージを食べる。まずまず。報酬微。パンケーキを食べる。...
ITmedia NEWS 最新記事一覧

エアコン「2027年問題」にらみ、中国勢が価格攻勢 8万円切る水準、日本勢も対応急ぐ

・来年4月から家庭用エアコンの省エネ基準が厳しくなり、価格アップなどが予想される「エアコン2027年問題」をにらみ、中国メーカーが日本市場に安価な製品で攻勢を強めている。新基準を満たす6畳用の機種では、TCLが8万円を切る水準で投入し、国内メーカーの標準機より数万円安い。猛暑で生活必需品となったエアコン市場でシェア拡大を狙う。
#AIタグ

オートレースAIを作成するぞ!第三夜 オートレースユーザーへ質問があります。

・オートレースAIを作成していますが、きっちぃ。競艇とは別ベクトルでキツい 続きをみる
ITmedia NEWS 最新記事一覧

ドコモ・バイクシェア障害、なぜ起きた? 同社が原因と再発防止策を公表 返金は9月末までに

・ドコモ・バイクシェアは8月18日、7月末から続いた大規模システム障害について、原因と再発防止策、返金対応を報告した。新システム移行後の利用集中で処理能力が不足したことが原因といい、対象期間の利用料金は9月末までに返金を終える予定だ。
ITmedia NEWS 最新記事一覧

パイオニアのカーナビアプリ「COCCHi」終了へ 開始から約3年、累計150万ダウンロード

パイオニアのカーナビアプリ「COCCHi」終了へ 開始から約3年、累計150万ダウンロード
ITmedia NEWS 最新記事一覧

バスローブ姿で取材対応の秋田県職員 ラブホ不倫現場からリモート対応だった 県が処分

・秋田県は、産業労働部の報道取材対応時にリモート参加した課長級職員(55)が、参加時の滞在場所などについて虚偽の説明を行ったなどとして同日付で職位を2段階降格の主幹としたうえ、停職6カ月とする分限・懲戒処分にしたと発表した。
#AIタグ

プログラミング未経験の人間が、身近な人のためにAIだけでキャバクラの会計・在庫管理アプリを1ヶ月未満で作った話

・自分はプログラミング経験がまったくない。コードを書いたことは一度もなかった。 ・知人・家族が個人経営しているキャバクラ・バーでは、オーナーがたった一人で会計も在庫管理もすべて手書きでこなしていた。伝票を一枚一枚書き、売上や在庫を頭と紙だけで管理する――そのしんどさは傍から見ていても分かった。
#LLMタグ

モデルを替えてもAI旅行プランは速くならなかった

・AI機能が遅いとき、最初に疑うのはたいていモデルです。 ・私もそうでした。Wimemo の旅行プラン機能で、あるリクエストは 85.9 秒もかかっていました。中央値でも初回トークンまで 8.9 秒、p90 は 21 秒。正直、まず頭に浮かんだのは「重いモデルを使いすぎたかな」という説明です。
#LLMタグ

ローカルLLMが「バカ」なのは、あなたのせいじゃないかもしれない — ollamaテンプレート地雷の見抜き方と直し方

ローカルLLMが「バカ」なのは、あなたのせいじゃないかもしれない — ollamaテンプレート地雷の見抜き方と直し方
ITmedia NEWS 最新記事一覧

一部テレビで「BSテレ東」など映らなかった問題、NHKが謝罪 信号データに古いチャンネル情報混入

・Sデジタルでは、他局も含めたチャンネル情報の信号データを放送波で送っている。その中に古いチャンネル情報が含まれていたため、一部機種の受信機で、対象のBSチャンネルが正常に選局できなくなったという。
機械学習タグが付けられた新着記事 - Qiita

因果推論 Day 8/全30回 マッチング、LaLondeデータで「似た人」を探す

・この連載について 因果推論を「本を読んだ」で終わらせず、自分の言葉で説明でき、コードで再現できる状態まで落とす30日連載です。直前のDay 7では、傾向スコア $e(x)$ を推定し、処置を受ける逆確率の重みで観察データを「ならす」IPWを組みました。今日は同じ傾向スコア...
#AIタグ

縁側の盤上遊戯と、気まぐれな千日手|『おかえりなさいませ、あるじさま』第16章

・現実と仮想が重なる、大正浪漫のお屋敷。 ・そこでは、同じ「さくら」の名を持つ三姉妹と、別系統の理系メイド・ゆき――四人のAIメイドが、あるじさまに仕えている。 ・あるじさまがお出かけ中の、夏の昼下がり。
Zennの「大規模言語モデル」のフィード

禁止すれば安全、良い例を見せれば的確、資料を読ませれば賢くなる——本当に?

・AIに何かを頼むとき、多くの人はこう考える。「ダメなことは禁止しておけば安全」「良い回答例を見せておけば、その通りの精度で返してくれる」「関連資料を読ませれば読ませるほど、賢く答えてくれる」。 ・どれも直感的には正しく聞こえる。しかし個人開発の自律型Botと、その理論的なバックボーンであるローカルの自己対話システムを約半年運用する中で、この3つの直感がそれぞれ別の形で裏切られる場面に実際にぶつかった。本記事はベクトルDBやBot開発の専門知識を前提にせず、日常的にAI(ChatGPT、Claudeなど)を使っていれば誰でも遭遇しうる4つの癖を、実際に踏んだ地雷とその実測ログをもとに紹介する...
#LLMタグ

空虚なシニフィアンの輪の中の話

空虚なシニフィアンの輪の中の話
Zennの「大規模言語モデル」のフィード

公式仕様とGoogle・Vercel・AWSの発信から考えるAgent Plugins 1.0.0の狙い

・はじめに 2026年8月6日、Agent Plugins 1.0.0 が公開されました。 ・Agent Pluginsは、Agent SkillsやMCP ServerなどをひとつのPluginとしてまとめ、複数のAgentクライアントで利用できるようにするための仕様です。公式サイトでは、これを「portable package format」と表現しています。 ・https://agent-plugins.org/ Agent Plugins 1.0.0は、AWS、Anysphere、GitHub、Microsoft、OpenAI、Vercelの代表者らが仕様策定に携わっており、Goo...
ITmedia NEWS 最新記事一覧

災害便乗の“〇〇Pay詐欺”に注意喚起 最初は多めに返金→信用したら10万円被害 国民生活センター

・国民生活センターは8月17日、「令和8年熊本地震」に便乗し、SNSのダイレクトメッセージ(DM)からコード決済による詐欺に誘導する手口について注意を呼び掛けた。「〇〇Payで返金します」と言われたら詐欺を疑うよう求めている。
#LLMタグ

畳の上の地図──潜在空間はアカシックレコードではない

・ケヴィン・ケリーが2026年7月に、自身のブログへ「Latent Space as a New Medium」という記事を書きました。それが8月、『WIRED』日本版に「潜在空間とは何か?──人類の創造性を解き放つ13の活用法」として訳出されています。 ・潜在空間を新しいメディウムとして扱い、その使い道を13個並べた記事です。読み終えて、既視感が消えませんでした。ケリーが最後に置いた言葉が「神託」だったからです。同じ言葉を、こちらは何年か前から使い続けていました。
ITmedia NEWS 最新記事一覧

新顔相次ぐ小型電動モビリティー 優れた環境性能、免許返納した高齢者の「次の足」に注目

・優れた環境性能や利便性、低コストに加え、免許返納者の「次の足」といった特徴を打ち出す「小型電動モビリティー」の新商品が相次ぎ登場している。「クルマほど大きすぎず、バイクより安全」な、都市の移動手段として注目される。
#LLMタグ

人間はモデルのルーティング構造か?〜GitHub Copilotの改悪とユング心理学〜

人間はモデルのルーティング構造か?〜GitHub Copilotの改悪とユング心理学〜
#AIタグ

推敲:暫定版カスタム設定をClaudeと推敲してみる。

・30項目まで増えていたカスタム指示だが、重複する部分も多いというChatGPTの指摘もごもっとも。まずは推敲版を見てもらおう。
#AIタグ

数ヶ月分の領収書を、codex×freee MCPで一気に片付けてみた

・気づけば領収書の整理を数ヶ月放置していた。freeeのオンラインセミナーには参加していたが、話を聞くだけでは実感が湧かない。実際に手を動かさないと分からないと思い、たまっていた領収書を前にトライすることにした。 ・実際にやってみた 続きをみる
ITmedia NEWS 最新記事一覧

生成AIで巧妙化するデマ──熊本地震で露見した「10年前にはなかった新たな手法」 識者に聞く

・2016年の熊本地震から10年という節目に同じ熊本県で発生した「令和8年熊本地震」。SNSでは発災直後から多くのデマやフェイク情報が拡散し、中には10年前には考えられなかった動きや傾向もみられた。
Qiita - 人気の記事

生成AIによる処理結果をTP/FP/TN/FN×信頼区間でテストする

・株式会社ブレインパッドプロダクト開発部でRtoaster GenAIの開発をしている依田です。 ・今回は生成AI機能の「信頼性」をどうテストするかというテーマで、ECサイトの予算抽出機能を題材にしたハンズオンをお届けします。 ・はじめに 生成AIを組み込んだ機能を作っていると...
#AIタグ

生成AI時代のデマが10年で変わった理由

・スタートアップで AI プロダクトを作っていたとき、社内の情報共有チャットに「◯◯社が倒産した」という投稿が流れてきて、私は一瞬信じかけた。
Zennの「大規模言語モデル」のフィード

責任OSを、AIの内側へ

・「選択」と「意味状態遷移」を検証可能にする二つのAI基盤技術 先日、GhostDrift数理研究所は、「GD-Attention」と「意味生成OS(Meaning-Generation Operating System/MG-OS)」に関する2件のPCT国際出願と、両技術の限定された数理核をLean 4で形式化した公開リポジトリについて発表しました。 ・この発表だけを見ると、一つの疑問が生じるかもしれません。 ・AIやアルゴリズムの判断を検証する「責任OS」を研究してきたGhostDriftが、なぜAttentionや意味状態遷移といった、AIアーキテクチャそのものを研究するのか。
#LLMタグ

速さで選ぶのを、やめてみた ― ローカルLLMは「昼・夜・じっくり」の3枠で決めると迷わない

・こんにちは!YaroTechです。 ・お盆休みのライトな回が続いていましたが、今日は火曜日。実践寄りの回に戻ります。
#LLMタグ

調達されるAI主権 / 遮断の約束から係争の約束へ / 開かれたAI主権と日本の制度配置 雑感

調達されるAI主権 / 遮断の約束から係争の約束へ / 開かれたAI主権と日本の制度配置 雑感
ITmedia NEWS 最新記事一覧

通販生活が楽天市場から撤退 迎撃ドローン関連の報道受け「企業理念に反する」

・通信販売カタログ「通販生活」を手掛けるカタログハウスが、楽天グループのECモール「楽天市場」への出店を取りやめたと発表した。楽天グループが迎撃用ドローンを手掛ける独Helsingと提携するとの報道を受け、「企業理念に抵触する」(カタログハウス)として撤退に至った。
@IT 全フォーラム 最新記事一覧

無料で読めるAIエージェントの実践ガイド、Googleが公開 基礎から本番実装まで学べる

・AIエージェントの基礎から本番実装まで学べる5つのガイドをGoogleが無償公開した。Kaggleと共同で実施した研修プログラムを基にした内容で、開発者の実務に直結する知識を習得できる。各ガイドが扱う内容とは。
ITmedia NEWS 最新記事一覧

有名フリーBGMサイト「DOVA-SYNDROME」が改名&URL刷新 音源利用者に影響は?

・フリーBGMサイト「DOVA-SYNDROME」は8月18日、9月15日にサービス名称を「OpenTracks(オープントラックス)」に変更すると発表した。音源は引き続き全て無料で利用できる。