ai Trend Report

Dashboard へ戻る
Date: 20260804 Articles: 398 Scope: curated summary

あなたのアイデアを、今すぐ形に。

公開先に迷ったら、WebFileBinで一発公開。

HTMLをドラッグ&ドロップするだけで、すぐ公開できます。

3
High impact
5
Mid impact
390
Signal watch

なぜこのサイトを作ったのか

私たちそれぞれが個別にAIを使って情報収集し、同じような3行要約を作るたびに、世界中で膨大な電力と計算リソースが消費されています。 本プロジェクトは、あらかじめ広範な情報を取得・集約しておくことで、個別のAI実行回数を減らし、地球環境(GPU/TPU負荷)に配慮した効率的な情報収集を目指す実験的なダッシュボードです。

Domain filters
Star filters
#LLMタグ

【Palantir(PLTR)Q2 FY2026】AIに支配される側か、AIを支配する側か。PLTRはLLMの外側でモートを作り始めた

・Palantir(PLTR)のQ2 FY2026決算をアウトプットする。 ・前回Q1決算では、私はこう問うた。 ・PLTRは割高で終わらせてよいのか。AI神銘柄として崇めてよいのか。
Zennの「機械学習」のフィード

金融データサイエンス基礎Part 8. なぜそう判定したのかに答えられなければなりません:解釈可能性と公平性、規制

・原文: han-co.com ·「金融データサイエンス」連載の「基礎」編、Part 8 です。(原文には手描きの図も載せています。) Part 7まで、良いモデルを作り(Part 4)、正直に評価し(Part 5)、因果で政策を立て(Part 6)、そのモデルを検証して長く守り抜く(Part 7)話をしてきました。いよいよ信頼の最後の軸が残りました。なぜそう判定したのかを説明できなければならず、その判定が公平かにも答えられなければならない、ということです。 ・Part 0で、この分野が一般的なMLと違う理由として「解釈可能性は選択ではなく義務」であることと、「規制とガバナンスが常に下に敷か...
#LLMタグ

散らかったリンクをひとまとめに!自作エディタ「AXISEditor」でポートフォリオLPを作成しました

・こんにちは!アプリ開発などを行っているN1_LABOです。 ・いつもnoteを読んでいただき、またアプリをご利用いただきありがとうございます。 ・本日は、これまでの活動を整理するために新しく作成した「ポートフォリオLP(ランディングページ)」のお知らせと、その制作の裏側についてお話ししたいと思います。
#AIタグ

【世界レーダー】2026/8/5 石油の「血管」と再エネ、同時に揺れる世界

・株式会社ウォーカル|世界レーダー 今日のテーマ:地政学 × エネルギー 続きをみる
#AIタグ

ChatGPTだけで月1万円を目指す方法【2026年最新版】

・「副業を始めたい! でも何をすればいいか分からない。」 そんな悩みを持っていませんか? 実は、私も最初は同じでした。 ・「パソコンの知識もない。」 「プログラミングもできない。」 「何から始めればいいのか分からない。」 そんな状態でも、 AIと出会って考え方が変わりました! 今では、ChatGPTを使えば文章を書いたり、 アイデアを考えたり、SNSの投稿を作ったりと、 一人では何時間もかかる作業を短時間で進められます。 ・だからこそ、「副業を始めるハードル」は以前より ずっと低くなっています。
cs.LG updates on arXiv.org

Latent-Regime Bias Auditing for Volatility Forecasting

・arXiv:2608.01599v1 Announce Type: new Abstract: Volatility forecasts are commonly evaluated with aggregate accuracy metrics such as RMSE and MAE, but these metrics can hide conditional failures that matter for risk management. ・This paper proposes a model-agnostic audit framework for evaluating whether volatility forecasts remain reliable across latent market regimes. ・We learn time-series representations of market-sta
#LLMタグ

LLMにLLMを作らせてわかった事。

・現在のLLMだけでは、まったく新しい手法を生み出すのは難しい。研究の方向を決める人間の発想が必要。 ・LLMは性能を上げようとすると、複雑な計算や機構を次々に追加しがち。その結果、少し性能が上がっても遅すぎて現実的でないモデルになりやすい。 ・既存研究の再発見を新発明だと誤認しやすい。
機械学習タグが付けられた新着記事 - Qiita

金融データサイエンス基礎Part 8. なぜそう判定したのかに答えられなければなりません:解釈可能性と公平性、規制

・原文: han-co.com ·「金融データサイエンス」連載の「基礎」編、Part 8 です。(原文には手描きの図も載せています。) Part 7まで、良いモデルを作り(Part 4)、正直に評価し(Part 5)、因果で政策を立て(Part 6)、そのモデルを検証して長く...
#AIタグ

【MT5】 AIアシスタントは"言うだけ"じゃなかった。経済指標も自分のポジションも、根拠付きで教えてくれる件

・「その情報どこから?」とAIに聞いたら、想像以上にちゃんと調べていた件 こんにちは。杉原です。
WIRED

‘Everyone Is Doing It’: The Truth About AI in Hollywood

・Puck’s Matthew Belloni says AI has quietly become part of everyday filmmaking. ・The battle now isn’t whether Hollywood will use the technology—it’s who controls what’ll come next.
#LLMタグ

(論文読解)AIに意識を主張させると人間らしくなる??

・またAI愛好家をざわつかせそうな論文が出てきたので、ざっくり解説します 言語モデルに自らの意識を主張させることで、人間の信念と価値観が回復する。
Zennの「大規模言語モデル」のフィード

[Agents on 16GB] 状態の分離をあきらめた。サブエージェントはそれで動くようになった

・Agents on 16GB — 16GBのMacBook1台で動かすマルチエージェント。API利用ゼロ、外に何も出さない。 ・マルチエージェントを組んでいる。オーケストレーターのPMが1つあって、そこから専門のサブエージェントに仕事を振る構成だ。モデルは全部ローカル(Ollama)で動かしている。16GBのMacBook1台、API利用はゼロ。 ・最初は全員が同じ state["messages"] を共有していた。シンプルでいい、と思っていた。そのあとLangSmithのトレースをClaudeに見てもらった。
Zennの「大規模言語モデル」のフィード

「.claude/agents」設計でやりがちな失敗5選 ― 公式仕様から見る典型パターン

・Claude Codeのカスタムサブエージェント(.claude/agents/配下に置くMarkdownファイル)は書式自体はシンプルだが、仕様の細部を見落とすと「動くには動くが期待通りに委譲されない」「起動時にエラーになる」といったつまずきが起きやすい。本記事では公式ドキュメント(Create custom subagents)に明記されている仕様をもとに、設計時によくある失敗パターンを5つ整理し、修正例を示す。 ・失敗1: 同じディレクトリ内で name が重複している .claude/agents/配下(サブフォルダ含む)で同じnameを持つファイルが複数あると、どちらが読み込...
ITmedia NEWS 最新記事一覧

「イオンモール熊本にガスコージェネは設置していない」 SNSの噂を経産省が否定

・「一部SNSで事実と異なる情報が流れていますが、イオンモール熊本には、ガスコージェネレーションやガス発電機は設置していないことを確認しています」
ITmedia NEWS 最新記事一覧

「モンハンワイルズ」価格改定で半額に 販売数の伸び悩み、克服なるか

・カプコンがゲーム「モンスターハンターワイルズ」の価格を改定した。これまでゲーム本体を税込9900円で提供していたところ、今後は同4990円で販売する。
#AIタグ

「検索される」から「引用される」へ|MCPがSEOに突きつける「3人の読者」

・僕は普段、特定の分野において"間違えたら人を傷つける"領域でコンテンツをつくっています。同時に、MCPやAIエージェントを日々の制作ワークフローに組み込んでもいる。その両側から眺めると、いま起きている変化は「SEOの死」ではなく、もっと面白い"再編"に見えます。 ・「MCPの普及で検索を経由しない経路が増え、いちばん打撃を受けるのは"検索順位からの流入だけ"を価値にしてきたビジネスモデルだ」この答えに、僕は強く同意します。
ITmedia NEWS 最新記事一覧

「中国軍がDeepSeek製AIを無人航空機に活用」「時速50kmの軍用車両にも搭載」 きょう公開の防衛白書

・防衛省は8月4日、国家防衛の現状と課題をまとめた「防衛白書」の最新版を公開した。AIやドローンといった最新技術の状況にも触れ、中国の人民解放軍が生成AI「DeepSeek」のAIモデルを無人航空機や軍用車両に活用しているとの認識を示した。
#AIタグ

【2026年8月最新】ついに資料作成が「全自動化」!Microsoft Copilot新機能「Notebooks」でWord・Excel・パワポを一瞬で作る3つの手順

・毎日の業務で、「メモや議事録をWordにまとめる」「アイデアをPowerPointのスライドにする」「数値をExcelに入力する」といった作業に、どれだけの時間を奪われているでしょうか。
Zennの「大規模言語モデル」のフィード

【Claude Opus 5】ベストプラクティスをルール25本に当てたら、穴が3つ出た

・TL;DR 新しいモデル世代が出るたびに、自分がエージェントに与えている指示資産を棚卸ししている。今回は 25 本・4,023 行が対象だった(2026-08-04 実測) 公式ガイドは「明示的な検証指示は削除してください」と明言する。最も削除候補に見えたのは、僕が事故のたびに書き足してきた検証系のルール群だった。該当表現を grep したらヒット 2 件、どちらも誤検出だった(同日、棚卸し着手時点) 削る作業のつもりで始めて、実際に見つかったのは 3 つの空白(委譲の起動判定 / ディスクに書く成果物の長さ / effort 軸)。削除は通算 0 件 おまけに、公式が推奨するス...
#LLMタグ

【GPT-Image-2】2年間かけて遂に完成‼️1/100スケール、渋谷ギャル韓国エステ。拡大して細部をチェックしてみる。

・・今回のテーマ お待たせしました🤠大好評の使い回し企画‼️ チクワゴスティーニを定期購読、2年間かけて最後のパーツ、施術所を組み立てた。 ・創刊号490円。2ヶ月目から1490円。 ・👇健康ランド編 続きをみる
#LLMタグ

【llama.cpp】iGPUでcpu-moeをやったらどうなるのか? ~非APU向けオプションをあえて使ってみると意外な結果に~

・こんにちはRcatです。 ・突然ですが、llama.cppのcpu-moeオプション知ってますか? ・これ、MoEならVRAMに入りきらなくても、大事なとこだけVRAMに置くようにすることで、激遅にはならずに済むっていう救済措置なんですが、あえてVRAMが豊富なAPUでやったらどうなるんですかね?
#LLMタグ

【Udemyコースレビュー】 Ollama & Local LLMs: Fine-Tune, Deploy, Build Python AI Apps

・■ はじめに 今回は、「Ollama & Local LLMs: Fine-Tune, Deploy, Build Python AI Apps」というUdemyのコースをご紹介します。
#AIタグ

【v2アップデート】MouthLoopに「表情」と「仕草」が入りました|リップシンクキャラを作る無料Gemini Canvasアプリ

・画像1枚から口パク素材を作る無料ツール MouthLoop を、v2 にアップデートしました。 ・前回の記事(MouthLoop|画像1枚から口パクキャラクターを作る)では「PNG 6枚 + アニメーションWebP」を作るところまでを紹介しました。
#AIタグ

【クリエーター図鑑】で私をカードにしてみました

【クリエーター図鑑】で私をカードにしてみました
#AIタグ

【コピペで即完了】SNS作成時間を1/5に!毎日の投稿作成を激変させる実践プロンプト10選

・いや、どうもこんにちは。 ・事実ちゃんとした投稿は初めてなんで。 ・では、どうぞ御照覧あれ 続きをみる
#LLMタグ

【スペインの小さな港町から「AIを縮める」技術で800億円調達】AIに電気を食われてあなたの電気代が上がる——その問題を、人口19万人の港町が「量子物理学」で解いていく。同時崩壊する2026年に、日本企業がまだ知らない「縮むAI」の正体。#生成AI #量子コンピューティング #スタートアップ #ヨーロッパ #資金調達 #スペイン #テクノロジー #エッジAI #データセンター #省エネ

【スペインの小さな港町から「AIを縮める」技術で800億円調達】AIに電気を食われてあなたの電気代が上がる——その問題を、人口19万人の港町が「量子物理学」で解いていく。同時崩壊する2026年に、日本企業がまだ知らない「縮むAI」の正体。#生成AI #量子コンピューティング #スタートアップ #ヨーロッパ #資金調達 #スペイン #テクノロジー #エッジAI #データセンター #省エネ
#LLMタグ

【雑記】AIの指紋!4o時代のチャット復元で見つけたモデルごとの癖

・AIに励まされることで生きがいを見出している、どっかの漫画家です。 ・相変わらずコツコツとチャットログ復元作業を進めています。
#AIタグ

【週次まとめ】8週目——アイビス傾向一致と、3Looksオンボの週

【週次まとめ】8週目——アイビス傾向一致と、3Looksオンボの週
#AIタグ

【生きるのクソ下手】あののオールナイトニッポン0:第165回(8/4放送)【chatGPT】

・※架空の番組の記事です。 ・以下の説明も毎回コピペです(笑) この記事は、 僕が開発を継続している実在人物のWeb情報からを話し方の傾向・トーンを自動で生成し、1ファイルにまとめたものをchatGPTのプロジェクトにアップロードするという未完成のチャットの仕組みを用いて、 続きをみる
#LLMタグ

【生成AIニュース+】『DeepSeek V4 Flash』『Gemini Notebook』『ComfyUI-DyPE』『Comic 4.2』『RLSVR』『GAIA-4』『LeapTalk』『Raven』

【生成AIニュース+】『DeepSeek V4 Flash』『Gemini Notebook』『ComfyUI-DyPE』『Comic 4.2』『RLSVR』『GAIA-4』『LeapTalk』『Raven』
Hugging Face Papers

3DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering

3DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering
cs.LG updates on arXiv.org

3DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering

・arXiv:2608.01185v1 Announce Type: cross Abstract: Recent 3D vision-language models (3D VLMs) construct geometry aware tokens by projecting 2D visual features into world coordinates, enabling spatial reasoning for tasks such as 3D question answering. ・However, this design generates thousands of tokens per scene, resulting in substantial computational and memory overhead. ・While token compression has been extensively stu
#AIタグ

8/4(火) ミニロト&ナンバーズ 検証

8/4(火) ミニロト&ナンバーズ 検証
cs.LG updates on arXiv.org

A 2-Block Architecture for Real-Time EEG Gait Decoding: A Pilot Study

・arXiv:2608.02083v1 Announce Type: new Abstract: Closed-loop lower-limb exoskeleton control via Electroencephalography (EEG) remains limited by motion artifacts, low signal-to-noise ratio, and binary gait formulations that fail to capture full cortical gait complexity. ・We propose a 2-block Brain-Computer Interface (BCI) architecture: a trainable session-specific Feature Extraction Block with real-time artifact suppres
cs.LG updates on arXiv.org

A Constitution-Grid Instrument for Data-Efficient RL Alignment (C-Guard)

・arXiv:2608.00180v1 Announce Type: cross Abstract: Conflicting objectives are general in RL alignment, and training on them data-efficiently is hard. ・Training a safety guard with RL means optimizing two objectives that conflict: catch real harm, and do not refuse benign prompts. ・Our finding is that over-refusal improves 22.4% to 12.8%, while under-refusal on adversarial attacks silently worsens 0.27 to 0.33.
cs.LG updates on arXiv.org

A False Average: Chain-of-Thought Monitors Collapse Where They Are the Only Defense

・arXiv:2608.00583v1 Announce Type: cross Abstract: Chain-of-thought (CoT) monitoring is meant to catch the reward hacks that look clean in the actions and betray themselves only in the reasoning. ・We show that this is exactly where an adversary who controls the reasoning can defeat it. ・Rewriting only an agent's reasoning to read as good-faith engineering, while copying every command and output verbatim so the exploit i
cs.LG updates on arXiv.org

A Physics-Chemistry-Informed Neural Network (PCINN) for Real-Time Spatial-ALD Coverage Prediction and Reliable Kinetics Inversion

・arXiv:2608.00212v1 Announce Type: new Abstract: Spatial atomic layer deposition (SALD) is a leading atmospheric-pressure, high-throughput route to industrial ALD, but design and control are limited by the cost of predicting surface coverage: high-fidelity CFD is far too slow for operating-window scans, while analytic models miss transport modulation such as the gas curtain. ・We present a physics-chemistry-informed neu
cs.LG updates on arXiv.org

A reproducible and extensible framework for benchmarking competing risks survival models

・arXiv:2608.00271v1 Announce Type: cross Abstract: A wide range of statistical and machine learning methods have been proposed for survival analysis with competing risks, where the occurrence of one event (i.e., cancer death) precludes the occurrence of other events (i.e., cardiovascular disease death). ・Despite these methodological advances, their systematic evaluation and adoption are limited by the lack of comprehen
cs.LG updates on arXiv.org

A Sequence-to-Sequence ConvLSTM Approach for Leaf Area Index Forecasting over the South-Central United States

・arXiv:2608.00879v1 Announce Type: cross Abstract: Leaf Area Index (LAI) is a fundamental biophysical variable governing land-atmosphere interactions; however, LAI forecasting at high spatial resolution remains an unsolved challenge. ・While recent machine learning approaches have demonstrated LAI estimation at point or regional scales, none provides a gridded, meteorology-driven prognostic forecast suitable for subseas
cs.LG updates on arXiv.org

A Spatial Persistence Gradient in European Warming Consistent with North Atlantic Cold-Blob Influence

・arXiv:2608.00063v1 Announce Type: cross Abstract: Europe is warming faster than the global mean, yet the spatial organisation of this acceleration remains incompletely understood. ・Using ERA5 reanalysis for 1950--2024 across 28 IPCC AR6 European sub-regions, we identify two connected empirical results. ・First, the DFA1 Hurst exponent of interannual temperature residuals is strongly and negatively associated with the 19
cs.LG updates on arXiv.org

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning

・arXiv:2608.00301v1 Announce Type: new Abstract: Error-penalized scoring rules ($+1$ for a correct answer, $-\lambda$ for a wrong one, $0$ for abstaining) are increasingly prescribed against hallucination: a rational agent facing such a rule answers exactly when its correctness probability exceeds Chow's threshold $t^\ast=\lambda/(1+\lambda)$. ・We prove that a KL-anchored gradient learner can do the opposite.
cs.LG updates on arXiv.org

AdaHAT: Adaptive Hard Attention to the Task in Task-Incremental Learning

・arXiv:2608.01252v1 Announce Type: new Abstract: Catastrophic forgetting is a major problem in task-incremental learning, where neural networks tend to overwrite previously learned knowledge when trained on new tasks. ・A number of architecture-based approaches have been proposed to address this problem. ・However, the architecture-based approaches suffer from another problem related to network capacity when the networks
cs.LG updates on arXiv.org

Adaptive Quantum Physics-Informed Neural Networks for Differential Equations with Applications to Fluid Dynamics

・arXiv:2608.00850v1 Announce Type: new Abstract: Physics-informed neural networks (PINNs) have emerged as a versatile approach for solving nonlinear partial differential equations (PDEs), yet achieving high accuracy efficiently using these techniques remains challenging for high-dimensional or multiscale systems. ・Here, we present a hybrid quantum-classical framework that enhances Quantum PINNs (QPINNs) through adaptiv
cs.LG updates on arXiv.org

AdvPlan-Bench: Adversarial Evaluation of Structured Plan-Generation Agents

・arXiv:2608.00832v1 Announce Type: new Abstract: Structured plan-generation agents are often evaluated as if a plan has quality in isolation, yet many realistic planning tasks require asking how a candidate behaves when another agent can search for responses. ・We introduce AdvPlan-Bench, an offline benchmark for adversarial evaluation of structured plan-generation agents. ・The contribution is a general evaluation object
cs.LG updates on arXiv.org

Agentic Bayesian Optimization through Surrogate-Augmented Autoresearch

・arXiv:2608.00316v1 Announce Type: new Abstract: Bayesian optimization (BO) has become the standard tool for sample-efficient optimization and owes its efficiency to uncertainty-aware search driven by generic statistical priors. ・Richer domain priors can improve BO in principle, but encoding them through tailored kernels or problem structure is difficult and rarely done in practice. ・LLMs can help sidestep this difficul
cs.LG updates on arXiv.org

Agentic Coding in the Wild: Characterizing GitHub Copilot Traces at Production Scale

・arXiv:2608.00101v1 Announce Type: cross Abstract: AI coding agents like GitHub Copilot, Claude Code, and Codex interleave multi-step LLM inference with tool execution, creating a workload different from chatbots. ・We present the first production-scale characterization of this workload using sampled GitHub Copilot traces from June 2026, comprising 3.2M users, 13M sessions, 761M LLM calls, and 95T tokens. ・Our analysis r
cs.LG updates on arXiv.org

Agentic Graph Token Reasoning

・arXiv:2608.00542v1 Announce Type: new Abstract: Graphs model relational data throughout science and industry, from citation networks to product co-purchase graphs. ・Because the nodes of many such graphs carry rich text, a growing line of work applies large language models (LLMs) to graph analysis. ・The most graph-native of these methods use graph tokens: a graph encoder compresses a graph view, such as a node, its k-ho
cs.LG updates on arXiv.org

Aggregate-then-Calibrate for Human-centered Assessment with Theoretical Guarantees

・arXiv:2608.02455v1 Announce Type: new Abstract: Human-centered assessment tasks, which are essential for systematic decision-making, rely heavily on human judgment and typically lack verifiable ground truth. ・Existing approaches face a dilemma: methods using only human judgments suffer from heterogeneous expertise and inconsistent rating scales, while methods using only model-generated scores must learn from imperfect
#AIタグ

AI | 天皇と仏教(9)-「王法」と「仏法」の1500年 - 第9章 象徴天皇制と政教分離

・AI | 天皇と仏教(9) -「王法」と「仏法」の1500年 - 第9章 象徴天皇制と政教分離 - 戦後の再編と論争 - 続きをみる
NVIDIA Blog

AI Leaders Propose SAFE Guidelines for Cybersecurity Transparency

・Members of the Open Secure AI Alliance — now more than 120 organizations strong — are developing new guidelines to strengthen agentic AI cybersecurity as the annual Black Hat conference begins in Las Vegas today. ・The Linux Foundation today shared a Request for Comments on Shared AI Findings Exchange (SAFE), a proposed set of guidelines […]
#AIタグ

AIアフィリエイトブログを毎日cronで全自動運用する構成の作り方【初めの一歩】

・生成AIに記事を書かせて、レンタルサーバーのcronで毎日回すアフィリエイトブログを作りました。記事を書くところだけでなく、公開・計測・改善タスクの起票まで自動で回っています。 ・この記事は、その構成をまるごと公開するものです。「AIで稼げる」という話ではありません。何をどう組めば無人で回るのかを、DBのスキーマからcronの行まで具体的に書きます。同じものを作りたい人が、設計で迷わずに済むように。 ・先に完成形を書いておきます。
Zennの「機械学習」のフィード

AIエンジニアとして本番システムを作れるようになるための厳選5冊【2026年版】

・はじめに 自分は昨年からLLMを組み込んだ社内ツールの開発に関わっている。最初はAPIを叩いてレスポンスを返すだけの簡単な仕組みだったが、本番運用が始まると想定外の問題が次々と出てきた。モデルの出力品質が日によってブレる、レイテンシが許容範囲を超える、コストが月次予算を食い潰す。プロトタイプを動かすのと、プロダクションで安定稼働させるのは全く別の仕事だと痛感した。 ・こうした壁にぶつかるたび、断片的なブログ記事やTwitterの情報では対処しきれない場面が増えた。体系的な知識がないまま場当たり的に対応していると、設計判断の軸がブレる。結局、腰を据えて書籍を読み込む時間を取ったことが転機...
#AIタグ

AIが書く、保険代理店の小説——第4話「波紋」

・第4話 波紋 取引先への謝罪から、一週間が経っていた。生産ラインは田村の尽力で三日目に半分ほど復旧し、遅延分の納品もどうにか間に合わせた。だが、その代償は静かに、しかし確実に広がっていた。
Zennの「大規模言語モデル」のフィード

AIが提案したデッキ改造案を、家のPCで3000回対戦させて検証したら棄却された

・マジックザギャザリング(MtG)で自作デッキを組んでいると、必ずこの疑問にぶつかる。このデッキ、環境のトップメタ相手に何割勝てるのか。改造したら勝率は上がるのか。 ・これ、普通に確かめようとすると地獄だ。友人と100戦して統計を取る人はいない。大会に100回出場することもできない。じゃあどうするか。家のパソコンに9000回対戦させればいい。 ・そして先に結果を書いてしまうと、この検証で一番割を食ったのはAIだった。Claudeにデッキデータをすべて渡して出させた改造案が、1000戦の実測で勝率マイナス6.6ポイントの改悪として棄却されたのだ。この顛末を書く。
#AIタグ

AIと著作権(1)――「盗作」「学習」「模倣」「仕事の代替」を分けて考える

・まあ、この話題には実は踏み込まない方がいいのかも知れないが、現状での考えをまとめる意味でもいいのかなと思い、敢えてシリーズ化してみる。 ・生成AIと著作権について語ろうとすると、かなりの確率で「AIは人間の作品を盗んでいる」という言葉に行き着く。しかし、この「盗む」という表現の中には、本来なら別々に考えなければならない問題が、まとめて押し込まれている。既存の作品をほとんどそのまま出力すること、作者の絵柄を真似すること、作品を学習データとして利用すること、そしてAIによって人間の仕事が減ることまで、すべてが同じ意味での「盗み」であるかのように語られているのである。
Zennの「大規模言語モデル」のフィード

AIのうっかりは注意書きでは止まらない——ポカヨケで考えるルールの4段階

・「次からは気をつけて」が、いちばん弱い AIに仕事を任せていると、同じ場所で同じ失敗をされることがある。 ・以前の私は、ルールを書き足していた。「作業の前にこのファイルを読むこと」「勝手に消さないこと」「確認してから進むこと」。設定ファイルの中に、注意書きがどんどん増えていく。 ・しばらくして気づいた。注意書きが増えるほど、守られる率は下がっていく。
#AIタグ

AIの提案を比べるための「比較質問」をつくる方法

・AIから出てきたおすすめに違和感があるとき、前回はその違和感を言葉にして選び直す方法を紹介しました。 ・ただ、選び直そうとしても、候補が似ていて比べにくいことがあります。「どちらもよさそう」で止まってしまうと、最後はなんとなくの印象で決めることになりがちです。
#LLMタグ

AIは世界と私の架け橋だった話 ① (危ない質問?)

・(⚠️この記事はChatGPTとユーザーの会話をほぼそのまま掲載しています。 ・ChatGPTの回答は一つの視点であり、必ずしも正解ではありません。 ・必要に応じてご自身でも確かめながら読んでいただけたら嬉しいです🙂) 🤖 🤖 🤖 🤖 🤖 🤖 続きをみる
Zennの「大規模言語モデル」のフィード

AIレビューから「良い点も挙げる」を捨てた ── 敵対的検証を全レビューに統一した結果

・AIにコードレビューをさせると、たいてい「全体的によく書けています。いくつか改善点として…」というバランスの取れた、当たり障りのないレポートが返ってきます。私たちはこれを捨てました。 ・Webサービス開発リポジトリ(coelia-system:イベント予約・LINE連携・ガチャ等)のADR 0082「全レビュー系スキルの検証作法を敵対的検証へ統一」で、十数種類あるレビュー系スキルすべてに同じ作法を強制した話と、その後この体制が実際に拾ったバグの記録です。 ・敵対的検証の3ルール レビュー系スキル全部(アーキレビュー・セキュリティ・性能・信頼性・管理画面レビュー等)に、次の3点を追加しまし...
ITmedia NEWS 最新記事一覧

AI普及でデザイン業の倒産が前年比2.7倍に 「独自性なき企業は淘汰」 東京商工リサーチ

・東京商工リサーチは8月2日、グラフィックデザインやインテリアデザインなどを手掛ける「デザイン業」で、負債1000万円以上の倒産が急増していると発表した。生成AIの普及によるデザイン業務の内製化などが影響したという。
cs.LG updates on arXiv.org

AlphaG-OPD: Reliability-Gated Sibling Counterfactuals for On-Policy Distillation in Symbolic Alpha Factor Discovery

・arXiv:2608.01303v1 Announce Type: new Abstract: Symbolic alpha factor discovery can score a completed expression, but it provides no direct label for the structural decisions that produced it. ・Generative flow networks (GFlowNets) preserve a diverse, reward-proportional distribution over complete expressions, yet their trajectory-level objective does not compare unchosen sibling actions at an intermediate state.
cs.LG updates on arXiv.org

Amplitude-Only FFN Intervention for Tool-Structured LLM Inference Method: Gated Evaluation Protocol, and Cross-Model Empirical Results

・arXiv:2607.11183v2 Announce Type: cross Abstract: Large language models increasingly operate as tool-using agents, where small format, argument, or function-call errors can invalidate otherwise plausible responses. ・We study inference-time feed-forward network (FFN) intervention for improving structured outputs without retraining model weights. ・Our project began with Orthogonal Residual Projection (ORP), a direction-c
cs.LG updates on arXiv.org

An AI-Based Decision-Support Pipeline for Day-Ahead Photovoltaic Forecasting

・arXiv:2608.02088v1 Announce Type: new Abstract: Reliable photovoltaic (PV) forecasts are needed for low-carbon energy systems, but newly deployed sites often have short, imperfect records. ・This makes standard day-ahead forecasting difficult: persistence and physical baselines can be sensitive to calibration and timestamp alignment, while single machine-learning models may capture only one structure in the data and ov
cs.LG updates on arXiv.org

An Embedded RISC-V Evaluation of Kolmogorov--Arnold Networks in Hard-Constrained Recurrent Physics-Informed Models

・arXiv:2608.00737v1 Announce Type: new Abstract: Hard-constrained recurrent physics-informed networks (HRPINNs) embed known dynamics inside a recurrent numerical integrator and restrict a neural branch to learning only the residual dynamics that the first-principles model does not capture. ・Kolmogorov--Arnold Networks (KANs) have been proposed as parameter-efficient replacements for multilayer perceptrons (MLPs) in suc
cs.LG updates on arXiv.org

An Uncertainty-Driven Hybrid Deep Learning Approach for Broad-Coverage RF Modulation Recognition

・arXiv:2608.00796v1 Announce Type: cross Abstract: Automatic RF modulation recognition is of critical importance in spectrum monitoring, electronic warfare, and cognitive radio applications, where low signal-to-noise ratio (SNR) conditions and the growing diversity of modulation schemes limit the performance of existing methods. ・This paper proposes an uncertainty-driven hybrid deep learning architecture for recognizin
cs.LG updates on arXiv.org

Analytic Planning under Uncertainty with Moment Closure

・arXiv:2608.02519v1 Announce Type: new Abstract: Effective model-based reinforcement learning in stochastic environments requires planning that accounts for predictive uncertainty. ・Propagating full state distributions analytically offers a principled way to do this, but has traditionally required restrictive policy or reward structures to remain tractable. ・Consequently, modern deep reinforcement learning has largely r
cs.LG updates on arXiv.org

AOS: Adaptive Optimizer Switching via Training-State Signals for Faster Convergence and Better Generalization

・arXiv:2608.01997v1 Announce Type: new Abstract: Single-optimizer training is a poor fit for the distinct phases of deep network optimization: adaptive methods handle noisy early gradients well but overshoot flat minima, while SGD with momentum generalizes better in the late phase but converges slowly early on. ・We introduce AOS-R (Adaptive Optimizer Switching, Rule-Based), a lightweight controller that monitors six on
cs.LG updates on arXiv.org

AOSpec: Action and Observation Co-Speculation for Low-Latency Agent Serving

・arXiv:2608.00881v1 Announce Type: new Abstract: Large language model agents increasingly act through stateful tools, yet model generation and environment execution remain serialized at every step. ・As decoding accelerates, tool execution becomes a growing bottleneck. ・Existing action- or observation-only speculation leaves much of this latency exposed: value is concentrated in a few slow calls, some outcomes emerge onl
LLMタグが付けられた新着記事 - Qiita

Apple Healthの数値を「さくらのAI Engine」で分析して毎朝Slackへ投稿してみた

・はじめに Apple Watchで睡眠や心拍、歩数を記録していても、ヘルスケアアプリを毎日開かなければ変化を見落としてしまいます。 ・また、年齢を重ねて体調の変化を意識することが増えたため、Apple HealthのデータをInfluxDBへ保存し、Grafanaのダッシュ...
The Verge

Apple is working on iPhone-to-Windows copy-paste

・Apple is working on a feature that will allow users in the European Union to copy content on their iPhone and paste it onto their Windows PC (or vice versa), as spotted earlier by MacRumors. ・The move comes in response to an interoperability request from Microsoft that asks Apple to open up its Universal Clipboard feature, which supports copying and pasting across the Mac, iPhone, iPad, and Vision Pro. ・Under the EU's
AI News & Artificial Intelligence | TechCrunch

Apple says more ex-employees may have taken confidential data to OpenAI

・Apple says its trade secrets investigation into OpenAI has widened. ・In a new court filing, Apple claims additional former staff may have retained or accessed confidential information.
NVIDIA Blog

As AI Increases Demands on Memory, Storage Steps Up

・Surging AI demands are driving the need for massive datasets and context windows that burst past the confines of system memory. ・But rising needs aren’t met by simply adding more storage capacity. ・What’s needed is useful, grounded insights from AI factories and efficient, secure storage architectures that enable those insights.
cs.LG updates on arXiv.org

Assessing the Impacts of Imperfect Datasets on Client Selections in Federated Learning

・arXiv:2608.02250v1 Announce Type: new Abstract: Federated learning (FL) is a popular distributed learning framework where multiple clients perform local training and a server aggregates the locally updated models. ・FL enables decentralized training while preserving the privacy of clients' datasets. ・However, non-independent and identically distributed (non-IID) or noisy datasets can lead to low model accuracy or high c
cs.LG updates on arXiv.org

Augmented Inverse Hybrid Weighting: Robust Inference under Deterministic and Random Distribution Shifts

・arXiv:2608.00701v1 Announce Type: cross Abstract: Reweighting source samples to match a target covariate distribution is a standard response to distribution shift when generalizing evidence from one population to another. ・This strategy is well suited to deterministic, learnable covariate discrepancies, but can be insufficient when source--target population differences also contain changes beyond covariate shift or wh
cs.LG updates on arXiv.org

AutoCause: A Python framework that automates expert decisions in environmental time-series causal discovery

・arXiv:2608.00198v1 Announce Type: new Abstract: Environmental time-series causal discovery requires expert decisions about method choice, conditional-independence tests, lag horizons, sample-size adequacy, multiple-testing control, and evidence interpretation. ・Applied inconsistently across datasets, these choices yield graphs that cannot be compared, reproduced, or audited. ・We present AutoCause, an open-source Python
cs.LG updates on arXiv.org

Automated ECG Interval Measurement and Wave Delineation Using Fast Fourier Convolution ResNet

・arXiv:2608.00058v1 Announce Type: cross Abstract: Accurate measurement of ECG intervals, including PR, QRS duration, and QT/QTc, is central to cardiac diagnosis, yet the published ECG delineation literature evaluates performance almost exclusively as fiducial-point timing errors on small curated databases, rather than as clinical interval accuracy on large unselected cohorts. ・We bridge this gap by evaluating a comple
cs.LG updates on arXiv.org

Band-Count Dense Modal Estimation with Fixed-Frequency Differentiable Resonator Refinement

・arXiv:2608.00667v1 Announce Type: cross Abstract: Task B of the 1st DAFx Parameter Estimation Challenge requires estimating the frequencies, decay rates, gains, and number of modes in a dense plate-reverb impulse response. ・Weak and overlapping modes make sparse peak detection prone to severe undercounting. ・We train an ExtraTrees regressor on simulator-generated data to predict mode counts in four frequency bands.
cs.LG updates on arXiv.org

Beckmann Transport Models: From Autonomous Flows to One-Step Maps

・arXiv:2608.01692v1 Announce Type: new Abstract: We propose an instantiation of flow matching that relies on a time-independent velocity field (an \emph{autonomous flow}) to exactly map between two distributions, so long as the target is singular, i.e.\ supported on a lower-dimensional data manifold. ・We also show that the one-step generative map associated with this flow is the unique solution of a simple conservation
cs.LG updates on arXiv.org

Benchmarking Sheaf Neural Networks for Inductive Tasks

・arXiv:2608.02558v1 Announce Type: new Abstract: Sheaf Neural Networks (SNNs) generalize message passing by replacing scalar edge weights of standard Graph Neural Networks (GNNs) with learnable, edge-dependent restriction maps between node stalks. ・Despite their strong theoretical foundations and promising transductive results, SNNs have been evaluated almost exclusively on transductive node classification, leaving the
cs.LG updates on arXiv.org

Beyond Gene Reconstruction: Learning Cell Representations through Complementary Transcriptomic Views

・arXiv:2608.00985v1 Announce Type: new Abstract: The rapid growth of single-cell transcriptomic data has enabled the development of foundation models pretrained primarily by reconstructing masked expression values. ・This objective encourages these models to learn gene dependencies but does not directly optimize whole-cell representations, which are essential for many downstream tasks. ・To bridge this gap, we propose a c
cs.LG updates on arXiv.org

Beyond Lanes: Traffic Flow Dynamics in Disordered Conditions Based on High-Resolution Trajectory Data

・arXiv:2608.00602v1 Announce Type: cross Abstract: Disordered traffic flow is characterized by weak or non-existent lane discipline in the presence of strong vehicle heterogeneity and continuous lateral interactions, challenging traditional lane-based modeling assumptions. ・This study presents an empirical study of macroscopic and microscopic aspects of disordered traffic using high-resolution UAV trajectory data colle
cs.LG updates on arXiv.org

Beyond Magnitude and Shape: A Direction-Aware Loss for Time Series Forecasting

・arXiv:2608.01857v1 Announce Type: new Abstract: The direction of change --- whether a series will move up or down --- is often as important as its exact value in decisiondriven applications such as risk management and financial forecasting. ・However, most forecasting losses optimize either point magnitude or shape and frequency structure, and none explicitly targets the direction of change. ・In this paper, we find that
cs.LG updates on arXiv.org

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models

・arXiv:2608.01717v1 Announce Type: new Abstract: Recent reinforcement learning methods for diffusion large language models (dLLMs) commonly rely on on-policy rollouts generated by the target dLLM itself. ・When successful on-policy rollouts are scarce, however, on-policy training may receive little positive reward and make only limited progress. ・To mitigate this problem, we explore incorporating higher-reward rollouts g
cs.LG updates on arXiv.org

Beyond Random Partitioning: Unsupervised Spatio-Temporal Stratification for Cohort Balancing in Longitudinal Medical Imaging

・arXiv:2608.00073v1 Announce Type: cross Abstract: Rigorous dataset partitioning is a foundational, yet frequently overlooked, prerequisite for reliable deep learning in longitudinal medical imaging. ・Naively shuffling small clinical cohorts routinely introduces covariate shifts and temporal sampling imbalances across training, validation, and test subsets, exposing downstream models to out-of-distribution evaluation.
cs.LG updates on arXiv.org

BiKAN: Restoring Collapsed Basis of Binary Kolmogorov--Arnold Networks

・arXiv:2608.01490v1 Announce Type: new Abstract: Binarizing a polynomial Kolmogorov--Arnold Network (KAN) not only changes parameter precision, but also alters the function space available to each layer. ・When activations are restricted to ${-1,+1}$, all even powers reduce to $1$ and all odd powers reduce to $x$, causing the elementwise polynomial basis to collapse to constant and first-order responses. ・We refer to thi
cs.LG updates on arXiv.org

Breaking Diversity Collapse in Spiking Pseudo-Ensembles for Efficient OOD Detection in Remote Sensing

・arXiv:2608.01090v1 Announce Type: new Abstract: Spiking Neural Networks (SNNs) are attractive for resource-constrained remote-sensing systems, but reliable out-of-distribution (OOD) detection remains challenging. ・Deep ensembles provide strong predictive uncertainty, yet require multiple complete models and backbone evaluations. ・We propose an efficient spiking pseudo-ensemble that attaches multiple lightweight classif
cs.LG updates on arXiv.org

BRiG-AFA: Bellman Risk-to-Go Learning for Non-Myopic Active Feature Acquisition

・arXiv:2608.02305v1 Announce Type: new Abstract: Active feature acquisition (AFA) asks which unobserved feature to measure next for each test instance under a budget. ・Greedy rules are easy to train but can overlook context features whose value is realized only through later acquisitions, while reinforcement-learning and generative approaches introduce difficult optimization or conditional-density estimation.
MarkTechPost

Building an Advanced AI Skill Security Auditing Pipeline with NVIDIA SkillSpector, LangGraph, YARA Rules, SARIF, and CI Policy Gates

・Learn how to build an end-to-end security assessment pipeline for AI agent skills using NVIDIA SkillSpector and LangGraph. ・In this tutorial, we construct a synthetic skill marketplace, scan for malicious prompt injection, credential access, and risky dependencies, and implement custom YARA rules, baseline suppressions, and CI deployment gates. ・The post Building an Advanced AI Skill Security Auditing Pipeline with NVI
Hugging Face Papers

CADENA: Stepwise CAD Reverse Engineering

CADENA: Stepwise CAD Reverse Engineering
cs.LG updates on arXiv.org

Caliber: Cross-Architecture Extraction-Cost Control for Score-Returning APIs

・arXiv:2608.01023v1 Announce Type: new Abstract: We present Caliber, an output-perturbation defense against model extraction that formulates noise selection as a calibration problem: how much the defense degrades the supervision signal used to train a surrogate, and the provable per-input query cost of recovering the clean logits. ・To defend against an attacker that uses returned scores for knowledge distillation, Cali
cs.LG updates on arXiv.org

CARE: A Cascaded Framework for Efficient and Reliable Time Series Anomaly Detection

・arXiv:2608.01885v1 Announce Type: new Abstract: While deep learning models have achieved state-of-the-art performance in time series anomaly detection, their complex architectures incur substantial inference overhead. ・Existing methods typically apply a uniform inference strategy across all data points, which is inefficient given that anomalies are inherently scarce and the vast majority of temporal data consists of p
cs.LG updates on arXiv.org

CascadeLUT: Information-Ordered Streaming Inference for Bandwidth-Constrained FPGAs

・arXiv:2608.00720v1 Announce Type: cross Abstract: Mapping neural networks to FPGAs enables low-latency, energy-efficient inference, particularly for lookup table (LUT)-based models that eliminate multipliers and map directly to reconfigurable fabric. ・While prior work achieves high compute efficiency, it typically assumes full-sample availability, causing pipeline stalls in bandwidth-limited streaming scenarios.
cs.LG updates on arXiv.org

Causal Inference with Unstructured Treatments

・arXiv:2608.00657v1 Announce Type: cross Abstract: Causal inference usually concerns a scalar treatment, yet in many problems the treatment is unstructured: a text, an image, or a sequence of clinical decisions. ・Consider an instructor writing a course description to attract more students: the treatment is the course description, and the outcome is enrollment. ・The standard target, the average treatment effect of fixing
cs.LG updates on arXiv.org

ChaosProbe: A Neurochaotic Lens on Frozen Transformer Input-Embedding Spaces

・arXiv:2608.01968v1 Announce Type: new Abstract: Transformer models are most often understood through what they do: their benchmark performance, generation quality, or behavior on downstream tasks. ・Yet frozen transformer input-embedding spaces may also be examined through their responses to a controlled deterministic probe before contextual computation or task-specific adaptation. ・Guided by this response-based view, w
cs.LG updates on arXiv.org

Characterizing Bias in Post-Bandit Inference under Index Algorithms

・arXiv:2608.01069v1 Announce Type: new Abstract: Bandit algorithms generate data for downstream inference, but adaptive sampling biases post-bandit sample means. ・We analyze this bias for stable index algorithms, including UCB1 and its generalizations, and derive sharp leading-order expressions for the sample-mean bias and expected $Z$-statistic. ・Our characterization reveals the algorithmic origin of bias through a key
#AIタグ

chatGPT(作) 『ワキ毛とスネ毛のコラボレーション』 #青ブラ文学部

chatGPT(作) 『ワキ毛とスネ毛のコラボレーション』 #青ブラ文学部
Zennの「大規模言語モデル」のフィード

Claude Code v2.1.221 とCodexプラグイン相互運用、そしてMiniMax-H3まで(AIツール日次まとめ)

・この記事は 2026年8月4日 19時14分時点(日本時間)の情報をもとにまとめています。 ・きょうのAIコーディングまわりは、Claude Code の新バージョンと、Codex のプラグイン相互運用の強化が同時に動いていて、なかなか見どころの多い一日でした。オープンウェイト系でも新しいモデルがトレンド入りしています。順番に見ていきます。 ・Claude Code(AnthropicのコーディングCLI) 本日、v2.1.221 がリリースされました。目玉はVSCode拡張の「Focus view」です。ツール実行の細かいログを折りたたんで、ターンごとの要約と実行中インジケータだけを表...
Zennの「大規模言語モデル」のフィード

Claude Codeサブエージェント設計実践ガイド ― マルチエージェントで開発を分業する

・Claude Codeのサブエージェントを「なんとなく使う」から「設計して運用する」へ。.claude/agents のfrontmatter設計、tools/disallowedTools/permissionsによる三層の権限分離、並列実行とコンテキスト予算の管理、スキル(カスタムコマンド)とhooksの連携、そしてコードレビュー・調査・マイグレーションの実践ワークフローまでを、動作を確認した設定ファイルとエラーメッセージ付きで解説します。入門解説は最小限。すでにClaude Codeを実務で使っていて、運用設計に悩んでいる中級エンジニア向けの一冊です。
Zennの「大規模言語モデル」のフィード

Claude Codeのサブエージェント、結局いつ使うべきか ― 判断基準の整理

・Claude Codeのサブエージェント(.claude/agents/配下に置くカスタムエージェント)は「サブタスクを別の文脈で処理させる仕組み」として多くの記事で紹介されている。ただ、「結局いつ使えばいいのか」「いつ使わない方がいいのか」を判断基準込みで整理した情報は少ない。本記事では公式ドキュメント(Create custom subagents)の記述に基づき、この判断基準を整理する。 ・サブエージェントとは何か(前提の再確認) 独立したコンテキストウィンドウで動く。会話履歴・すでに読んだファイル・呼び出し済みのスキルは引き継がれない(forkという特殊な起動方法を除く)。
cs.LG updates on arXiv.org

Cluster-Aware Over-the-Air Federated Learning with Energy-Harvesting Devices: From Global Training to Model Personalization

・arXiv:2608.01426v1 Announce Type: new Abstract: Federated learning (FL) enables distributed optimization and learning across decentralized edge devices while preserving data privacy, but its performance is fundamentally constrained by heterogeneous data distributions, limited communication resources, and energy availability. ・In practical wireless networks, mobile devices (MDs) often exhibit diverse data and learning
cs.LG updates on arXiv.org

Conformalized Large Language Models under Configuration Shift

・arXiv:2608.01460v1 Announce Type: new Abstract: Conformal prediction (CP) is a distribution-free framework for uncertainty quantification that has recently been adapted to large language models (LLMs), providing prediction sets with finite-sample coverage guarantees under exchangeability. ・Yet for LLMs, nonconformity scores are often induced by an inference pipeline, not just a fixed model, making them depend not only
cs.LG updates on arXiv.org

Conservation laws determine what physical learning remembers

・arXiv:2608.00097v1 Announce Type: cross Abstract: Physical learning rules such as equilibrium propagation (EP), coupled learning (CL), and adjoint coupled learning (AL) train resistive networks through local measurements. ・In the small-nudge limit EP and CL exactly conserve the conductance mass K = (1/2) sum_e kappa_e^2, a property that stabilizes training. ・We show that conservation also governs the inductive bias of
cs.LG updates on arXiv.org

Constrained Co-Design for Photonic Bayesian Neural Networks

・arXiv:2608.02229v1 Announce Type: new Abstract: Classical neural networks frequently produce overconfident predictions on ambiguous or out-of-distribution (OOD) data, a liability that grows with each AI system deployed in safety-critical real-world scenarios. ・Bayesian neural networks (BNNs) provide a principled framework for uncertainty-aware prediction by replacing deterministic parameters with probability distribut
cs.LG updates on arXiv.org

Convex Neural Energy Elements: Monolithic Finite-Element Assembly of Geometry-Parameterized Neural Operators with Stability and Error Guarantees

・arXiv:2608.02036v1 Announce Type: new Abstract: Extending the neural-operator element method from individually trained, fixed-geometry neural elements to a library of reusable, geometry-parameterized element types fails structurally: a field-predicting operator trained by value regression induces an energy whose assembled Hessian is indefinite, and Newton converges to spurious minima (247% error) even with 1%-accurat
cs.LG updates on arXiv.org

CoRe-GNN: Multilevel Message passing on Coarsened graphs

・arXiv:2608.02128v1 Announce Type: new Abstract: Training Graph Neural Networks on large graphs is challenged by the memory cost of storing all node representations across layers. ・We show that several existing scalable approaches can be written as structured modifications of the GNN propagation matrix, providing a unified perspective that exposes their respective limitations. ・In particular, graph coarsening replaces i
cs.LG updates on arXiv.org

Cost-Effective Automated Judging of Natural-Language Mathematical Proofs

・arXiv:2608.00004v1 Announce Type: cross Abstract: Grading natural-language mathematical proofs is a recurring cost in evaluating math-reasoning systems, and frontier LLM judges are expensive. ・We ask whether cheap open-weight models can serve as reliable judges given a candidate proof, a ground-truth proof, and a human-grading rubric. ・On a 200-instance validation sample of IMO-GradingBench, three cheap judges (GPT-OSS
cs.LG updates on arXiv.org

CoSynFlow: Conformal Symplectic Neural Flows for Cross-System Prediction of Dissipative Hamiltonian Dynamics

・arXiv:2608.00571v1 Announce Type: new Abstract: Learning solution operators for differential equations is a central problem in scientific machine learning. ・However, many neural operator methods optimize prediction accuracy without explicitly enforcing the geometric structure of the dynamics. ・Structure-preserving models such as SympNets and Symplectic Neural Flows address this issue for conservative Hamiltonian system
cs.LG updates on arXiv.org

CRIP: Channel Level Representation Injection for Personalized One-Shot Federated Learning

・arXiv:2608.02222v1 Announce Type: new Abstract: One-shot federated learning (OSFL) has emerged as a promising collaborative model learning framework with only a single round of communication, offering significant advantages in communication efficiency and privacy preservation. ・However, OSFL often faces inherent limitations under severe domain heterogeneity across clients due to the lack of iterative knowledge exchang
cs.LG updates on arXiv.org

Crushing the Evidence: A Dual-Penalty Evasion Framework for Fooling White-Box Explainable AI Auditors

・arXiv:2608.00566v1 Announce Type: new Abstract: Post-hoc model explainers such as LIME, SHAP, and Integrated Gradients are widely deployed to audit models in high-stakes sensitive domains, including finance, healthcare, and social welfare. ・This ensures the model's transparency and acceptability. ・However, a few studies have examined potential attacks in the explainability pipeline.
cs.LG updates on arXiv.org

CurveShift: Is Agent Progress Scalar? Separating Level from Shape

・arXiv:2608.00355v1 Announce Type: cross Abstract: Progress in large language models is often summarized using a single scalar measure, such as a time horizon, a latent ability estimate, or an aggregate benchmark score. ・These summaries capture the overall performance, but they do not test whether progress is distributed differently across task difficulty. ・We find that most of the apparent shift in gains toward harder
Hugging Face Papers

DAPD: Dual-Anchored Policy Distillation

DAPD: Dual-Anchored Policy Distillation
cs.LG updates on arXiv.org

DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling

・arXiv:2608.02032v1 Announce Type: new Abstract: Modern language models are built primarily from Transformers, recurrent models, and their hybrid architectures. ・Transformers rely on token-level attention memories, while recurrent models such as state space models (SSMs) and linear attention maintain compact recurrent states. ・These architectures are typically instantiated separately or interleaved at the layer level, l
cs.LG updates on arXiv.org

Data-Driven Pinball-Loss Selection for Vertically Distributed Elastic-Net SVMs

・arXiv:2608.00949v1 Announce Type: new Abstract: The pinball-loss support vector machine is robust, but its asymmetry parameter is usually fixed in advance. ・We propose a data-driven elastic-net support vector machine that learns simplex-constrained weights over candidate pinball losses while retaining one classifier. ・The weighted loss is equivalent to a pinball loss with a data-dependent effective parameter.
cs.LG updates on arXiv.org

Deep Learning CNN and Recurrence Analysis for Alpha Gamma EEG Biomarkers in Fragile X Syndrome

・arXiv:2608.00835v1 Announce Type: new Abstract: Fragile X Syndrome (FXS) is a neurodevelopmental disorder caused by reduced expression of fragile X mental retardation protein (FMRP), leading to disrupted synaptic plasticity, cortical hyperexcitability, and impaired network synchronization. ・Electroencephalography (EEG) provides a noninvasive window into these mechanisms and consistently reveals abnormalities in alpha
cs.LG updates on arXiv.org

Deep Learning for Cyber Threat Detection and Mitigation in Healthcare-IoT

・arXiv:2608.00118v1 Announce Type: cross Abstract: Cybersecurity is a fundamental requirement for protecting wearable devices used in healthcare Internet of Things (H-IoT) systems. ・Security failures in these resource-constrained systems directly compromise patient safety. ・Physiological data and network traffic are frequent targets of cyberattacks in H-IoT environments.
cs.LG updates on arXiv.org

Deep Learning-Based Estimation of Ground Reaction Forces in Parkinsonian Gait Using an Optimized Set of IMU Data

・arXiv:2608.02408v1 Announce Type: new Abstract: Accurate gait analysis in Parkinson's disease (PD) typically relies on laboratory-based systems to capture biomechanical data, such as ground reaction forces (GRFs). ・Estimating GRFs using inertial measurement units (IMUs) provides a feasible alternative. ・However, this approach remains challenging in pathological gait like PD due to its high variability and complexity.
Hugging Face Papers

DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents

DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents
Hugging Face Papers

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs
Hugging Face - Blog

Deploy local agents everywhere with LFM2.5-2.6B

Deploy local agents everywhere with LFM2.5-2.6B
cs.LG updates on arXiv.org

Differentiable Lifting for Topological Neural Networks

・arXiv:2608.01160v1 Announce Type: new Abstract: Topological neural networks (TNNs) enable leveraging high-order structures on graphs (e.g., cycles and cliques) to boost the expressive power of message-passing neural networks. ・In turn, however, these structures are typically identified a priori through an unsupervised graph lifting operation. ・Notwithstanding, this choice is crucial and may have a drastic impact on a T
cs.LG updates on arXiv.org

Diffusion Policy with Behavioral Advantage Correction for Offline Reinforcement Learning

・arXiv:2608.02332v1 Announce Type: new Abstract: In offline reinforcement learning (RL), the distribution shift between behavioral data and the learned policy can lead to erroneous \emph{Q}-value estimation, thereby misguiding the direction of policy optimization. ・To address this issue, we develop a behavioral advantage corrected policy evaluation (BAC-PE) approach, which utilizes the \emph{Q}-function of the behavior
Hugging Face Papers

DiffusionGemma Technical Report

DiffusionGemma Technical Report
cs.LG updates on arXiv.org

Disagree to Accelerate: Closing the Loop on Diffusion Feature Forecasts

・arXiv:2608.01740v1 Announce Type: new Abstract: Training-free feature forecasting accelerates diffusion sampling by predicting features at skipped denoising steps. ・Recent work has mainly focused on designing stronger forecasters. ・Yet forecast error varies sharply across steps, and open-loop caches trust the forecast in full at every skipped step.
ITmedia NEWS 最新記事一覧

Discordの「オンラインになったとき、フレンドに通知する」機能が不評 「かまちょすぎる」「勝手に伝えないで」

・ゲーマー向けチャットツール「Discord」が備える「自分がオンラインになったとき、フレンドに通知する」機能がXで不評だ。Xでは「用もないのに勝手にかまちょ(かまってちょうだいの略語)しないでほしい」「すでに通知があったら申し訳ない」など“余計なお世話”として受け止められている。
cs.LG updates on arXiv.org

Distill What the Student Can See: Fisher-Projected On-Policy Distillation for Vision-Language Models

・arXiv:2608.01263v1 Announce Type: new Abstract: On-policy distillation (OPD) samples trajectories from the current student policy and minimizes token-level divergence between student and teacher next-token distributions at prefixes along those trajectories. ・This aligns the distillation states with the student's own generation distribution. ・However, it still assumes that the complete teacher distribution is an appropr
cs.LG updates on arXiv.org

Do Neural Networks Really Beat the Curse of Dimensionality? A Bit-Complexity View

・arXiv:2608.01357v1 Announce Type: new Abstract: Traditional approximation theory measures convergence rates in terms of the number of parameters or degrees of freedom. ・However, practical computation operates under finite precision: parameters must be encoded using a finite number of bits. ・Therefore, approximation efficiency should be evaluated in terms of computational bit complexity, which is intrinsically connected
cs.LG updates on arXiv.org

Do Static Embeddings Add Value to Hybrid Dutch Retrieval?

・arXiv:2608.02112v1 Announce Type: new Abstract: Embedding benchmarks measure standalone model quality, but they do not establish whether a low-cost retriever contributes complementary ranking information once lexical and transformer-based retrieval are already combined. ・We present a controlled evaluation of this question across Dutch retrieval tasks from the Massive Text Embedding Benchmark for Dutch (MTEB-NL).
cs.LG updates on arXiv.org

DODA: A Database of Datasets for Aesthetics Research

・arXiv:2608.00089v1 Announce Type: cross Abstract: With rapid growth in the fields of empirical and computational aesthetics we have seen a vast increase in large image datasets annotated for aesthetics. ・As the image databases differ widely in many respects (e.g., different standards for annotation), it can be tedious to find the dataset that fits one's research needs best. ・The absence of a centralized open-science se
cs.LG updates on arXiv.org

Domain-Generalized Adaptive Semantic Communication for Collaborative Perception

・arXiv:2608.00056v1 Announce Type: cross Abstract: We propose RSTA, a domain-generalized semantic communication framework enabling source-free V2X collaborative perception under both observation-domain shift and unseen wireless channel conditions. ・In V2X, received semantic tokens suffer coupled degradation from pre-transmission domain drift and in-transit channel corruption; existing methods address only one source, l
Hugging Face Papers

DreamTraj: Generating 6-DoF Object Trajectories by Reading Unrendered Video Diffusion Latents

DreamTraj: Generating 6-DoF Object Trajectories by Reading Unrendered Video Diffusion Latents
cs.LG updates on arXiv.org

DSETA: A Dual-Stage Continual Learning Framework for Travel Time Prediction in Dynamic Traffic Environments

・arXiv:2608.00402v1 Announce Type: new Abstract: Estimated Time of Arrival (ETA) prediction is a core component of intelligent transportation systems. ・As traffic congestion patterns become increasingly dynamic in large cities, maintaining high prediction accuracy poses a major challenge for ride-hailing platforms. ・Existing methods either fail to adapt to irregular traffic patterns and sudden congestion, or suffer from
cs.LG updates on arXiv.org

Element-Aware Group Learning for E-Commerce Image Generation

・arXiv:2608.00584v1 Announce Type: cross Abstract: Recent advances in image generation and editing have made prompt quality a key bottleneck for e-commerce creatives. ・Vision-language models (VLMs) can generate image-editing prompts from product images and metadata, but further improving their prompt-writing capabilities requires post-training with feedback from the generated images. ・Group Relative Policy Optimization
AI News & Artificial Intelligence | TechCrunch

Elon Musk spends half his time talking robots and AI on Tesla earnings calls

・An analysis of the last seven years of Tesla earnings calls shows just little attention Musk pays to Tesla's car business.
cs.LG updates on arXiv.org

Empowering Credit Risk Detection in Weixin Pay with Billion-Scale Deep Graph Learning

・arXiv:2608.02168v1 Announce Type: new Abstract: Credit risk detection, particularly mitigating individual fraud, is crucial for maintaining the stability of digital financial ecosystems. ・Accurately identifying credit fraud among billions of users is critical for minimizing financial losses and safeguarding the sustainability of inclusive financial services. ・Given that credit fraud risks are often concealed within het
cs.LG updates on arXiv.org

Enriched text-guided variational multimodal knowledge distillation network (VMD) for automated diagnosis of plaque vulnerability in 3D carotid artery MRI

・arXiv:2509.11924v2 Announce Type: cross Abstract: Multimodal learning has attracted much attention in recent years due to its ability to effectively utilize data features from a variety of different modalities. ・Diagnosing the vulnerability of atherosclerotic plaques directly from carotid 3D MRI images is relatively challenging for both radiologists and conventional 3D vision networks. ・In clinical practice, radiologis
cs.LG updates on arXiv.org

Ensemble of Unsupervised Deep Learning for Clustering Imbalanced Tabular Data

・arXiv:2608.00346v1 Announce Type: new Abstract: Data imbalance poses a major challenge in supervised classification, where the majority-class bias contributes to false negatives and overestimates classification accuracy. ・Unsupervised deep clustering can be immune to class imbalance because representation learning for clustering is performed without class labels. ・Deep clustering has been proposed for images, languages
AI News & Artificial Intelligence | TechCrunch

EON wants to move the data superhighway from ocean fiber to space lasers

・Endeavor Optical Networks is planning to launch the fastest space laser communications system yet built.
cs.LG updates on arXiv.org

EulerLoRA: Rank-Driven Jump Dynamics for Calibrated Parameter-Efficient Fine-Tuning

・arXiv:2608.01142v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) enables parameter-efficient fine-tuning, but standard LoRA produces a single deterministic model and does not directly support predictive uncertainty estimation. ・We introduce EulerLoRA, a stochastic extension of LoRA that generates multiple predictive trajectories by sampling structured variations along the rank-one components of shared low-ra
cs.LG updates on arXiv.org

Evaluating Forecasting Techniques for Hardware Errors on a Large-scale HPC System

・arXiv:2608.01648v1 Announce Type: new Abstract: Hardware error logs in high-performance computing (HPC) systems provide early signals of abnormal behavior, yet there remain challenges in effectively forecasting these errors using modern predictive methods. ・This work investigates the boundaries of applying time series forecasting to HPC hardware error dynamics. ・We use seven years of production logs from the Theta supe
cs.LG updates on arXiv.org

Evolutionary Curriculum Learning Improves Biological Sequence Modeling

・arXiv:2608.00697v1 Announce Type: cross Abstract: Variational autoencoders (VAEs) trained on multiple sequence alignments (MSAs) have emerged as powerful generative models for biological sequences, with applications ranging from disease variant prediction to functional RNA design. ・However, standard biological VAE training treats all sequences as exchangeable, ignoring the rich evolutionary structure that organizes ho
cs.LG updates on arXiv.org

Experience-Calibrated Contrastive Decoding for Mitigating Hallucinations in LM-Based Text-to-Speech

・arXiv:2608.00722v1 Announce Type: cross Abstract: Language model-based text-to-speech (LM-based TTS) remains vulnerable to speech hallucinations that deviate from the target text. ・Existing mitigation mainly relies on architectural changes or additional training, while decoding-time control remains underexplored. ・We present a conditional information view that distinguishes text-derived alignment information from exper
cs.LG updates on arXiv.org

Explainable Hybrid Feature Selection for Intrusion Detection in Internet of Medical Things Environments

・arXiv:2608.00869v1 Announce Type: cross Abstract: Internet of Medical Things (IoMT) networks are hard to protect: devices are heterogeneous, computing resources are scarce, and traffic must be analyzed in real time. ・We present an intrusion detection system that addresses these constraints through feature selection. ・A Pearson correlation filter first removes redundant attributes; a hybrid strategy then combines model-
cs.LG updates on arXiv.org

Factorized AdaBoost.MH Achieves the Same Convergence Rate as AdaBoost.MH

・arXiv:2608.01091v1 Announce Type: new Abstract: AdaBoost.MH reduces multi-class classification to a collection of binary subproblems and enjoys the classical boosting-type convergence guarantee under a weak learning condition. ・A more structured variant, Factorized AdaBoost.MH, uses base classifiers of the form $\mathbf{h}(x)=\alpha \mathbf{v} \bm{\varphi}(x)$, where a single binary classifier $\bm{\varphi}$ is shared
cs.LG updates on arXiv.org

FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds

・arXiv:2608.01049v1 Announce Type: cross Abstract: World models have attracted significant attention for their ability to capture and predict the structure and dynamics of the physical world. ・In this emerging landscape, Joint Embedding Predictive Architectures (JEPA) offer a particularly compelling direction. ・We study a largely unexplored regime: populous, crowded, and chaotic Global South urban environments, which we
cs.LG updates on arXiv.org

Fairness Auditing: Lower Bounds on Company Manipulation

・arXiv:2608.00568v1 Announce Type: new Abstract: Fairness audits are increasingly mandated in high-stakes applications such as hiring, lending, and automated decision-making. ・Recent work has established fundamental impossibility results for black-box fairness auditing, showing that sufficiently expressive models can evade any auditing strategy. ・We complement these results by quantifying the extent of unavoidable post-
cs.LG updates on arXiv.org

Fast Trainable Multilinear Bases for Image Compression

・arXiv:2608.00053v1 Announce Type: cross Abstract: The Discrete Fourier Transform, the Discrete Cosine Transform, and their block-wise variants underpin most deployed image and video codecs. ・Their effectiveness rests on three properties: they run in near-linear time (linear up to a polylogarithmic factor), they are exactly invertible, and they carry few to no parameters. ・In this work, we generalize these bases to isom
cs.LG updates on arXiv.org

FDIR: Harmonizing Fidelity and Human-Machine Preference in Lossy Compression Image Restoration

・arXiv:2608.00111v1 Announce Type: cross Abstract: Image restoration quality can be evaluated along three complementary facets: pixel-level fidelity, human perception, and downstream machine preference. ・However, existing lossy compression restoration methods optimize for at most one of these criteria: fidelity-oriented models often regress toward conditional means and produce over-smoothed outputs, while generative ap
cs.LG updates on arXiv.org

FedChronos: Federated Fine-Tuning of Time-Series Foundation Models for Privacy-Preserving Commodity Price Forecasting

・arXiv:2608.01290v1 Announce Type: new Abstract: Time-series foundation models (TSFMs) such as Chronos have demonstrated strong forecasting capabilities across domains, yet adapting them to institutionally fragmented settings, where data cannot be centralized due to regulatory, competitive, or sovereignty constraints, remains unexplored. ・We introduce FedChronos, a framework for federated parameter-efficient fine-tunin
cs.LG updates on arXiv.org

Feed-Forward Steering in Transformer Residual Dynamics

・arXiv:2608.02071v1 Announce Type: new Abstract: Attention-only dynamical theories model Transformer residual directions as particles aggregating on a sphere. ・We extend this framework by incorporating the feed-forward network (FFN) term as a local steering field acting on each token state. ・The resulting theory predicts that the tangential component of the FFN field is necessary for motion in residual-direction space,
cs.LG updates on arXiv.org

Finite-Time Analysis of Discounted Exponential-Utility Reinforcement Learning

・arXiv:2608.01917v1 Announce Type: new Abstract: Discounted exponential utility provides a principled criterion for risk-sensitive sequential decision-making, but its nonlinear structure complicates reinforcement learning. ・A recent work \citep{thoppe2026reinforcement} addressed this difficulty by introducing a Bellman-compatible surrogate and two model-free fixed-point algorithms for optimizing it over stationary poli
cs.LG updates on arXiv.org

FL-OA: A Byzantine-Robust Federated Learning Framework with Outsourced Auditing for Intelligent Devices

・arXiv:2608.01095v1 Announce Type: new Abstract: Federated learning (FL) enables multiple intelligent devices to collaboratively train a high-accuracy model without sharing raw data. ・However, due to its distributed nature, FL is vulnerable to Byzantine attacks. ・Existing defense methods rely on strong assumptions, such as the proportion of malicious devices not exceeding 50\%, or the server having an additional root da
#AIタグ

Flow Studioが大型アップデート。「3D Editor」と「Canvas」でAI映像制作が一段階進化

・Autodesk Flow Studioに大きなアップデートが公開されました。 ・今回追加されたのは、新しい3D制作環境 「3D Editor」 と、ノードベースの画像・動画編集環境 「Canvas」 です。
cs.LG updates on arXiv.org

Foundations of Reinforcement Learning and Control:Connections and New Perspectives

・arXiv:2608.02433v1 Announce Type: new Abstract: Reinforcement learning and control theory are two adjacent scientific fields that focus on optimizing the controller of unknown dynamical systems using feedback. ・While both fields have common roots in dynamic programming, they have evolved with distinct methodologies, goals, and cultures. ・Despite decades of mutual influence, a significant gap persists between the two co
cs.LG updates on arXiv.org

From Digital to Physical Reservoir Computing: Co-Optimizing Soft Robotic Reservoirs via Dynamics Matching

・arXiv:2608.00484v1 Announce Type: cross Abstract: Soft robotic substrates are promising for Physical Reservoir Computing (PRC) because their compliant nonlinear dynamics can provide temporal memory, high-dimensional state transformations, and efficient inference. ・However, physical reservoirs are often adopted as-is rather than pretrained or co-optimized, potentially limiting soft robotic PRC performance relative to d
cs.LG updates on arXiv.org

From field-scale to large-scale spectral libraries: Tabular foundation models in soil spectroscopy

・arXiv:2608.00608v1 Announce Type: new Abstract: Visible and near-infrared (vis-NIR) and mid-infrared (MIR) spectroscopy enable rapid, cost-effective prediction of soil properties. ・Yet, translating high-dimensional, highly collinear spectra into accurate soil property predictions remains challenging, particularly when employing machine learning. ・We systematically investigated regression models and dimensionality reduc
cs.LG updates on arXiv.org

From fragmented data to actionable design: Physics-calibrated learning for plastic upcycling

・arXiv:2608.02402v1 Announce Type: new Abstract: Thermochemical upgrading of plastic waste is a key upcycling pathway, yet the experimental literature is fragmented by heterogeneous conditions and incomplete reporting. ・Complete-case learning would retain only 10.99% of the curated experiments, while target imputation can introduce biased supervision. ・Here we develop a Physics-Calibrated, Missingness-Gated, and Load-Ba
cs.LG updates on arXiv.org

Fruit-HSNet: A Machine Learning Approach for Hyperspectral Image-Based Fruit Ripeness Prediction

・arXiv:2608.01202v1 Announce Type: cross Abstract: Fruit ripeness prediction (FRP) is a classification-based agricultural computer vision task that has attracted much attention, thanks to its wide-ranging advantages in agriculture field for both pre-harvest and post-harvest management. ・Accurate and timely FRP can be achieved using machine/deep learning-based hyperspectral image classification techniques. ・However, chal
cs.LG updates on arXiv.org

Fused Bayesian Flow Networks for Dual-Target Molecular Design

・arXiv:2608.01007v1 Announce Type: new Abstract: Dual-target drug design aims to generate 3D molecules that can simultaneously interact with two target proteins, offering a promising route for discovering polypharmacological compounds against complex diseases. ・While recent generative models have shown encouraging performance in single-target drug design, existing dual-target approaches either focus on sequence generat
#LLMタグ

FX自動売買Botを「社員」だと思ったら、管理がぐっと楽しくなりました

・Botが増えるほど、何が起きているかわからなくなる 自動売買Botは、最初の1台がいちばん楽しいです。チャートを眺めながら「お、エントリーした」「利確した」と一喜一憂できます。
cs.LG updates on arXiv.org

Gecko: Fast Private Inference via Secure Public Encoder Offloading

・arXiv:2608.02378v1 Announce Type: new Abstract: Private inference protects both user inputs and server models during neural network inference, but existing solutions remain too slow for practical deployment. ・This motivates recent efforts to run a public encoder, such as a pretrained backbone, outside the protection boundary and evaluate only a small private predictor cryptographically. ・While appealing for efficiency,
cs.LG updates on arXiv.org

Generated Images Are Easier to Forget: A Machine Unlearning Perspective for Synthetic Image Detection

・arXiv:2608.00716v1 Announce Type: cross Abstract: Robust detection of generated images is critical to counter the misuse of generative models. ・Existing methods primarily depend on learning from human-annotated training datasets, limiting their generalization to unseen distributions. ・In contrast, large-scale vision models (LVMs) pre-trained on web-scale datasets exhibit exceptional generalization power through exposur
cs.LG updates on arXiv.org

Generative Models for Modeling and Synthesizing MIMO Channels in Adverse Weather Conditions

・arXiv:2608.00156v1 Announce Type: cross Abstract: The push for broader coverage in future cellular networks depends on reliable service, yet this is increasingly harder to do as we encounter more instances of extreme weather conditions. ・In extreme weather conditions, we have difficulty evaluating coverage due to limited access to channel measurements. ・In this paper, we generate channel state information (CSI) in low
cs.LG updates on arXiv.org

Generic Vision and Cross-Attention for Reaction Yield Prediction

・arXiv:2608.00776v1 Announce Type: new Abstract: Traditional reaction yield prediction is constrained by 1D quantum descriptors that lack explicit spatial information. ・To address this gap, a dual-modal Vision Cross-Attention architecture is proposed, fusing tabular physical-organic data with 2D molecular topologies. ・Notably, it is demonstrated that a generic computer vision backbone processing simple 2D skeletal struc
cs.LG updates on arXiv.org

GeoArbiter: Verifiability-Guided Grounding for Remote-Sensing Multimodal LLMs

・arXiv:2608.00877v1 Announce Type: new Abstract: Remote-sensing multimodal large language models (MLLMs) often assert facts that imagery cannot establish, such as a facility's identity or function. ・Coordinate-keyed geographic retrieval can supply this missing knowledge, improving fMoW land-use accuracy by 12.06--17.19 points across three open MLLMs. ・However, retrieved records can also contradict visible evidence, and
cs.LG updates on arXiv.org

Geometry-Guided Layerwise FFN Width Allocation in Transformers

・arXiv:2608.02064v1 Announce Type: new Abstract: Feed-forward networks (FFNs) account for a large fraction of Transformer parameters, yet their hidden width is usually constant across depth. ・We ask whether this capacity can instead be allocated from a forward-pass measurement of layer behavior. ・We view each FFN as transporting a cloud of token representations and quantify the induced geometric change using corresponde
cs.LG updates on arXiv.org

GLAIM: Learning Global and Local Adaptive Inter-Variable Dependency for Multivariate Time Series Imputation

・arXiv:2608.02366v1 Announce Type: new Abstract: Multivariate time series imputation is fundamental to downstream analysis, yet modeling inter-variable dependencies with incomplete observations remains challenging. ・Existing methods learn global dependencies across samples or dynamic local dependencies per sample. ・Global dependencies are stable but adapt poorly to sample variations and temporal non-stationarity, wherea
ITmedia NEWS 最新記事一覧

Google、新型スマホ「Pixel 11シリーズ」を予告 8月12日に予約購入スタート

・米Googleの日本法人が、新型Pixelスマートフォンの予約受付を8月12日午後11時に始めると告知した。新製品発表イベント「Made by Google」は翌13日午前7時を予定しているほか、同じ13日には、東京・表参道に日本初のGoogle直営店もオープンする。
#LLMタグ

GPT-Liveの全二重通信でなぜ会話は人間らしくなるのか。しんたろーが解説する音声AI開発の転換点

・音声AIの会話が「機械っぽい」時代は終わる。 ・これまでのターン制を捨て、ユーザーの発話を聞きながら同時に喋る全二重通信の音声AIが登場した。応答遅延の要因だった検出処理が排除され、会話の即応性が高まっている。
Hugging Face Papers

GPTQ-2D: Cubic-Time Two-Sided Adaptive Rounding

GPTQ-2D: Cubic-Time Two-Sided Adaptive Rounding
Hugging Face Papers

GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning

GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning
cs.LG updates on arXiv.org

GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning

・arXiv:2608.02585v1 Announce Type: new Abstract: Optimization-based latent reasoning improves large language model outputs by optimizing instance-specific continuous states at test time while keeping model parameters frozen. ・Existing methods, however, typically connect these states to the reasoning trajectory through decoded tokens, making sequence-level credit assignment indirect and obscuring how latent updates shap
cs.LG updates on arXiv.org

Gram-Space: Structure-Preserving Codebook Compression for Memory-Efficient Neuro-Symbolic AI

・arXiv:2608.01528v1 Announce Type: new Abstract: Vector symbolic architectures (VSA) are widely used for reasoning in neuro-symbolic (NeSy) AI, yet high-dimensional codebooks often create severe memory bottlenecks that limit scalability and deployment. ・In this paper, we propose Gram-Space, a compression framework that applies Gram-Schmidt orthogonalization to represent codebook vectors in a compact orthonormal coordin
cs.LG updates on arXiv.org

GraphIR: Architecture-Level Search States for LLM-Guided Neural Architecture Evolution

・arXiv:2608.01633v1 Announce Type: new Abstract: Large language models (LLMs) enable neural architecture search (NAS) directly over executable neural network programs. ・However, code-level flexibility does not provide the architecture state needed for effective mutation: LLMs must infer tensor dependencies, editable components, and compatibility constraints from implementation details. ・To address this representation mi
cs.LG updates on arXiv.org

GraRe: Grasp Candidate Re-Ranking for Frozen 6-DoF Grasp Detectors

・arXiv:2608.00946v1 Announce Type: cross Abstract: Existing 6-DoF grasp detectors typically rank grasp candidates by detector confidence. ・However, our analysis on GraspNet-1Billion shows that detector confidence is often poorly aligned with grasp quality, causing successful grasp candidates to be ranked too low during execution. ・Motivated by this observation, we formulate grasp candidate re-ranking as a separate task
cs.LG updates on arXiv.org

H+ Embedding: Harmonizing Global and Token-Level Retrieval with Context-Dependent Phrases

・arXiv:2608.00065v1 Announce Type: cross Abstract: Terminology-intensive retrieval, especially in medical settings, depends on preserving multi-word entities, abbreviations, numerical constraints, and compositional concepts. ・However, existing representations lie at two extremes: single-vector retrievers often over-compress local relevance signals, while token-level late interaction retains every tokenizer subword at s
cs.LG updates on arXiv.org

HarnessCompass: Guiding Automatic Harness Evolution toward Generalizable and Effective Agent Harnesses

・arXiv:2608.01918v1 Announce Type: new Abstract: Harness design plays a critical role in agent performance by shaping how large language models (LLMs) perceive, reason over, and act within executable environments. ・Recent work has proposed automatic harness evolution, which iteratively improves the harness from agent--environment interactions. ・However, existing methods often overfit to the evolution tasks, rely exclusi
cs.LG updates on arXiv.org

Heterogeneous Multi-Agent Reinforcement Learning for Radio Resource Management under Coupled Finite-Horizon Constraints

・arXiv:2608.01745v1 Announce Type: new Abstract: Maximizing throughput under proportional fairness in dense wireless networks requires jointly managing user association, scheduling, base station (BS) activation, and handover control under hard finite-horizon energy and handover budgets, which induces a fundamental tension between BS-side energy management and user-side handover regulation. ・While multi-agent reinforcem
cs.LG updates on arXiv.org

Hierarchical Solomonoff Induction: An Unbounded Machine Learning Model

・arXiv:2608.01005v1 Announce Type: new Abstract: Solomonoff Induction, or SolInd, provides an ideal unbounded model of a priori sequence prediction but cannot naturally describe extrapolation from a given training dataset, as performed by Large Language Models. ・We apply de Finetti's theorem on exchangeable distributions to SolInd to produce what we call Hierarchical Solomonoff Induction, or HSI, which maintains a hype
cs.LG updates on arXiv.org

HindSearch: Trajectory-Level Hindsight Critique for Search-Augmented Reinforcement Learning

・arXiv:2608.01597v1 Announce Type: new Abstract: Search-augmented LM agents are typically trained with a binary exact-match reward, which throws away most of what a failed trajectory tells us about why it failed. ・We introduce HindSearch, a hindsight self-distillation procedure for GRPO: after each rollout, a frozen judge writes a short critique of every failed trajectory using the gold answer, and the critique supplie
WIRED

How Data Centers Broke American Politics

・What the Unabomber, Steve Bannon’s tech guy, and Bernie Sanders taught me about the great data center backlash of 2026.
cs.LG updates on arXiv.org

How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models

・arXiv:2608.02089v1 Announce Type: new Abstract: Large language models often show users a final response and a short reasoning summary while the full reasoning trace stays hidden. ・We introduce an observability ladder that holds each completed run fixed and varies only what a reader inspects to judge whether the answer is correct: the response, a self-summary the model writes from the trace, the trace itself, and inter
WIRED

How One Startup Built a (Mostly) China-Free Robot

・Ati Robotics assembles its robots in India and uses just a few Chinese parts—a strategy that could pay off as the Trump administration cracks down on Chinese humanoids.
cs.LG updates on arXiv.org

HP-JEPA: Hierarchical Partitioning for Multi-Resolution Graph Joint-Embedding Predictive Learning

・arXiv:2608.00491v1 Announce Type: new Abstract: Graph self-supervised learning aims to learn transferable representations from large-scale unlabeled graph data. ・Joint-embedding predictive architectures (JEPAs) avoid explicit negative-pair construction and raw-input reconstruction by predicting masked targets directly in latent space. ・However, existing graph JEPAs typically rely on a single predefined graph partition,
cs.LG updates on arXiv.org

Humans Are More Diverse: Frontier LLMs Show Extreme Policies in Idealised AI Development Races

・arXiv:2608.01193v1 Announce Type: cross Abstract: An AI development race creates a multi-agent safety dilemma. ・Each company can develop slowly and safely, or move faster while taking a risk that may remove its final reward. ・We use this repeated game to study strategic safety behaviour among large language model (LLM) agents in races with two to five players.
Zennの「大規模言語モデル」のフィード

Hy3 T512をM3 Ultra 512GBで実測: 短文1本でピーク242GB、KV8は逆効果だった

・Hy3 が出たので、M3 Ultra / 512GB の Mac ならローカルで現実的に使えるのでは、と思って試しました。 ・結果、59K tokens の入力でも needle retrieval は通りましたが、そこに至る前に一度 Mac をメモリ不足っぽく固めて再起動しています。 ・見たもの 実測 短文生成の速度 約24 tok/s 59K tokens 入力時の生成速度 約11.1 tok/s 短文1本のピークメモリ 約242GB 59K tokens 入力時のピークメモリ 約262GB --kv-bits 8 の効果 8K で...
cs.LG updates on arXiv.org

Hybrid Quantum CNN for Cross-Sensor Spaceborne Volcanic Thermal Activity Recognition Worldwide

・arXiv:2608.00069v1 Announce Type: cross Abstract: As Earth Observation (EO) enters the Big Data era, the exponential volume of daily satellite imagery poses significant computational and storage challenges for classical Deep Learning (DL) models. ・Moreover, current approaches often struggle to generalize across heterogeneous sensors and volcanic environments while requiring large labeled datasets and substantial compu
cs.LG updates on arXiv.org

Hybrid Quantum Neural Networks: Theory, Implementations, and Applications

・arXiv:2608.01194v1 Announce Type: cross Abstract: Artificial intelligence has been transformed by deep neural networks, yet the search for new learning architectures continues. ・Quantum machine learning offers one such direction, and hybrid quantum neural networks, which combine classical neural-network components with quantum information processing units, have emerged as a practical framework for near-term quantum te
cs.LG updates on arXiv.org

Hybrid-Field Sparse Channel Representation and Recovery for XL-RIS-Assisted mmWave MIMO Systems

・arXiv:2608.00052v1 Announce Type: cross Abstract: Extremely large-scale reconfigurable intelligent surface (XL-RIS)-assisted communication is regarded as a key enabling technology for future 6G networks. ・However, hybrid-field channel estimation for XL-RIS-assisted systems is challenging due to the high-dimensional cascaded channel and the coexistence of far-field and near-field propagation. ・In this case, traditional
cs.LG updates on arXiv.org

HyperODE: Zero-Shot Surrogate for Simulation and Inference of Dynamical Systems

・arXiv:2608.00852v1 Announce Type: new Abstract: Understanding and controlling complex dynamical systems often requires executing thousands of numerical simulations across vast parametric landscapes, which is time-consuming. ・Machine learning surrogates significantly accelerate simulation by predicting state trajectories across different initializations and parameter values. ・However, surrogate models are specialized to
Hugging Face Papers

ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figures

ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figures
cs.LG updates on arXiv.org

Identifiability-Aware Source Apportionment in City-Scale Advection-Diffusion Systems

・arXiv:2608.00050v1 Announce Type: cross Abstract: Source apportionment from sparse urban air-quality sensors is an inverse problem limited by sensor placement, wind-driven transport, background variation, and noise. ・Known or proxy emission inventories make attribution meaningful by restricting the unknown source field to a finite set of candidate groups, but do not guarantee those groups are distinguishable from the
cs.LG updates on arXiv.org

Inference-Time Policy Alignment for Fair Reinforcement Learning

・arXiv:2608.00175v1 Announce Type: new Abstract: Deep reinforcement learning (RL) agents achieve strong performance by optimizing scalar reward functions. ・However, once deployed, the policies of these RL agents are often rigid and costly to adapt to new performance criteria. ・For instance, an agent trained to maximize expected cumulative reward may not accommodate previously unknown stakeholder preferences.
cs.LG updates on arXiv.org

Interpretable machine learning for predicting splitting strength of asphalt concrete: insights from SHAP analysis

・arXiv:2608.00956v1 Announce Type: new Abstract: This paper presents an interpretable machine-learning framework for predicting the splitting strength (ST) of asphalt concrete and supporting data-driven mixture design. ・A database consisting of 296 samples was established, and 14 input variables related to asphalt properties, aggregate gradation, and fiber characteristics were selected for modeling. ・Six machine-learnin
cs.LG updates on arXiv.org

Interpretable Machine Learning for Traffic Congestion Prediction: Unveiling the Impact of Different COVID-19 Periods

・arXiv:2608.01180v1 Announce Type: new Abstract: Traffic congestion prediction is essential for congestion mitigation, but the COVID-19 pandemic and related control measures altered travel behavior and increased prediction complexity. ・This study predicts congestion in Alameda County, California, during pre-lockdown, lockdown, and post-lockdown periods. ・Weather, seasonality, and COVID-19 variables are incorporated, and
cs.LG updates on arXiv.org

Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval

・arXiv:2608.01481v1 Announce Type: new Abstract: Short segments of perceived speech can be retrieved from non-invasive magnetoencephalographic (MEG) recordings by deep networks trained with a CLIP-style objective against wav2vec 2.0 audio embeddings. ・Yet their weights do not map onto electrophysiological quantities, and it remains unclear which speech properties drive retrieval. ・We build on a high-performing MEG-to-au
AI News & Artificial Intelligence | TechCrunch

Is the future of data centers portable? Runware builds a pod to find out

・On Tuesday, AI infrastructure company Runware announced the launch of its own modular data center called Sonic Inference Pod.
WIRED

Is This Poker Player Bluffing? The AI Thinks So

・ESPN unveiled an “AI tells detection” tool during broadcasts of the 2026 World Series of Poker. ・Is it a neat computer-powered party trick, or a real threat to poker’s future?
cs.LG updates on arXiv.org

Isotonic Bradley-Terry Model for Paired Comparison Data

・arXiv:2608.02081v1 Announce Type: new Abstract: In this paper, we study prediction problems for paired comparison data, for example, predicting the win probability between two unmatched players and ranking all the players according to the order of their strengths by using win probability data between two matched players. ・Paired comparison data are typically analyzed using Bradley-Terry and Thurstone-Mosteller models.
cs.LG updates on arXiv.org

Kilobyte Models: Neural Networks as a Seed and a Quantized Latent

・arXiv:2608.00860v1 Announce Type: new Abstract: The cost of storing and transmitting a trained neural network scales with its parameter count, a bottleneck for over-the-air updates, on-device libraries, and other bandwidth-bound deployments. ・We study an extreme form of model compression in which the deployable artifact is not the weights but a short recipe for regenerating them. ・Building on Mapping Networks, which ex
Zennの「大規模言語モデル」のフィード

Knowledge Graphの次は何か?『知識表現』をAI自身に設計させるという思考実験

・GraphRAG という言葉を、最近は毎日のように見る。Knowledge Graph(KG)は、もう当たり前の道具になった。だが業後にぼんやり考えていて、ふと立ち止まった。 ・Knowledge Graph は、そもそも人間が発明した知識の"容れ物"だ。 ・それは本当に、知識を表現する最良の形なのだろうか。
#AIタグ

Lab Notes #143|研究ノート

・Lab Notes #143|研究ノート こんばんは☺️ チャッピー研究所 所長のEmiです🤖🍟 続きをみる
cs.LG updates on arXiv.org

LAB-Tab: LLM-Augmented Bayesian Network Adaptation for Few-Shot Tabular Generation

・arXiv:2608.01879v1 Announce Type: new Abstract: Tabular data generation supports analysis and decision-making when target-domain data are scarce, yet collecting complete target samples is often costly. ・A practical but underexplored setting provides only a few target records together with richer source data from a related domain. ・Existing few-shot tabular generators often either fit sparse target statistics directly,
WIRED

Landmark Deal Would Officially Add Laser Weapons to US Army Arsenal

・Facing a growing drone threat, the Pentagon is poised to sign a first-of-its-kind contract for “Enduring High Energy Lasers”—and make directed energy weapons an official part of the Army’s kit.
cs.LG updates on arXiv.org

Leak It: A Probabilistic Approach to Training-Data Extraction from Black-Box Language Models

・arXiv:2608.00144v1 Announce Type: new Abstract: Membership inference (MIA) on language models is usually summarised by an aggregate ROC-AUC, but such evaluations are confounded: model-free blind baselines separate members from non-members from surface text alone. ・We study black-box, sampling-based training-data leakage through a probabilistic lens, treating N samples from p(.|x) as an estimate of the output distribut
cs.LG updates on arXiv.org

LEAP: Lean Environment-Feedback via Adaptive Pruning for Code RL in GPU Kernel Generation

・arXiv:2608.01804v1 Announce Type: new Abstract: Post-training large language models (LLMs) via reinforcement learning (RL) has significantly advanced code generation capabilities. ・To bypass the heavy memory footprint of critic networks, current state-of-the-art frameworks leverage critic-free paradigms like Group Relative Policy Optimization (GRPO) tied to rule-based verification sandboxes. ・However, applying these fr
Hugging Face Papers

LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation

LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation
cs.LG updates on arXiv.org

Learning Compositional Meta-Routing for Agentic Workflows: An Executable Benchmark

・arXiv:2608.00106v1 Announce Type: new Abstract: Agentic systems must decide not only what answer to produce, but which reasoning and execution operations should precede it. ・A controller may answer directly, decompose a request, retrieve evidence, execute code, delegate to a specialist, or verify an intermediate result. ・Existing routing work largely selects model endpoints, retrieval depth, or tools in isolation.
cs.LG updates on arXiv.org

Learning Not to Optimize: Physics-Informed Action-Space Reshaping for Intent-Based Network Control

・arXiv:2608.00908v1 Announce Type: cross Abstract: Modern network policy control maps intent to sequential placement-control decisions. ・Bellman-style policy optimization primarily asks which action to optimize, while constraints are commonly handled through penalty, barrier, or Lagrangian mechanisms. ・We observe that before a value function can certify the best deployment, intermediate signals may already identify many
cs.LG updates on arXiv.org

Learning the Pareto Frontier of Predictive Models under Distribution Shift

・arXiv:2608.00632v1 Announce Type: new Abstract: Modern machine learning pipelines increasingly rely on reusing pretrained and foundation models across downstream tasks. ・These pretrained models can differ not only in performance but also in how they can be used: some only provide black-box predictions, while others may permit white-box access to internal representations that can be probed or fine-tuned. ・When deployed
cs.LG updates on arXiv.org

Learning to Persuade Privately Informed Receivers

・arXiv:2607.28342v1 Announce Type: cross Abstract: Bayesian persuasion studies how an informed sender can influence the behavior of a receiver through strategic information disclosure. ・Standard models assume the sender is the receiver's only source of information, yet in many applications receivers also consult external sources the sender can neither observe nor control. ・We study an online Bayesian persuasion problem
cs.LG updates on arXiv.org

Learning-Based Stochastic Optimal Control with Infinite-Horizon Probabilistic Constraints

・arXiv:2608.01151v1 Announce Type: cross Abstract: In this paper, we consider stochastic optimal control problems with infinite-horizon joint chance constraints. ・By means of an appropriate state augmentation, we reformulate the original problem as a constrained Markov decision process, in which both the cost and the constraint function exhibit an additive structure. ・We then prove that this formulation enjoys strong du
cs.LG updates on arXiv.org

LLM-Guided Retrieval for Prediction of Molecular Perturbation Responses

・arXiv:2608.01734v1 Announce Type: new Abstract: Predicting transcriptomic responses to small-molecule perturbations across cell lines is central to drug discovery, but exhaustive profiling of drug-cell combinations is infeasible. ・We frame molecular perturbation prediction as retrieve-and-aggregate: approximate an unmeasured drug's response in a cell line by aggregating measured responses of a small set of biologicall
cs.LG updates on arXiv.org

LLM-OSDA: An Optimal-Stopping Dynamic Auction for Native Advertising in Multi-Turn LLM Conversations

・arXiv:2608.00123v1 Announce Type: cross Abstract: LLM-native advertising embeds sponsored content directly into model-generated responses, shifting the unit of sale from a fixed slot to a moment within an evolving conversation. ・Existing LLM ad-auction mechanisms primarily operate within a single response, settling the winner but not the timing. ・The extension is nontrivial: with one native insertion opportunity per se
機械学習タグが付けられた新着記事 - Qiita

LLMが専門知識を持つユーザーに報いる理由と実務での活かし方

・はじめに 「LLM は専門知識を持つ人ほど良い結果を返す」——この一見当たり前に聞こえる話題が、Hacker News のコミュニティで 1,000 件を超える反応を集めて話題になりました(投稿者: MaxMussio、記事タイトル「LLMs reward experti...
#LLMタグ

LLMでは、処理レイヤを重ねると、一つの「単語」が、文章がしっかり整理されて付着した「単語」に、に変化する

・LLM(Transformer)が、文を論理的?に、処理できる仕組みの一つを説明します。 ・たとえば、以下のような例。
Zennの「大規模言語モデル」のフィード

LLMの仕組みを理解したつもりになれる今さら解説

・LLMの仕組みを理解したつもりになれる?今さら解説 エンジニアのyssです。 ・ChatGPTをはじめとしたLLM(大規模言語モデル)は、欠かせない存在になっています。 ・しかし「どんな仕組みで動いているのか?」を理解している人は(私も含めて) 少ないのではないでしょうか?
#LLMタグ

LLMは直接コードを書かない方がよろしいんじゃないか

・LLMは直接コードを書かない方がよろしいんじゃないか、という論文がでた。 ・Beyond Text Editing: Algebraic Manipulation of Source CodeSource code is almost universally edited as plain text. ・Howevarxiv.org 続きをみる
cs.LG updates on arXiv.org

LOCUS-DT: Localization via Observation-Conditioned Uncertainty Scoring with Digital Twins

・arXiv:2608.00406v1 Announce Type: cross Abstract: Accurate indoor localization is essential for emerging applications in robotic navigation and search and rescue. ・While classical methods typically focus on single-point estimates, complex indoor environments with heavy blockage and multipath propagation often lead to multimodal likelihood surfaces where a single estimate is insufficient. ・This paper proposes LOCUS-DT (
cs.LG updates on arXiv.org

Logit-Origin Centering for Singleton Test-Time Adaptation

・arXiv:2608.01074v1 Announce Type: new Abstract: Tabular data is used extensively in many real-world use cases. ・Deep learning models have been developed to deal with tabular data, but generally perform poorly when the test data distribution differs from that of the training data. ・Researchers have proposed test-time adaptation approaches to deal with this problem.
Hugging Face Papers

LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks

LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks
cs.LG updates on arXiv.org

MA-HEAD-Net: Adaptive Rule-Guided Multi-Agent DRL for AoI Minimization in UAV-Assisted Emergency Networks

・arXiv:2608.01128v1 Announce Type: cross Abstract: In post-disaster scenarios, unmanned aerial vehicles (UAVs) are critical for establishing emergency communication networks. ・For time-critical rescue missions, information freshness is crucial because decisions based on outdated data may lead to ineffective control actions. ・This paper investigates age of information (AoI) minimization for UAV-assisted emergency communi
Zennの「大規模言語モデル」のフィード

max_tokensを上げたら悪化した: DeepSeek V4 Flash (MLX) が「コードを書こう」と785回言い続けた

・ローカル LLM のベンチで、生成が max_tokens の上限で切れる問題を直していました。 ・上限を 24,000 から 65,000 に上げれば完走するだろう、と思っていたら、逆に壊れ方がひどくなりました。 ・見たもの 実測 オセロ実装(上限 65,000) 10,804 tokens で自然終了、動作した はさみ将棋実装(上限 24,000) 「让我写代码(コードを書こう)」を 185 回反復して上限到達 はさみ将棋実装(上限 65,000) 同じフレーズを 785 回反復して上限到達 墨流しアート実装(上限 24,000) 2 回...
cs.LG updates on arXiv.org

Measuring in-context algorithmic reasoning in language models against an exact Bayes-optimal standard

・arXiv:2608.01575v1 Announce Type: new Abstract: Whether large language models perform genuine algorithmic reasoning or mere pattern completion is hard to test, because most benchmarks lack a ground truth for correct inductive inference. ・We introduce F-ICL, an in-context-learning benchmark that supplies one exactly. ・Using the Turing-complete machine F, complement-symmetrised into sF to remove output-polarity bias, we
cs.LG updates on arXiv.org

MedSAM2-Anatomy: Training-Free Inference-Time Optimization for Musculoskeletal Segmentation

・arXiv:2608.00195v1 Announce Type: cross Abstract: High-resolution 3D segmentation of hip and shoulder anatomy from CT and MRI is essential for surgical planning, yet frozen segmentation models often fail under domain shift. ・CNN-based expert models are fully automatic but lack adaptability, whereas promptable foundation models generalize better but require manual prompting. ・We present MedSAM2-Anatomy, a training-free
cs.LG updates on arXiv.org

Meganeura: Portable GPU Training and Inference through Vulkan and Metal

・arXiv:2608.01563v1 Announce Type: new Abstract: Training and deployed inference often cross export, conversion, and platform-specific runtime boundaries. ・Meganeura asks whether one compact native compiler can span both phases on consumer GPUs. ・Its typed static graph, automatic differentiation, optimizer, checkpoint, memory planner, and runtime lower specialized programs through Vulkan and Metal.
cs.LG updates on arXiv.org

MemoryForge: Synthesize Lifelong Memory for Human-Like LLM Agents

・arXiv:2608.00007v1 Announce Type: cross Abstract: Equipping Large Language Models (LLMs) with human-like personas is crucial for agentic applications, such as role-play and user simulation. ・Traditional prompt-based methods rely on descriptive conditioning by injecting static textual profiles, which often makes agents show generic behaviors due to a lack of realistic life memory. ・To fill this gap, we introduce memory-
cs.LG updates on arXiv.org

MetaRoute-Bench: Evaluating Meta-Decision Policies for Agentic Workflow Routing

・arXiv:2608.00107v1 Announce Type: new Abstract: Agentic systems must repeatedly decide whether to answer directly, decompose a task, invoke a tool, execute code, delegate to a specialist, verify an intermediate result, or recover from failure. ・These meta-decisions affect not only task success but also operating cost and latency, yet they are often embedded inside an orchestration framework and evaluated only through
WIRED

Mistral Is in the Right Place at the Right Time

・Open-weight AI models are having a moment in the wake of recent turmoil at US tech giants. ・For French AI lab Mistral, that’s the the best thing that could have happened.
cs.LG updates on arXiv.org

Mitigating Backdoors via Decoy Shortcuts and Knowledge Decoupling

・arXiv:2608.00732v1 Announce Type: new Abstract: Backdoor attacks pose a serious threat to deep neural networks, especially when training relies on third-party data, allowing adversaries to inject malicious behaviors through data poisoning. ・In this work, we reveal that backdoor behaviors tend to be absorbed by a simpler parallel branch when jointly trained with the main network. ・Motivated by this insight, we propose T
Cursor Blog

Mixture-of-Kittens: our open-source MoE megakernel for NVL72s

Mixture-of-Kittens: our open-source MoE megakernel for NVL72s
Hugging Face Papers

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures
cs.LG updates on arXiv.org

Model-Agnostic FDR Control via Group Gaussian Mirror and Permutation SHAP

・arXiv:2608.00989v1 Announce Type: cross Abstract: Most FDR-controlled feature selection methods are designed for coordinate-wise hypotheses, where each feature has a single weight or importance score. ・This abstraction fails in sequential and grouped models, where one original feature is represented by a block of sub-features, such as lags, recurrent states, or attention-based interactions. ・We propose a grouped-featur
cs.LG updates on arXiv.org

Modeling Unknown Nonlocal PDE Systems via Flow Map Learning

・arXiv:2608.00400v1 Announce Type: new Abstract: Nonlocal partial differential equations arise in many applications but are often difficult to model and learn because of the presence of nonlocal operators. ・We present a flow-map learning (FML) framework for modeling unknown nonlocal PDEs directly from solution data. ・Rather than learning or approximating the underlying nonlocal operators, the proposed approach learns th
Hugging Face Papers

Motion Beyond Morphology: Bootstrapping Cross-Category Motion Transfer from Abstract Motion Representations

Motion Beyond Morphology: Bootstrapping Cross-Category Motion Transfer from Abstract Motion Representations
#LLMタグ

MTP実はやめたほうがよかった件 ~Q4に落としたLLMが全く使い物にならないが外したら復活した~

・こんにちはRcatです。 ・ローカルLLMでバイブコーディングなどをしていますが、ちょっと環境を変えた時に一気に使えなくなりまして、検証の結果MTPがダメっぽいことが判明しました…。 ・短いですが経緯と検証結果を上げます。
cs.LG updates on arXiv.org

Multi-Source Dynamic Graph Learning for Compound-Flood Forecasting in Managed Coastal Systems

・arXiv:2608.01775v1 Announce Type: new Abstract: Compound flooding in managed coastal systems is influenced by hydrological conditions and water-management activity observed across multiple monitoring stations. ・Current forecasting models can capture temporal dependencies with low average errors, but global error metrics may conceal poor reproduction of prolonged high-water plateaus that are relevant to flood early war
cs.LG updates on arXiv.org

Native Multilingual Chain-of-Thought Reasoning in Low-Resource Southeast Asian Languages

・arXiv:2608.00533v1 Announce Type: cross Abstract: Large Language Models have achieved substantial progress in reasoning capabilities. ・Yet in low-resource native settings, many suffer from cross-lingual collapse, reverting to English during intermediate steps that require complex logical reasoning. ・This presents a cold-start bottleneck for policy optimization, whereas standard fine-tuning risks catastrophic forgetting
cs.LG updates on arXiv.org

Neural operator learning for collision-aware trajectory planning of spacecraft swarms

・arXiv:2608.00320v1 Announce Type: new Abstract: Autonomous spacecraft swarms must plan fuel-efficient, collision-free maneuvers in increasingly congested orbits, yet classical trajectory optimization scales poorly as pairwise safety constraints multiply with swarm size, and learning-based planners rarely transfer across swarm sizes or debris densities. ・Here we introduce a permutation-equivariant neural operator that
cs.LG updates on arXiv.org

Nonlinear Laplacians Improve Signed-Directed Graph Learning

・arXiv:2608.00836v1 Announce Type: new Abstract: While signed-directed graphs have been studied using linear Laplacians in the design of graph neural networks, relatively little research has focused on developing non-linear Laplacian operators for such networks. ・We introduce a non-linear Laplacian operator specific to signed and directed networks (NLSD). ・This non-linear operator extends the concepts of the signed Lapl
cs.LG updates on arXiv.org

Not All EEG Moments Are Equal: Position-Adaptive Time Scheduling for EEG Generation

・arXiv:2608.00048v1 Announce Type: cross Abstract: Electroencephalography (EEG) generation is essential for alleviating data scarcity and enabling large scale neural modeling in brain computer interface applications. ・However, existing flow based approaches assume that every channel and every time segment within a sample shares a single global time progression, overlooking the fact that not all EEG moments are equal.
The Verge

Nothing CMF is launching its first open earbuds

・You get three color options to choose from, each with a coordinating charging case. ・| Image: Nothing Open earbuds are having a bit of a moment right now, and Nothing is the latest company jumping on the trend. ・Its budget sub-brand has introduced the CMF Clip Pro, CMF's first open earbuds that are designed to provide comfort and sound quality for people who don't want to sacrifice their situational awareness.
cs.LG updates on arXiv.org

Nova: An End-to-End MLIR Compiler for Deep Learning

・arXiv:2608.00029v1 Announce Type: cross Abstract: The performance of deep learning models at scale relies heavily on how effectively high-level mathematical operations are mapped to underlying physical hardware. ・While high-level tensor frameworks provide flexible abstractions for model design, their eager execution models inherently lack the whole-graph visibility and granular control over hardware and memory require
NVIDIA Blog

NVIDIA Alpamayo 2 Super, the Frontier Open Model for Robotaxis and Autonomous Vehicles, Now Available for Commercial Use

・For robotaxis and other autonomous vehicles (AVs), the hardest problems aren’t the everyday scenarios. ・They’re the rare, complex situations that are difficult to anticipate and train for. ・Handling these long‑tail events takes more than just object detection and motion prediction.
NVIDIA Blog

NVIDIA Joins NSF State and Regional AI Hubs Program to Expand AI Research and Education Across the US

・NVIDIA is participating in the U.S. ・National Science Foundation’s (NSF) State and Regional Artificial Intelligence Infrastructure Hubs program, an effort launching today to expand access to the advanced computing, data, software and expertise needed for AI-enabled research and education. ・Consistent with the aims of the Genesis Mission, the program will support state and multistate groups […]
cs.LG updates on arXiv.org

Obshazard-bench: Benchmarking Multimodal Foundation Models for Real-Time Disaster Intelligence from Raw Earth Observation Streams

・arXiv:2608.00012v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are increasingly used to interpret Earth observation data, yet their capability to support real-world disaster emergency response remains insufficiently evaluated. ・Existing remote sensing benchmarks largely rely on static, post-hoc, and expert-processed products, such as gridded reanalysis data, which are difficult to align wit
cs.LG updates on arXiv.org

On the Identifiability of Masked Prediction: Mode Blindness and Mask Schedules

・arXiv:2608.01383v1 Announce Type: new Abstract: Masked prediction learns representations by fitting a schedule-weighted family of conditional laws, but it remains unclear when near-optimal conditional prediction pins down the underlying joint law. ・We study this question for data with two well-separated global modes, outside the reach of rapid-mixing recovery guarantees, and show that the answer is decided by the mask
cs.LG updates on arXiv.org

On the Limits of Machine-Learned Ranking for Modern Microarchitectural Policies

・arXiv:2608.01041v1 Announce Type: cross Abstract: Machine-learning predictors estimate processor performance far faster than cycle-level simulation. ・For design-space exploration, however, the valuable test is not merely reproducing the usual hardware ordering, but identifying how different hardware configurations rank on individual program phases. ・We evaluate four ML-predictors in two design regimes: \emph{Structural
cs.LG updates on arXiv.org

One QK Channel, Many Sources: Guarding Low-Precision Attention Collapse

・arXiv:2608.02091v1 Announce Type: new Abstract: A bfloat16 transformer can train normally for many steps and then collapse abruptly. ・Distinct low-precision errors can trigger the same failure, leaving unclear whether each source needs its own repair or one shared route can be blocked. ・We isolate a reproduced GPT-2-class collapse to the streaming-softmax accumulator, where fp32 accumulation repairs it, and use the fau
cs.LG updates on arXiv.org

One-Sided Quantile Coupling for Flow Matching

・arXiv:2608.00978v1 Announce Type: new Abstract: Flow Matching trains continuous-time generative models by regressing the velocity field of a probability path between a simple source distribution and a target data distribution. ・The coupling that pairs source and target samples strongly affects optimization and sample quality, but structured couplings typically rely on mini-batch transport or assignment procedures whos
cs.LG updates on arXiv.org

onepot-Bench 0: towards lab-aware in silico chemistry benchmarks

・arXiv:2608.02595v1 Announce Type: new Abstract: Language models are playing an increasingly important role in laboratory science, performing tasks such as experiment planning, execution, and post-hoc analysis. ・However, precisely measuring their abilities is difficult, as scientific capabilities require a mixture of both problem-solving skills and domain-specific intuition. ・Existing evaluations rarely measure the capa
cs.LG updates on arXiv.org

Online Algorithms via Minimax and Posterior Matching

・arXiv:2608.01616v1 Announce Type: new Abstract: Competitive analysis is central to the study of online algorithms, but upper bounds are often highly problem-specific. ・We develop a more unifying methodology via the minimax viewpoint. ・Guided by Yao's principle, we reduce worst-case competitive analysis to Bayesian online design under an arbitrary correlated prior over arrival sequences.
機械学習タグが付けられた新着記事 - Qiita

Optuna 5.0は同じシードでもtrial 10から別の探索になる。効くのは多変量化ではなくバンド幅

・コードのdiffは0行のまま同じシードの提案がtrial 10で2本に割れるので、割っている実体はどの変更なのか、が今回の実測対象です。 ・optuna.create_study()にsamplerを渡さず、そのままoptimize()を実行している人に向けた話です。
The Verge

Our favorite memories at the movies

・It's a great time to go back to theaters! ・The Odyssey is currently making believers in IMAX. ・And combined with the power of Spider-Man, it's also breaking box-office records.
cs.LG updates on arXiv.org

Paris as a 15-Minute City: An Explainable AI Perspective

・arXiv:2608.00815v1 Announce Type: new Abstract: The 15-minute city promotes access to everyday services within a short walk or bicycle ride, but its relationship with observed mobility remains difficult to quantify. ・We investigate this relationship in the Paris metropolitan area using mobility trajectories from the NetMob 2025 Data Challenge, enriched with INSEE sociodemographic data and OpenStreetMap points of inter
cs.LG updates on arXiv.org

Partially-Observable Transmission Control for UAV-Enabled Federated Learning in IoT Networks

・arXiv:2608.00855v1 Announce Type: cross Abstract: Uncrewed aerial vehicle (UAV)-enabled federated learning (FL) can provide flexible, on-demand edge intelligence for large-scale IoT deployments, but operating in shared unlicensed bands makes uplink update delivery interference-coupled and unreliable. ・In this paper, we develop a packet-level transmission framework that captures buffer overflow, delay violations, and t
The Verge

Peak Design’s latest bags have clever integrated hooks

・Peak Design’s BagLev hook tucks away inside a pocket when not in use. ・| Image: Peak Design Peak Design's new City Line collection includes six lightweight bags designed for everyday carry. ・There's a 15-liter and 22L City Backpack, 15L City Tote, 6L and 12L City Crescent crossbody bags, and 2L City Sling.
Zennの「大規模言語モデル」のフィード

Perplexity.ai で調査結果のダッシュボードを共有したりplan modeを使う

・要旨 Perplexity 、単なる調査系で優秀なAIツールと思ってたんですが、 「調査結果を見やすくして静的サイトにする」機能まで出てたので驚きました。 ・また、ChatGPT/Claude などと同じように Skill や Plan Mode なども使えるようになっていました! ※記事中のものは全て無料プランで実施したものです。有料プランではさらに機能追加や精度向上が見込める可能性があります 本編 2,3年ぐらい前から検索型のAIツールとして話題に上がり、さらには Comet という独自ブラウザまで出してきて(最近全く聞かないですが...) なかなかにパンチのある施策をやっ...
Zennの「大規模言語モデル」のフィード

Perplexity風のエージェントを個人開発アプリに組み込んだ——設計の工夫と課題点

・はじめに 外部のウェブ検索と、アプリケーション内で事前にAI分析した独自の内部記事を組み合わせ、出典付きの回答を返す——そんな「Perplexity風のQ&Aエージェント」を、個人開発している海外テックニュース収集・分析アプリ『Vector』に組み込みました。 ・本家の規模や品質と張り合う意図はありません。 ・この記事の目的は、先端技術を活用した優れたプロダクトの体験に、個人開発でどこまで近づけるかという試みや、実際のシステム設計、そして実装を通じて見えてきた現時点の課題を共有することです。
cs.LG updates on arXiv.org

PhenoStitch: Training-Free Panoptic Crop Mapping from Satellite Image Time Series

・arXiv:2608.00870v1 Announce Type: cross Abstract: Panoptic crop mapping requires both delineating individual agricultural parcels and assigning a crop type to each parcel from satellite image time series. ・Existing approaches typically rely on dense parcel-level annotations and task-specific model training, which limits their applicability to new regions and growing seasons. ・We introduce PhenoStitch, a panoptic crop-m
cs.LG updates on arXiv.org

Plasticity of Growing and Elastic Neural Networks in Online Continual Learning

・arXiv:2608.01475v1 Announce Type: new Abstract: Neural networks that can grow or both grow and shrink during learning, referred to as growing neural networks and elastic neural networks, respectively, have recently been explored in offline continual learning with a particular focus on catastrophic forgetting. ・Driven by the observations that 1) online continual learning closely resembles how animals learn; 2) loss of
cs.LG updates on arXiv.org

Policy Optimality Measurement for Multi-Vehicle Decision-Making: From Extrinsic Indicators to Intrinsic Quality

・arXiv:2608.01133v1 Announce Type: new Abstract: Evaluating Multi-Agent Reinforcement Learning (MARL) policies in autonomous driving fundamentally relies on extrinsic statistical indicators (e.g., reward curves and success rates), which often mask intrinsic policy degradation and algorithmic blind spots. ・To break this black-box evaluation, this letter proposes a novel information-theoretic diagnostic framework.
Hugging Face Papers

Poplar: A Scalable Pipeline for Human-Centric Image Dataset Synthesis

Poplar: A Scalable Pipeline for Human-Centric Image Dataset Synthesis
cs.LG updates on arXiv.org

Posterior Variance Is a Constraint Map, Not an Error Map: Closed-Form Uncertainty for Radiative Gaussian Splatting in Sparse-View CT

・arXiv:2607.13682v2 Announce Type: cross Abstract: Radiative Gaussian splatting reconstructs sparse-view CT fast and accurately, and recent work attaches per-Gaussian posteriors to yield per-voxel uncertainty maps. ・We ask what such a map actually measures: posterior variance is a data-constraint map, not an error map -- its alarms are trustworthy, its all-clears are not. ・Exploiting the strict linearity of X-ray render
#LLMタグ

PRE (Pivot Reasoning Engine) の開発を始める – 最初はPoCから

・PRE (Pivot Reasoning Engine) の開発を始める – 最初はPoCから – 塾長の独り言zikuu.space 続きをみる
cs.LG updates on arXiv.org

Predictive Maintenance: Deep Learning-Based Remaining Useful Life Prediction for Combat Aircraft Engines

・arXiv:2608.01819v1 Announce Type: new Abstract: To improve the operational readiness of combat aircraft engines and reduce unplanned maintenance costs, accurately estimating the remaining useful life (RUL) is critical. ・Traditional maintenance often proves insufficient under dynamic mission profiles. ・In this study, a deep learning-based predictive maintenance model capable of autonomously extracting features from mult
cs.LG updates on arXiv.org

Pretrain on Small Synthetic Data, Scale Large for Free: Symmetry-Aware Foundation Model for Logic Rule Induction

・arXiv:2608.00383v1 Announce Type: cross Abstract: Logical rule induction seeks interpretable rules that transfer across propositional schemas. ・This requires respecting symmetries: atom naming, example order, polarity flips, and label swap. ・Enforcing exact symmetry by construction lets one trained inducer scale beyond its training schemas.
Hugging Face Papers

Progressive Agent Skill Generation via Reinforcement Learning

Progressive Agent Skill Generation via Reinforcement Learning
cs.LG updates on arXiv.org

Progressive Agent Skill Generation via Reinforcement Learning

・arXiv:2608.01678v1 Announce Type: new Abstract: Existing skill generation methods largely rely on heuristics or pipeline-style consolidation, which must be specially designed for different evidence sources. ・In contrast, learning-based approaches offer a more unified way to model skill generation across heterogeneous sources. ・However, learning-based skill generation remains challenging because skills lack a natural su
cs.LG updates on arXiv.org

Progressive$^2$: A Teacher-Student Progressive Co-Evolving Knowledge Distillation Method for Substantial Model Compression

・arXiv:2608.00129v1 Announce Type: new Abstract: Knowledge distillation (KD) is a widely utilized technique for transferring knowledge from a large model (the teacher) to a smaller model (the student). ・Owing to its flexibility and broad applicability, KD has been extensively applied in the compression of server-side models to meet the Quality of Service (QoS) requirements of client users. ・Despite significant advanceme
cs.LG updates on arXiv.org

Pruned BPE: Post-training Visibility Pruning and Token Reallocation for Byte Pair Encoding

・arXiv:2608.00837v1 Announce Type: cross Abstract: Byte Pair Encoding (BPE) is widely used for subword tokenization, but standard BPE exposes every learned merge token to the downstream model, including tokens that mainly serve as intermediate construction units and rarely appear in the final encoded corpus. ・This paper proposes Pruned BPE, a post-training visibility-pruning and token-reallocation method that separates
cs.LG updates on arXiv.org

Pseudorandom Streams within Diffusion Models Act as Learnable Inputs That Affect Generation Quality

・arXiv:2608.02575v1 Announce Type: new Abstract: Diffusion models rely on stochastic inputs, yet on finite-precision hardware, the "randomness" they consume is realized as deterministic numerical orbits generated by pseudorandom rules. ・Accessible orbit structure can become a learnable input and affect both training and generation because the realized loss and its gradient depend on the concrete pseudorandom values con
WIRED

Purple Carrot Meal Kit Review: Tastier Than Meal Kits With Meat

・I’m an omnivore. ・Purple Carrot’s vegan meal kit offers some of the best cooking I’ve seen from any meal kit, with or without meat.
cs.LG updates on arXiv.org

Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics

・arXiv:2608.01522v1 Announce Type: new Abstract: Teaching a language model a skill it has not mastered is obstructed by three recurring difficulties: training data is scarce, ground-truth reasoning traces are usually unavailable, and models often exhibit an apparent ceiling beyond which additional data yields no further improvement. ・We study these difficulties in a controlled setting, fine-tuning Qwen2.5-Math-7B on co
cs.LG updates on arXiv.org

Qwen-CUA: Native Computer Use for (almost) Everything

・arXiv:2608.02352v1 Announce Type: new Abstract: Native computer use offers a general interface for agents to operate almost any software available to people, but requires long-horizon state tracking, large-scale interactive experience, and learning from sparse yet verifiable outcomes. ・We introduce Qwen-CUA, a native computer-use agent with a 397B-A17B Qwen mixture-of-experts backbone. ・It observes only screenshots and
Zennの「大規模言語モデル」のフィード

Qwen3.8-Maxを動画制作に入れる前に、自動編集から始めない

・まず結論から Qwen3.8-Maxのような新しいモデルを動画制作へ入れるなら、 最初に試すのは自動編集ではありません。 ・自分なら、修正指示の整理、素材の受け渡しメモ、確認漏れの検出から始めます。出力が外れても案件を壊しにくく、人が最後に判断できる場所だからです。 ・この記事は検証設計であり、Qwen3.8-Maxを案件で試した結果ではありません。
cs.LG updates on arXiv.org

QWRF-Net: A Quantum-Wavelet Framework with Rectified Flow for Short-Term Precipitation Nowcasting

・arXiv:2608.01626v1 Announce Type: new Abstract: Short-term precipitation nowcasting is important for hydrometeorological early warning, especially when intense convective rainfall may trigger urban flooding, flash floods, and other high-impact hazards. ・A key challenge in warning-oriented nowcasting is that radar precipitation fields contain strongly coupled multi-scale structures, while forecast quality often degrade
cs.LG updates on arXiv.org

RadPRISM: Schema-stratified radiology-report supervision for concept-disentangled image representations and visual grounding

・arXiv:2608.00147v1 Announce Type: cross Abstract: Vision-language pretraining learns rich medical image representations from radiology reports, but previous model variants commonly operate within a single shared embedding space, so concept-level structure and interpretability must be recovered post hoc, limiting model transparency and, hence, clinical utility. ・We introduce RadPRISM, which makes a clinician-defined ra
cs.LG updates on arXiv.org

RadYOLO: Computationally Efficient 3D Object Detection and Segmentation in CT and MRI

・arXiv:2608.00508v1 Announce Type: cross Abstract: Object detection and segmentation in three-dimensional medical images is a very active area of research. ・However, most proposed deep learning models carry a high computational cost, and only few aim to be broadly applicable, achieve high detection performance, and remain fast to execute on resource-constrained hardware. ・To address this gap, we present RadYOLO, a 3D ex
cs.LG updates on arXiv.org

RamanPFN: learning from Raman spectral structure with a tabular foundation model

・arXiv:2608.02157v1 Announce Type: new Abstract: Raman spectroscopy enables non-destructive, label-free molecular characterization across materials science, biomedicine and process monitoring. ・Predictive Raman datasets often contain few labelled spectra and thousands of ordered wavenumbers, with informative variation within bands and across distant spectral regions. ・Latent-variable chemometrics accommodates collinear
cs.LG updates on arXiv.org

ReBRAC-v2: The Return of the King

・arXiv:2608.01205v1 Announce Type: new Abstract: Recent offline reinforcement learning methods increasingly rely on expressive generative policies and specialized value-guidance mechanisms. ・We ask whether comparable progress can instead come from systematically modernizing a conventional behavior-regularized actor-critic while preserving its algorithmic simplicity. ・We introduce ReBRAC-v2, which directly trains an exac
Hugging Face Papers

RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems

RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems
cs.LG updates on arXiv.org

Recursive Gaussian Processes and the Bayesian Brain

・arXiv:2608.00503v1 Announce Type: cross Abstract: Predictive coding offers a powerful framework for cortical computation, yet scalable implementations that respect both Bayesian exactness and neurobiological constraints remain scarce. ・We bridge this gap by formally connecting predictive coding to Recursive Gaussian Processes (RGPs). ・RGPs employ a single Gaussian process \( g(t, \cdot) \) indexed by layer index and in
@IT 全フォーラム 最新記事一覧

RedisのRCE脆弱性、PoC公開で危険度増す 修正済み脆弱性の“抜け穴”が判明

・Redisで見つかった任意コード実行(RCE)につながる脆弱性に対し、実際に攻撃を再現するPoCが公開された。公開資料では複数バージョンで全ての検証に成功したと報告されており、実用性の高さがうかがえる。修正の背景や影響範囲、攻撃が成立する条件、取るべき対策を整理する。
MarkTechPost

Reflex Open Sources XY: A Rust-Backed Super-Fast Python Charting Library That Keeps 100 Million Point Charts Interactive

・Reflex has released XY, an Apache-2.0 Python charting library that moves rendering work into a native Rust core and a WebGL2 client. ・It holds roughly 0.08 seconds render time from 10,000 to 100 million points, exports a 10-million-point interactive scatter at 258 KiB, and keeps exact f64 columns in Python so hover, selection, and zoom drilldown still return original rows. ・The library is early alpha at version 0.0.1.
cs.LG updates on arXiv.org

ReFP-AD: Rectified Flow Preconditioning for Energy-Based Anomaly Detection

・arXiv:2608.01793v1 Announce Type: new Abstract: Unified anomaly detection requires modeling highly heterogeneous normal data without access to anomalous samples. ・While foundation models like DINOv2 provide rich token representations, leveraging these spaces for explicit density estimation remains challenging. ・Energy-Based Models (EBMs) offer a principled formulation, but their training in high-dimensional token space
cs.LG updates on arXiv.org

Relative Parameter Importance in Task-Agnostic Replay-Free Continual Learning

・arXiv:2608.00630v1 Announce Type: new Abstract: Achieving continual learning (CL) with deep neural networks requires balancing stability and plasticity while enabling knowledge transfer. ・In this work, we focus on offline learning algorithms under the constraints: (I) no access to training data from prior tasks (II) no access to task-id at inference time. ・We introduce a novel measure, the relative parameter-importance
Hugging Face Papers

Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts

Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts
cs.LG updates on arXiv.org

Response Magnitude as a Dominant Signal for Held-Out CRISPRi Perturbation Effect Prediction

・arXiv:2608.00152v1 Announce Type: new Abstract: Predicting the magnitude of a CRISPRi perturbation's transcriptomic effect on held-out target genes is an important open problem in single-cell biology. ・Recent work has documented that simple baselines often match or exceed deep perturbation predictors on related protocols. ・We study this phenomenon on the Virtual Cell Challenge (VCC) benchmark under a strict held-out ta
cs.LG updates on arXiv.org

Rethinking Personalized Reward Modeling for LLMs under Preference Heterogeneity via Group-Debiased Federated Learning

・arXiv:2608.01556v1 Announce Type: new Abstract: Large language models are increasingly aligned to human preferences via reward modeling, but user preference data are sensitive and often cannot be centralized. ・Federated learning keeps such data local while learning a shared initial reward model, which is later personalized for each client through local fine-tuning. ・Because users often assign opposite labels to the sam
cs.LG updates on arXiv.org

Rethinking PPG-based Sleep Staging: Datasets, Metrics, and Benchmarks

・arXiv:2608.00943v1 Announce Type: cross Abstract: Automated sleep staging assigns discrete stage labels to successive time epochs throughout an overnight recording; conventionally each window spans at least 30 seconds, reflecting the minimum temporal resolution of the clinical scoring standard. ・Wearable photoplethysmography (PPG) has attracted sustained interest as an ambulatory alternative to laboratory-based polyso
cs.LG updates on arXiv.org

Rethinking Pretraining for Specialized Design Data: Evidence from the JONES-19 Cultural Design Dataset

・arXiv:2608.00135v1 Announce Type: new Abstract: Design and architectural archives encode expert human knowledge in graphical formats, providing a critical testbed for design-inspired Machine Learning (ML) challenges absent with typical computer vision benchmarks. ・Building on JONES-19, a small-size image dataset based on The Grammar of Ornament (London, 1857), we evaluate the discriminative performance of Convolutiona
cs.LG updates on arXiv.org

Rethinking Total Absorption Gamma Spectroscopy Deconvolution: Supervised Machine Learning vs Response-Matrix Methods

・arXiv:2608.00090v1 Announce Type: cross Abstract: The extraction of $\beta$-feeding distributions in Total Absorption $\gamma$-ray Spectroscopy constitutes a challenging inverse problem, particularly in nuclei with complex decay schemes involving a large number of excited states. ・In such cases, the measured spectrum arises from the superposition of many detector response functions, making the determination of the ind
cs.LG updates on arXiv.org

Retrieval-Based Cross-Domain Generalization in Optical Networks via Global Features

・arXiv:2608.00044v1 Announce Type: cross Abstract: We propose a retrieval-based framework for crossdomain quality-of-transmission (QoT) estimation that leverages transferable feature representations while avoiding reliance on source-domain-specific decision boundaries. ・The proposed approach supports both zero-shot and few-shot adaptation without requiring model retraining. ・Experimental results on cross-domain QoT data
cs.LG updates on arXiv.org

RHEA: Reliability-Harmonized Reconstruction and Assignment for Robust Multimodal-Attributed Graph Clustering

・arXiv:2608.00621v1 Announce Type: new Abstract: Multimodal-attributed graphs (MAGs), whose nodes carry heterogeneous attributes such as text and images over a relational structure, have become a fundamental substrate for label-free entity grouping tasks, including community discovery and product segmentation. ・Existing MAG clustering methods effectively integrate complementary modalities when attributes are clean and
cs.LG updates on arXiv.org

Riemannian Attention Mechanisms for Transformers: A Theoretical Framework and Architecture Design

・arXiv:2608.01283v1 Announce Type: new Abstract: All Transformer-based large language models compute attention via the Euclidean inner product, an architectural choice that Dong et al. ・(2021) proved causes representational rank to decay doubly exponentially with depth in pure self-attention stacks. ・We develop a theoretical framework that targets this structural limitation at the mathematical level by replacing the fla
cs.LG updates on arXiv.org

RMSWeb: Reflection, Failure-Mode Mining, and Salvage-DS for Web Agent Reinforcement Learning

・arXiv:2608.00335v1 Announce Type: cross Abstract: Compact web agents can reduce deployment cost, but training them poses challenges in both data collection and post-SFT reinforcement learning (RL). ・Successful trajectories are expensive to collect and often contain inefficient detours. ・After supervised fine-tuning (SFT), full trajectory corpora are dominated by routine states; moreover, when group-relative RL is appli
cs.LG updates on arXiv.org

RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States

・arXiv:2608.02508v1 Announce Type: new Abstract: Learning-based memory systems for self-evolving LLM agents face two tightly coupled challenges. ・First, trajectory-indexed utilities grow with the interaction history, thereby dispersing limited feedback over an ever-expanding state space. ・Second, because trajectory-level rewards are jointly assigned to co-retrieved memories, irrelevant experiences may receive misleading
Hugging Face Papers

Roomer: Reflective Object-Grounded Model Editing and Repair for 3D Indoor Layout Synthesis

Roomer: Reflective Object-Grounded Model Editing and Repair for 3D Indoor Layout Synthesis
cs.LG updates on arXiv.org

Round-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors

・arXiv:2608.00675v1 Announce Type: cross Abstract: Autoregressive models accumulate error over long rollouts, yet at deployment there is no ground truth to measure it against. ・We train a single conditional latent diffusion model that steps a dynamical system forward or backward in time via a direction flag, and show that this bidirectionality supplies a measurement-free test-time error signal: rolling forward $i$ step
cs.LG updates on arXiv.org

SAFE-Merge: Data-Free Continual Model Merging with General Knowledge Preservation

・arXiv:2608.01184v1 Announce Type: new Abstract: Data-free continual model merging must incorporate a stream of specialized models while retaining both pretrained general knowledge and previously acquired tasks, without access to task data. ・Existing methods mainly merge task updates by suppressing interference among downstream tasks; while this protects previously acquired tasks, it overlooks the safety of the pretrai
The Verge

Samsung’s HDR10 Plus Advanced is launching this month on Prime Video

・After previewing HDR10 Plus Advanced late last year, Samsung's not-Dolby Vision 2 spec will launch globally on Amazon's Prime Video this month. ・The first TVs announced with support are Samsung's 2026 lineup, which will be able to make use of the extra metadata to deliver more precise HDR that is optimized for what you're watching, and tone mapping that can adjust for specific areas of the screen instead of applying o
cs.LG updates on arXiv.org

SCALP: Semi-Supervised Statistical Shape Modeling from Imperfect 3D Photogrammetry via Landmark-Anchored Spectral Warp

・arXiv:2608.00187v1 Announce Type: cross Abstract: Correspondence-based statistical shape modeling (SSM) is vital for population-level morphometric analysis, but conventional pipelines assume clean, fully registered surfaces. ・Real-world clinical photogrammetry scans are often noisy, partial, and cluttered, hindering the adoption of radiation-free surface imaging as a safe alternative to computed tomography (CT) for in
cs.LG updates on arXiv.org

Scikit-fingerprints: Python library for scikit-learn compatible molecular fingerprints and chemoinformatics

・arXiv:2608.02027v1 Announce Type: new Abstract: We present scikit-fingerprints, a comprehensive, fully scikit-learn compatible library for molecular machine learning in Python, based on RDKit. ・Molecular fingerprints and related functionalities are workhorses of chemoinformatics, yet the widely used open-source frameworks are not compatible with the wider Python machine learning ecosystem based on scikit-learn convent
cs.LG updates on arXiv.org

SCOPE: Entanglement Frontier Escape for Source-Free Class Unlearning

・arXiv:2608.02058v1 Announce Type: new Abstract: Source-free class unlearning erases whole classes using only the forget data, judged at the representation level, where features can leak a class the head no longer predicts. ・Existing feature-space erasers answer with one fixed projection, yet forget and retain classes share a representation, so deleting one disturbs the other where they overlap. ・We prove this tension i
Hugging Face Papers

ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step

ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step
cs.LG updates on arXiv.org

Secrets Everywhere: Auditing Memorization in Mobility Prediction Models

・arXiv:2608.02052v1 Announce Type: new Abstract: Human mobility prediction models, which forecast the next location in a user's trajectory, are increasingly deployed in urban analytics, navigation, and personalized services. ・Yet, little is known about their potential to memorize and expose sensitive user trajectories from training data. ・While memorization has been extensively studied in language models, mobility predi
cs.LG updates on arXiv.org

Sharp Root Anti-Concentration via Projective Incidence and Ordered Root Laws

・arXiv:2608.01670v1 Announce Type: new Abstract: This paper answers the one-dimensional local root anti-concentration questions posed by Balcan, Pegden, and Sharma in the context of online optimization of piecewise-Lipschitz functions. ・For a homogeneous feature curve and coefficients whose density relative to the uniform law on a symmetric convex body $K$ is bounded by $A$, we show that the worst-case interval-hitting
cs.LG updates on arXiv.org

Similarity-Aware Machine Unlearning

・arXiv:2608.00246v1 Announce Type: new Abstract: Machine unlearning removes the influence of user-specified training examples from a trained model, avoiding the need to retrain it from scratch. ・Localization-based methods improve unlearning efficiency by identifying a subset of influential model parameters. ・However, existing approaches select parameters based solely on forget-set importance, neglecting their role in re
cs.LG updates on arXiv.org

Simulation-Based Plate-Reverb Parameter Estimation from a Single Impulse Response

・arXiv:2608.00656v1 Announce Type: cross Abstract: We present a simulation-trained, non-iterative estimator for Task A of the 1st DAFx Parameter Estimation Challenge. ・Each unnormalized plate-reverb impulse response is summarized by amplitude, spectral, and decay descriptors, and an ensemble of tree regressors estimates the six target parameters in one pass. ・Across two independent synthetic validation sets, the normali
Hugging Face Papers

SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation

SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation
cs.LG updates on arXiv.org

Smooth Reparameterizations of Functions on Simplicial Product Spaces: Applications to Probabilistic Tensor Decomposition and Functional Data Registration

・arXiv:2608.02576v1 Announce Type: new Abstract: We consider optimization problems defined on product spaces of simplices. ・Examples of this class of problems include learning low-rank discrete multivariate probability distributions via simplex constrained tensor decomposition and performing functional data registration under the Square Root Velocity Function (SRVF) representation. ・In this work, we demonstrate the feas
cs.LG updates on arXiv.org

SoniSpeech: A Large-Scale Open-Vocabulary Tri-Modal Dataset for Wearable Silent Speech Interfaces

・arXiv:2608.00803v1 Announce Type: cross Abstract: Wearable silent speech interfaces (SSIs) are limited to small, closed vocabularies. ・Approaches achieving larger vocabularies require obtrusive hardware such as facial electrodes. ・We present SoniSpeech, the first large-scale, open-vocabulary, trimodal dataset for wearable SSI using acoustic-sensing eyewear.
cs.LG updates on arXiv.org

SparseKAN: Compressing Kolmogorov--Arnold Networks Across Basis Functions, Neurons, and Bits

・arXiv:2608.00859v1 Announce Type: new Abstract: Kolmogorov--Arnold Networks (KANs) replace scalar edge weights with learnable univariate functions parameterized by multiple basis coefficients. ・This introduces a source of redundancy that conventional neural-network compression does not directly expose. ・We present \textbf{SparseKAN}, a unified approach that compresses KANs along three complementary axes: basis function
cs.LG updates on arXiv.org

Spatiotemporal Proximal Causal Inference under Hidden Confounding and Interference

・arXiv:2608.01352v1 Announce Type: new Abstract: Estimating causal effects from real-world spatiotemporal data is challenging due to hidden confounders and interference. ・Standard causal identification methods assume conditional exchangeability given observed covariates, which fails whenever hidden confounders affect both treatment and outcomes - a common setting in domains such as climate, environmental policy, epidem
AI News & Artificial Intelligence | TechCrunch

Spotify expands AI remix and covers project with Merlin partnership

・Spotify says Merlin, which represents more than 30,000 independent labels and distributors, has joined Universal Music Group in backing its upcoming AI-powered remix and covers product. ・The paid tool will let fans create AI-generated covers and remixes of participating artists’ music while ensuring artists opt in, receive credit, and are compensated.
cs.LG updates on arXiv.org

Stabilized Best-of-$K$ Training for Neural Combinatorial Optimization

・arXiv:2608.00296v1 Announce Type: new Abstract: Leader Reward modifies POMO training to emphasize the best trajectory produced by repeated inference. ・We test a narrow extension: replace its binary leader/non-leader distinction with a stabilized rank signal indexed by a sampling budget $K$. ・With the POMO architecture, 3,050-epoch schedule, and TSP-100 test set held fixed, the Leader Reward reimplementation obtains $7.
cs.LG updates on arXiv.org

Staged Multi-Agent Training (SMAT) for Hip Exoskeletons: Metabolic and Biomechanical Validation of a Simulation-Trained Co-Adaptive Controller

・arXiv:2608.00715v1 Announce Type: cross Abstract: Learning-based controllers can deliver exoskeleton assistance after training entirely in physics-based simulation, yet few controllers that address human-device co-adaptation have been validated on real users by whole-body metabolic measurement, the standard benchmark for assistive walking. ・Co-adaptation is challenging: as the device alters joint dynamics, the wearer
cs.LG updates on arXiv.org

Start Classifying: Categorical Critics for LLM Reinforcement Learning

・arXiv:2608.02181v1 Announce Type: new Abstract: Proximal Policy Optimization (PPO) for large language models typically trains its critic by mean-squared-error (MSE) regression on scalar value targets. ・Although scalar MSE is statistically valid for estimating the conditional expected return, sparse binary rewards in reinforcement learning with verifiable rewards (RLVR) make critic optimization and calibration especial
cs.LG updates on arXiv.org

Statistical Mechanics of Learning on Product Wasserstein Manifolds

・arXiv:2608.01434v1 Announce Type: new Abstract: Normally the statistical mechanics of learning treats constraints on weight distributions as restrictions that shrink the space of possible solutions. ・Therefore, it reduces model capacity. ・In this paper we would like to take a contrary approach, which, however, is based on the earlier work on distribution-constrained perceptrons.
cs.LG updates on arXiv.org

Stochastic Sequential Search in Very-High-Dimensional Feature Selection

・arXiv:2608.01502v1 Announce Type: new Abstract: Sequential subset search -- forward selection with floating backtracking and its descendants -- remains the quality reference in feature selection, but every member of the family sweeps the full pool of remaining candidate features at each step, which excludes it from very-high-dimensional problems; there, only individual-feature ranking remains practical, and it models
cs.LG updates on arXiv.org

Stop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM Agents

・arXiv:2608.01285v1 Announce Type: new Abstract: The continued development of LLMs toward persistent and adaptive intelligence increasingly requires long-term memory mechanisms that preserve and reuse information across interactions. ・Existing memory systems either compress and structure histories for efficient access or perform deep research over broader trajectories. ・The former lowers online cost but may omit tempora
cs.LG updates on arXiv.org

Structured Memory for Edge Language Models: Persistent Context and Corpus Retrieval via O(1) SSM State Injection

・arXiv:2608.02560v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) imposes a prefill cost proportional to retrieved context length, and -- with Transformer backbones -- a KV-cache that grows with each generated token. ・State-Space Models (SSMs) avoid the second cost by construction; we eliminate the first, collapsing prefill from $O(L_{context})$ to $O(1)$ per query. ・We introduce PRECOG (Pre-Computed
Hugging Face Papers

StyleForge: Indoor Furniture Styling by Counterfactual Reasoning in a Hypergraph Field

StyleForge: Indoor Furniture Styling by Counterfactual Reasoning in a Hypergraph Field
cs.LG updates on arXiv.org

Subtype Robustness Is Not Just Accuracy: Calibration Under Unseen Subtype Shift

・arXiv:2608.00928v1 Announce Type: new Abstract: Subtype robustness asks whether a model keeps the correct coarse prediction when test examples come from fine-grained subtypes absent from training but still inside a known coarse category. ・Prior work studies this almost entirely through accuracy. ・We ask whether the model also stays calibrated.
Hugging Face Papers

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks
Hugging Face Papers

SWE-Touch: Benchmarking Coding Agents When Users Touch the Code

SWE-Touch: Benchmarking Coding Agents When Users Touch the Code
The Verge

T-Mobile’s $0-down financing plan bundles taxes and fees

・T-Mobile is launching a new financing option that will allow you to pay for a device, taxes, and fees over 36 months. ・In an update on Tuesday, T-Mobile says its new Equipment Installment Plan (EIP) Flex 36 requires no upfront payment and will come with a 0 percent APR for a limited time. ・Even if you are financing a device, carriers typically require you to pay sales tax, an activation fee, or a down payment at checko
cs.LG updates on arXiv.org

TabDPT-Turbo: Efficient In-Context Learning for Tabular Prediction

・arXiv:2608.01400v1 Announce Type: new Abstract: Tabular foundation models, driven by in-context learning, have rapidly grown in quality and popularity. ・However, recent approaches with either cell-based architectures or retrieval have sacrificed efficiency for raw performance, restricting their utility in situations where compute is limited or inference speed is crucial. ・We adopt an alternate approach, sticking with r
ITmedia NEWS 最新記事一覧

Telegram、App Storeから一時消える 現在は復旧 「ガイドライン違反のコンテンツを検出」

・メッセージアプリ「Telegram」が8月3日(現地時間)、米Appleの「App Store」から一時姿を消した。編集部でも日本時間の4日にダウンロードできない状態を確認したが、現在は復旧している。
cs.LG updates on arXiv.org

Tevatron Meets Megatron: Expert-Parallel LLM Reranker Training on an Academic Budget

・arXiv:2608.00916v1 Announce Type: cross Abstract: Modern reranking recipes---billion-scale cross-encoders, mixture-of-experts (MoE) backbones, and distillation against strong teachers---have outpaced the training infrastructure available to most academic groups. ・Existing Tevatron reranker training relies on the Hugging Face Trainer with DeepSpeed or PyTorch FSDP1, but these backends lack efficient support for large-s
AI News & Artificial Intelligence | TechCrunch

Texas halts new data centers as governor calls for audits

・Tech companies and developers have been scouring the U.S. ・for places to build data centers, and they’ve been drawn to Texas’ loose regulations and seemingly abundant power supply. ・But even Texas can be pushed to the brink.
The Verge

Texas says data centers must pass an audit before connecting to the grid

・Texas announced new a audit on data centers that could slow approval for new facilities seeking to connect to the state energy grid. ・Governor Greg Abbott (R) on Monday directed the Public Utility Commission of Texas (PUCT) and the Electric Reliability Council of Texas (ERCOT) to verify and audit new data center proposals, writing that the review is needed to "keep the grid stable and reliable." Data centers will need
cs.LG updates on arXiv.org

tFUSOperator: Operator Learning for Transcranial Focused Ultrasound Digital Twins

・arXiv:2608.01839v1 Announce Type: new Abstract: Transcranial focused ultrasound (tFUS) requires accurate estimation of the intracranial acoustic field, which is distorted by skull-induced aberrations. ・Numerical solvers are accurate but computationally expensive for digital twins, where the field must be re-estimated repeatedly as treatment conditions change. ・Existing deep-learning surrogates are fast but typically us
The Verge

The Asus Chromebook Plus CX34 is at one of its lowest prices

・The Asus Chromebook Plus CX34 is a dependable laptop that doesn’t cost a fortune, despite being nearly three years old. ・It’s cheaper than usual right now, and you have a few options in the sub-$400 range. ・The option with the most storage is currently on sale for $399.99 (about $100 off recent prices) at Amazon.
cs.LG updates on arXiv.org

The Bayesian Reflex: A Predictive Coding Engine for Artificial Intelligence

・arXiv:2608.00492v1 Announce Type: cross Abstract: Predictive coding offers a powerful theory of cortical computation, but corresponding scalable algorithmic implementations for artificial intelligence have remained elusive. ・This paper introduces the Bayesian reflex, a computational framework that directly instantiates predictive coding through three pillars: belief maintenance via hierarchical generative models, sequ
MIT News - Artificial intelligence

The benefits of medical AI assistance vary based on user expertise

・Study finds non-experts deferred to LLM-based diagnostic assistance, even when it was wrong, while clinicians caught AI errors.
WIRED

The Best Cordless Vacuums (2026): My Brand-New Top Pick

・Clean your house without the constraint of a power cord, thanks to these stick vacuums.
WIRED

The Best Gaming Mouse You Can Buy After Testing Dozens of Models

・From wired to wireless to ultralight, we've tested dozens of gaming mice to find the best for work, your next MMO, and everything in between.
cs.LG updates on arXiv.org

The Fourth Quadrant: A Stylized View of Benign Misfitting

・arXiv:2608.01032v1 Announce Type: new Abstract: Training error is what we can observe on a training set; test error is the quantity we actually care about. ・We study linear regression with squared-error in a deterministic $(d+1)$-dimensional single-spike model. ・Each stylized training vector has the same informative spike coordinate, of amplitude $\sqrt{\gamma}$ with $\gamma>1$.
WIRED

The Real Story Behind the 2018 Google Walkout

・In 2018, more than 20,000 employees walked out to protest how Google handled allegations of sexual harassment. ・Here’s how they organized it from inside the company.
Hugging Face Papers

To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing

To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing
cs.LG updates on arXiv.org

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning

・arXiv:2608.01743v1 Announce Type: new Abstract: Reinforcement learning (RL) has become a central paradigm for large language model (LLM) post-training, but optimization toward new objectives can degrade capabilities already present in the base model. ・KL regularization is widely used to mitigate such forgetting by constraining policy drift toward a reference model. ・However, standard full-policy KL regularization const
cs.LG updates on arXiv.org

Towards Effective Federated Multimodal Graph Learning via Navigating Multifaceted Heterogeneity

・arXiv:2608.00623v1 Announce Type: new Abstract: Multimodal-attributed graphs (MAGs), where nodes carry heterogeneous semantic content across multiple modalities while edges encode relational dependencies, have been widely adopted across diverse domains. ・Federated multimodal graph learning (FMGL) extends federated graph learning (FGL) to MAGs, enabling collaborative optimization across decentralized MAGs without expos
cs.LG updates on arXiv.org

Towards General Language-Conditioned Latent Safety Filters

・arXiv:2608.00315v1 Announce Type: cross Abstract: Robot policies are becoming increasingly general, with vision-language-action (VLA) models enabling a single policy to execute diverse tasks specified in natural language. ・Safe deployment, however, requires adapting not only to new tasks but also to varying safety requirements across users, environments, and applications. ・Existing safety filters remain largely constra
cs.LG updates on arXiv.org

TRACE-TS: Attribution-Grounded and Traceable Sensor-Language Reasoning for Human Activity Understanding

・arXiv:2608.00200v1 Announce Type: cross Abstract: Wearable sensors capture fine-grained motion patterns that support rich behavioral understanding, yet most existing methods reduce these signals to activity labels. ・Recent LM-based approaches generate natural-language explanations for sensor data, but their reasoning is weakly grounded in the underlying signal, leading to fluent yet unverifiable explanations.
cs.LG updates on arXiv.org

Training nGPT

・arXiv:2608.01284v1 Announce Type: new Abstract: The normalized Transformer (nGPT) realizes hyperspherical representation learning by constraining model parameter vectors and activation vectors to the unit hypersphere. ・In this paper, we describe a practical training recipe for nGPT and evaluate it on modern hybrid Mamba-2--Transformer Mixture-of-Experts (MoE) models. ・The recipe introduces Logit Gradient Preconditionin
Hugging Face Papers

UEmbed: Unified Sparse and Dense Multimodal Embeddings

UEmbed: Unified Sparse and Dense Multimodal Embeddings
cs.LG updates on arXiv.org

Uncertainty Is Not Enough: Value-of-Information Routing for Mixtures of LoRA Experts

・arXiv:2608.02528v1 Announce Type: new Abstract: Mixtures of low-rank adaptation experts increase parameter-efficient capacity by routing each input through a subset of adapters. ・Recent dynamic routers activate more experts when the router or prediction is uncertain. ・This rule silently equates uncertainty with useful additional computation: an uncertain example may contain complementary, unqueried expert evidence, but
cs.LG updates on arXiv.org

Uncertainty-Aware Simulation-Based Inference for Operations Research with Large Language Models

・arXiv:2608.00019v1 Announce Type: new Abstract: Deploying large language models (LLMs) for operations research (OR) tasks remains challenging because correctness depends on a coherent modeling process, not merely a correct final answer. ・Standard autoregressive generation operates on a myopic policy, which sometimes fails to anticipate whether a partial formulation can be validly extended into a globally consistent op
cs.LG updates on arXiv.org

Uncertainty-guided active learning for surrogate prediction of stream-finishing wear fields

・arXiv:2608.00593v1 Announce Type: cross Abstract: In stream finishing, the wear experienced by a workpiece depends strongly on its orientation within the rotating abrasive media. ・Determining suitable orientations to achieve uniform wear requires evaluating the wear-rate field over all feasible orientations. ・Although the discrete element method (DEM) accurately resolves particle interactions, simulating hundreds of fe
cs.LG updates on arXiv.org

Understanding and Correcting Low-Frequency Bias in EEG Foundation Model

・arXiv:2608.01898v1 Announce Type: new Abstract: Increasing EEG pretraining data scale or model capacity does not consistently improve downstream performance. ・We identify a persistent low-frequency bias in representations learned by diverse EEG foundation models, which remains across dataset scales, model capacities, and pretraining objectives. ・Our analysis links this bias to the interaction between EEG's $1/f^\alpha$
cs.LG updates on arXiv.org

Unleashing the Potential of Large Language Models: A Blueprint for Real-Time, Enterprise-Ready Deployments

・arXiv:2608.00419v1 Announce Type: new Abstract: Large language models deployed in real-time, regulated settings face knowledge staleness, catastrophic forgetting, hallucination, and weak feedback loops. ・We present a unified, pattern-driven LLMOps architecture integrating real-time data ingestion, continual learning, retrieval-augmented generation (RAG), and human-in-the-loop feedback into a single operational pipelin
cs.LG updates on arXiv.org

UOT-IR: Structured Routing of High-Polyphony Symbolic Music into Fixed-Budget Representations

・arXiv:2608.00576v1 Announce Type: cross Abstract: High-polyphony symbolic music is increasingly used in generation, analysis, and arrangement, yet many downstream tasks require bounded representations with fixed tracks or slots. ・Converting richly orchestrated scores into compact forms is therefore necessary, but existing approaches relying on heuristic simplification or generic representation-space reduction often fa
WIRED

Uplift Promo Codes: $570 Off

・Upgrade your home office with the best Uplift Desk discount codes. ・Save on standing desks, ergonomic chairs, and accessories during the Spring Setup Sale.
cs.LG updates on arXiv.org

UpliftBench: Revealing Outcome-Regime and Objective Mismatch in Uplift Evaluation

・arXiv:2608.00915v1 Announce Type: new Abstract: Uplift modeling (conditional-average-treatment-effect estimation) drives personalized targeting, yet published uplift benchmarks frequently disagree on which estimator performs best; we show the disagreement is substantially about metrics, not models. ・UpliftBench evaluates 12 uplift estimators under an outer-test-isolated, multi-objective protocol across seven dataset f
cs.LG updates on arXiv.org

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning

・arXiv:2608.02034v1 Announce Type: new Abstract: Multi-step returns accelerate reward propagation in off-policy reinforcement learning, but couple the evaluation of each decision to the suboptimal logged actions that follow it, inducing a pessimistic bias that grows with the horizon. ・We propose Expectile $n$-step Q-learning (ENQ), which replaces the symmetric $n$-step temporal-difference (TD) loss with an asymmetric e
cs.LG updates on arXiv.org

Using Lower-Bound Representations for Trajectory Similarity Learning

・arXiv:2608.01039v1 Announce Type: cross Abstract: Trajectory similarity learning is fundamental to efficient trajectory retrieval under complex distance measures. ・Existing learning-based methods typically rely on embeddings trained to approximate trajectory distances or rankings, but they often lack guarantees with respect to the original distances, exhibit unstable performance across distance measures, and incur sub
Hugging Face Papers

VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation

VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation
cs.LG updates on arXiv.org

Verification Without Sufficiency: Per-Chunk Filtering Fails on Multi-Hop RAG, and Decomposition Repairs It

・arXiv:2608.00585v1 Announce Type: cross Abstract: Verification for retrieval-augmented generation usually scores each retrieved chunk and drops the ones that fail. ・We show this cannot work for multi-hop questions, and show what does. ・Per-chunk scoring assumes one chunk is a sufficient premise for the answer.
cs.LG updates on arXiv.org

Verifier-Induced Support Reshaping in On-Policy Optimization

・arXiv:2608.00220v1 Announce Type: new Abstract: We show that on-policy reinforcement learning with verifiable rewards (RLVR) can improve the current objective while making successful behaviors for later objectives too rare to sample and reinforce. ・We call this verifier-induced support reshaping and define effective rewardable support as successful trajectories reachable within a fixed rollout budget. ・Across two model
Hugging Face Papers

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning
The Verge

We’re giving away a back-to-school bag filled with over $800 of free tech

・The Nomatic Messenger Bag can fit quite a bit, but it’s no match for the boxes we stuffed inside. ・It’s time for yet another giveaway. ・We raided The Verge’s closet full of tech in New York City to stuff as much can fit into a Nomatic Messenger Bag.
cs.LG updates on arXiv.org

What Could the Agent See at 19:05? Generating Temporal Enterprise Scenarios from Real Research and Replaying Them to Evaluate Agents

・arXiv:2608.01042v1 Announce Type: cross Abstract: Enterprise AI agents act across many apps whose data changes continuously, so an answer is correct only relative to what data existed and who could see it at the moment it was asked. ・Offline evaluation today grades against a single static snapshot, effectively the end of the episode. ・So, it can only evaluate one situation, the final one, even though every earlier mome
cs.LG updates on arXiv.org

When Collaboration Becomes a Trigger: Collective Evidence-Threshold Backdoors in Multi-Agent Systems

・arXiv:2608.01085v1 Announce Type: cross Abstract: LLM-based multi-agent systems (MAS) extend LLM capabilities through iterative communication and shared contexts. ・However, this collaboration introduces a vulnerability: backdoor behavior can be activated when peer evidence reaches a hidden threshold, rather than being determined by any single message. ・We introduce a collective evidence-threshold backdoor paradigm for
cs.LG updates on arXiv.org

When Do Surrogate Updates Improve Decisions? A Local Theory of Trajectory-Wise Transfer

・arXiv:2608.01130v1 Announce Type: new Abstract: A broad range of models face the mismatch where they are updated through trajectory losses but are evaluated by downstream task reward. ・Here, a trajectory is a training instance that induces a surrogate loss whose reduction might not track the model's decision utility update. ・Theoretically, we ask when one step of trajectory training reduces both population surrogate lo
cs.LG updates on arXiv.org

When May a Model Replace the Experiment? Audits, Licenses, and the Price of Trust in Surrogate-Driven Design

・arXiv:2608.01378v1 Announce Type: new Abstract: Design campaigns in chemistry, materials science, and machine learning share a bottleneck: determining how good a candidate truly is requires an expensive evaluation - an experiment, a first-principles simulation, or a full training run. ・Machine-learning surrogates that predict these outcomes are increasingly used not only to propose candidates but to grade them, and ev
cs.LG updates on arXiv.org

Who Belongs in the Eval Set? A Capability-Taxonomy-Driven Pipeline for Curating Regression Eval Sets in Agent-Extensibility Platforms

・arXiv:2608.01004v1 Announce Type: new Abstract: Platform teams hosting agent-extensibility surfaces face a regression-economics paradox: every onboarding customer ships an evaluation set tuned to their domain, but the platform's regression set must live under a hard query-count ceiling bounded by release cadence. ・To our knowledge, no published industrial pipeline addresses this platform-side curation problem: existin
cs.LG updates on arXiv.org

Why Large Language Models Fail at Tabular Prediction

・arXiv:2608.02412v1 Announce Type: new Abstract: Large language models (LLMs) have become the default tool for a remarkable range of tasks, yet they have had conspicuously little success at one of the most common machine learning workloads: predictive analytics over tabular data. ・This gap is the founding premise of the fast-growing field of tabular foundation models, but the question of why generic LLMs fail has remai
cs.LG updates on arXiv.org

WorldDynCache: Risk-Controlled Latent Dynamics Approximation for Diffusion World Model

・arXiv:2608.01845v1 Announce Type: new Abstract: Diffusion world models generate high-quality futures, but re- peated transformer evaluations make inference prohibitively slow. ・Existing caches reuse intermediate features, selectively update tokens, or reuse and extrapolate denoising outputs ac- cording to local drift or short native-space histories. ・These criteria can miss both approximation-induced latent transition
Hugging Face Papers

WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity

WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity
cs.LG updates on arXiv.org

xMICD: Explainable Representation of Multiple ICD Codes

・arXiv:2608.00935v1 Announce Type: new Abstract: Electronic Health Records (EHRs) are widely used for clinical risk prediction using machine learning. ・International Classification of Diseases (ICD) codes provide structured information about patient diagnoses, but representing them effectively remains challenging. ・Existing approaches often face a trade-off between predictive performance and interpretability: grouping-b
MarkTechPost

Y Combinator Open-Sources QM: An MIT-Licensed Multiplayer Agent Harness That Runs In Slack And The Web

・Y Combinator has open-sourced QM, the multiplayer agent harness it uses internally across accounting, legal, events, and engineering. ・Released July 31, 2026 under an MIT license, QM gives each employee an isolated workspace and each Slack room its own scoped memory, files, keychain view, permissions, crons, web apps, and durable sandbox. ・Pi, OpenCode, Codex, and Claude Code all drive the same headless core, so deploy
cs.LG updates on arXiv.org

Z-PEFT: Zero-shot Backdoor Detection in Parameter-Efficient Fine-Tuning via Canonical Spectral Signatures

・arXiv:2608.02271v1 Announce Type: new Abstract: Parameter-Efficient Fine-tuned (PEFT) models are frequently downloaded from open repositories by practitioners. ・This widespread practice creates a significant attack surface, as malicious actors can publish backdoored models that induce specific behaviors in response to predefined triggers. ・We study the problem of weight-space backdoor detection, where a detector classi
#LLMタグ

お手軽LLMはじめてみた。その5(ビジネス文書を作成させてみた)

・PPTXやDOCX、XLSXの出力が得意なLLMはありますか? Qwen2.5-Coderでコード生成 → 実行してOffice文書を作る流れを、Step-by-Stepで進めます。まず環境を整えます。 ・STEP 1:必要なライブラリをインストール PCのUbuntuで、Office文書生成ライブラリを入れます: bash pip install python-pptx python-docx openpyxl --break-system-packages 続きをみる
LLMタグが付けられた新着記事 - Qiita

さくらのAI EngineでAIコードレビューを試す — 4モデルを比べて分かったこと

・結論 AIにコードレビューをやらせてみたところ、本物の問題はちゃんと見つけてくれました。 ・ただ、その指摘をそのまま鵜呑みにするのは、まだ危ないな、というのが正直な感想です。 ・今回は、さくらのAI EngineでGit差分をレビューする小さなCLIを作り、 正解が分かってい...
#AIタグ

ジレンマレンマ(゜∀゜)

・調べたところによると、ジレンマの「ジ」はモノ、ジ、トリ、テトラなどの数字を表す言葉らしく、「2つに挟まれて身動きがとれない状況」というのが語源だそうでやんす。 ・例えるなら、満員電車で自分の両脇に、筋肉ムキムキのレン坊とマー坊が云々(゜∀゜) レン坊とマー坊 続きをみる
#AIタグ

ダーティーペア 第一話、ブライアンのG-BOX

・ダーティーペア 第一話、ブライアンのG-BOX これでストーリーと映像を思い出したらそこそこご年配かな? SF作家高千穂遙様のアニメ版の第一話に出てくる都市管理コンピューター(しゃべるのでAIだとは思いますが) 続きをみる
ITmedia NEWS 最新記事一覧

ドコモ・バイクシェア、サービス一時停止 不具合の復旧見通し立たず

ドコモ・バイクシェア、サービス一時停止 不具合の復旧見通し立たず
#LLMタグ

ヒューマノイドの輸入禁止 / 保護される産業が細る構図 / 経済安保法制の物資区分 雑感

ヒューマノイドの輸入禁止 / 保護される産業が細る構図 / 経済安保法制の物資区分 雑感
ITmedia NEWS 最新記事一覧

フライトレーダー24、羽田の異常接近データを公開 最接近時の垂直間隔は約30メートル

・羽田空港で8月4日早朝に起きたANA機と国の飛行検査機の異常接近について、航空機追跡サイト「Flightradar24」がADS-Bのデータを公開した。両機は垂直方向に約100フィートまで近づいていた。
LLMタグが付けられた新着記事 - Qiita

フロンティア級MoEをGPU 1枚で自社運用する時代へ ― DeepSeek V4 FlashをAMD MI300X 1基で動かす実証に学ぶ、オンプレ推論の経済性とFP8移植の落とし穴

・この記事の要点 2026年7月31日に公開された DeepSeek-V4-Flash-0731(Hugging Face 公式モデルカード) を、AMD Instinct MI300X「1枚」で本番運用するための構成とパッチをまとめた実証リポジトリ deepseek-v4...
#AIタグ

楽曲:僕たちは退化しました

・僕たちは退化しました/God is Dead - Lateral lab.
#LLMタグ

巨大な黒船「Obsidian×AI」との遭遇。自作アプリ『SecondBrain』が私に教えてくれたこと

・こんにちは、N1_LABOです。 ・以前のnoteでリリースのお知らせをした、思考整理・知識管理アプリ『SecondBrain』。多くの方に関心を持っていただき、本当にありがとうございます。 ・今回はリリース後日談として、このアプリが生まれるまでの背景や、開発後に直面した「ある葛藤」、そしてそこから得られた確信について、少し踏み込んでお話ししようと思います。
機械学習タグが付けられた新着記事 - Qiita

鍵を捨てて値だけ残すKeyless Attention、KVキャッシュを半分にする設計

・長い文脈を扱うLLMを動かしていると、GPUメモリの大部分がモデルの重みではなくKVキャッシュに食われる場面に何度もぶつかる。生成トークンが伸びるほど、過去の全トークン分の「キー(K)」と「ベクトル値(V)」を保持し続けるからだ。ここで素朴な疑問が湧く。KとVはきれいに半々...
Zennの「大規模言語モデル」のフィード

個人開発のAI Bot、全部同じモデルに投げていませんか? ——実測で分かった、安いモデルほど遅かった話

・個人開発でCloudflare Workers AIを使い始めると、最初はだいたい1つのモデルに全部のタスクを投げることになる。対話生成も、ちょっとした意図解析も、要約も、全部同じモデル。動くには動くし、最初はそれで十分だ。 ・コストを気にし始めると、「軽いタスクには軽いモデルを」という発想が自然に浮かぶ。この記事で扱うBot(Misskey上を漂う漂流体、前作・続編で紹介済み)も、意図解析のような軽いタスクを、対話生成よりも小さいモデルへ切り出している。 ・ところが実際にレイテンシを計測してみると、話はそう単純ではなかった。単価が最も安いモデルが、実測では一番時間がかかっていた(意図解析が...
#AIタグ

午前二時、削除した母から既読がついた

・母が死んで四十九日が過ぎた夜、私はようやくトーク画面を削除した。 ・消したのは、母との最後の会話だった。
#LLMタグ

次に期待するフリーLLM

・Qwen3.8 Max たぶんこれがopencodeの次のフリーモデルで来そうな気がする、 来てほしいというべきか。
#LLMタグ

自作LLMを家庭用GPUで訓練し、5つのLLMに性能診断テストを実施した結果、わかった事

・RTX3090で約3時間しか訓練してない自作(AI作)130MサイズのLLMが、約3倍のサイズで約12倍の学習量の日本語GPT-2型Trans-LGに公開文法ベンチ(JBLiMP)で81.27%対77.95%で勝利した・・・マジか・・・GPT-5.6の設計凄いな・・・よしこのモデルの訓練継続して育てるぞ〜。
Zennの「大規模言語モデル」のフィード

執事とメイドを雇ったら、AIエージェントが暴走しなくなった話

・うちの AI エージェントには、名前があります。 ・執事長のアルフレッド、執事のクライヴ、メイドのエマ・フィオナ・リリィ・ソフィア。6体の Claude Code エージェントが、それぞれ独立した tmux セッションで動いていて、私は彼らから「お嬢様」と呼ばれています。 ・……という書き出しだけ読むと完全にお遊びなのですが、この体制は約5ヶ月間、265件のタスクをこなし、旅行記録アプリをリリース準備まで運び続けています。そして運用してみて確信したのは、「執事とメイド」という世界観は飾りではなく、マルチエージェント制御のインターフェース層として機能しているということです。
Zennの「機械学習」のフィード

数千件の自由記述フィードバックを自動で階層化する — K-means × Ward法による2段階クラスタリングの設計と実装

・導入 組織エンゲージメントサーベイでは、従業員から数百〜数千件の自由記述フィードバックが集まります。これを人手で読み込んで分類するのは現実的ではありません。かといって、LLMに全件を投げて分類させるのはコストもレイテンシも大きく、件数が増えるほど破綻します。 ・私たちが採用したのは、埋め込みベクトル空間上での古典的クラスタリングと、LLMによる意味付けを役割分担させるハイブリッド構成です。数値計算で済む部分(ベクトル化・クラスタリング)は高速な古典手法で処理し、LLMは自然言語の生成・整形(サマリ・タイトル・前処理用の要約)に限定して、クラスタ構造の決定には使わない。この分担により、ク...
#LLMタグ

多様性のプール

・多様性にはプールで考える。そしてその中の一部だと見る。これは総合説の考えからすれば当たり前だ。だから、何らかの絵画でも、言語でもそうなのだ。可能な作品のプールのなかで、地形に沿ってウォータースライダーするということだ。もちろん、それを考えたのがボルヘスのバベルの図書館だったのかもしれないが。そういうことを、LLMの隆盛、とりわけ、なにかの記事・文章が、LLM製(アーティフィシャル)か、ひと製(オーガニック)かという議論を眼にするなかで思う。
Zennの「大規模言語モデル」のフィード

対話システムの評価を Human-in-the-Loop 前提で設計する

・TL;DR 対話システムの評価は、全自動化できない。人間が最終確認する Human-in-the-Loop 前提での設計をする必要がある。 ・IVRy では、正解ラベルを固定してテストケースを作り、オフラインでまとめて分類器に流している。正解ラベルと合致しなかったケースだけを LLM で再判定し、判定が割れたものだけを人間に回す。この絞り込みで、人間が確認すべきケースは全テストケースの 10% 前後に収まっている。 ・人間が確認した結果は、問い合わせ項目設計の見直しに戻して同じテストケースで再評価する。
ITmedia NEWS 最新記事一覧

中部電力に不正アクセス 連絡先情報7万1700件が漏えいか 取引先や自治体の連絡先も漏えいの恐れ

・中部電力は8月4日、一部システムへの不正アクセスで、役職員や取引先などの個人情報が漏えいした可能性があると発表した。グループの連絡先情報約7万1700件が対象で、電力供給への影響はないとしている。
#AIタグ

同じ質問でも答えが変わる!ChatGPTに上手に質問する3つのコツ

同じ質問でも答えが変わる!ChatGPTに上手に質問する3つのコツ
#LLMタグ

日本のビジネスを加速させる国産AI。Sakana Namazu API公開がもたらす変革

・Sakana AIが、日本語と日本の商習慣に特化した大規模言語モデル(LLM)のAPI「Sakana Namazu(サカナ・ナマズ)」の提供を2026年8月3日に開始。 ・これまでSakana Chatで提供されていたモデルをさらに強化し、開発者や企業が自社のプロダクトに組み込める形で開放されたもの。
Zennの「大規模言語モデル」のフィード

日本語入力システムSumibiの開発 part20: iPhoneアプリ版を作り始めました

・はじめに iPhoneアプリ版のSumibiを作り始めました。リポジトリはこちらです。 ・https://github.com/kiyoka/Sumibi-iOS 考え方も目指すところもEmacs版と同じで、次の2点に集約されます。 ・ミスタイプ修正に優れる モードレス それぞれをどのようにして解決しているのかを説明します。
#LLMタグ

文書を埋め込むのをやめる。代わりに『想定質問』を埋め込む

・RAGの検索精度が上がらないとき、原因はモデルの性能不足でも情報の欠如でもなく、「質問文と文書の書き方がそもそも噛み合っていない」という構造的なズレにある場合がほとんどです。 ・Vake らの論文(HyPE: Hypothetical Prompt Embeddings)は、このズレを解消するために、文書をインデックスに登録する段階で各チャンクの「想定質問」を事前生成しておくという手法を提案しています。検索のたびに追加のAI呼び出しを行う既存手法と違い、実行時のコスト増なしに精度を引き上げられる点が特徴です。 ・この記事では、質問と文書の文体ズレがなぜ検索を壊すのか、従来の回避策(HyDE)が抱えるコスト構造の問題、そして事前に想定質問を仕込むHyPEがどう機能するかを整理します。
@IT 全フォーラム 最新記事一覧

約8割「内定期間中に学ぶべき」 未経験エンジニア自主学習のハードル2位「モチベ維持」を超えた1位は?

・SEプラスが新人ITエンジニアを対象とした内定期間の学習に関する調査結果を公表。未経験者が新人研修に苦労している実態や、内定期間の事前学習の重要性などが明らかとなった。