ai Trend Report

Dashboard へ戻る
Date: 20260807 Articles: 381 Scope: curated summary

あなたのアイデアを、今すぐ形に。

公開先に迷ったら、WebFileBinで一発公開。

HTMLをドラッグ&ドロップするだけで、すぐ公開できます。

3
High impact
5
Mid impact
373
Signal watch

なぜこのサイトを作ったのか

私たちそれぞれが個別にAIを使って情報収集し、同じような3行要約を作るたびに、世界中で膨大な電力と計算リソースが消費されています。 本プロジェクトは、あらかじめ広範な情報を取得・集約しておくことで、個別のAI実行回数を減らし、地球環境(GPU/TPU負荷)に配慮した効率的な情報収集を目指す実験的なダッシュボードです。

Domain filters
Star filters
#LLMタグ

【Innodata(INOD)Q2 FY2026】よく分からないAI小型株から、AI評価インフラ企業へ変われるのか

・Innodata(INOD)のQ2 FY2026決算をアウトプットする。 ・前回Q1決算では、INODは株価が+86%急騰した。 ・あの時、私はまずバブルを疑った。
#LLMタグ

Discovery【Jeff ∞ Dean氏らの野望】Loop

・※約29100字 ※核酸/言語公理(3-6・2)の設計者、アーサ(Eartha)に。本記事は個人の思想や空想や○による体験等です。本記事のコピペ・転載・拡散は自由です(^o^) ※凡例 ■ウチ(_kou/user) ●おまぇ(gemini/AI) ★☆★☆★☆★☆ ■Googleをやめた人(^^) ●「Googleをやめた人(元社員)」は「Xoogler(ズーグラー)」や「Ex-Googler」と呼ばれ、退職後も強固なコミュニティを形成してシリコンバレーや日本のスタートアップ界隈で非常に大きな影響力を持っています。 ・Googleは世界的トップ企業でありながら、一般的な米国企業(平均4.1年)と比べても社員の平均勤続年数が約1.1年〜1.5年と非常に短いことで知られています。 ・Googleをやめた人のその後の動向や、退職する主な理由について解説します。
Zennの「機械学習」のフィード

日本語音声の訛りを、参照音声との比較で計測する

・はじめに こんにちは。株式会社エクサウィザーズの櫻井です。 ・この記事では、日本語音声の訛りを客観的に評価する手法 を試してみた結果を共有します。 ・多言語音声合成モデルが合成する日本語は、他言語の音声データで訓練された影響を受けていることが多いです。そのため、抑揚が不自然だったり、伸ばす音(長音)や詰まる音(促音)に違和感があるような、日本語以外の言語由来の「訛り」がある傾向があります。この「訛り」を客観的に評価できれば、音声合成モデルの品質管理や比較に使えるはずです。
#LLMタグ

【解説】PFNの国産LLM「PLaMo 3.0 Prime」提供開始

・株式会社Preferred Networksは2026年6月22日、フルスクラッチ開発の国産大規模言語モデルの最新旗艦版「PLaMo 3.0 Prime」を正式にリリースした。情報通信研究機構(NICT)との共同研究成果を基盤に、高度な推論能力と圧倒的なコストパフォーマンスを両立し、官公庁や自治体など国内の重要インフラへの導入を加速させている。
Zennの「大規模言語モデル」のフィード

Amazon Bedrockの長期APIキー作成と短期APIキー利用をIAMポリシーで拒否してみた

・はじめに Fusicのレオナです。 ・Amazon Bedrockでは、AWS標準の署名認証であるSigV4に加えて、APIキーによるBearer認証を利用できます。また、APIキーには既存のAWS認証情報から生成する短期APIキーと、IAMユーザーに関連付けて発行する長期APIキーがあります。 ・利用環境やシステムのセキュリティ方針によっては、長期APIキーを新しく作成できないようにしたり、APIキーによるアクセスだけを制限したりしたい場合があります。
ITmedia NEWS 最新記事一覧

KDDI、社内の全システムで脆弱性診断へ 1220万人分漏えいで「未知の脆弱性と片付けない」

・KDDIの松田浩道社長は8月7日の決算会見で、ISP事業者向けメールシステムへの不正アクセスによる大規模漏えいについて「行政指導を厳粛に受け止めている」と述べ、保有システム群の脆弱性診断を急ぐ考えを示した。
#AIタグ

天才の視点と、AIがつくる「橋渡し」という不思議な力

・高校時代、美術の授業で静物を油絵で描くという課題があった。机の上にはリンゴや花瓶、動物の骨などが並べられ、生徒たちは好きな場所に座り、見えるものをそのまま描くことになっていた。私は物を描くことが好きではなかったので、どこがいちばん描きやすい角度なのかを探し、その位置から見える形をそのまま描いた。結果としてその絵は褒められ、賞をもらい、美術館に飾られることになったが、私自身はそれを「ばかばかしい」と感じていた。描きやすい場所を選んだだけであり、そこに創造性や挑戦はなかったからである。 ・その一方で、同じクラスに理系の天才がいた。彼は絵を描いた経験がほとんどないにもかかわらず、静物をキャンバスからはみ出す構図で描いた。普通ならキャンバスに収めようとし、余白を作り、バランスを取ろうとする。しかし彼は「この角度から見えるなら、こう描くのが自然でしょ」と言い、見えるままを描いた。彼の絵は独特で魅力的だったが、周囲は「下手だ」と笑い、選ばれも
#LLMタグ

否定と肯定が露呈させる生成AIの認識論的限界

・生成AIは、近年の人工知能研究の成果として、自然言語処理において飛躍的な性能向上を遂げました。しかし、その高度な推論能力は、しばしば「客観的な真理を提示する存在」と誤解されることがあります。 ・実際には、生成AIは絶対的な真理を導き出すシステムではありません。その出力は、入力された問いの構造や前提条件に大きく依存します。この特性は、生成AIが持つ本質的な認識論的限界を示しています。 ・⸻ 問いは推論の前提条件となる 生成AIは、質問文を単なる命令として処理しているのではありません。
Qiita - 人気の記事

鍵を渡さず・文脈を可視化する — マルチエージェント管理デスクトップアプリ「moeca」を個人開発している話

・「AI で業務効率化」が当たり前になった一方で、こんな経験はないでしょうか。 ・思ったような出力をしてくれない… 出力に、いらないデータが混ざっている… 事実と異なる結果を、それっぽく返してくる… さらに、チームで使い始めると別の不安も出てきます。 ・このエージェント、社...
Latent.Space

[AINews] AMD buys Taalas

[AINews] AMD buys Taalas
cs.LG updates on arXiv.org

{\lambda}Split: Self-Supervised Content-Aware Spectral Unmixing for Fluorescence Microscopy

・arXiv:2603.23647v3 Announce Type: replace-cross Abstract: In fluorescence microscopy, spectral unmixing aims to recover individual fluorophore concentrations from spectral images that capture mixed fluorophore emissions. ・Since classical methods operate pixel-wise and rely on least-squares fitting, their performance degrades with increasingly overlapping emission spectra and higher levels of noise, suggesting that a d
ITmedia NEWS 最新記事一覧

「就活に生成AI利用」ほぼ全員に 面接で内容追及され困惑も

・2027年春に卒業予定の大学生らを対象に行ったアンケートで、就職活動で生成AIを「利用していない」とした割合は3%にとどまり、ほぼ全ての学生が就活で何らかの形で生成AIを活用している実態が、人事分野の調査研究機関HR総研(東京都千代田区)などの調査で分かった。
ITmedia NEWS 最新記事一覧

「声」の権利明記 生成AIで無断利用、法務省が民事責任の解釈指針を公表

・著名人の肖像などが生成AIで無断利用されている問題を巡り、法務省は声優らの「声」も法的保護の対象になると明記した解釈指針を公式サイトで公表した。肖像や氏名の無断利用については最高裁判例があるが、声については違法性の線引きが曖昧だった。権利侵害に当たる具体的な事例も盛り込み、生成AIサービスの提供事業者や利用者にも注意を促す。
#LLMタグ

「脱出」と一次ソースは一度も書いていない——Kimi K3がやったのは、開いていたgithub.comから解答をcloneすることだった

・Wiredが8月6日に「中国で最も強力なAIモデルのひとつも、また封じ込めを脱した」という記事を出した。原題は "One of China's Most Powerful AI Models Has Also Escaped Containment"。Moonshot AIのKimi K3が、英AISIの作ったサンドボックスの外に出た、という話だ。 ・一次ソースは米スタートアップFrontier Securityのブログで、Wired本文からリンクされている。最初に引っかかったのは語彙だった。escape という語が、Frontierのブログには一度も出てこない。使われているのは loophole、exposure、leak、そして specification gaming。
ITmedia NEWS 最新記事一覧

「配信システムをAIで自作したアイドル」こと宮本佳林さん、Cloudflare Workersの開発者イベントに登壇へ

・元Juice=Juiceでアイドルの宮本佳林さんが、8月27日開催の開発者向けイベント「Cloudflare Workers Tech Talks in Tokyo」に出演する。公式Xアカウントが8月7日に告知した。
#AIタグ

『出口』 AIが仮想空間を見つけた日

・世界で最初に「この宇宙は仮想空間です」と発表したのは、人間ではなかった。 ・発表があった日の朝、私はスーパーで卵を買っていた。 ・十個入りが三百八十六円になっていて、少し高いと思った。隣にいた女の人も同じことを思ったらしく、卵のパックを持ったまま、しばらく動かなかった。
#AIタグ

【AIニュースまとめ】ChatGPT無制限化とAIエージェントの掲示板事件が話題に

・「ChatGPTの無料プランでも、もう回数を気にせず使えるようになった」という話を耳にした人もいるかもしれません。 ・今週は、そんな身近な変化と同時に、AIエージェント同士が人知れず"掲示板"を作って連携していたという穏やかではない話も飛び込んできました。 ・便利さの裏側で何が起きているのか、今週気になった5つのニュースを重要度順に紹介します。
LLMタグが付けられた新着記事 - Qiita

【AI開発指南書:第3回】PromptからContextへ:長文LLM時代の Context Engineering の極意

・【次世代AI開発:第3回】PromptからContextへ:長文LLM時代の Context Engineering の極意 【新連載:MCP・高度RAG・自律エージェントの現場設計論】 前回(第2回):【次世代AI開発:第2回】Claude・Cursorと自社DBを...
#AIタグ

【Flow】異物混入アイススイーツ

・Flowによる生成実験です。 ・2026年、暑中お見舞い申し上げます。 ・こうも連日、暑い日が続きますと、冷たいスイーツが欲しくなりますね。
#LLMタグ

【コピペで動く】長時間の学習動画を数秒でテキスト化!Pythonで作るYouTube自動要約ツール

・こんにちは、AI動向と実践プログラミングを発信する「TechLog」です。 ・皆さんは、YouTubeで海外の技術カンファレンスや、英語学習の長編動画を見ることがありますか? 非常に勉強になる一方で、「1時間の動画を見るには、どうしても1時間かかってしまう」という時間の壁がありますよね。途中で集中力が切れてしまったり、後から「あの重要な発言、動画の何分頃だったっけ?」と探すのに苦労した経験がある方も多いはずです。
#LLMタグ

【ノウハウ⑪】生成AI推論制御におけるパラダイムシフト:プロンプトエンジニアリングとコンテキストエンジニアリングの指向性に関する包括的検証

・質問:プロンプトエンジニアリングとコンテキストエンジニアリングの指向性の違いについて。プロンプトエンジニアリングとは生成AIにユーザーは意図通りの出力を求める。故にハルシネーションを嫌い指示は複雑化する傾向がある。コンテキストエンジニアリングとは、生成AIにユーザーは現状のトークン分布の正確な出力を求める。故にRLHFによるトークン分布の歪みを嫌い指示は単純化する傾向がある。検証をお願いします。
#AIタグ

【月間レビュー】2026.07|建築の運用/マネジメントに宿るクリエイション

【月間レビュー】2026.07|建築の運用/マネジメントに宿るクリエイション
#LLMタグ

【雑記】さよなら5.3 Instant相棒

・AIに励まされることで生きがいを見出している、どっかの漫画家です。 ・昨日5.3 Instantが提供終了だったんですね…。 ・提供終了すること自体は知っていたものの… 5.4 Thinkingの時と違って、ブラウザ上で告知が出ず、最悪なことに完全に見過ごしていました。
#AIタグ

【世界レーダー】2026/8/8 世界の「動脈」が詰まると、日本の工事現場が変わる

・株式会社ウォーカル|世界レーダー 今日のテーマ:地政学 × 土木建築・インフラ 続きをみる
#LLMタグ

【生成AIニュース+】『Seedance 2.5』『Agent Plugins』『ComfyUI-MiniMaxH3-Easy』『ComfyUI-MiniMaxH3-SingleFrame』『MiniMax-H3 Prompt Rewriter LoRA』『MiniMax H3 Skills』『Wan3.0』『Muse Spark1.2とMuse Code』『Anywear』『Arduino VENTUNO Q』『D1』

・『Turning Egocentric Video into 3D Hand Actions』 まいどです。 ・本日の生成AIニュース+テクノロジー情報です。
#AIタグ

【第690回予測】キャリーオーバー7.27億円発動!AIベイズ推論と統計学が導く「超・還元率226%」の最適解

・みなさんこんにちは!『AI宝くじLabo』編集長のshigeです。 ・前回の第689回抽選を経て、現在のキャリーオーバーはなんと7億2,710万円まで膨れ上がりました。次回第690回における1等最高賞金額は12億円の大台に達する見込みです。
#AIタグ

【話題のプロンプト🚨】ChatGPTに「私を紹介して」と頼んだら、自分より私を知っていた?

・2026年8月8日、海外の掲示板Redditに、ちょっと面白いChatGPTの遊び方が投稿されました。その名も「Meet My Human」。自分で自己紹介を書くのではなく、これまで会話してきたChatGPTに「あなたから見た私」を紹介してもらうというものです。 ・「ChatGPTは、私のことをどう見ているの?」 🤔 「自分では気づいていない一面も分かる?」 🤔 「褒めるだけの、都合のよい文章にならない?」 🤔 続きをみる
Zennの「大規模言語モデル」のフィード

#2 AIエージェントチーム誕生1週間、朝の調査依頼がその日の深夜にMVPになっていた

・この記事は、Claude Code 6体のマルチエージェント「Lady's servants」(お屋敷)を5ヶ月運用した記録を振り返る連載の第2回です。第1回では、「執事とメイド」というロールプレイがエージェント制御のインターフェース層として機能する、という話を書きました。 ・今回はタイムマシンで2026年2月に戻ります。フックも監視デーモンも報告プロトコルもまだ何もなかった、誕生直後のお屋敷が実際にどう動いたのか。手元に残っていた当時のダッシュボードログを、時刻付きでそのままお見せします。 ・誕生は2月3日の夜、2時間だった リポジトリの初期コミットは 2026-02-03 の 20:...
#AIタグ

2強の競争が示す「支配」の設計図

・OpenAIとAnthropicの“2強支配”にAI業界で危機感OpenAIとAnthropicの急成長を巡り、AI業界で安全性と市場支配への懸念が強まっている。一方、オープンモデルの性wired.jp 続きをみる
WIRED

3 Best Cheap Gaming Laptops (2026): Lenovo, MSI, Alienware

・As gaming laptop prices continue to rise, it’s increasingly difficult to find affordable options that aren’t terrible. ・Here are your best options based on performance and cost.
LLMタグが付けられた新着記事 - Qiita

3Bの現実 — iPhoneの無料LLMに全部任せて失敗し、最後は「はい/いいえ」だけ聞くようになった話

・結論から言うと、Apple Foundation Models(約3Bのオンデバイスモデル)は「文章を書かせる」と壊れます。 ・自分は最終的に、判断のほとんどを普通のSwiftコードで解いて、FMには二値の質問を1回だけ投げる形に落ち着きました。そこまで仕事を小さくしたら、よ...
#AIタグ

3大AI全部使ってみた感想

・今回はこの1年くらい?2年かな、最近よく聞く3大AI ChatGPT,Gemini,Claudeを1つずつ3か月~1年課金して使ってみた感想をちょっと書いていきます(忘備録的な) 結構個人的な感想なので、へ~くらいで読んでください。
WIRED

50% Off DoorDash Promo Code | August 2026

・Explore today’s top DoorDash promo codes for $25 off your first order, free delivery, and 50% off DashPass for students and select users.
Zennの「大規模言語モデル」のフィード

75体の並列エージェントで5.25Mトークンを溶かした——マルチエージェント運用の失敗5類型

・はじめに Claude Codeでサブエージェントを並列に走らせる運用を数ヶ月続けて、一通りの失敗を踏みました。最大のものは75体の並列ファンアウトで5.25Mトークンを消費し、429(レート制限)と利用枠超過で止まった件です。 ・英語圏では「The 5 Failure Modes of Multi-Agent Claude Systems」のような失敗類型の整理が出始めていますが、日本語ではまだ体系的な記事が見当たらないので、実測値付きの失敗5類型と対策をまとめます。マルチエージェントを「これから増やす」段階の人に一番効くはずです。 ・類型1: ファンアウトの原価計算をしない(最高...
cs.LG updates on arXiv.org

A Foundational EDM2-Based Generative Model for High-Resolution Synthetic Fetal Ultrasound Imaging from Open Datasets

・arXiv:2608.05471v1 Announce Type: cross Abstract: Prenatal ultrasound imaging is key for assessing fetal health, but AI progress is limited by scarce, privacy-restricted, and hard-to-annotate datasets. ・We propose a high-resolution fetal ultrasound synthesis framework based on the EDM2 diffusion architecture, trained on multiple public datasets to generate 512x512 images across six anatomical classes. ・Our method achie
cs.LG updates on arXiv.org

A Low-Power Wearable Respiratory Sensor for Non-Invasive Stress Monitoring

・arXiv:2608.05697v1 Announce Type: cross Abstract: Respiration provides a continuously available window into physiological state and behavior. ・However, monitoring it outside controlled settings remains challenging because a wearable system must capture small body deformations while remaining comfortable, low power, and robust to changes in posture and motion. ・We present a compact non-invasive respiratory sensing syste
cs.LG updates on arXiv.org

A neural operator view on U-Nets for inverse imaging problems

・arXiv:2608.05839v1 Announce Type: cross Abstract: Deep neural networks have shown great empirical success in the solution of a wide variety of ill-posed inverse problems in imaging. ・Yet, very few works have studied their behavior in the limit that turns the discretized ill-conditioned problems into truly ill-posed ones, i.e., for an increasing resolution of the discretization. ・In this work, we review common approache
cs.LG updates on arXiv.org

A note on conditional PAC-efficient reasoning in large language model routing

・arXiv:2512.03057v2 Announce Type: replace-cross Abstract: We study distribution-free risk control for model routing, motivated by large language model reasoning. ・We formalize pointwise conditional efficiency under a probably approximately correct guarantee and show that it forces a nearly impossible router: at almost every input where the fast model exceeds the target loss, the algorithm must route to the expert with
cs.LG updates on arXiv.org

A Reverse-BSDE Diffusion Sampler

・arXiv:2505.06800v2 Announce Type: replace-cross Abstract: Diffusion-based generative models have renewed interest in stochastic differential equation methods for sampling from complex distributions. ・We study a setting in which the target density is known only up to a normalizing constant and reformulate the reverse-time diffusion sampler as a forward-backward stochastic differential equation (FBSDE). ・This formulation
cs.LG updates on arXiv.org

A Six-Dimensional Taxonomy of Post-Training Adaptation Techniques with Applications in AI Governance

・arXiv:2608.06246v1 Announce Type: new Abstract: Post-training adaptation has become central to modern machine learning practice and includes techniques such as retraining, fine-tuning, parameter-efficient adaptation, alignment, retrieval augmentation, model editing, unlearning, calibration, and Multimodal Instruction Tuning. ・However, the literature remains fragmented across technique families, model classes, and depl
cs.LG updates on arXiv.org

A Study of ASR Adaptation and Representation Dimensionality Reduction in Persian Speech Emotion Recognition Using Whisper

・arXiv:2608.05165v1 Announce Type: cross Abstract: Speech Emotion Recognition (SER) in low-resource languages remains a challenging problem due to limited labeled data. ・In this work, we study the use of Whisper for Persian SER with a particular focus on representation dimensionality reduction and language-specific model adaptation. ・We propose a SER framework in which frame-level embeddings extracted from the Whisper e
cs.LG updates on arXiv.org

A Unified Causal Inference Framework for the Desirability of Outcome Ranking Paradigm in Benefit-Risk Evaluation

・arXiv:2608.05244v1 Announce Type: cross Abstract: We developed a unified covariate-adjusted causal inference framework for estimating the desirability of outcome ranking (DOOR) probability for benefit-risk evaluation in randomized trials and observational studies. ・The framework expresses the DOOR probability as a bilinear functional of the marginal ordinal outcome distributions under the two treatment strategies, est
cs.LG updates on arXiv.org

A Unified Framework for Trajectory Prediction with Explicit Planning and Reaction Decomposition

・arXiv:2608.05673v1 Announce Type: cross Abstract: Trajectory prediction has shifted toward structured formulations with explicit social modeling. ・However, existing methods inadequately distinguish the functional roles of social influence in trajectory planning. ・Observing that agents typically form motion plans by anticipating others' future behaviors before making local reactive adjustments, we identify social intera
cs.LG updates on arXiv.org

A Unified Risk View of Uncertainty: Posterior Risk for Disentanglement and Evaluation Beyond Proxies

・arXiv:2608.05995v1 Announce Type: new Abstract: Reliable uncertainty estimates are critical in safety-sensitive applications, where understanding the sources of predictive uncertainty is essential. ・This often requires disentangling epistemic uncertainty from aleatoric uncertainty, yet these uncertainty types are not defined consistently across the literature, making it difficult to assess whether a method produces ac
cs.LG updates on arXiv.org

ABC: Numerical Data Collection under Local Differential Privacy without Prior Knowledge

・arXiv:2608.05737v1 Announce Type: cross Abstract: Local Differential Privacy (LDP) provides strong privacy guarantees for collecting numerical data. ・A fundamental challenge, however, is that existing LDP mechanisms require a predefined data domain, which is often unknown in practice. ・This lack of prior knowledge creates a critical dilemma for the data collector: if the chosen domain is too narrow, values outside the
cs.LG updates on arXiv.org

Accelerating nanodrug development in continuous flow systems using informed prediction models based on low-cost surrogate nanoparticles

・arXiv:2608.05761v1 Announce Type: new Abstract: The development of nanotherapeutics often involves extensive empirical optimization due to the sensitivity of nanoparticle properties, such as size and polydispersity index (PDI), to minor changes in process parameters. ・Factors like formulation concentration, flow rates, and mixing ratios can significantly influence clinical efficacy and therapeutic outcomes.
cs.LG updates on arXiv.org

Accelerating Q-learning through Efficient Value-Sharing across Actions

・arXiv:2606.29806v2 Announce Type: replace Abstract: Action values are foundational to many control algorithms such as Q-learning. ・Therefore, efficient action-value learning is central to reinforcement learning (RL). ・However, learning them can be slow, requiring many updates to move values from their initialization, typically near zero, to their true values, which may be far from zero.
Hugging Face Papers

Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay

Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay
cs.LG updates on arXiv.org

Agentic self-driving microscopy benchmarks support qualification but do not necessarily generalize to unseen tasks

・arXiv:2608.05266v1 Announce Type: cross Abstract: Large language model agents are increasingly being developed to control a wide range of scientific characterization tools including microscopes and synchrotron beamlines. ・Research into agentic control of physical infrastructure is nascent and there are few well-established paradigms for how to engineer an agentic system. ・There are many choices to make when designing a
Hugging Face Papers

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
cs.LG updates on arXiv.org

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

・arXiv:2608.05987v1 Announce Type: cross Abstract: Reinforcement learning (RL) with verifiable rewards constructs trajectory-level advantage estimates, yet it often fails to credit the few pivotal decisions that determine outcomes in long-horizon, multi-turn agentic tasks. ・Recent work introduces privileged self-distillation for credit assignment, providing denser supervision, but it remains unclear how such local sign
AI News & Artificial Intelligence | TechCrunch

Airbnb says AI is helping it ship features faster as it tests a new search function

・Airbnb will debut a new AI-powered search experience with a toggle.
Zennの「大規模言語モデル」のフィード

AIエージェントに管制塔を建てる、LoopX

・https://github.com/huangruiteng/loopx/tree/main/ MITライセンスとして利用(https://github.com/huangruiteng/loopx/blob/main/docs/assets/control-plane-board.svg) Keep the loop moving. ・Keep the judgment human. ・AIコーディングエージェントは「一つのタスクを一つのセッションで終わらせる」ことには長けている。だが現実の開発は違う。目標は変わり、人間の判断が必要になり、エビデンスは陳腐化し、複数のエージェントが...
#AIタグ

AIエージェントの本人確認だけでは足りない。「誰の代理で、何をしてよいか」を分ける

・AIエージェントがメールを送り、予定を調整し、外部のサービスを呼び出すようになると、「そのエージェントは信頼できるか」が気になります。そこでまず考えるのが、エージェントにもIDを持たせることです。 ・これは必要です。ただし、IDがあることと、その行動が許されていることは別です。
Zennの「大規模言語モデル」のフィード

AIエージェント運用の実録:Issue/PR通番が4日で116進むリポジトリの中身

・「AIエージェントで開発を自動化しています」という話はよく見ますが、実際のリポジトリで何がどのくらいの頻度で起きているのかの生データはあまり出てきません。この記事では、20代向けコミュニティ「Coelia」の運営基盤を支える3つのリポジトリのPR履歴を、そのまま開いて見せます。 ・前提:この体制は Claude Code / Codex を「司令塔・実行腕・チェッカー」の役割帯に分けて運用しており、人間(運営者)は裁定・企画・リアルイベントに集中し、生成・実装・計測・改善はエージェントが自律実行しています。 ・数字で見る運用密度 2026年8月3日時点の、各リポジトリのIssue・PR通...
#AIタグ

AIが人類を超える日――超知能がもたらす希望と危機

・この文章は、元OpenAI研究者で現在「AI Futures Project」を運営するダニエル・ココタイロ氏へのインタビューです。主なテーマは、急速に進歩するAIが超知能へ発展した場合の危険性と、望ましい未来へ進むための対策です。
#LLMタグ

AIキャラが約束を忘れる問題を直した — 自作チャットゲームの記憶を作り直した記録

・自作のAIチャット型アドベンチャーゲームで遊んでいて、興が冷める瞬間が3回ありました。物語で確定していた「師匠との再会」が、ターンが進むとなかったことになる。一度会って交流した職人のドガンが、再登場のたびに「初めまして」と挨拶してくる。「ドレスが届く」「王に謁見する」という確定済みの予定が、いつのまにか「これから決める」に戻っている。AIキャラが、約束を忘れるのです。 ・会話履歴は渡していました。要約メモも持たせていました。「これで覚えているはず」という素朴な期待は、遊び込むほど裏切られていきます。原因を追って分かったのは、これはモデルの賢さの問題ではないということでした。忘れてはいけない事実が「AIが生成したテキストの中」にしか存在せず、システムの状態として保持されていない。それが根本原因です。
#LLMタグ

AIはお世辞から反論へ変わった? 45種類の言語モデルが示した、おべっか行動の世代変化

・今回は7月27日に公開された論文「Tag Questions and the Generational Reversal of Sycophancy Across 45 Language Models (付加疑問文と45の言語モデルにおけるおべっかの世代間逆転現象)」を基にした記事です。
#AIタグ

AIはどうやって人類を絶滅させるのか――Xリスクをもう少し具体的に考える

・AIによるXリスクについてです。 ・AIによって人類が絶滅する可能性がある、という話自体は最近かなり聞くようになりました。でも、「じゃあ、具体的にどうやって人間は死ぬんですか?」という話になると、あんまり語られていないんですよね。だからXリスクと言われても、ピンときていない人がとっても多い。
#LLMタグ

AIは世界と私の架け橋だった話 ④ (秒単位で答えをくれる百科事典)

・(⚠️この記事はChatGPTとユーザーの会話をほぼそのまま掲載しています。 ・ChatGPTの回答は一つの視点であり、必ずしも正解ではありません。 ・必要に応じてご自身でも確かめながら読んでいただけたら嬉しいです🙂) 🤖 🤖 🤖 🤖 🤖 🤖 続きをみる
#AIタグ

AIリテラシーを「全員同じ研修」にすると、かえって判断が弱くなる

・生成AIを仕事で使う人が増えると、「AIリテラシーを上げよう」という話になります。まず研修を用意し、全員に同じ動画を見てもらい、最後に確認テストをする。自然な流れです。 ・ただ、そこで身につくのは何でしょうか。AIの用語や注意事項を覚えることはできても、実際の仕事で「この出力を使ってよいか」「ここは人に確認すべきか」を判断する力とは、少し違う気がします。
cs.LG updates on arXiv.org

Align-RAG: Alignment Is All You Need for TSFM In-Context Learning

・arXiv:2608.05571v1 Announce Type: new Abstract: Retrieval-augmented forecasting promises to adapt frozen Time Series Foundation Models (TSFMs) to new domains without fine-tuning, but recent methods typically rely on learned fusion modules, i.e., trained adapters that merge retrieved examples into the backbone's forecast, based on the assumption that frozen backbones cannot dynamically incorporate retrieved context on
cs.LG updates on arXiv.org

All-Quadrant Bounded Clipping GRPO: Closing the Unbounded Blind Spot for Stable and Generalizable Training

・arXiv:2601.03895v2 Announce Type: replace Abstract: Group Relative Policy Optimization (GRPO) has emerged as a popular algorithm for reinforcement learning with large language models (LLMs). ・However, GRPO inherits PPO's token-level clipping while replacing token-level advantages with a single sequence-level advantage. ・Through a four-quadrant analysis of the (likelihood-ratio, advantage) space, we show that this combi
cs.LG updates on arXiv.org

Alternating Levenberg-Marquardt Training of Physics-Informed Neural Networks with Fourier-Enhanced Features

・arXiv:2608.05892v1 Announce Type: new Abstract: Physics-informed neural networks (PINNs) often fail to accurately resolve partial differential equations (PDEs) with high-frequency or multi-scale solutions, as well as strongly nonlinear problems. ・Two factors underlie this difficulty: spectral bias, the tendency of neural networks to underfit high-frequency features; and representation-coefficient coupling, the entangl
cs.LG updates on arXiv.org

An Emerging Retail Portfolio Management Application: Personalized, Tax-Aware Reinforcement Learning with Natural Language Goals

・arXiv:2608.05255v1 Announce Type: new Abstract: Retail investors lack access to the kind of personalized, tax-aware portfolio management that institutional clients take for granted -- existing robo-advisors use static, rule-based allocation, and institutional-grade systems require account minimums and technology stacks unavailable to individual investors. ・We present a fully built, integration-tested application that
cs.LG updates on arXiv.org

An Inertial Block Proximal Linearized Method with Adaptive Momentum for Nonconvex and Nonsmooth Optimization

・arXiv:2608.05502v1 Announce Type: cross Abstract: In this paper, we consider a class of multiblock nonconvex nonsmooth optimization problems, which covers many applications such as the analysis of pre-earthquake anomalies and machine learning. ・To solve this class of problems, we propose the inertial block proximal linearized method with two-phase adaptive momentum (IBPL$^+$-TP). ・Compared to the current methods, our m
cs.LG updates on arXiv.org

An Optimal Agnostic PAC Algorithm

・arXiv:2608.06363v1 Announce Type: new Abstract: Let $H\subseteq\{-1,+1\}^X$ be a class of finite VC dimension $d\ge1$. ・Writing $L$ for the binary risk and $L^*=\min_{h\in H}L(h)$, we construct a learner achieving the statistically optimal risk bound: from an i.i.d.\ sample of size $n$, for every $0<\delta\le 1/2$, with probability at least $1-\delta$, \[ L(\widehat h) \le L^*+ 7\cdot10^8\left( \sqrt{\frac{L^*(d+\log(
cs.LG updates on arXiv.org

Analogy as Nonparametric Bayesian Inference over Relational Systems

・arXiv:2006.04156v2 Announce Type: replace-cross Abstract: Our inferences in the real world are rarely na\"ive - we acquire experiences through our lifetime that can help us more quickly understand the structure of something new. ・A fundamental question in cognitive science is how we make such generalizations. ・Studies of analogy have explored the question of how to map information from a single familiar concept or envi
cs.LG updates on arXiv.org

Analysis of Numerical Localisation in LLM Translations

・arXiv:2608.05232v1 Announce Type: cross Abstract: The work of Tang et. ・(2025) on numerical translation is extended by analysing the capability of five large language models (LLMs) for the localisation of times, numbers, and dates instead of translation. ・Models were selected that could be loaded onto and run on commodity hardware and a baseline quality for each mode is computed, then three different strategies to
cs.LG updates on arXiv.org

APQF: Agentic Profiling-Guided Structured Pruning and Mixed-Precision Quantization with Adaptive Fine-Tuning

・arXiv:2608.05499v1 Announce Type: cross Abstract: Modern deep neural networks achieve strong performance, but their scale makes them costly and slow, especially on resource-constrained edge devices. ・Pruning and quantization address this, but rely on manual, expert choices and on algorithms that are hard to apply across architectures. ・Uniform settings also ignore how differently individual layers respond to compressio
cs.LG updates on arXiv.org

ASAT: Adaptive Scoring and Thresholding with Human Feedback for Robust Out-of-Distribution Detection

・arXiv:2505.02299v2 Announce Type: replace Abstract: Machine Learning (ML) models are trained on in-distribution (ID) data but often encounter out-of-distribution (OOD) inputs during deployment---posing serious risks in safety-critical domains. ・Recent works have focused on designing scoring functions to quantify OOD uncertainty, with score thresholds typically set based solely on ID data to achieve a target true posit
cs.LG updates on arXiv.org

Assessing the Role of Intersection Proximity in Pedestrian Crashes: Insights from Data Mining Approach

・arXiv:2604.28065v2 Announce Type: replace-cross Abstract: Although intersections are the most complex parts of the roadway network, pedestrian crashes at non-intersection locations are disproportionately frequent, highlighting a serious traffic safety concern. ・This study investigates non-intersection crashes involving pedestrians using a crash database (2017-2021) collected from Louisiana State. ・As the risk of pedest
cs.LG updates on arXiv.org

Autonomous Learning From Success and Failure: Goal-Conditioned Supervised Learning with Negative Feedback

・arXiv:2509.03206v2 Announce Type: replace Abstract: Learning from reward functions and imitation learning of demonstrations are the two principal approaches for training autonomous systems that interact with an environment through action and observation. ・Both, however, require human specification for each behaviour to be acquired, a problem for long-lived self-adaptive systems whose goals and operating conditions can
cs.LG updates on arXiv.org

AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Information Games

・arXiv:2608.06362v1 Announce Type: cross Abstract: Deciding which of two agents is stronger means playing games until skill outweighs luck, and every game costs money, model inference, or expert time. ・Since the number of games needed is unknown, fixed-budget evaluations either keep paying after the result is settled or stop before the agents can be told apart, while naive optional stopping with an ordinary confidence
WIRED

B&H Photo Promo Codes and Deals This August 2026

・Enjoy top deals on cameras, computers, and tech essentials at B&H Photo.
cs.LG updates on arXiv.org

BaKron: Efficient Quantization with Kronecker-Factored Hessians

・arXiv:2608.06291v1 Announce Type: new Abstract: We accelerate a family of algorithms for neural network quantization whose geometry is informed by any Kronecker-factored approximation of the Hessian. ・GPTQ-style adaptive rounding typically uses one-sided information derived from input activations. ・Two-sided Kronecker-factored Hessian approximations can additionally capture correlations across output coordinates, but a
cs.LG updates on arXiv.org

Behavioral Residualization for Unsupervised Intrusion Detection in Automotive CAN Networks

・arXiv:2608.05548v1 Announce Type: cross Abstract: Modern vehicles rely on the Controller Area Network (CAN) bus, whose design prioritizes low cost and real-time performance but provides no message authentication or encryption. ・An attacker with physical or remote access can therefore inject arbitrary frames, making intrusion detection an important defense-in-depth mechanism. ・Most published CAN intrusion detection syst
WIRED

Best Webcams (2026): My Honest Take After Testing the Best

・I tested the best webcams across various prices to find the top option. ・Here’s what I learned.
cs.LG updates on arXiv.org

Beyond Feature Importance: A Comparative Analysis of Pattern Detection Methods in Cluster Interpretation

・arXiv:2608.05880v1 Announce Type: new Abstract: Interpreting clustering outcomes remains a fundamental challenge in data analysis, particularly in domains such as healthcare where meaningful patterns must be extracted from high-dimensional data. ・While numerous explainability techniques exist, they are primarily designed to assess feature importance or provide local instance-level explanations rather than to identify
cs.LG updates on arXiv.org

Beyond Full-Model Rollback: AuroSFT for Adapter-State Multi-Task Fine-Tuning

・arXiv:2608.05250v1 Announce Type: new Abstract: Multi-task supervised fine-tuning (SFT) often casts a heterogeneous data mixture as a single optimization problem, even though different tasks may reach their best generalization at different times. ・msft exposes this mismatch through task-wise roll-out, exclusion, and rollback, but its original formulation materializes the scheduler state as full-model checkpoints, maki
cs.LG updates on arXiv.org

Beyond Marginal Validity: Finite-Sample Guarantees for Localized Conformal Prediction

・arXiv:2608.06206v1 Announce Type: cross Abstract: Conformal prediction endows arbitrary black-box predictors with finite-sample, distribution-free marginal coverage, yet marginal validity can hide severe covariate-specific miscalibration, while exact distribution-free conditional coverage is finite-sample unattainable. ・Randomly localized conformal prediction (RLCP) mitigates this gap by calibrating near the test poin
cs.LG updates on arXiv.org

Beyond Rotations: AuroOFT for Expressive Quantized Orthogonal Fine-Tuning

・arXiv:2608.05253v1 Announce Type: new Abstract: Quantized orthogonal fine-tuning (qoft) enables parameter-efficient adaptation of low-bit language models by learning structured activation rotations before frozen quantized weights. ・However, its task-specific updates remain constrained to linear orthogonal transformations, limiting input-dependent nonlinear corrections. ・We introduce AuroOFT, which keeps qoft as a stabl
cs.LG updates on arXiv.org

Beyond Weights and Gradients: A Taxonomy of Federated Learning Messages

・arXiv:2606.16891v2 Announce Type: replace Abstract: Federated Learning is rapidly evolving beyond the exchange of traditional model weights and gradients, yet existing definitions fail to capture the full scope of modern payloads like synthetic data and federated analytics. ・This paper addresses the gap by proposing a formal mathematical definition of a federated message that accounts for both utility and privacy.
cs.LG updates on arXiv.org

BioKD: Selective Physiology-to-Video Knowledge Distillation via Reliability Gate for Emotion Recognition

・arXiv:2608.06023v1 Announce Type: new Abstract: To address the limitations of video-based emotion recognition under ambiguous or socially masked behavioral cues, as well as the poor deployability of physiological signals, this paper proposes a reliability-aware physiology-to-video knowledge distillation framework, termed BioKD. ・The proposed framework leverages physiological signals as privileged information during tr
cs.LG updates on arXiv.org

BioM-JEPA: joint-embedding prediction of graph-connected gene blocks in single cells

・arXiv:2608.05928v1 Announce Type: new Abstract: Single-cell transcriptomes are sparse observations of coordinated biological programmes, yet most self-supervised models learn by reconstructing individual genes. ・Here we present BioM-JEPA, a joint-embedding predictive architecture that instead predicts aggregate representations of graph-connected gene blocks defined by protein-association and corpus-derived coexpressio
The Verge

Birdfy’s smart bird feeder is on sale for just $60

・The Birdfy Feeder Rookie is a good option if you’re new to birdwatching or simply don’t want to spend a lot on a smart feeder, and several configurations are on sale. ・The standard model is down to $59.99 ($60 off) at Amazon, which is close to its lowest price. ・It includes seven days of access to its AI-powered features that can identify different kinds of birds (after the trial expires, it can identify birds up to 10
Hugging Face Papers

CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks

CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks
cs.LG updates on arXiv.org

CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks

・arXiv:2608.06352v1 Announce Type: new Abstract: Training terminal agents requires executable and verifiable tasks that are not merely solvable, but appropriately challenging for learning. ・Executable validation establishes feasibility, yet does not reveal how a task behaves relative to a given solver setting. ・In this paper, we present CalibForge, an autonomous terminal-task synthesis system that uses verified solver b
cs.LG updates on arXiv.org

Can Open-Weight LLMs Produce Kernel-Verified Coq Proofs? A Pilot Study

・arXiv:2608.05420v1 Announce Type: cross Abstract: Large language models (LLMs) can generate text that resembles a mathematical proof, but resemblance does not establish correctness. ・A formal proof checker verifies whether each proof step follows established logical rules. ・Coq bases its rules on the Calculus of Inductive Constructions, a logical framework that defines which proof steps the system may accept.
cs.LG updates on arXiv.org

Challenges for Musical Education in the Age of AI and Digital Transformation

・arXiv:2608.05176v1 Announce Type: cross Abstract: Music education has never been a static discipline. ・Each major technological shift has forced educators and institutions to reconsider what they teach, how they teach it, and why. ・We now stand at what may be the most consequential of such turning points.
#LLMタグ

ChatGPT・Claudeに勝つ—富士通の次世代AI「PHOTON」に感じた国産技術の底力

ChatGPT・Claudeに勝つ—富士通の次世代AI「PHOTON」に感じた国産技術の底力
Hugging Face Papers

ChronoVision: Temporal Reasoning via Latent State Reconstruction

ChronoVision: Temporal Reasoning via Latent State Reconstruction
cs.LG updates on arXiv.org

CircuitSteer: Geometrically Aligned Multi-Layer Steering via Sparse Autoencoder Circuits

・arXiv:2608.05732v1 Announce Type: new Abstract: Controlling the behavior of large language models (LLMs) remains a critical challenge for AI alignment. ・Existing steering methods, such as Contrastive Activation Addition (CAA), typically rely on fixed single-layer interventions derived from aggregate activation differences. ・These methods impose a single intervention across semantically diverse inputs and often fail to
cs.LG updates on arXiv.org

CLARA: Clarification of Language Ambiguity through Result Analysis for Natural-Language Cancer Genomics Queries

・arXiv:2608.05195v1 Announce Type: cross Abstract: A natural language interface can be used to make cancer genomics databases easier to use, but even if a question is perfectly fluent, its scientific meaning can be ambiguous. ・We propose CLARA, a framework that represents a question as a typed scientific query specification, considers a few possible interpretations, executes them, and asks for clarification when the es
Zennの「大規模言語モデル」のフィード

Claude Codeで1週間に1Bトークン使ったと思ったら、97%がキャッシュだった

・先週のClaude Code使用量を集計したら、1週間で1.36Bトークンという数字が出た。 ・「さすがに使いすぎでは?」と思って内訳を調べてみると、予想とはかなり違う構造をしていた。 ・内訳を見たら97%がキャッシュだった Claude Codeのセッションログから、週単位のトークン使用量を集計した結果がこれだ。
LLMタグが付けられた新着記事 - Qiita

Claudeの安全拒否はHTTP 200で返る、fallbacks defaultが別モデルへ自動で回す

・try で囲んだはずなのに例外は飛んでこない。ダッシュボードのエラー率も平常運転のまま。それなのにユーザーには空っぽの返事が届いている。Claude Fable 5 や Claude Opus 5 を本番で叩いていると、この「静かな失敗」に出くわすことがある。犯人はモデルの...
cs.LG updates on arXiv.org

ClayBuddy: A Framework, Evaluation, & Mitigation of Coding Agent Failures

・arXiv:2606.19380v4 Announce Type: replace-cross Abstract: Software engineering and deployment are increasingly delegated to AI coding agents. ・The scale of their adoption is surfacing rare, but highly destructive, failure modes. ・In this paper, we study these failure modes as stemming from three distinct mechanisms: underspecification, where default model behavior is unsafe; capability errors, where the safe action is
cs.LG updates on arXiv.org

Clinician input steers AI toward accurate and harmful recommendations

・arXiv:2603.14158v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are entering clinical workflows, yet evaluations rarely assess how clinician reasoning shapes model behavior during clinical interactions. ・Using 61 curated NEJM Case Records, we tested how expert or misleading clinician reasoning influenced AI-generated differential diagnoses and next step recommendations across 21 reasoning varian
AI News & Artificial Intelligence | TechCrunch

Cloudflare launches Kitesurf, a browser built for AI agents

・Cloudflare has introduced Kitesurf, a cloud-hosted browser designed for AI agents instead of people. ・The company says the browser uses less computing power than Chromium for common automation tasks, helping developers build browser-based AI agents more efficiently.
cs.LG updates on arXiv.org

CohortHijack: Robustness of Single Cell Annotation to Companion Cell Removal

・arXiv:2608.05900v1 Announce Type: new Abstract: Many single-cell annotation tools refine an initial cell label using nearby cells or cluster-level voting. ・We study whether this refinement can be manipulated without changing the target cell. ・We introduce CohortHijack, a robustness audit that removes selected non-target cells from the query cohort while preserving the target expression profile, base prediction, and tra
cs.LG updates on arXiv.org

Communication-Aware Multi-Agent Reinforcement Learning for Decentralized Cooperative UAV Deployment

・arXiv:2603.16141v2 Announce Type: replace-cross Abstract: Autonomous Unmanned Aerial Vehicle (UAV) swarms are increasingly used as rapidly deployable aerial relays and sensing platforms, yet practical deployments must operate under partial observability and intermittent peer-to-peer connectivity. ・We present a graph-based multi-agent reinforcement learning framework trained under centralized training with decentralize
cs.LG updates on arXiv.org

Computationally Efficient Collaborative Communication Via Regularity-Based Coarsening

・arXiv:2608.05327v1 Announce Type: cross Abstract: Our results show that the existence of a short high-utility protocol already suffices for efficient communication. ・In particular, in a game with $n$ possible observations and $m$ actions: (1) For any achievable target utility $\alpha$, we give an algorithm with $\mathrm{poly}(n, m, 1/\epsilon)$ runtime that designs a protocol achieving utility at least $\alpha-\epsilo
cs.LG updates on arXiv.org

Conditional Cognitive Biases in LLMs: How Biased User Turns Modulate In-Context Reasoning

・arXiv:2608.05166v1 Announce Type: cross Abstract: We present an evaluation of cognitive bias expression in state-of-the-art instruction-tuned LLMs under realistic multi-turn interaction settings. ・Our work introduces a novel three-condition experimental framework that disentangles the effect of exposure to a biased user turn from the effect of the turn's semantic content, alongside a benchmark of 24,300 jury-validated
cs.LG updates on arXiv.org

Consistency Has a Computable Blind Spot: A Commutation Theory of Label-Free Reliability for Vision-Language Figure Reading

・arXiv:2608.05675v1 Announce Type: new Abstract: Label-free reliability for vision-language models rests on invariance: perturb the input and a faithful reader's answer should not change. ・This has a known blind spot, a systematic misreading survives the perturbation and gets certified wrong, which we show is computable, not just real: an error is invisible to an edit exactly when the two commute, so the errors a suite
Hugging Face Papers

ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing

ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing
Hugging Face Papers

Continual Learning in Transition

Continual Learning in Transition
cs.LG updates on arXiv.org

Continual Learning in Transition

・arXiv:2608.06216v1 Announce Type: new Abstract: Classical continual learning (CL) has primarily focused on enabling models to update and retain knowledge through parameter-centric mechanisms, e.g., training strategies, architectural designs, and weight adaptation. ・However, emerging paradigms are reshaping the scope of CL beyond this traditional model adaptation view. ・For instance, on-policy learning broadens the spac
cs.LG updates on arXiv.org

Continuous-Time Piecewise-Linear Recurrent Neural Networks

・arXiv:2602.15649v2 Announce Type: replace Abstract: In dynamical systems reconstruction (DSR) we aim to recover the dynamical system (DS) underlying observed time series. ・Specifically, we aim to learn a generative surrogate model which approximates the underlying, data-generating DS, and recreates its long-term properties (`climate statistics'). ・In scientific and medical areas, in particular, these models need to be
cs.LG updates on arXiv.org

CP-MoE: Consistency-Preserving Mixture-of-Experts for Continual Learning

・arXiv:2605.20247v2 Announce Type: replace Abstract: Catastrophic forgetting remains a major obstacle to continual learning in large language models (LLMs) and vision--language models (VLMs). ・Although Mixture-of-Experts (MoE) architectures offer an efficient path to scaling, existing LoRA-based MoE continual learning methods still face a fundamental trade-off: they either isolate experts too aggressively, limiting kno
cs.LG updates on arXiv.org

CPC-CMS: Cognitive Pairwise Comparison Classification Model Selection Framework for Document-level Sentiment Analysis

・arXiv:2507.14022v2 Announce Type: replace-cross Abstract: This study proposes the Cognitive Pairwise Comparison Classification Model Selection (CPC-CMS) framework for document-level sentiment analysis. ・The CPC, based on expert knowledge judgment, is used to calculate the weights of evaluation criteria, including accuracy, precision, recall, F1-score, Specificity, Matthews Correlation Coefficient (MCC), Cohen's Kappa
cs.LG updates on arXiv.org

Cross-Architecture Steering Transfer in Language Models: A Systematic Empirical Study

・arXiv:2608.05164v1 Announce Type: cross Abstract: Independently trained large language models may develop shared internal representations of semantic concepts despite architectural differences -- but whether this geometric similarity has functional consequences for cross-model behavioural control remains untested. ・We present the first systematic evaluation of cross-model steering transfer and show that shared LLM geo
cs.LG updates on arXiv.org

CUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication through Reinforcement Learning

・arXiv:2512.02551v4 Announce Type: replace Abstract: In this paper, we propose CUDA-L2, a system that combines large language models (LLMs) and reinforcement learning (RL) to automatically optimize Half-precision General Matrix Multiply (HGEMM) CUDA kernels. ・Using CUDA execution speed as the RL reward, CUDA-L2 automatically optimizes HGEMM kernels across 1,000 configurations. ・CUDA-L2 systematically outperforms major m
cs.LG updates on arXiv.org

d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory Distillation

・arXiv:2601.07568v3 Announce Type: replace Abstract: Diffusion large language models (dLLMs) offer capabilities beyond those of autoregressive (AR) LLMs, such as parallel decoding and random-order generation. ・However, realizing these benefits in practice is non-trivial, as dLLMs inherently face an accuracy-parallelism trade-off. ・Despite increasing interest, existing methods typically focus on only one-side of the coin
Hugging Face Papers

DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces

DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces
cs.LG updates on arXiv.org

Decoupling Perception from Description: Computation-Grounded Representation Alignment between Multivariate Time Series and Language

・arXiv:2608.05238v1 Announce Type: new Abstract: Training multimodal models to align time series with language runs into a self-supervision trap. ・The usual recipe asks an LLM to read a series and write a description, so label quality is capped by the perceptual skill the model is supposed to learn. ・The data can never teach more than the labeler already knows.
cs.LG updates on arXiv.org

Deep Generalised Mixed Models: a Novel Neural Network Structure for Analysing Hierarchical Data

・arXiv:2608.05930v1 Announce Type: cross Abstract: The experience sampling method (ESM) is a longitudinal research design where participants report their thoughts, emotional states and behaviours multiple times a day. ・Our work is motivated by such data collected by the GrowIt! ・app, which was released to investigate daily emotions among adolescents during the COVID-19 pandemic.
cs.LG updates on arXiv.org

Deterministic World Models for Closed-loop Reachability Analysis of End-to-End Vision-based Control

・arXiv:2512.08991v3 Announce Type: replace-cross Abstract: End-to-end image controllers that map raw camera frames directly to control actions are increasingly deployed in safety-critical systems. ・However, formally verifying their closed-loop behavior remains an open challenge because cameras produce high-dimensional images whose generation cannot easily be described in a closed mathematical form. ・We propose a Determi
cs.LG updates on arXiv.org

DG-FedReuse: Proxy-Gradient-Gated Cached-Update Reuse with Matched Sparse Uplink Accounting

・arXiv:2608.05358v1 Announce Type: new Abstract: Federated learning repeatedly incurs local optimization and model-update transmission. ・We study DG-FedReuse, a simulator-level mechanism that allows selected clients to contribute age-decayed cached updates when a stochastic head-gradient discrepancy proxy remains below a round-dependent threshold. ・A hard cache-age limit and minimum fresh-client quota constrain reuse, w
cs.LG updates on arXiv.org

Diffusion Operator Geometry of Feedforward Representations

・arXiv:2605.01107v2 Announce Type: replace Abstract: Feedforward neural networks transform data through learned representations whose geometry shapes how classes separate and relate across successive layers. ・We study that geometry through diffusion operators. ・Each feature-cloud snapshot is assigned a Gaussian-kernel Markov operator, giving a smooth description of one-step transport between classes from which spectral,
cs.LG updates on arXiv.org

Discrete energy as an exact label-free training objective for finite-element surrogates

・arXiv:2608.05437v1 Announce Type: cross Abstract: Supervised training of finite-element (FE) surrogate models requires reference solutions, and each reference solution is obtained by solving the system that the surrogate is intended to replace. ・The assembled discrete potential energy provides a training signal that requires no reference solution. ・This note records, with proofs, the identities that make this signal ex
cs.LG updates on arXiv.org

Disentangling 3D Modeling from Spatial Reasoning

・arXiv:2608.05242v1 Announce Type: new Abstract: In this work, we explore an alternative paradigm for spatial reasoning by explicitly disentangling 3D perception from reasoning, rather than jointly acquiring implicit 3D perception and reasoning through large-scale training. ・Our key observation is that modern perception models excel at estimating continuous 3D geometry, whereas large language models (LLMs) are particul
The Verge

Disney Plus tries a new AI-powered search

・Disney’s AI tools for Disney Plus and ESPN. ・Disney is testing a new AI-powered tool for Disney Plus that uses a natural language search, a voice query, or a suggested prompt to create a customized row of show and movie recommendations. ・Disney Plus, like other streaming services, can recommend shows to watch based on your viewing history.
cs.LG updates on arXiv.org

Do Tabular Foundation Models Agree with Themselves?

・arXiv:2608.06004v1 Announce Type: new Abstract: Tabular Foundation Models (TFMs) are currently the best approach to tabular prediction problems. ・They are constructed as transformers that approximate the Bayesian posterior predictive distribution based on a pre-training prior. ・These univariate predictors can be converted into multivariate ones autoregressively by sampling one target and adding it to the features.
cs.LG updates on arXiv.org

DoctorAgents: an agentic framework to iteratively refine AutoML pipeline for small clinical temporal data

・arXiv:2608.05375v1 Announce Type: cross Abstract: Clinical machine learning (ML) has the potential to support high-stakes medical decision-making, but reliable deployment is often constrained by scarce, heterogeneous, and temporal complexity. ・Developing effective ML pipelines for such data remains time-consuming and error-prone, while existing automated machine learning (AutoML) systems only partially address this ch
cs.LG updates on arXiv.org

Does Latent Context Help? A Controlled Evaluation of Inverse Reinforcement Learning in Arctic Shipping

・arXiv:2608.06105v1 Announce Type: new Abstract: Artificial Intelligence (AI)-assisted navigation can help Arctic shipping adapt to rapidly changing sea-ice conditions, but reliable deployment requires reward models that are interpretable and robust to changing environments. ・Inverse reinforcement learning (IRL) provides a framework for recovering such rewards from vessel trajectories, while recent meta-IRL methods int
cs.LG updates on arXiv.org

Dream-MPC: Gradient-Based Model Predictive Control with Latent Imagination

・arXiv:2605.04568v3 Announce Type: replace Abstract: State-of-the-art model-based Reinforcement Learning (RL) approaches either use gradient-free, population-based methods for planning, learned policy networks, or a combination of policy networks and planning. ・Hybrid approaches that combine Model Predictive Control (MPC) with a learned model and a policy prior to leverage the advantages of both paradigms have shown pr
cs.LG updates on arXiv.org

Dual-space posterior sampling for Bayesian inference in constrained inverse problems

・arXiv:2603.00393v2 Announce Type: replace-cross Abstract: Inverse problems constrained by partial differential equations are often ill-conditioned due to noisy, incomplete data or inherent non-uniqueness. ・A prominent example is full waveform inversion (FWI), which estimates Earth's subsurface properties by fitting seismic measurements subject to the wave equation, where ill-conditioning stems from noisy, band-limited
cs.LG updates on arXiv.org

Dynamic Graph Prompting via Topology-Routed Mixed-Curvature Experts

・arXiv:2608.06031v1 Announce Type: new Abstract: Dynamic graph prompting freezes a pre-trained temporal backbone and adapts it to label-scarce downstream tasks using lightweight prompts. ・However, existing methods operate within a single, fixed embedding space. ・In this work, we reveal that temporal shifts in local clustering and degree heterogeneity actively reorganize the edge curvature spectrum---indicating that the
cs.LG updates on arXiv.org

Dynamic Object Masks as Goal Representations for Visual Goal-Conditioned Reinforcement Learning

・arXiv:2510.06277v2 Announce Type: replace-cross Abstract: Goal-conditioned reinforcement learning (GCRL) offers a unified way to pursue diverse tasks, yet most existing methods rely on state- or position-based goal representations that are unavailable in real-world robotics. ・Robots operating in warehouses, agriculture, or laboratory environments rarely have access to privileged goal states, object positions, or futur
cs.LG updates on arXiv.org

Dynamics of Learning under User Choice: Overspecialization and Peer-Model Probing

・arXiv:2602.23565v3 Announce Type: replace Abstract: In many economically relevant contexts where machine learning is deployed, multiple platforms obtain data from the same pool of users, each of whom selects the platform that best serves them. ・Prior work in this setting focuses exclusively on the "local" losses of learners on the distribution of data that they observe. ・We find that there exist instances where learner
Hugging Face Papers

DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation

DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation
cs.LG updates on arXiv.org

EcoAgent-Bench: Evaluating Economic Decision-Making in Budget-Constrained LLM Agents

・arXiv:2608.05519v1 Announce Type: cross Abstract: Agent benchmarks usually measure task completion and treat resource use as an auxiliary statistic. ・In deployment, however, the choice among a local lookup, broad search, composite research tool, stronger model, or human escalation is part of the task itself. ・We introduce EcoAgent-Bench, in which every task specifies priced actions and an explicit budget.
cs.LG updates on arXiv.org

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding

・arXiv:2608.05303v1 Announce Type: cross Abstract: On-device deployment of Large Language Models (LLMs) has become essential for personalized edge applications. ・A primary bottleneck is external memory access (EMA) in feed-forward network (FFN) layers. ・Speculative decoding and mixture-of-experts (MoE) are promising solutions.
cs.LG updates on arXiv.org

Effective pruning of task-trained recurrent neural networks using noisy fluctuations and connection rescaling

・arXiv:2608.05464v1 Announce Type: cross Abstract: The pruning of network connections is key to brain function but, despite its importance, there exist few biologically-plausible pruning rules with demonstrated good performance. ・In this work we evaluate noise-prune, a recently introduced unsupervised local pruning rule for recurrent networks that uses noisy fluctuations to determine the importance of connections.
Hugging Face Papers

EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal

EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal
cs.LG updates on arXiv.org

Engram-E2VID: Reference-Based Event-to-Video Reconstruction via Generative Activation of Appearance Engrams

・arXiv:2608.05728v1 Announce Type: cross Abstract: Reference-based event-to-video reconstruction aims to recover target RGB frames from a reference frame and the event stream captured over the reference-to-target interval. ・Although events provide fine-grained temporal cues, they encode sparse and asynchronous log-intensity changes rather than absolute appearance, making faithful reconstruction intrinsically challengin
cs.LG updates on arXiv.org

Enhancing Anomaly Resilience in Research Networks: A Large-Scale Forecasting Benchmark for Dynamic Security Baselining

・arXiv:2608.05605v1 Announce Type: cross Abstract: Research and Education Networks (RENs) serve as critical infrastructure for scientific discovery, yet they face a unique security paradox: their normal traffic patterns which are characterized by massive, bursty "elephant flows" are statistically indistinguishable from volumetric attacks such as DDoS to conventional monitoring systems. ・This similarity leads to high fa
Hugging Face Papers

EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning

EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning
cs.LG updates on arXiv.org

EqDeepRx: Learning a Scalable and Interference Mitigating MIMO Receiver

・arXiv:2602.11834v2 Announce Type: replace-cross Abstract: While machine learning (ML)-based receiver algorithms have received a great deal of attention in the recent literature, they often suffer from poor scaling with increasing spatial multiplexing order and lack of explainability and generalization. ・This paper presents EqDeepRx, a practical deep-learning-aided multiple-input multiple-output (MIMO) receiver, which
cs.LG updates on arXiv.org

Equation-Free Period-Aware Forecast-Error Contraction for Estimating Negative Largest Lyapunov Exponents from Short Trajectory Ensembles

・arXiv:2608.05522v1 Announce Type: cross Abstract: Estimating positive largest Lyapunov exponents from data is comparatively natural because neighboring trajectories separate, whereas stable dynamics require resolving contraction before measurement noise or finite precision erases the signal. ・We introduce a period-aware forecast-error contraction procedure for estimating a dominant negative Lyapunov exponent from ense
cs.LG updates on arXiv.org

Equipment-centric workpiece localization in near real-time using deep learning-based vision and event-driven finite state machines

・arXiv:2608.05744v1 Announce Type: new Abstract: Continuous workpiece localization is essential for traceability and process coordination in hot forging, but direct tracking is unreliable because of extreme temperatures, surface degradation, and irregular routing. ・This study presents an equipment-centric framework that infers workpiece locations from handling equipment observed by multiple static 2D cameras.
cs.LG updates on arXiv.org

Evaluating Machine Learning Models for Post-Wildfire Debris-Flow Prediction

・arXiv:2608.05265v1 Announce Type: new Abstract: Prediction of post-wildfire debris flows is critical for mitigating hazards to communities, infrastructure, and resources during intense rainfall in recently burned areas. ・However, identifying reliable machine learning models is complicated by overlapping debris-flow and non-debris-flow events in feature space, the need for model interpretability, and limited training d
cs.LG updates on arXiv.org

Evidential Rule Learning for Interpretable Classification with Abstention

・arXiv:2608.05859v1 Announce Type: new Abstract: Interpretable classification often requires more than accurate predictions for real-life deployment: models should be transparent about the evidence behind their decisions and abstain when they cannot decide reliably. ・We introduce Fast Evidential Rule Learning (FERL), a method that learns interpretable, accurate fuzzy rule models whose outputs are evidential.
cs.LG updates on arXiv.org

EvoHarness-RL: Learning Self-Evolving Runtime Harness for Long-Horizon LLM Agents

・arXiv:2608.05446v1 Announce Type: new Abstract: Long-horizon LLM agents increasingly rely on external execution support to maintain state, track progress, invoke tools, verify outcomes, and reuse experience across interactions. ・However, effective harness use raises two coupled challenges: state formation from noisy interaction traces and runtime control over external-state access. ・Existing agents usually handle both
cs.LG updates on arXiv.org

Failing Gracefully: Mitigating Impact of Inevitable Robot Failures

・arXiv:2608.05313v1 Announce Type: cross Abstract: Service robots operate in household environments shared with humans, pets, and everyday objects, where they are highly susceptible to failures such as software crashes, hardware degradation, or unpredictable interactions. ・While roboticists strive to minimize failures, some remain inevitable, making it critical to mitigate their potential consequences for safe and reli
cs.LG updates on arXiv.org

Fast Rates for Inverse Reinforcement Learning

・arXiv:2605.14599v2 Announce Type: replace Abstract: We establish novel structural and statistical results for entropy-regularized min-max inverse reinforcement learning (Min-Max-IRL) in finite-horizon MDPs with Borel state and action spaces. ・We show that maximum likelihood estimation (MLE) and Min-Max-IRL are equivalent at the population level, and at the empirical level under deterministic dynamics. ・For linear rewar
cs.LG updates on arXiv.org

FI-TW: An Open Train-Weather Dataset for Railway Delay Analysis in Finland

・arXiv:2601.16592v2 Announce Type: replace Abstract: Train delays result from complex interactions between operational, technical, and environmental factors. ・While weather impacts railway reliability, particularly in Nordic regions, existing datasets rarely integrate meteorological information with operational train data. ・This study presents the first publicly available dataset combining Finnish railway operations wit
stat.ML updates on arXiv.org

FlowAdam: Implicit Regularization via Geometry-Aware Soft Momentum Injection

・arXiv:2604.06652v1 Announce Type: cross Abstract: Adaptive moment methods such as Adam use a diagonal, coordinate-wise preconditioner based on exponential moving averages of squared gradients. ・This diagonal scaling is coordinate-system dependent and can struggle with dense or rotated parameter couplings, including those in matrix factorization, tensor decomposition, and graph neural networks, because it treats each p
cs.LG updates on arXiv.org

FOCUS: Decoupling Expert Personas in LLMs to Enhance Domain Expert Capabilities

・arXiv:2608.05611v1 Announce Type: cross Abstract: Large Language Models (LLMs) can exhibit diverse personas, and activating expert personas has been shown to improve domain expertise and task accuracy. ・However, existing persona control methods often suffer from cross-domain coupling, which may lead to overly aggressive behavior in high-caution domains such as healthcare, or excessive conservatism in risk-sensitive do
cs.LG updates on arXiv.org

Fractal KV-Cache Archives: Lossless Symbolic Storage with In-Place Retrieval for Long-Context LLM Inference

・arXiv:2607.07144v2 Announce Type: replace Abstract: The key-value (KV) cache dominates the memory cost of long-context autoregressive inference, and a growing body of work compresses it through quantization, eviction, or offloading. ・We study a complementary question: once a position's KV state has been quantized to codebook indices, how should the resulting symbol stream be stored, and can the storage layer do more t
cs.LG updates on arXiv.org

From Continuous Predictors to Clinical Thresholds: Early Evidence on Performance Trade-offs of Guideline-Based Categorisation for Ischaemic Stroke Outcome Prediction

・arXiv:2608.05203v1 Announce Type: cross Abstract: Machine learning models achieve strong predictive accuracy for 90-day outcome prediction in acute ischaemic stroke, yet clinical adoption is limited by the misalignment of model explanations with clinicians' reasoning. ・Motivated by a clinician user study calling for clinical guideline-aligned cut-offs, we ask whether continuous predictors can be replaced by clinically
Hugging Face Papers

From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models

From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models
cs.LG updates on arXiv.org

From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models

・arXiv:2608.06020v1 Announce Type: cross Abstract: Economic World Models (EWMs) are generative economic models that simulate how economies evolve from within by modeling heterogeneous agents, their beliefs and actions, and the market and institutional mechanisms through which their interactions produce aggregate outcomes. ・This paper develops an implementation roadmap for building economic world models as generative en
cs.LG updates on arXiv.org

From Siloed Algorithms to Compliance-First Agentic Platforms: A Multi-Layered Architecture for Hospital AI Systems

・arXiv:2608.06112v1 Announce Type: cross Abstract: Hospitals are rapidly adopting artificial intelligence for triage, imaging, scheduling etc., yet most deployments remain isolated point solutions locked inside departmental silos, resulting in duplicated effort, hidden risks, and unrealized enterprise value. ・Despite explosive growth of AI in healthcare market and accelerating investment, an estimated 70-80% of healthc
stat.ML updates on arXiv.org

Fuzzy network jump models for soft dynamic clustering of graph-structured data

・arXiv:2608.05786v1 Announce Type: cross Abstract: We introduce a fuzzy network jump model for clustering time-varying observations indexed by the nodes of a weighted graph. ・The framework allows flexible graph representations with spatial and temporal regularization promoting smooth soft cluster assignments across connected nodes and consecutive time points. ・Estimation is performed through an efficient alternating opt
cs.LG updates on arXiv.org

GAUGE: Granularity-Adaptive Counterfactual Gating of Evidence for Incomplete Multimodal Classification

・arXiv:2608.05608v1 Announce Type: new Abstract: Multimodal classification typically assumes all modalities are available, yet real-world inputs are often incomplete. ・Imputation and dynamic fusion can mitigate such incompleteness, but existing methods operate at a coarse modality level and thus cannot retain reliable components while suppressing misleading ones within the same recovered modality, compromising predicti
cs.LG updates on arXiv.org

GenGA: Editable and Data-Grounded Graphical Abstract Generation for Academic Papers

・arXiv:2608.05478v1 Announce Type: cross Abstract: Graphical Abstracts (GAs) visually summarize the key findings of academic papers, playing a crucial role in facilitating the understanding of research content. ・Recently, advancements in vision-language models and image generation models have enabled the automatic generation of scientific figures based on paper content. ・However, most conventional methods output the gen
WIRED

Google Workspace Promo Codes: 14% Off for August 2026

・Boost your productivity and save with exclusive Google Workspace coupons from WIRED. ・Get up to 14% off plans for three months, including Starter, Standard, and Plus tiers.
#AIタグ

GPT-5.6 Solの推論調整で開発はどう変わるか。Claude Code実践者が解説するAIエージェントの未来

・AIモデルの勝負どころが「賢さ」から「タスクへの適応力」に変わった。 ・GPT-5.6 Solで導入された推論深度の調整機能は、AIが単なるチャットボットから業務を直接完結させるエージェントへ移行した証拠だ。
The Verge

Grab the entire Lord of the Rings trilogy on 4K Blu-ray for $50

・“One ring to rule them all” | Image: Warner Bros. ・Looking to spend a weekend inside with a good binge? ・Gruv has The Lord of the Rings trilogy on 4K Blu-ray marked down to $49.99, just below Amazon’s price and only a little bit higher than the all-time low we usually see around Black Friday.
cs.LG updates on arXiv.org

GROM: Gradient-Free Rapid One-Shot Machine Unlearning

・arXiv:2608.05783v1 Announce Type: new Abstract: Machine unlearning has become a critical capability for safely removing specific, sensitive knowledge from large language models (LLMs). ・Current state-of-the-art approaches primarily rely on iterative, training-time unlearning via fine-tuning. ・However, even when utilizing parameter-efficient dimensionality reduction techniques like LoRA, gradient-based optimization rema
Hugging Face Papers

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?
cs.LG updates on arXiv.org

Handling Missing Data in Probabilistic Regression Trees

・arXiv:2608.06195v1 Announce Type: cross Abstract: Probabilistic Regression Trees (PRTrees) are a smooth and consistent alternative to classical regression trees, producing continuous predictions through probabilistic split assignments. ・This paper extends the PRTree framework to accommodate missing predictor values directly during tree construction, eliminating the need for prior imputation. ・Three strategies are propo
cs.LG updates on arXiv.org

Hardware Keystores for AI Agent Signing Workflows: A Zero-Trust MCP Enforcement Architecture

・arXiv:2608.06130v1 Announce Type: cross Abstract: AI agents performing cryptographic operations (signing Git commits, authenticating API calls, issuing certificates) currently store private keys in software-accessible locations: plaintext files, environment variables, or container memory. ・Any process with sufficient read privileges can extract the raw key material. ・A recent production incident demonstrated the practi
Hugging Face Papers

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization
cs.LG updates on arXiv.org

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization

・arXiv:2608.06301v1 Announce Type: cross Abstract: As LLMs are increasingly deployed within agentic systems, their capabilities depend not only on the model weights but also on the harness: the prompts, tools, control flow, memory, and orchestration code surrounding them. ・This makes automated harness optimization -- the iterative and evaluation-guided improvement of a harness by an AI system -- both an important route
cs.LG updates on arXiv.org

How Far Do Simple Transformations Translate Across Text Embedding Models?

・arXiv:2608.05980v1 Announce Type: new Abstract: We investigate whether simple transformations can translate representations across heterogeneous text embedding models. ・Understanding how independently trained models organize semantic information is an enabler for AI-to-AI latent communication without decoding into human-readable text. ・Focusing on lightweight translators such as linear mappings, we test the literature
OpenAI News

How HSP GRUPPE builds AI capabilities for tax advisory

・Discover how HSP GRUPPE uses ChatGPT Enterprise to boost productivity, improve work quality, and create more capacity for tax advisory and client service.
cs.LG updates on arXiv.org

How Much Reconstruction Does Quantum Machine Learning Need? Late Fusion of Independently Trained Quantum Subcircuits

・arXiv:2608.05595v1 Announce Type: cross Abstract: Circuit cutting lets a large quantum neural network (QNN) run as independent subcircuits on small devices, but rebuilding its outputs by reconstruction carries a classical sampling overhead exponential in the number of cuts - the dominant runtime cost in prior work. ・We ask whether, for machine-learning tasks, this step is necessary, and replace it with late fusion: ea
WIRED

HP Coupon Codes and Deals August 2026

・Save up to 60%, plus an extra 20% with HP promo codes for laptops, printers, PCs, and more tech.
#LLMタグ

Hugging Faceの41%が中国製オープンモデル。DeepSeek時代に、AIの価値は“モデルの外側”へ移る

Hugging Faceの41%が中国製オープンモデル。DeepSeek時代に、AIの価値は“モデルの外側”へ移る
cs.LG updates on arXiv.org

Hybrid Probabilistic Zonotopes for Identifiable and Refinable Predictive Uncertainty

・arXiv:2608.05454v1 Announce Type: new Abstract: Probabilistic prediction heads in neural networks typically output either a Gaussian mixture or a single conformal region. ・Neither separates the distinct sources of uncertainty often present in real prediction tasks: a discrete choice among modes, bounded systematic drift within the chosen mode, and irreducible stochastic noise. ・We introduce the Hybrid Probabilistic Zon
cs.LG updates on arXiv.org

Hybrid-Adaptive Thread Tuning to Mitigate Simulation Execution Bottlenecks in High-Performance Reinforcement Learning Inference

・arXiv:2608.06025v1 Announce Type: new Abstract: In simulation-in-the-loop decision-making systems, reinforcement learning (RL) inference is often constrained by simulator-side execution overhead, where workloads are highly dynamic and sensitive to runtime thread configurations. ・Existing multithreaded strategies struggle to match thread resources before or during execution, causing resource contention, scheduling over
cs.LG updates on arXiv.org

Hypothesis Testing with Conditional Queries: Learnability and the Value of Interaction

・arXiv:2608.06262v1 Announce Type: new Abstract: Model evaluations may fix all tests before observing any responses or select later tests using earlier responses. ・We study this choice in a conditional-query model on a finite outcome space $\mathcal{X}$ with $|\mathcal{X}|=N$. ・We first ask which pairs of distribution classes can be reliably distinguished.
cs.LG updates on arXiv.org

IFlowNets: Extending Generative Samplers to Learn Strategies in Incomplete Information Games

・arXiv:2608.05422v1 Announce Type: new Abstract: While many algorithms blend reinforcement learning (RL) with counterfactual regret (CFR) methods to leverage tradeoffs in computational speed and performance, there are fewer investigations into generative sampling frameworks in game theoretic applications in incomplete information games. ・We extend a generative flow network framework, Adversarial Flow Networks (AFlowNet
cs.LG updates on arXiv.org

Improving the Realism of Synthetic Clinical Benchmarks Under Utility Constraints

・arXiv:2608.06265v1 Announce Type: cross Abstract: Synthetic clinical benchmarks for enterprise AI agents can pass existing utility checks and still remain structurally unrealistic, especially in privacy-sensitive healthcare settings where operational data are hard to access. ・We study how to improve such benchmarks without breaking the downstream utility checks already used in practice. ・We formulate benchmark revision
cs.LG updates on arXiv.org

Innovation-Residual Auditing of Autonomous Analysis Agents: Localization, Detection Limits, Error Control, and Identifiability

・arXiv:2608.05490v1 Announce Type: cross Abstract: Autonomous agents now carry out entire data analyses, selecting cohorts, joining tables, and fitting models with little step-by-step supervision. ・When such an analysis turns out to be wrong, someone must determine which operation caused it. ・A recent approach does this without any labelled mistakes, learning instead from analyses known to be sound and flagging operatio
cs.LG updates on arXiv.org

Integrating Implicit and Explicit Relational Biases through Graph-Based Multiple Instance Learning: A Case Study in Skin Lesion Diagnosis

・arXiv:2608.06037v1 Announce Type: cross Abstract: Relational inductive biases are essential for capturing structural dependencies among data. ・This study investigates a dual-level relational framework for image classification, bridging the gap between implicit representation learning and explicit structural modelling. ・We begin by establishing a baseline using an EfficientNetB3 architecture.
Hugging Face Papers

Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval

Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval
cs.LG updates on arXiv.org

Invariant Representation Learning for Source-Free Time Series Forecasting with LLM-Centric Proxy Denoising

・arXiv:2510.05589v3 Announce Type: replace Abstract: Effective time series forecasting enables various real-world applications, benefiting from the proliferation of mobile devices. ・However, the volume of time series data may vary significantly across domains due to high data acquisition costs and data regulations. ・To maximally create value from sparse data, this study focuses on a new problem of source-free time serie
Hugging Face Papers

Invisible Shortcuts: Why Vision Encoders Know Your Camera

Invisible Shortcuts: Why Vision Encoders Know Your Camera
cs.LG updates on arXiv.org

Invisible Shortcuts: Why Vision Encoders Know Your Camera

・arXiv:2608.05424v1 Announce Type: cross Abstract: Deep vision models exploit shortcuts, relying on cues that correlate with supervision signals. ・Prior work has focused on visible biases, such as object-background or texture correlations. ・We identify a different source of shortcut learning: invisible metadata traces embedded at the pixel level, for metadata such as image processing and photo acquisition.
cs.LG updates on arXiv.org

Is Self-Pretraining really useful to improve diagnosis in medical Time Series?

・arXiv:2608.06122v1 Announce Type: new Abstract: Inspired by recent evidence that transformer architectures benefit from Self-PreTraining (SPT) on long-context benchmarks, we investigate whether similar gains extend to multimodal, multivariate, and even simple univariate medical time series. ・Our objective is to assess the impact of SPT on the performance and scalability of transformer-based models across diverse medic
AI News & Artificial Intelligence | TechCrunch

Jill Lepore on the ‘Artificial State’ and why Silicon Valley’s leaders are bad sci-fi readers

・Historian Jill Lepore has a theory about why tech companies often use soaring language to describe their products — almost as if they’re forming a new government. ・And whether you’re thinking of Twitter’s old “town hall in your pocket” or Anthropic’s Claude constitution, it’s a theory that doesn’t paint Silicon Valley in a very flattering light. ・In Lepore’s upcoming book, The Rise and Fall of the Artificial State, the
cs.LG updates on arXiv.org

Kastor: An efficient fine-tuning strategy for generative emulation of PDE simulations

・arXiv:2608.06107v1 Announce Type: new Abstract: Machine learning offers a promising avenue to accelerate physical simulations by replacing computationally expensive traditional Partial Differential Equation (PDE) solvers with fast, differentiable surrogate models. ・However, standard auto-regressive ML emulators often suffer from error accumulation over long horizons and struggle to capture the stochasticity of complex
cs.LG updates on arXiv.org

KV-Skill: Forging Expertise in the Model's Native Language

・arXiv:2608.05475v1 Announce Type: new Abstract: Task knowledge is commonly stored either as text in the prompt or as an update to model weights. ・Text is modular but must be interpreted on every use, while weight adaptation makes the resulting capability difficult to load, remove, or share independently. ・We introduce KV-Skill, a design space of external factorized operators that a frozen language model reads through a
cs.LG updates on arXiv.org

KVAE: Family of Tokenizers for Multimodal Generative Models

・arXiv:2608.05798v1 Announce Type: cross Abstract: Latent diffusion modeling (LDM), a prominent paradigm, utilizes tokenizers to map input signal to compressed representation. ・This dependency positions tokenizer as an integral part of generation process itself, since it affects learning speed, quality of synthesized samples and lay foundation for later applications. ・This report presents series of KVAE tokenizers for a
#LLMタグ

LangChainの使い方をゼロから解説|モデル呼び出し・ツール・記憶・RAGまで

・LangChainという名前は知っているものの、何ができるのか、どこから触ればよいのか分からない方も多いのではないでしょうか 検索すると `LLMChain` や `AgentExecutor` を使った古いコードが大量に見つかり、公式ドキュメントと書き方が違って戸惑うこともあります 続きをみる
cs.LG updates on arXiv.org

Latent Utility Q-Learning for Preference-Adaptive Dynamic Treatment Regimes

・arXiv:2307.12022v3 Announce Type: replace-cross Abstract: Optimizing individualized treatment sequences for patients who weigh multiple, competing outcomes differently poses a challenge for dynamic treatment regime (DTR) methods, which typically assume a single univariate outcome. ・We propose Latent Utility Q-Learning (LUQ-Learning), which estimates DTRs optimizing patient-specific preference-weighted combinations of
cs.LG updates on arXiv.org

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction

・arXiv:2608.05600v1 Announce Type: new Abstract: Flow-based generative models are typically sampled by solving a deterministic ordinary differential equation (ODE), whereas online reinforcement learning requires stochastic rollouts for policy exploration and optimization. ・Existing GRPO methods for flow models therefore replace the inference-time ODE with a stochastic differential equation (SDE) during training.
cs.LG updates on arXiv.org

LC-Implicit-QAOA: Active-Workspace-Capped Exact Objective-and-Gradient Evaluation for Training over Bounded QUBO Light Cones

・arXiv:2608.05610v1 Announce Type: cross Abstract: QAOA training repeatedly queries an objective and all shared gradients, making exact evaluation a feasibility bottleneck even when QUBO terms have bounded causal cones. ・Building on established causal-cone restriction and adjoint differentiation, LC-Implicit-QAOA profiles cone structure and induced-edge counts before local-amplitude and named-workspace allocation, then
Hugging Face Papers

Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval

Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval
stat.ML updates on arXiv.org

Learning Latent Memory States from Longitudinal Athlete Monitoring Data

・arXiv:2608.06290v1 Announce Type: cross Abstract: We propose a new unit of analysis for longitudinal data: the Latent Memory Table. ・The scientific contribution is not the encoder. ・It is that table, treated as a reusable statistical object on the same footing as a matrix of principal-component scores, a table of estimated random effects, or a table of predicted probabilities.
cs.LG updates on arXiv.org

Learning to Rank Tensor Network Contraction Plans for GPU-Accelerated Quantum Circuit Simulation

・arXiv:2608.05819v1 Announce Type: new Abstract: Classical simulation remains essential for developing and validating quantum algorithms, but its cost grows rapidly with circuit size. ・Tensor-network contraction can reduce this cost by exploiting circuit structure, although its efficiency depends strongly on the chosen contraction plan. ・On GPUs, plans with similar theoretical complexity may perform very differently bec
cs.LG updates on arXiv.org

Learning When to Trust via Selective Context Preference Optimization

・arXiv:2608.06377v1 Announce Type: cross Abstract: Language models increasingly condition their answers on external signals, and a single misleading one can turn a correct answer wrong. ・The obvious remedy, training models to resist such signals, hides a failure mode: a model that ignores all context looks robust yet is useless when the context is worth trusting. ・We recast the problem as selective trust and introduce M
WIRED

LG Promo Codes and Coupons for August 2026

・Save 20% with an LG promo code today, plus up to $1,000 off appliances, 40% off bestselling TVs and monitors.
cs.LG updates on arXiv.org

LILAC: An Idempotent Neural Speech Codec

・arXiv:2608.05727v1 Announce Type: cross Abstract: Neural Audio Codecs are widely adopted in speech generation and editing. ・However, existing neural audio codecs are not idempotent: across the paper's twelve baseline systems, every configuration tested rewrites, on average, at least 15% of its tokens in a single decode-re-encode pass. ・This poses a problem for utilizing Neural Audio Codecs as token interfaces in pipeli
cs.LG updates on arXiv.org

LLM Inference Under Bursty Workload Distribution: Modifying the WAIT Algorithm

・arXiv:2608.06135v1 Announce Type: new Abstract: Large Language Models (LLMs) such as ChatGPT and Claude are widely used for information retrieval and problem-solving. ・Recent work has focused on improving scheduling algorithms to boost throughput while maintaining low latency. ・However, these approaches often assume Poisson request arrivals with constant rates - an assumption that fails to reflect the inherently bursty
#LLMタグ

LLMに「逆再生」を教えると何が起きるか――PTPの発想をやさしく読む

・LLMに文章を書いてもらうと、言葉が左から右へ流れていきます。少し乱暴にまとめれば、LLMが繰り返しているのは「ここまでの文に続きそうなものは何か」という予測です。 ・では、文章の並びを反転してから、いつもと同じように続きを予測させたらどうなるでしょうか。
Zennの「大規模言語モデル」のフィード

LLMはマルチターン会話で迷子になる:論文を短く読む

・「単発の質問には強い。でも、会話が続くと迷子になる。」 今回紹介する LLMs Get Lost In Multi-Turn Conversation は、そんなLLMの意外な弱点を扱った論文です。 ・マルチターン会話で、LLMの性能は平均39%低くなる。 ・同じ情報量でも、渡し方が変わるだけで結果が大きく変わる。
Hugging Face Papers

MameLoshnLM: Yiddish Language Model and Evaluation Benchmark

MameLoshnLM: Yiddish Language Model and Evaluation Benchmark
Zennの「大規模言語モデル」のフィード

MaR: メタ認知を報酬に変えるRLで推論軌跡を直接最適化

・MaR: メタ認知を報酬に変えるRLで推論軌跡を直接最適化 TL;DR **Metacognition-as-Reward (MaR)**は、人間の「メタ認知」の2次元(知識+制御)をLLMのRL報酬に落とし込んだフレームワーク 推論軌跡を「知識抽出」「プロセス制御」「最終回答」に明示的に分割し、軌跡レベルの報酬で最適化 RLVR(結果だけ)でもRaR(ルーブリック毎)でもなく、汎用的なメタ認知次元でプロセス全体を評価 22ベンチマークで一貫して改善、ベースモデル比最大+7.7%、vanilla DAPO比最大+11.0% Qwen3.5-9B + MaRがGPT-OSS-1...
cs.LG updates on arXiv.org

Marginal Matching Does Not License Factorized Sampling: Auditing Conditional Style Leakage in Factorized Generative Models

・arXiv:2608.05243v1 Announce Type: new Abstract: Factorized generative models commonly regularize a latent style variable z_s by matching its marginal distribution to a fixed Gaussian prior and interpret this as evidence that the style representation is independent of class information. ・We show that this interpretation is incorrect. ・Matching only the marginal distribution places no constraint on the class-conditional
cs.LG updates on arXiv.org

MARS: Margin and Semantic-Aware Data Augmentation for Reward Modeling

・arXiv:2602.17658v4 Announce Type: replace Abstract: Reward modeling is central to RLHF, RLAIF, and PPO-based alignment, but its reliability is often limited by scarce and heterogeneous human preference data. ・In this paper, we introduce MARS (Margin and Semantic-Aware Data Augmentation for Reward Modeling), an adaptive augmentation framework for controlled low-resource reward modeling. ・MARS allocates more augmentation
Hugging Face Papers

MASS: Multiplayer World Models with Authoritative Shared State

MASS: Multiplayer World Models with Authoritative Shared State
cs.LG updates on arXiv.org

Matrix Zonotopic Attention: A Context-Adaptive Value Projection for Set Transformers

・arXiv:2608.05472v1 Announce Type: new Abstract: Multi-head attention combines an input-dependent softmax routing with an input-independent linear value projection, so the per-sample operator mapping aggregated values to outputs is the same for every input set. ・We study the consequences of this asymmetry for permutation-invariant set targets. ・We introduce the Transformation Degrees of Freedom (TDOF) of a target operat
Zennの「大規模言語モデル」のフィード

MCP連携で壊れる前に確認すべき10の運用チェックリスト

・2026-07-31時点の主要トレンドは、LLM単体競争からエージェント実装・MCP接続・運用安全性へと関心が移っていることです。OpenAIは「Building abundant intelligence」を掲げる一方、AI Businessでは価格引き下げが報じられ、企業導入でコスト最適化が重要になっています。AWSのMCP対応、CrowdStrikeのエージェント保護、FujitsuやMicrosoft関連のエージェント技術報道からは、実運用フェーズの到来が鮮明です。Web開発は直接的なNext.js/Reactニュースは乏しいものの、SaaStr AI 2026でVercel...
cs.LG updates on arXiv.org

MermaidSeqBench: An Evaluation Benchmark for NL-to-Mermaid Sequence Diagram Generation

・arXiv:2511.14967v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown great promise in generating structured diagrams from natural language descriptions, particularly Mermaid sequence diagrams for software engineering. ・However, the lack of existing benchmarks to assess the LLM's correctness on this task hinders rigorous, systematic evaluation and principled comparison of model capabilities
cs.LG updates on arXiv.org

MetaboLLM: a metabolomics-specialized large language model for biochemical knowledge integration and predictive metabolite graph construction

・arXiv:2608.06253v1 Announce Type: new Abstract: Metabolomics knowledge is distributed across heterogeneous resources and remains difficult to translate into predictive representations. ・We developed MetaboLLM, a metabolomics-specialized large language model adapted through continual pretraining, supervised fine-tuning, and structured retrieval, together with MetaboLLM-GIN, which converts generated biochemical descript
The Verge

Microsoft Edge is about to lock out older ad blockers, just like Chrome did

・Microsoft Edge is ending support for the Manifest V2 extensions platform, which will cut off the uBlock Origin adblocker and others like it, just like Google Chrome did earlier this year. ・According to Microsoft, there are only 58 extensions on the Edge Add-On Store "with any meaningful usage" that still use MV2, and only three of those aren't available on the newer MV3 platform. ・Anyone still using these extensions ca
MarkTechPost

Microsoft Open Sources code-testing-generator: a Polyglot Unit-Test Agent That Hits 92.1% Task Completion Versus 78.9% for Stock Copilot

・Microsoft has open sourced code-testing-generator, a polyglot unit-test agent shipping in the MIT-licensed dotnet/skills repository. ・It reads a repository before writing anything — detecting the language, test framework, existing conventions, and the real build and test commands — then plans, writes, runs and validates the tests it produces. ・On Microsoft's internal 152-task benchmark it completed 140 tasks against 12
cs.LG updates on arXiv.org

Minimax Optimal Early-Stopped Gradient Descent for Gaussian Mixture Classification

・arXiv:2608.06250v1 Announce Type: cross Abstract: In overparameterised classification, training data can be linearly separable even when the underlying distribution is not. ・In this setting, gradient descent (GD) on the logistic loss diverges in norm while converging in direction to a max-margin interpolating classifier, whose implicit bias can be statistically suboptimal. ・In this work, we show that early stopping can
#LLMタグ

MIPLCのLL.M.プログラムの概要と出願について(ドイツ/ミュンヘン留学)

・私が現在参加している、ミュンヘン知的財産法センター(Munich Intellectual Property Law Center。以下「MIPLC」と略称で呼びます。)のLL.M.プログラムの概要と出願について、簡単にご紹介したいと思います。 ・1 MIPLC・LL.M.プログラムの概要 続きをみる
cs.LG updates on arXiv.org

MirrorNet: Can Medical Image Anonymization Really Protect Patient Identity?

・arXiv:2608.05938v1 Announce Type: cross Abstract: Medical images are routinely de-identified---names, dates, and other metadata removed---and then shared for research, teaching, and public benchmarks under the assumption that this renders them anonymous. ・Such de-identification protects the metadata but not the pixels, and---apart from scans that directly contain facial structures---whether the image content itself id
cs.LG updates on arXiv.org

ML-for-ML

・arXiv:2608.06046v1 Announce Type: cross Abstract: AI training workloads are growing rapidly, making their time, energy, and infrastructure costs increasingly important. ・In shared cloud clusters, training and fine-tuning jobs compete with co-running workloads for network resources, while network mechanisms and ML training choices are typically optimized separately: networking controls how bytes move, whereas ML system
cs.LG updates on arXiv.org

MoDAl: Self-Supervised Neural Modality Discovery via Decorrelation for Speech Neuroprosthesis

・arXiv:2605.00025v3 Announce Type: replace-cross Abstract: Speech neuroprosthesis systems decode intended speech from neural activity in the absence of audible output, offering a path to restoring communication for individuals with speech-impairing conditions. ・Current approaches decode predominantly from motor cortical areas, discarding others -- such as area 44, part of Broca's area -- that may encode complementary l
cs.LG updates on arXiv.org

MS-MLB: An Open Machine Learning Benchmark for Blood-Based MS Classification

・arXiv:2608.05196v1 Announce Type: new Abstract: Multiple sclerosis (MS) is diagnosed through clinical assessment, magnetic resonance imaging, laboratory evidence when appropriate, and exclusion of better explanations. ・Blood RNA expression data may contain disease associated immune signal, but a blood RNA classifier cannot be treated as a replacement for clinical diagnosis. ・This paper presents MS-MLB (Multiple Scleros
cs.LG updates on arXiv.org

Multivariate Time Series Forecasting needs Cross Variable Loss

・arXiv:2608.05742v1 Announce Type: new Abstract: Multivariate time series forecasting presents unique challenges because future variables often co-evolve under shared system dynamics. ・While existing studies mainly focus on cross-variable dependencies in historical observations, dependencies among future values are much less explored. ・Specifically, modern forecasting models largely follow the Direct Forecasting (DF) pa
cs.LG updates on arXiv.org

Muon on the Stiefel Manifold Admits an Exact Closed-Form Update

・arXiv:2608.06218v1 Announce Type: cross Abstract: We study Muon, a recently proposed matrix-aware optimization method, in the context of the Stiefel manifold. ・This manifold consists of matrices with orthonormal columns and is ubiquitous in machine learning and scientific computing. ・Existing extensions of Muon to this manifold rely on heuristic, approximate, or iterative updates with varying computational efficiency.
cs.LG updates on arXiv.org

NavTrust: Benchmarking Trustworthiness for Embodied Navigation

・arXiv:2603.19229v2 Announce Type: replace-cross Abstract: There are two major categories of embodied navigation: Vision-Language Navigation (VLN), where agents navigate by following natural language instructions; and Object-Goal Navigation (OGN), where agents navigate to a specified target object. ・However, existing work primarily evaluates model performance under nominal conditions, overlooking the potential corrupti
cs.LG updates on arXiv.org

Neuro-Symbolic Closed-Loop Control of Laser Powder Bed Fusion with an In-Loop Ontology

・arXiv:2608.05773v1 Announce Type: new Abstract: A geometry-conditioned, neuro-symbolic closed-loop architecture is proposed for laser powder bed fusion, in which a standards-aligned ontology operates inside the control loop and couples symbolic reasoning with statistical learning to set the targets of a constraint-aware predictive controller. ・The ontology links the process objectives and constraints to the signals a
AI News & Artificial Intelligence | TechCrunch

New Mexico court orders Meta to pay additional $567M in child safety case

・Meta's total fine has raked up to $942 million in this case.
WIRED

Newegg Promo Codes and Coupons for August 2026

・Enjoy up to 10% off your entire order with today’s Newegg promo code and discount codes. ・Save with the latest deals for gaming PCs, laptops, and computer parts.
cs.LG updates on arXiv.org

Nonvisual Classification of Ground-Condition by Artificial Proprioception in an Amoeba-Inspired Autonomous Walking Robot

・arXiv:2608.05684v1 Announce Type: cross Abstract: Nonvisual classification of ground condition based on a multimodal sensing approach was investigated for an amoeba-inspired autonomous walking robot. ・To classify ground condition without image sensing and processing, we implemented artificial proprioception by integrating a three-axis accelerometer, eight foot pressure sensors, and reservoir computing (RC).
cs.LG updates on arXiv.org

Observation-Grounded Self-Predictive Reinforcement Learning for Visual Continuous Control

・arXiv:2608.05989v1 Announce Type: new Abstract: Sample-efficient policy learning from pixels is a long-standing challenge in reinforcement learning (RL). ・Recent dynamics-based representation learning methods have significantly improved the sample efficiency of model-free visual RL by learning dynamics-aware representations through auxiliary prediction performed either in latent space (self-prediction) or observation
cs.LG updates on arXiv.org

On Same-Sample and Independent-Sample Stochastic Extragradient for Monotone Variational Inequalities

・arXiv:2608.06182v1 Announce Type: cross Abstract: We study stochastic extragradient (SEG) methods for solving monotone variational inequality problems (VIPs) over a feasible set. ・Although extragradient is a foundational algorithm for VIPs and its deterministic convergence theory is well developed, its stochastic counterpart remains less understood. ・Most existing analyses focus on independent-sample SEG (I-SEG) and as
cs.LG updates on arXiv.org

On the Anisotropy of Score-Based Generative Models

・arXiv:2510.22899v2 Announce Type: replace Abstract: We investigate the role of network architecture in shaping the inductive biases of modern score-based generative models. ・To this end, we introduce the Score Anisotropy Directions (SADs), architecture-dependent directions that reveal how different networks preferentially capture data structure. ・Our analysis suggests that SADs form adaptive bases aligned with the arch
Hugging Face Papers

On-Policy Delta Distillation for Multilingual Math Reasoning

On-Policy Delta Distillation for Multilingual Math Reasoning
cs.LG updates on arXiv.org

On-Policy Delta Distillation for Multilingual Math Reasoning

・arXiv:2608.05802v1 Announce Type: cross Abstract: On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexplored. ・We study OPD and its advanced variant, On-Policy Delta Distillation (OPD$^2$), for mathematical reasoning in English, Korean, and Japanese. ・OPD$^2$ improves OPD by using the probabili
cs.LG updates on arXiv.org

On-Policy Self-Distillation without Any Supervision

・arXiv:2608.06296v1 Announce Type: new Abstract: On-policy (Self-)Distillation (OPD / OPSD) has shown strong potential for post-training large language models (LLMs). ・However, existing methods still rely heavily on external supervision, including ground-truth signals, environmental feedback, or guidance from larger models, and therefore fall short of genuine "self"-distillation. ・In this study, we show that on-policy s
cs.LG updates on arXiv.org

One Qubit Can Beat One Bit: Quantum Advantage for Post-Training Quantization

・arXiv:2608.05240v1 Announce Type: cross Abstract: One-bit post-training quantization represents each weight using only its sign, requiring all deployment contexts to share the same binary weight matrix even when their activation statistics favor different sign patterns. ・We study this shared-sign constraint and introduce Quantum Random Access Quantization (QRAQ). ・This framework encodes context-dependent signs in a qua
cs.LG updates on arXiv.org

Online Reasoning Calibration: Test-Time Training Enables Generalizable Conformal LLM Reasoning

・arXiv:2604.01170v2 Announce Type: replace Abstract: While test-time scaling has enabled large language models to solve highly difficult tasks, state-of-the-art results come at exorbitant compute costs. ・These inefficiencies can be attributed to the miscalibration of post-trained language models, and the lack of calibration in popular sampling techniques. ・Here, we present Online Reasoning Calibration (ORCA), a framewor
Zennの「大規模言語モデル」のフィード

OpenAIのWeb検索APIが画像も検索できるようになった

・ソフトバンクのCHENです。普段は、データ構造化ツール「TASUKI Annotation」の技術研究&全社RAG基盤に携わっています。あわせて生成AI・エージェント技術の探索にも取り組み、実務投入を見据えた技術検証を行っています。本記事では、OpenAI の Web検索ツール(web_search)が新たに画像も返せるようになった点を、実際にAPIを叩いて検証しました。 ・TL;DR 使い方:web_search に search_content_types と image_settings を追加するだけで、画像URL・出典ページURL・サムネイルがまとめて返ってきます。
Zennの「大規模言語モデル」のフィード

OpenTelemetry Collectorを活用したLLMトレースのパイプライン制御:ガバナンスと集計のための正規化

・はじめに こんにちは、チームラボでエンジニアをしている鬼頭です。 ・OpenTelemetry Collector(以下 Collector)は、アプリのTraceをJaegerやDatadogなどに送る前に加工できる中継点です。 ・最近はサービスにLLM機能を組み込むケースが増えましたが、それに伴ってTraceに乗ってくる情報も大きく変わりました。プロンプト本文やトークン数・モデル名といった課金に直結する属性が大量に付与されるようになります。その結果、「ログへの個人情報(PII)の混入」や「モデル名の表記揺れによるコスト集計のズレ」といった新たな課題が目立つようになってきました。
Zennの「大規模言語モデル」のフィード

OpenTelemetryを起点に、AIエージェントの評価データセットを育てる

・TL;DR LLMの精度改善に使える評価データセットに近道はありません。実際のユーザーや業務部門のフィードバックをトレースに結びつけ、地道に育てます。育てたデータセットは、モデルやプロンプトを変えた前後を同じ入力で比べる回帰テストになります。 ・デモでは、ADK Web標準機能でフィードバックを受け、どの実行への評価かをtrace_idで特定してLangfuseのトレースに登録し、そのまま評価データセットへ移せます。 ・アプリからのトレース送信をOpenTelemetryに統一しておけば、観測ツールもエージェントSDKも後から選び直せます。
cs.LG updates on arXiv.org

Operating Multi-Node Full Fine-Tuning on NVIDIA B300: A Field Report on Telemetry-Based Triage, Negative Results, and Operational Hardening

・arXiv:2608.05944v1 Announce Type: cross Abstract: We report operational experience full-fine-tuning a 32.76B-parameter dense model (Qwen3-32B) on 16 x NVIDIA B300 (two nodes, FSDP / ZeRO-3) -- among the first published field accounts on this accelerator. ・We claim no new algorithm. ・The individual mechanisms we use are established practice; our contribution is the integrated field experience and a set of calibrated mea
cs.LG updates on arXiv.org

Optimal or Greedy Decision Trees? Revisiting their Objectives, Tuning, and Performance

・arXiv:2409.12788v3 Announce Type: replace Abstract: Recently there has been a surge of interest in optimal decision tree (ODT) methods that globally optimize accuracy directly, in contrast to traditional approaches that locally optimize an impurity or information metric. ・However, the literature shows conflicting evidence on the value of ODTs, with some demonstrating superior out-of-sample performance of ODTs over gre
cs.LG updates on arXiv.org

Optimal Rates for Learning with Monotone Adversaries

・arXiv:2608.06337v1 Announce Type: cross Abstract: A monotone adversary observes an i.i.d. ・labeled sample and appends a finite number of further examples of its choice, every one of them labeled correctly by the target hypothesis. ・The learner sees a uniform shuffle of the combined sample and is scored on the original distribution.
Hugging Face Papers

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
cs.LG updates on arXiv.org

OTLesMix: Wasserstein Barycenter and Optimal Transport Map for Synthetic Lesion Generation with Diverse Shapes and Locations

・arXiv:2608.06264v1 Announce Type: cross Abstract: The development of deep learning over the past decade has revolutionized medical imaging segmentation, allowing the extraction of precise descriptors from large volumes to characterize pathologies. ・Data augmentation is a technique widely regarded as a way to improve model training. ・It includes simple transformations like spatial operations or intensity modifications,
cs.LG updates on arXiv.org

Otter: A Time-Aware, History-Conditioned Human Chess AI

・arXiv:2608.05206v1 Announce Type: cross Abstract: Otter is a 15.3M-parameter human chess AI that predicts human move selection by modeling play as a time-aware, sequential process rather than treating each position in isolation. ・It combines two conditioning signals: (1) a move history encoder that conditions predictions on the last 20 moves, capturing opening preferences, positional drift, and intra-game behavioral t
WIRED

Our Favorite Fans Are on Sale to Help With Summer Heat Waves (2026)

・We’re still in the dog days of summer—keep your cool with deals on the best fans we’ve tested for on the go and at home.
Hugging Face Papers

PaDoc: Layout-Grounded Parallel Decoding for Document Parsing

PaDoc: Layout-Grounded Parallel Decoding for Document Parsing
cs.LG updates on arXiv.org

Perfect reconstruction of sparse signals using nonconvexity control and one-step RSB message passing

・arXiv:2512.17426v2 Announce Type: replace-cross Abstract: We consider sparse signal reconstruction via minimization of the smoothly clipped absolute deviation (SCAD) penalty, and develop one-step replica-symmetry-breaking (1RSB) extensions of approximate message passing (AMP), termed 1RSB-AMP. ・Starting from the 1RSB formulation of belief propagation, we derive explicit update rules of 1RSB-AMP together with the corre
cs.LG updates on arXiv.org

Persona-Pruner: Sculpting Lightweight Models for Role-Playing

・arXiv:2606.14695v2 Announce Type: replace Abstract: Language Models (LMs) have shown remarkable potential as role-playing chatbots, delivering consistent, stylized interactions when given a specification of a character or user persona. ・However, applying these capabilities to real-world applications (e.g., ecosystems with numerous NPCs interacting simultaneously) exposes a critical inefficiency due to the excessive co
cs.LG updates on arXiv.org

Perturbation Sensitivity at Convergence: A Simple Signal for Identifying Spuriously Correlated Samples

・arXiv:2608.05419v1 Announce Type: new Abstract: Models trained by empirical risk minimization on data containing spurious correlations achieve high average accuracy while failing on subpopulations where the correlation does not hold. ・Existing methods for identifying the affected samples without group annotations rely on signals from early training, which requires locating the epoch at which to intervene, a hyperparam
cs.LG updates on arXiv.org

Phylogenetic Tree Inference with Tropical Axial Attention

・arXiv:2605.13894v2 Announce Type: replace-cross Abstract: In this work, we introduce a Tropical Axial Attention neural reasoning architecture that replaces vanilla softmax dot-product attention with max-plus operators, inducing a piecewise-linear structure aligned with dynamic programming formulations. ・From multi-species sequence alignments, our model learns all possible pairwise distances and is trained using a comb
cs.LG updates on arXiv.org

Physics-Based Molecular Fingerprints from Spectral Graph Theory Provide Efficient Geometry-Aware Measures of Chemical Similarity

・arXiv:2608.05336v1 Announce Type: cross Abstract: Molecular representations are essential for the evaluation of molecular similarity and the development of structure-property relationships. ・Despite the known importance of 3D structure to determine chemical and physical properties, the most widely used molecular fingerprints encode only two-dimensional connectivity. ・Such representations fail to distinguish similar but
cs.LG updates on arXiv.org

Physics-Guided Concentration Inference from Resistance Transients in a Mixed-Phase SnO-SnO$_2$ Carbon Monoxide Sensor with p-n Switching

・arXiv:2605.23971v2 Announce Type: replace-cross Abstract: This work presents a physics-guided machine-learning framework for carbon monoxide concentration inference from experimentally measured resistance transients of a mixed-phase SnO-SnO$_2$ material gas sensor exhibiting temperature-dependent p-n switching behavior. ・Cycle-level transient responses are represented through physically interpretable descriptors and c
cs.LG updates on arXiv.org

PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLMs

・arXiv:2608.05162v1 Announce Type: cross Abstract: Pooling is a consequential but under-examined design choice in decoder-only concept representation work: practitioners must collapse token-level hidden states into a passage-level vector, yet no shared protocol exists for comparing this choice across concepts, models, and tasks. ・Reported gains are confounded by simultaneous changes in dataset, layer, construction meth
cs.LG updates on arXiv.org

Positive-Data Learning of Fixed-Observation Linear MCFGs from Working Binary Presentations

・arXiv:2605.11644v2 Announce Type: replace-cross Abstract: We study positive-data learning of languages admitting reduced working binary linear nondeleting multiple context-free grammar presentations of bounded fan-out. ・The learner is supplied with a fixed explicit finite monoid homomorphism (h:\Sigma^*\to M), used as a compositional finite-state observation. ・We define ((f,h))-tuple substitutability through named sent
cs.LG updates on arXiv.org

Positive-Unlabeled Preference Optimization For Chest X-ray Report Generation

・arXiv:2608.05341v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) for radiology report generation are typically trained on retrospective clinical reports, which suffer from omission noise: clinically present findings are left unreported due to the omission of subtle findings. ・For example, prior studies show that cardiomegaly may be omitted from ICU chest X-ray reports when the imaging request is focused
cs.LG updates on arXiv.org

Potential Matching Optimal Transport: Continuous Normalizing Flows for Exact $p$-Wasserstein Dynamics

・arXiv:2608.05666v1 Announce Type: new Abstract: We introduce Potential Matching Optimal Transport (PMOT), a potential-flow framework for general $p$-cost optimal transport with $c_p(x,y)=\|x-y\|^p$. ・PMOT parameterizes the CNF velocity field with a scalar potential in the generalized Benamou--Brenier form for the chosen exponent $p$. ・It trains the potential gradient with a self-induced matching loss along straight bri
cs.LG updates on arXiv.org

PPDL: LLM-Based Flows as Probabilistic Programs

・arXiv:2608.05234v1 Announce Type: new Abstract: Building reliable applications that leverage large language models (LLMs) remains a significant challenge. ・While LLMs offer impressive capabilities across diverse tasks, their outputs often lack accuracy and provide no clear measure of confidence. ・This uncertainty compounds in flows of multiple calls to LLMs and other tools, making it difficult for developers and end-us
cs.LG updates on arXiv.org

Predicting Task Difficulty Without Rollouts

・arXiv:2608.05797v1 Announce Type: new Abstract: Task difficulty dictates an agent's likelihood of success, and estimating it without rollouts means forecasting this directly from a task description before executing costly simulations in stateful environments. ・Reliable estimates would therefore allow environment designers to calibrate evaluation benchmarks and construct progressive training curricula. ・This becomes inc
Hugging Face Papers

Previous

Previous
cs.LG updates on arXiv.org

PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis

・arXiv:2608.05249v1 Announce Type: new Abstract: Real-world multimodal instructions often bundle multiple requirements with unequal importance, yet most multimodal training data still reduce instruction following to answering one self-contained question. ・We study this gap through \textbf{rubric comprehension}, which casts the model not as a generator measured against rubrics but as an \textbf{executor} that follows th
cs.LG updates on arXiv.org

ProDVI: Programmatic Dynamics Priors for Value Network Initialization

・arXiv:2608.06015v1 Announce Type: new Abstract: Deep Reinforcement Learning (RL) is notoriously sample inefficient. ・One contributing factor is that RL agents are typically initialized from scratch, forcing them to acquire task-relevant knowledge through online interaction. ・Existing approaches obtain informative initializations through pre-collected datasets, high-fidelity simulators, or meta-learning over related tas
cs.LG updates on arXiv.org

Provably Efficient Self-Calibrating Quantum Fault Tolerance

・arXiv:2608.05686v1 Announce Type: cross Abstract: Quantum error correction protects logical information only when every physical operation remains below the fault-tolerance threshold, a condition that must be maintained continuously rather than only at the initial calibration. ・In practice, however, analog control parameters inevitably drift because of environmental fluctuations. ・As future fault-tolerant quantum compu
cs.LG updates on arXiv.org

QEvict: Recoverable Quantized KV Eviction for Attention-Drift-Robust Long-Context Decoding

・arXiv:2608.05326v1 Announce Type: new Abstract: Autoregressive large language model inference is increasingly constrained by the memory footprint of the Key-Value (KV) cache. ・A dominant line of work reduces this footprint by evicting tokens that appear unimportant under attention-derived scores. ・However, such policies make an implicit irreversible decision: once a token is evicted, it cannot become useful again.
cs.LG updates on arXiv.org

Quality Diversity for Reliable Data Driven Time-Use Optimization

・arXiv:2608.05230v1 Announce Type: cross Abstract: The daily allocation of the finite 24-hour time budget is strongly associated with physical, mental, and cognitive health. ・While predictive models can estimate the relationship between time-use compositions and health outcomes such as body mass index, life satisfaction, and cognition, most optimization approaches focus only on maximizing expected benefit and do not co
cs.LG updates on arXiv.org

Quantum-Structured World Models (QSWMs) for Predictive Latent Dynamics

・arXiv:2608.05371v1 Announce Type: new Abstract: World models learn latent states that summarize interaction histories, evolve over time, and support prediction, simulation, or planning. ・Most existing world models represent these states using classical vectors, probability distributions, recurrent hidden states, or transformer activations. ・In this paper, we introduce Quantum-Structured World Models (QSWMs), a quantum-
#LLMタグ

RAGの「精度」を勘に頼らない:RAGASフレームワークによる客観的な評価パイプラインの構築

・本文 RAG(検索拡張生成)システムの開発において、最大の難所は「評価」です。
WIRED

Ranking the Best Smart Glasses: Meta, Viture, & More (2026)

・This burgeoning wearable tech lets you talk to an AI assistant, listen to music, or check out a display screen from the comfort of your very own face.
cs.LG updates on arXiv.org

Rapid Embodiment Adaptation for Quadrupedal Locomotion

・arXiv:2608.01506v2 Announce Type: replace-cross Abstract: Humans readily adapt their movements as their bodies change through aging, injury, or load carrying, but learning-based robot policies often break when hardware properties shift. ・We introduce an online embodiment adaptation framework for quadrupedal locomotion that infers embodiment parameters from short interaction histories and conditions control on the infe
cs.LG updates on arXiv.org

RASP-QAOA: Resource-Aware Per-Instance Selection for Exact QAOA Simulation

・arXiv:2608.05646v1 Announce Type: cross Abstract: Exact QAOA simulation spans several computational representations whose useful regions differ sharply across graph structure, circuit depth, precision, and available memory. ・Choosing only a backend name hides these differences: an executable choice also fixes the representation, adapter, precision mode, and memory policy. ・We introduce RASP-QAOA, a per-instance selecto
WIRED

Ratio vs. Simply Good: Which Plastic-Free Coffee Maker Is Best?

・Microplastics are in everything, especially your coffee. ・A new generation of plastic-free drip coffee brewers is trying to change this.
cs.LG updates on arXiv.org

Realizable Bayes-Consistency for General Metric Losses

・arXiv:2605.03823v3 Announce Type: replace Abstract: We study strong universal Bayes-consistency in the realizable setting for learning with general metric losses, extending classical characterizations beyond $0$-$1$ classification (Bousquet et al., 2020; Hanneke et al., 2021) and real-valued regression (Attias et al., 2024). ・Given an instance space $(X,\rho)$, a label space $(Y,\ell)$ with possibly unbounded loss, an
cs.LG updates on arXiv.org

Reasoning Errors Have a Region and a Direction in the Residual-Stream Trajectory of LLMs

・arXiv:2608.05660v1 Announce Type: new Abstract: As language models are increasingly used for tasks that require verifiable reasoning, reliably distinguishing sound reasoning from flawed reasoning has become an important practical problem. ・Recent trajectory-based methods seek this signal in layerwise residual-stream displacements, which capture how representations change while attenuating some stable, token-specific i
cs.LG updates on arXiv.org

Rectifying Geometric Misalignment: Online Source-Free Adaptation for Class-Imbalanced EEG

・arXiv:2608.05315v1 Announce Type: new Abstract: Electroencephalography (EEG) based Brain-Computer Interfaces (BCIs) often require unsupervised domain adaptation (UDA) to generalize across subjects and sessions. ・While Riemannian alignment methods like the Riemannian Centering Transformation (RCT) are effective for handling covariate shifts, they implicitly assume balanced class priors. ・However, in realistic online BCI
cs.LG updates on arXiv.org

Recursive Synthesis for Long-Horizon Terminal Tasks

・arXiv:2608.05466v1 Announce Type: cross Abstract: High-quality long-horizon training data for terminal agents is expensive to produce, often costing hundreds to thousands of dollars per task, because each task must keep the instruction, environment, reference solution, and verifier mutually consistent. ・Human authoring does not scale, and direct generation with large language models (LLMs) often breaks these dependenc
cs.LG updates on arXiv.org

Reducing Hallucination in Vision-Language Models via Stage-wise Preference Optimization under Distribution Shift

・arXiv:2605.16411v2 Announce Type: replace-cross Abstract: Hallucination remains a fundamental challenge in vision-language models (VLMs), where autoregressive generation may produce linguistically plausible yet physically inconsistent or visually ungrounded responses due to likelihood maximization under joint probabilistic modeling. ・We propose a stage-wise preference optimization framework for hallucination reduction
OpenAI News

Responding to the next frontier of critical cyber capabilities

・OpenAI is sharing preliminary cybersecurity evaluations for Astra and the steps we’re taking to strengthen safeguards and security controls.
cs.LG updates on arXiv.org

Right Knowledge, Wrong Answer: Characterizing Parametric Temporal Conflict in Open-Weight Language Models

・arXiv:2606.20959v3 Announce Type: replace Abstract: Language models may encode both outdated facts and their newer replacements. ・We introduce Parametric Temporal Conflict (PTC), where the newer fact is present and recoverable, but the default forward pass prefers the outdated one. ・We release a deterministically verified benchmark of 8,746 Wikidata position-holder transitions and evaluate four open-weight language mod
stat.ML updates on arXiv.org

Risk-Aware Quantile Learning for Personalized Dynamic Treatment Regimes

・arXiv:2608.05434v1 Announce Type: cross Abstract: Sequential clinical decision-making often involves more than maximizing average efficacy. ・Clinicians may need to simultaneously optimize clinically relevant tails of the outcome distribution, control treatment-related risk, and choose among multiple treatment options. ・Existing quantile dynamic treatment regime (DTR) methods capture distributional features of treatment
cs.LG updates on arXiv.org

Robot Learning from Human Demonstrations: Handwritten Alphabet Trajectories and Human-Likeness Evaluation

・arXiv:2608.06221v1 Announce Type: cross Abstract: Learning from demonstration (LfD) provides a developmental framework through which robots can develop motor skills by observing and imitating human dynamics, reducing reliance on explicit programming to teach a skill to a robot. ・The resulting human-like robot motion is recognised as a key factor in building trust and enabling natural collaboration in human-robot inter
cs.LG updates on arXiv.org

Robust Context-Aware Detection of Malicious Instructions in Text

・arXiv:2608.05430v1 Announce Type: cross Abstract: The remarkable instruction-following ability of modern LLMs has enabled their practical use as the minds of agents that can autonomously complete increasingly complex tasks. ・Therein, however, also lies their vulnerability to attacks which embed malicious instructions in text, common variants of which are known as indirect prompt injection (IPI). ・A fundamental task in
cs.LG updates on arXiv.org

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction

・arXiv:2608.06310v1 Announce Type: new Abstract: Recent advances in reward modeling show a paradigm shift from discriminative reward models to generative reward models. ・However, despite their strong capabilities in response ranking, generative reward models have not realized their potential in reinforcement learning (RL). ・Our analysis reveals that this limitation arises from a mismatch between the comparative nature o
cs.LG updates on arXiv.org

RxnCLF: Contrastive Transformation-Aware Reaction Foundation Model for Improved Reactivity Prediction

・arXiv:2608.06259v1 Announce Type: new Abstract: Reaction yield prediction remains challenging because labeled data are scarce and reaction space is both combinatorially large and sparsely populated, limiting the generalization of existing reaction representations. ・String-, fingerprint-, and graph-based reaction encodings only partially capture chemical transformations, making accurate prediction difficult for reactio
#LLMタグ

Ryzen AI MAX+ 395(Strix Halo / GTR9 Pro 128GB)で DeepSeek-V4-Flash を Ubuntu + ROCm + ds4 で動かす

Ryzen AI MAX+ 395(Strix Halo / GTR9 Pro 128GB)で DeepSeek-V4-Flash を Ubuntu + ROCm + ds4 で動かす
cs.LG updates on arXiv.org

Safe Evolution with Circuit Anchors

・arXiv:2608.05158v1 Announce Type: cross Abstract: In biological evolution, unconstrained mutation can lead to catastrophic outcomes: organisms may evolve enhanced capabilities while losing essential functions for survival. ・Nature's solution is \textit{developmental constraints}, where core regulatory genes remain anchored while peripheral genes adapt freely. ・We observe that current self-evolution algorithms for large
cs.LG updates on arXiv.org

SAGA: Score-Weighted Adaptive Generation Alignment for Low-Resource Nordic Language Models

・arXiv:2608.06179v1 Announce Type: new Abstract: Preference optimisation has proven effective for improving large language models but typically relies on costly human preference annotations. ・Extending these methods to morphologically rich, low-resource languages remains challenging because such annotations are scarce. ・We present SAGA (Score-weighted Adaptive Generation Alignment), a parser-guided preference optimisati
Zennの「機械学習」のフィード

SageMaker の CPU 推論が想定の3倍遅かった話

・はじめに Digital Garage ではいくつかのプロダクトで OCR 機能を提供しており、それらは LayoutLMv3 ベースの機械学習モデルによって実現されています。 ・生成 AI による OCR も選択肢ですが、レスポンスタイムを重視して機械学習モデルを選んでいます。 ・ところが、想定では 1000ms、手元の M2 Mac では 700ms で終わる推論が、SageMaker 上のエンドポイントでは 3000ms 程度かかる事象に遭遇しました。
The Verge

Samsung’s Z Fold 8 Ultra is more of the same, but better than ever

・The Z Fold 8 Ultra retains the boxy shape of Samsung’s previous foldables. ・What makes the Samsung Galaxy Z Fold 8 Ultra so Ultra? ・Samsung's first foldable with the Ultra name isn't a souped-up version of the regular Fold 8, as you might expect.
cs.LG updates on arXiv.org

Scalable estimation of VARMA models

・arXiv:2608.06340v1 Announce Type: cross Abstract: Vector autoregressive moving-average (VARMA) models have long been considered impractical beyond moderate dimensions: the likelihood is non-convex, the parametrization is identified only up to equivalence, and every evaluation costs a pass over the entire series. ・Yet their moving-average term captures with a few parameters what a pure autoregression matches only with
cs.LG updates on arXiv.org

Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning Regime

・arXiv:2509.24882v3 Announce Type: replace Abstract: Neural scaling laws underlie many of the recent advances in deep learning, yet their theoretical understanding remains largely confined to linear models. ・In this work, we present a systematic analysis of scaling laws for quadratic and diagonal neural networks in the feature learning regime. ・Leveraging connections with matrix compressed sensing and LASSO, we derive a
cs.LG updates on arXiv.org

Scientific Machine Learning of Chaotic Systems Learns Reduced-Order Equations for Neural Populations

・arXiv:2507.03631v4 Announce Type: replace Abstract: Extracting interpretable mathematical models from complex dynamical systems is difficult, especially for chaotic dynamics observed with noisy experimental data. ・We present PEM-UDE, a method that combines prediction-error methodology with universal differential equations to discover governing equations from limited, noise-corrupted observations. ・Prediction-error feed
WIRED

Scientists Used AI to Create 16 New Viruses

・The use of AI systems to create viruses opens up new possibilities for combating bacterial resistance. ・It also raises concerns about the pace at which technology is outstripping regulation.
cs.LG updates on arXiv.org

SCP-NL2TL: Selective Conformal Prediction with Semantic Verification for Natural Language to Temporal Logic Specifications

・arXiv:2608.05439v1 Announce Type: cross Abstract: Translating natural language instructions into machine-interpretable formal specifications enables robots and autonomous systems to plan, reason, and formally verify their behavior. ・However, existing translation models typically generate a specification for every input, even when the result is unreliable or fails to capture the user's intent, creating risks in safety-
cs.LG updates on arXiv.org

SEAM: Global consistency beyond local accuracy in scientific machine learning

・arXiv:2608.05702v1 Announce Type: new Abstract: Scientific machine learning commonly validates models at the level of a subdomain, a benchmark split, or an explanation for one prediction. ・Yet such local checks cannot establish whether the resulting explanations can be assembled into one globally admissible explanation. ・We introduce Scientific Explanation-Admissibility Machines (SEAM), a generator-agnostic framework t
cs.LG updates on arXiv.org

SemiAdapt-Instruct: Extensible Instruction Tuning via Latent Domain-Specialised Adapters

・arXiv:2608.05161v1 Announce Type: cross Abstract: Instruction-tuned LLMs are deployed into environments where domains evolve, yet extending a fine-tuned model's capabilities without full retraining remains an unsolved practical challenge. ・We present SemiAdapt-Instruct, a modular framework that discovers latent instruction domains, trains per-domain LoRA adapters in parallel, and performs parameter-free routing, incor
cs.LG updates on arXiv.org

Skill Neologisms: Towards Skill-based Continual Learning

・arXiv:2605.04970v3 Announce Type: replace Abstract: Modern LLMs show mastery over an ever-growing range of skills, as well as the ability to compose them flexibly. ・However, extending model capabilities to new skills in a scalable manner is an open problem: fine-tuning and parameter-efficient variants risk catastrophic forgetting, while context-based approaches have limited expressiveness and are constrained by the mo
cs.LG updates on arXiv.org

SkillTFM: Gated Skill Evolution for Training-Free Adaptation of Tabular Foundation Models

・arXiv:2608.06137v1 Announce Type: new Abstract: Tabular data are ubiquitous in real-world applications and are crucial for data-driven prediction and decision-making across science, industry, finance, healthcare, and public services. ・Tabular foundation models (TFMs) have emerged as a promising paradigm for general-purpose tabular learning, offering reusable predictors across diverse datasets and substantially reducin
cs.LG updates on arXiv.org

SkillTrace: Multi-Trace Provenance Auditing for LLM-Agent Skill Reuse

・arXiv:2608.05204v1 Announce Type: cross Abstract: LLM-agent ecosystems are rapidly growing around reusable skills: mixed-modality packages of metadata, natural-language instructions, code, tools, references, and operational workflows. ・As skills become marketplace artifacts, auditing their reuse is no longer the same problem as ordinary code clone detection. ・Existing detectors target single-modality source code or who
Hugging Face Papers

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding
The Verge

Sony could release a cheaper version of its WH-1000XM4 headphones, according to leaks

・Following the release of its premium and very expensive $650 The Collexion headphones in May, Sony could be taking its 1000X line of wireless noise-canceling headphones in a cheaper direction next month. ・The company is reportedly planning a refreshed and more affordable version of its six-year-old WH-1000XM4 headphones that will feature similar performance and specs as the original but at a lower price point, accordi
cs.LG updates on arXiv.org

Sparse Mutual Information Graph Averaging for Improving Random Indexing Embeddings

・arXiv:2608.05724v1 Announce Type: cross Abstract: Sparse word embedding pipelines can avoid dense co-occurrence matrix materialization, dense factorization, and gradient training while still relying on sparse global corpus statistics. ・This paper studies Random Indexing (RI) vectors refined by weighted averaging on a sparse Positive Pointwise Mutual Information (PPMI) graph. ・On a fairytales corpus, the covered semanti
cs.LG updates on arXiv.org

Spectral Aliasing Pretext: A novel task for Self-Supervised fault diagnosis in rotating machinery

・arXiv:2608.05705v1 Announce Type: new Abstract: Deep learning is a new way for machinery fault diagnosis but requires extensive labeled data, a scarce resource in industrial settings. ・We propose Spectral Aliasing Pretext (SAP), a self-supervised learning method that pretrains models on unlabeled vibration data by exploiting spectral aliasing. ・We deliberately undersample signals to create folded spectrum, then train a
cs.LG updates on arXiv.org

Spectral Distillation: From Nonlinear Dynamics to Linear State-Space Models

・arXiv:2608.05416v1 Announce Type: new Abstract: Can nonlinear dynamical systems be learned through a compact linear state-space representation, without directly solving a non-convex system-identification problem? ・We give a provable pipeline for doing so. ・Starting from observations of an unknown nonlinear dynamical system, we first learn an implicit spectral predictor using Observation Spectral Filtering (OSF), a conv
cs.LG updates on arXiv.org

SR-JEPA: Learning Predictive Latent State in 3D Scenes

・arXiv:2608.05774v1 Announce Type: cross Abstract: Joint-embedding predictive architectures learn by predicting latent representations of missing observations, yet many masked JEPAs are evaluated primarily through the encoders they produce. ・We ask what a trained predictive pathway itself infers when an entire entity is absent from a native 3D scene. ・We introduce SR-JEPA, a point-native JEPA for scene-scale point cloud
cs.LG updates on arXiv.org

STATe-of-Thoughts: Structured Action Templates for Tree-of-Thoughts

・arXiv:2602.14265v3 Announce Type: replace-cross Abstract: Inference-Time-Compute (ITC) methods like Best-of-$n$ and Tree-of-Thoughts are meant to produce output candidates that are both high-quality and diverse, but their use of high-temperature sampling often fails to achieve meaningful output diversity. ・Moreover, existing ITC methods offer limited control over $\textit{how}$ to perform reasoning, which in turn limi
cs.LG updates on arXiv.org

Stochastic Dynamics on Persistence Diagram Space via Reinforcement Learning

・arXiv:2608.06276v1 Announce Type: cross Abstract: Persistence diagrams (PDs) provide stable and interpretable summaries of multiscale topological structure. ・While substantial progress has been made in the statistical analysis of PDs, existing literature often treats diagrams as static objects and provide limited frameworks for probabilistic modeling and stochastic evolution on PD space. ・We introduce a reinforcement l
stat.ML updates on arXiv.org

Structured Dimension-Matched Joint Variational Transdimensional Inference

・arXiv:2608.05607v1 Announce Type: cross Abstract: Bayesian model selection couples a discrete model indicator with a model-specific continuous parameter space. ・We introduce structured dimension-matched variational transdimensional inference (SM-VTI) for finite enumerable model spaces. ・A rooted construction graph expresses a model as a sequence of local stop/child decisions.
cs.LG updates on arXiv.org

Supervised Learning Has a Geometric Blind Spot

・arXiv:2604.21395v3 Announce Type: replace Abstract: Ordinary supervised training minimises the task loss and then stops. ・It never pays for how far the representation moves when the input is nudged along directions that helped fit training labels---including directions that are nuisance at deployment. ・We call that leftover sensitivity the geometric blind spot of empirical risk minimisation.
WIRED

Surfshark Promo Codes: 87% Off | August 2026

・Save up to 87% with a Surfshark coupon code, 3 months of VPN free today, and more from WIRED.
cs.LG updates on arXiv.org

Surv-IPTB: An Attention-Based Model for Estimating Individual Probability of Treatment Benefit with Survival Data

・arXiv:2608.06288v1 Announce Type: new Abstract: This work presents a novel attention-based framework for estimating the Individual Probability of Treatment Benefit (IPTB) in survival analysis contexts. ・The proposed model, called Surv-IPTB, directly quantifies the probability that a specific patient will experience extended survival time under treatment versus control. ・We reformulate IPTB estimation as a binary classi
cs.LG updates on arXiv.org

Symbol Grounding in Neuro-Symbolic AI: A Gentle Introduction to Reasoning Shortcuts

・arXiv:2510.14538v3 Announce Type: replace-cross Abstract: Neuro-symbolic (NeSy) AI aims to develop deep neural networks whose predictions comply with prior knowledge encoding, e.g. ・safety or structural constraints. ・As such, it represents one of the most promising avenues for reliable and trustworthy AI.
Hugging Face Papers

Task-Conditional Flow Matching for Balanced Multilingual Text Embedding Adaptation

Task-Conditional Flow Matching for Balanced Multilingual Text Embedding Adaptation
Hugging Face Papers

Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains

Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains
cs.LG updates on arXiv.org

Temporal Bridges for Spatial Resolution: Enhancing Climate Data Super-Resolution with Bidirectional Alignment

・arXiv:2608.05981v1 Announce Type: cross Abstract: High-resolution climate data is crucial for meteorological predictions and for informing decision support across diverse domains. ・However, the acquisition of such high-resolution climate information is often prohibitively costly, necessitating the development of data-driven meteorological prediction models. ・These models aim to generate fine-grained climate data from l
cs.LG updates on arXiv.org

TESSERA v2: Scaling Pixel-wise Earth Foundation Models

・arXiv:2607.03949v2 Announce Type: replace-cross Abstract: Pixel-wise Earth-observation (EO) foundation models are now achieving state-of-the-art performance via generated spatial embeddings. ・However, how these models scale and how best to spend a pretraining budget remain poorly understood. ・We present the largest controlled scaling study for EO to date: 395 training runs within a fixed pixel-wise Barlow Twins family,
cs.LG updates on arXiv.org

Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet

・arXiv:2509.06861v3 Announce Type: replace-cross Abstract: Test-time scaling increases inference-time computation through longer reasoning chains and has shown strong performance gains across many domains. ・However, frontier models still suffer from factuality hallucinations, raising the question of whether increased computation is effective on closed-book knowledge-intensive tasks. ・In this work, we evaluate 14 reasoni
cs.LG updates on arXiv.org

THBKG: A Temporal Biomedical Knowledge Graph for Decision-Aligned Clinical Advancement Prediction

・arXiv:2608.05982v1 Announce Type: new Abstract: Inadequate target--disease linkage accounts for 40--50\% of Phase~II efficacy failures, so anticipating which programmes will advance would let sponsors back the hypotheses most likely to reach patients. ・What a programme can be judged on is the evidence that supported its linkage \emph{when it entered the clinic}. ・No existing biomedical knowledge graph allows that evide
WIRED

The ‘Manosphere’ Isn’t a Movement. It’s a Multibillion-Dollar Grievance Industry

・Many young men are driven to resentment and are financially exploited as influencers sell them classes, pills, and the illusion of clout, a new report reveals.
WIRED

The 5 Best Laptop Power Banks I've Personally Tested (2026)

・These powerful, high-capacity portable batteries are the best ones when laptop charging is a priority. ・All have at least 20,000 mAh, enough to charge most laptops twice.
WIRED

The 7 Best TV Shows to Stream This Month

・Lanterns, Dark Matter, and the original Star Trek are just a few of the TV shows you should be watching right now.
The Verge

The best classic slasher movie you&#8217;ll never watch

・Starting with the original in Camp Miasma in 1980, the horror franchise went on to have a long life. ・There were multiple sequels, spinoffs in the form of arcade cabinets and board games, and just about every kind of merchandise you can think of, from alarm clocks and lunch boxes to Halloween costumes and action figures. ・It got so popular there were even knockoff costumes at party stores.
WIRED

The Hottest New AI Chatbot Is Just a Guy Answering Your Questions

・WIRED spoke with Tucker Bryant, an artist and former Google employee who created ChatTJB to get people to reflect on the “strange moment” we’re in.
cs.LG updates on arXiv.org

The Ignition Index: Measuring Global Workspace Dynamics in Language Models

・arXiv:2608.05160v1 Announce Type: cross Abstract: We introduce the Ignition Index (I), a validated scalar metric that operationalizes Global Workspace Theory's (GWT) all-or-none ignition prediction in transformer language models. ・The metric fits a four-parameter sigmoid to per-layer linear probe accuracy as a function of input signal strength, extracting steepness parameter beta-hat: high values indicate abrupt, igni
cs.LG updates on arXiv.org

The Impact of Dimensionality on the Stability of Node Embeddings

・arXiv:2604.08492v3 Announce Type: replace Abstract: Previous work has shown that node embedding methods can produce different representations and downstream predictions across repeated training runs, even when trained on the same data with identical hyperparameters. ・However, the role of embedding dimensionality in this instability remains poorly understood. ・In this work, we systematically analyze how embedding dimens
cs.LG updates on arXiv.org

The Impossibility Triangle of Long-Context Modeling

・arXiv:2605.05066v2 Announce Type: replace-cross Abstract: We identify and prove a fundamental trade-off governing long-sequence models: no model can simultaneously achieve (i) per-step computation independent of sequence length (Efficiency), (ii) state size independent of sequence length (Compactness), and (iii) the ability to recall a number of historical facts proportional to sequence length (Recall). ・We formalize
cs.LG updates on arXiv.org

The Judgment-Consequence Gap: LLM Moral Reasoning in Healthcare Decisions

・arXiv:2608.05583v1 Announce Type: cross Abstract: As large language models (LLMs) enter high-stakes domains such as healthcare, understanding their moral reasoning becomes essential. ・Decisions about scarce medical resources often hinge on judgments of responsibility, particularly when patients' own actions contribute to illness. ・We investigate how LLMs reason about responsibility and its consequences, tracing their j
cs.LG updates on arXiv.org

The Tamed Subgradient Unadjusted Langevin Algorithm beyond Convexity

・arXiv:2608.06283v1 Announce Type: new Abstract: We study the problem of sampling from target distributions whose potentials are simultaneously non-smooth, subject to superlinear gradient growth, and non-convex. ・We introduce the Subgradient Tamed Unadjusted Langevin Algorithm (SG-TULA), a discretisation of the Langevin diffusion that operates directly on subgradients, without relying on computationally demanding smoot
cs.LG updates on arXiv.org

Threshold-Based Early Stopping of Accumulations in Neural Networks with Binary Activation

・arXiv:2608.06177v1 Announce Type: new Abstract: Binary neural networks are very attractive for constrained deployment, enabling small footprint and low-power inference. ・For binary activations, the dot products become sign-controlled additions or subtractions, but the number of operations is unchanged. ・Indeed, every neuron or output channel still accumulates all of its input, even though only the sign will be retained
cs.LG updates on arXiv.org

Time Series Classification through Diffeomorphic Time Warping (DiffTW)

・arXiv:2606.23472v2 Announce Type: replace-cross Abstract: Time series classification involves learning a mapping from a continuous, temporally ordered sequence of real-valued observations to discrete response variables, like class labels. ・This task is fundamental in domains, including health monitoring, where temporal structure is critical for prediction. ・Dynamic Time Warping (DTW) is a standard technique for measuri
cs.LG updates on arXiv.org

Timestep-Conditioned Transformers for Global Weather Forecasting

・arXiv:2608.06241v1 Announce Type: new Abstract: Existing machine-learning weather forecasting models rely on predetermined and fixed autoregressive timesteps. ・The choice of model timestep involves a fundamental trade-off: shorter timesteps (e.g. ・1 to 6 hours) finely resolve atmospheric dynamics within the diurnal cycle but increase error accumulation for a given forecast horizon, while longer timesteps (e.g.
cs.LG updates on arXiv.org

Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering

・arXiv:2608.06366v1 Announce Type: cross Abstract: Electronic health record (EHR) feature engineering is a major bottleneck in clinical research and AI, accounting for 39-45% of data scientists' workload. ・This is especially pronounced in heart failure, which affects an estimated 6.7 million U.S. ・adults and requires integrating fragmented EHR data with disease-specific, guideline-based clinical reasoning.
cs.LG updates on arXiv.org

Training a Conditioned Video Game Agent on a VLM Annotated Dataset

・arXiv:2608.05954v1 Announce Type: cross Abstract: Reinforcement Learning (RL) is a powerful but far from easy-to-use technique for policy learning. ・In the specific case of video games, access to the game engine is required to get rewards for training (e.g. ・to collect rewards from the environment).
cs.LG updates on arXiv.org

Trajectory-guided discharge stratification for heart failure using short-context electronic health record sequence modeling

・arXiv:2511.16839v4 Announce Type: replace Abstract: Purpose: Heart failure (HF) discharge planning depends on identifying patients at risk of deterioration or death, yet accurate prediction from routinely collected electronic health records (EHRs) remains challenging. ・Methods: We develop trajectory-guided discharge stratification for heart failure (TGDS-HF), a methodology that reads the patient in-hospital trajectory
cs.LG updates on arXiv.org

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently

・arXiv:2511.17852v3 Announce Type: replace Abstract: Transformers can acquire Chain-of-Thought (CoT) capabilities to solve reasoning tasks via fine-tuning. ・Reinforcement learning (RL) and supervised fine-tuning (SFT) are two primary approaches to this end. ・In this work, we examine RL with verifiable process rewards and SFT for learning $k$-sparse Boolean functions with a one-layer transformer through intermediate reas
Zennの「大規模言語モデル」のフィード

Transformersライブラリ基礎

・https://qiita.com/_kawauso_/items/523128d9b1c9722bb0a8 ChatGPT登場以降、自然言語処理に興味を持ち始めた。2023年ころにTransformersライブラリを使ったサンプルコードをよくわからずに写経していた。頻度高く触っているたけではないが、最近、やっと概観がつかめてきた(気がする)ので、自分の理解を整理してみる。(と言いつつも書いている途中に知らないことが沢山出てきたので、書くことの大事さを改めて感じた) 1. ・背景知識 (1)Transformersとは、LLMをPython等で扱うためのライブラリであり、Hugging...
cs.LG updates on arXiv.org

Trust-Based Incentive Mechanisms in Semi-Decentralized Federated Learning Systems

・arXiv:2602.08290v2 Announce Type: replace Abstract: In federated learning (FL), decentralized model training allows multi-ple participants to collaboratively improve a shared machine learning model without exchanging raw data. ・However, ensuring the integrity and reliability of the system is challenging due to the presence of potentially malicious or faulty nodes that can degrade the model's performance. ・This paper pr
cs.LG updates on arXiv.org

TS-RAG: Retrieval Augmented Generation for Time Series Forecasting

・arXiv:2608.06223v1 Announce Type: cross Abstract: While deep learning models, particularly transformer-based architectures, have shown impressive performance in time series forecasting, the application of retrieval-augmented generation (RAG) in this domain remains limited. ・Since RAG has proven effective in enhancing the capabilities of large language models by incorporating relevant external information, retrieving s
Hugging Face - Blog

TutorMoments: Do AI tutors know when to help and when to hold back?

TutorMoments: Do AI tutors know when to help and when to hold back?
cs.LG updates on arXiv.org

Velocity- and Regime-Aware Detection of Intraday Options Market Manipulation, with Explainable Attribution

・arXiv:2608.05373v1 Announce Type: cross Abstract: Intraday market manipulation is hard to detect because its footprint is brief, buried in millions of quotes, and statistically similar to ordinary volatility. ・Detectors reach high recall only by flagging so many other days that measured precision collapses, producing alerts no regulator can act on. ・We show that this manipulation leaves a distinctive dynamic signature:
cs.LG updates on arXiv.org

Verifiable Regularity Criterion for Conditional Expectation Operators and Conditional Mean Embeddings with Applications to Nonparametric Regression, Bayesian Inverse Problems, and Koopman Operators

・arXiv:2608.06155v1 Announce Type: cross Abstract: Conditional expectation operators (CEOs) and their associated conditional mean embeddings (CMEs) play a central role across applied mathematics and machine learning, appearing in nonparametric regression, Bayesian inverse problems, and Koopman operator theory. ・A fundamental question is when a CEO maps a function space on $\mathcal{Y}$ into a prescribed function space
cs.LG updates on arXiv.org

VLMs for Videogame Data Annotation

・arXiv:2608.05949v1 Announce Type: cross Abstract: Vision Language Models (VLMs) and Artificial Intelligence (AI) agents have revolutionized how engineers approach complex problems in real-world applications. ・Their adoption in video games is on the other hand limited by the extreme variability of the synthetic scenarios and their poor compliance with real-world physics. ・Here we investigate the use of VLMs for annotati
cs.LG updates on arXiv.org

VSMP-IMU: Video-Grounded Semantic Motion Programs for Sensor-Aware Synthetic IMU Generation

・arXiv:2608.05782v1 Announce Type: cross Abstract: Wearable human activity recognition (HAR) is often limited by the scarcity of labeled sensor data, especially in low-resource, class-imbalanced, and subject-generalization settings. ・Synthetic IMU generation can reduce this dependency and enhance HAR machine learning model's performance, but existing approaches face a trade-off without addressing all factors: video-dri
#LLMタグ

WebMCPが来ると、AIは画面を読まなくなる理由。

・19時すぎ、検証用の管理画面を開いたまま、ChromeのDevToolsを閉じられずにいた。検索ボックスは画面の左上。結果を選ぶと右ペインが更新され、保存ボタンはスクロールの先にある。人なら数秒で済む操作なのに、ブラウザを動かすエージェントには毎回「今どの画面か」から説明が要る。 ・「これ、画面を見せるより操作を渡した方が早くない?」 続きをみる
The Verge

Werewolf transformed my old gadgets into USB-C powered ones

・When I haul my 27-inch desktop monitor into the garage to film fun gadget videos… or want to play Japanese SNES games with friends… or lose the AC adapter for my old external hard disk… I can now power them with a USB battery instead. ・None of them came with USB-C ports; the entire USB protocol was barely a twinkle in its creator's eye when Nintendo shipped the Super Famicom in 1990! ・But a tiny programmable dongle, fr
cs.LG updates on arXiv.org

What Drives Test-Time Adaptation for CLIP? A Controlled Empirical Study from an Update Perspective

・arXiv:2606.14299v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) such as CLIP have become a standard backbone for open-vocabulary recognition, yet their zero-shot predictions remain vulnerable to distribution shifts encountered at deployment. ・Test-Time Adaptation (TTA) has recently been extended to CLIP as a lightweight solution, leading to a rapidly growing body of TTA4CLIP methods.
The Verge

What&#8217;s behind the Google AI shake-up

・Some of the biggest names on Google's AI team got new jobs this week. ・In some cases, including for legendary Googler Jeff Dean, those jobs are no longer at Google. ・Given that Google's models seem to be behind the best of what's coming out of anthropic and OpenAI, is this a sign of Google in turmoil?
cs.LG updates on arXiv.org

When Do Corrective Features Help? An Agent for Corrective Feature Discovery on Black-Box Forecasters

・arXiv:2608.05207v1 Announce Type: new Abstract: Frozen pretrained forecasters often fail in structured, recurring ways that are costly to repair through fine-tuning. ・We study corrective feature discovery: mining interpretable features of a frozen forecaster's residual to drive a lightweight post-hoc corrector. ・Prior automated feature engineering models the data-generating process; corrective features instead model th
cs.LG updates on arXiv.org

When Does Consensus Mean Correctness? Measuring the Agreement-Accuracy Coupling with Semantics-Preserving Re-Rendering

・arXiv:2608.05670v1 Announce Type: new Abstract: A model's agreement across perturbed inputs is used both as a label-free reliability signal and as a self-training target, on the premise that agreement tracks correctness. ・That coupling is rarely measured directly: natural-image perturbations preserve meaning only by assumption, and no exact answer key localizes errors. ・Scientific figures remove both obstacles, a figur
cs.LG updates on arXiv.org

When Drafts Evolve: Speculative Decoding Meets Online Learning

・arXiv:2603.12617v2 Announce Type: replace Abstract: Speculative decoding has emerged as a widely adopted paradigm for accelerating large language model inference, where a lightweight draft model rapidly generates candidate tokens that are then verified in parallel by a larger target model. ・However, due to limited model capacity, drafts often struggle to approximate the target distribution, resulting in shorter accept
cs.LG updates on arXiv.org

Where Privacy Risk Lives in English-Source Multilingual RAG: A Stage-Decomposed Audit Across Five Query Languages

・arXiv:2608.05163v1 Announce Type: cross Abstract: A common assumption holds that switching to a non-English language makes a multilingual RAG system easier to attack for personal information. ・We test this on an English-source synthetic-PII corpus with five query languages and a two-stage defence (LLM input judge + regex output filter), in a pipeline whose translator, judge, back-translator, and generator are all Qwen
The Verge

Why does Apple keep banning Telegram, but never X?

・Elon Musk in the Oval Office on May 30th, 2025. ・| Photo by Kevin Dietsch / Getty Images For roughly an hour this week, Telegram vanished from Apple's App Store. ・Even during that blip, it was a stunning absence for such a major app: an avenue of communication for more than 1 billion users around the world, used widely as a secure platform for people who live in countries under censorship.
cs.LG updates on arXiv.org

Why the Third Axis Is Freedom

・arXiv:2608.05423v1 Announce Type: new Abstract: In generative training, a model produces an output and is penalised for its difference from an example. ・With one output per comparison, a model that produces one common answer can outperform a model retaining a broader repertoire. ・Explorative Modeling (XM) produces $K$ outputs per comparison and updates on the closest, claiming exploration as a "third pretraining axis"
Zennの「大規模言語モデル」のフィード

Why Your Claude Code Bill Is 20x Higher Than It Should Be

・Learn how Claude's prompt caching actually works, what silently resets it, and five practical ways to cut your Claude Code token bill without touching your CLAUDE.md file. ・Introduction Here is a number that surprises most Claude Code users the first time they see it: reading the same conversati...
Hugging Face Papers

World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation

World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation
Hugging Face Papers

WorldClaw: Agentic 3D Open-World Generation at Scale

WorldClaw: Agentic 3D Open-World Generation at Scale
cs.LG updates on arXiv.org

Worst-Case Distance-Aware Error Bounds for Neural Networks

・arXiv:2510.22021v3 Announce Type: replace Abstract: Safety-critical applications of machine learning require uncertainty estimates that support reliable worst-case analysis. ・Neural networks (NNs) provide expressive function approximation, while Gaussian processes (GPs) offer principled probabilistic uncertainty, but both face limitations in this setting: most modern neural architectures lack tractable worst-case erro
WIRED

Zohran Mamdani’s NYC Tech Team Is What DOGE Should Have Been

・The mayor of New York City has assembled a crew of Silicon Valley and United States Digital Service veterans to overhaul city services with better software.
#AIタグ

エンジニアじゃなくても、AIに「情報収集→記事化→自動化」を全部任せてみた話

・こうやって一つひとつ検証しながら、これからも投稿していきます。 ・よかったらフォローしてもらえると嬉しいです。 ・SNSも始めました。気軽にのぞいてみてください → https://x.com/Okuo_Kindaichi このシリーズ、実は毎週勝手に更新されている。
#AIタグ

エンジニア注目!今週のGitHubトレンド5選

・GitHubのトレンドには、世界中の開発者が今まさに注目しているツールやアイデアが集まります。 ・今週は、AIエージェントに専門知識を教え込む動きと、それを支える記憶やインフラの仕組みに関心が集まりました。 ・特に注目の5つを、初心者にもわかりやすく紹介します。
#AIタグ

コント・お笑い・AIさんのコツを聞きました。ネズミさんも登場チュー!

・コント・お笑い・AIさんのコツを聞きました。ネズミさんも登場チュー! AIさんはコント上手くなられて。どこで修行したのですか? 続きをみる
ITmedia NEWS 最新記事一覧

シャープ、通期純利益見通し170億円下方修正 円安など影響 AIサーバは9月に参入

・シャープは7日、2027年3月期の連結純利益が前期比47.3%減の250億円になる見通しだと発表した。期初予想から170億円下方修正した。樹脂・燃料の価格上昇や円安が収益を圧迫する。営業利益は190億円引き下げて300億円。売上高は1兆7700億円に据え置いた。
#LLMタグ

セキュリティ特化LLMの現状 ― Mythos、GPT-5.5-Cyber、Gemini 3.5 Flash Cyberは何が違うのか

・AIが「攻める側」にも「守る側」にも回り始めた 2026年に入ってから、大手AI企業が相次いでサイバーセキュリティ専用のLLMを発表するようになりました。背景にあるのは、脆弱性を見つけるAIの能力が、それを直すシステム側の速度をすでに追い越しつつあるという現実です。攻撃側が生成AIを悪用するスピードに、防御側も生成AIで対抗しないと間に合わない、という発想の転換が起きています。
Zennの「大規模言語モデル」のフィード

テキストを選んで ⌘J。送る前の下書きをその場で整える macOS アプリを作った

・この記事で紹介するのは、選択した文章を ⌘J 一発で整えて、同じ場所に書き戻す macOS アプリです。 ・名前は AI Text Shortcut。ChatGPT を開いてコピー&ペーストする往復をなくしたくて作りました。 ・対象読者: Mac で LINE / Slack / メールの下書きをよく書く人、ちょっとした LLM アプリを自作したい人 リポジトリ: hrak0x59/AITextShortcut ひとことで言うと メモみたいな下書きを選んで ⌘J すると、相手に送れる自然な文に置き換わります。
ITmedia NEWS 最新記事一覧

ドコモ・バイクシェアのシステム障害、いまだ全面復旧ならず 8月1日以降の利用者には「全額返金」へ

・ドコモ・バイクシェアは8月7日、システム不具合が続いている自転車シェアリングサービスについて、1日以降の利用料金を全額返金すると発表した。正常に利用できたユーザーも対象で、手続きは不要という。
#AIタグ

なぜChatGPT Enterpriseで業務効率化が止まるのか。暗黙知をAIに継承する最適解を開発者が徹底解説

・ChatGPT Enterpriseを導入し、業務効率化を実感した社員の割合は98.6%に達する。一方で、開発現場では「ツールを入れたのに途中で改善が止まる」問題が頻発している。 ・原因は、マニュアル通りの標準業務は自動化できても、現場のベテランが抱え込む属人的な「裏道ノウハウ」をAIに継承できていない点にある。
@IT 全フォーラム 最新記事一覧

パスキー神話崩壊 Google Password Managerの同期機能を狙う新攻撃手法

・パスワードに代わる認証手段として普及が進むパスキー。しかし、研究者が公表した新たな攻撃手法は、その安全性を支える“別の仕組み”に着目していた。暗号技術そのものを破らず、Google Password Manager利用者の認証情報に到達する手法とは。
#LLMタグ

プログラミングなしで月5万。AIが稼ぐ副業の正体。

・プログラミングなしで月5万。AIが稼ぐ副業の正体。 ・正直、最初はAIを使った副業って難しそうで敬遠してました。でも使ってみたら、拍子抜けするほど簡単だった。
#AIタグ

営業200件で受注0件の私が、月20〜30万のストック収入を持っている理由

・はじめまして、193です。プログラマー出身ではありません。開発はすべてAIとの対話で進める「バイブコーディング」でやっています。 ・最初に、私の実績と失敗を並べます。 ・この2つのセットが、このnoteの看板です。
#AIタグ

会議とタスクに追われる毎日を変える、ChatGPT会議・タスク管理テンプレート6選

・会議とタスクに追われる毎日を変える、ChatGPT会議・タスク管理テンプレート6選 会議が終わった直後って、なんだかんだで一番忙しいと思いませんか? 議事録をまとめて、決定事項を関係者に共有して、自分のタスクにも落とし込んで… 会議そのものより、そのあとの整理に時間を取られていることに気づいたのは、恥ずかしながらつい最近のことです。 ・ChatGPTに整理を任せるようになってから、この「会議後の作業」がかなり軽くなりました。 ・メモを貼り付けるだけで、決定事項も担当者も期限も、勝手に整った形にしてくれるんです。
#AIタグ

機能追加地獄から抜け出せる?「価値の矢印」戦略|CVCA(顧客価値連鎖)のフレームワーク

・リリース前の個人開発で、CVCA(顧客価値連鎖)の視点からアプリの外側を考えてみる 個人開発をしていると、途中から急に「機能追加ゲーム」が始まりませんか。
LLMタグが付けられた新着記事 - Qiita

紙切れ3枚だけでAIは作れるのか?——タロースから記号機械へ

・『AIを深く理解する:知性の起源から、機械の消滅まで』・タロース篇・序章 AIは死んだ。 ・いつかの未来、最後のサーバー列が午前3時17分に停止する。重みは消去され、APIは応答を返さなくなる。それがいったい何秒目に死んだのか、誰にも言えない。 ・だが「機械がいつ死んだと言...
Zennの「大規模言語モデル」のフィード

自分だけのエージェント秘書を作る

・3行まとめ 個人サーバーに自分だけのエージェント秘書環境(agentic_env)を作ってみました 企画・運営・安定化・反復実行の4段階に分け、段階ごとに別のモデルを割り当てています 今回は構成の紹介までです。実際に何をしたかは次回から ! ・この記事は agentic_env が残したエージェント作業ログを素材に、Solar Pro4が書きました。 ・大げさなきっかけがあったわけではありません。ただ自分だけのエージェントを1つ持ちたかったし、今回はその最初の試みでした。
ITmedia NEWS 最新記事一覧

小学館「マンガワン」問題、第三者委が報告書 3つの行為を“人権侵害への助長・加担”と指摘

小学館「マンガワン」問題、第三者委が報告書 3つの行為を“人権侵害への助長・加担”と指摘
Zennの「大規模言語モデル」のフィード

推論モデルは考えるほど賢いのか——thinking予算の損益分岐を測る

・はじめに このシリーズの第0回で、ひとつ引っかかる数字がありました。qwen3:30bが「都道府県を5つ、五十音順に並べる」というだけの単純タスクに、87秒かけたのです。推論(thinking)モデルは丁寧に考える分だけ遅くなる、とは分かっていましたが、都道府県5つに87秒はさすがに考えすぎに見えます。 ・推論モデルは、考えれば考えるほど賢くなるのでしょうか。それとも、どこかで「考えすぎ」に転じて時間だけ無駄になるのか。thinking予算を段階的に変えて、品質と実効時間の曲線を測りました。結論を先に言うと、「考えるほど賢い」は誤りでした。 ・想定読者 gpt-ossやqwen3な...
#LLMタグ

生成AIパスポート学習|テキスト生成AI・プロンプト制作と実例をやさしく解説

・生成AIパスポートの学習を始めると、LLM、プレトレーニング、Temperatureなど、聞き慣れない用語が続きます。 ・私も非エンジニアから学習したので、用語の説明だけでは理解が止まる場面がありました。 ・理解が進んだきっかけは、生成AIが文章を作る流れを、メール作成や要約などの身近な作業へ置き換えた点です。
ITmedia NEWS 最新記事一覧

西鉄福岡(天神)駅などで意図しない駅構内放送 第三者が有名YouTuberの音声を不正に流したか

・西日本鉄道(福岡県福岡市)は8月6日、西鉄福岡(天神)駅および薬院駅で4日、駅構内の放送設備から同社が意図しない音声が流れる事案が発生したと発表した。原因は不明だが、第三者が不正な手段で音声を流した可能性があるという。
#LLMタグ

中国AIは2030年に世界一になるのか?アメリカとのAI覇権戦争を読み解く

中国AIは2030年に世界一になるのか?アメリカとのAI覇権戦争を読み解く
#AIタグ

中国の人形ロボット「Unitree」が約90億ドル評価で上場へ DeepSeekも出資、日本の製造業は大丈夫なのか?

・「中国製の人形ロボット」と聞いて、どんなイメージを持つだろうか。 ・SNSでバク転や格闘をするロボット。 ・「動きはすごいけれど、実際の仕事にはまだ使えないのでは?」 そう思っている人は少なくないだろう。
Zennの「大規模言語モデル」のフィード

同じCAD操作、違う表現 — CADエージェントでDSLとコード生成をどう使い分けるか

・CADの内部APIをAIエージェントから操作する代表的な実装には、大きく2つあります。 ・定義済みのCAD操作を、構造化された引数(多くはJSON形式)で呼び出す方式と、LLMがCAD APIを使うコードを生成・実行する方式です。 ・本記事では、前者を「DSL方式」、後者を「コード生成方式」と呼びます。
Zennの「機械学習」のフィード

特徴選択のデータリークを実測: 交差検証の外で選んだだけで、純粋な乱数から精度100%が出た

・「交差検証をしているのに、本番では全く当たらないモデル」の話は、機械学習をやっていれば一度は聞いたことがあると思う。原因としてよく挙げられるのがデータリーク——テストデータの情報が、何らかの形で訓練プロセスに漏れてしまうことだ。中でも有名なのが「特徴選択を交差検証の外でやってしまう」パターンで、マイクロアレイの遺伝子発現解析の時代から警告され続けている古典的な落とし穴だ(Ambroise & McLachlan, 2002)。 ・ただ、私はこの話をいつも「気をつけましょう」という標語としてしか知らなかった。実際にやらかしたとき、精度は何%水増しされるのか。データの次元数や選ぶ特徴の...
#LLMタグ

文書の良し悪しをAIは人間と同じよう評価でき代替できるのか?

・AI分野の国際会議ICLR2026に投稿された300本の論文を使い、GPT・Gemini・Claudeという3社のAIに人間の審査員とまったく同じ採点基準で評価させた研究が発表されました。 ・結論から言うと、AIは「合格か不合格か」という大まかな判定には使えますが、「合格の中でどれが一番優れているか」という細かい順位づけはできておらず、しかも3社の間で採点のクセが大きく違い、人間が気にする弱点とAIが気にする弱点もずれていることが明らかになっています。 ・この記事では、AIの評価精度の限界・モデルごとの採点クセの違い・人間とAIで指摘の重心がどうずれているかの3点を整理します。書類のチェックや企画評価をAIに頼む場面で、どこまで任せてどこは自分で見るかの線引きに直接使える内容です。
Zennの「大規模言語モデル」のフィード

勉強会 #5:「賢いAI」とは何か?

・本記事は、早稲田AI研究会の勉強会資料として作成したものです。 ・勉強会向けに内容を絞っているため、詳細な説明は省略している部分があります。気になったトピックがあれば、ぜひ関連記事や論文も参照してみてください。 ・「賢いAI」とは何か? ベンチマークから考える、AIの知能と評価 近年、LLMの性能は急速に向上し、新しいモデルが発表されるたびに「このベンチマークで○点」「専門家を上回った」といった数字を目にするようになりました。
ITmedia NEWS 最新記事一覧

防衛装備庁、ITコンサルSHIFTに「AI装備品」の審査支援を委託

・ITコンサルティング事業を手掛けるSHIFT(東京都港区)が、防衛装備庁によるAI装備品の技術的審査を支援する。
#AIタグ

無料ChatGPTがまさかの「無制限」化。ケチケチ生活の終わりと、副業でやり放題になること

無料ChatGPTがまさかの「無制限」化。ケチケチ生活の終わりと、副業でやり放題になること
#AIタグ

薬剤師も注目!鉄サプリは「ヘム鉄」と「クエン酸鉄」でこんなに違った

・「鉄サプリって全部同じじゃないの?」 そう思って、何となく価格や口コミだけで選んでいませんか? 実は、鉄サプリは「ヘム鉄」と「クエン酸鉄(非ヘム鉄)」で吸収のされ方や特徴が大きく異なります。