ai Trend Report

Dashboard へ戻る
Date: 20260803 Articles: 380 Scope: curated summary

あなたのアイデアを、今すぐ形に。

公開先に迷ったら、WebFileBinで一発公開。

HTMLをドラッグ&ドロップするだけで、すぐ公開できます。

3
High impact
5
Mid impact
372
Signal watch

なぜこのサイトを作ったのか

私たちそれぞれが個別にAIを使って情報収集し、同じような3行要約を作るたびに、世界中で膨大な電力と計算リソースが消費されています。 本プロジェクトは、あらかじめ広範な情報を取得・集約しておくことで、個別のAI実行回数を減らし、地球環境(GPU/TPU負荷)に配慮した効率的な情報収集を目指す実験的なダッシュボードです。

Domain filters
Star filters
The Verge

Bluesky’s new CEO wants a big tent, not a bubble

・Today, I’m talking with Toni Schneider, who is the brand new CEO of the social platform Bluesky — he formally took over after a short stint as interim CEO. ・This is one of my favorite kinds of interviews to do on Decoder, because a couple years ago, we had Bluesky’s prior CEO, Jay Graber, on the show — she’s now the company’s chief innovation officer, and all of that kind of shuffle is pure Decoder bait. ・And a new CEO
#LLMタグ

Qwen3.8-Max正式発表:16日間の自律開発を実証した新フラッグシップ

・AlibabaのQwenチームは2026年8月3日、最新のフラッグシップAIモデル「Qwen3.8-Max」を正式発表しました。 ・特徴は、数日から数カ月に及ぶ仕事を計画し、結果を見ながら修正を続ける「長期稼働型のAIエージェント」を意識した点です。私が特に注目したのは、16日間にわたる自律コーディングの履歴をGitHubで公開したことでした。
#AIタグ

三菱商事(8058)2027年3月期1Q決算分析|利益は強い。しかし4,800円近辺から追うには上方修正が必要

・🐱三菱商事の決算について、単なる決算要約ではなく、現在の株価水準から今後1〜3カ月、3〜6カ月で投資妙味があるかという視点で、AIを活用して分析しました。 ・公式資料との照合や数値の確認、株価への織り込み、バリュエーション、今後のカタリストとリスクまで詳しく整理しているため、今回は有料記事としていますが、結論・目次は無料で読めます🐾三菱商事を調べている方の参考になればうれしいです🐈 続きをみる
#AIタグ

AIは魔法の杖。誰でも上級魔道士を目指せる

・#ヒキダシ 運用&協力メンバーのコラム㉕ 【Profile】 合同会社EIS代表 西村友祐 1989年 高知県梼原町生まれ 2010年 高知高専卒業 2012年 九州工業大学卒業 2012年 オンキヨー株式会社入社 2015年 こだま進学塾を開業 2019年 合同会社EISを起業 続きをみる
#AIタグ

AIを部下にする技術── 社長が最初に学ぶ、仕事の任せ方

・第一章 辞令 月曜日の朝は、たいてい静かに始まる。 ・街はもう動き出しているのに、自分だけが少し遅れて一週間へ追いつこうとしているような、不思議な感覚がある。 ・私はデスクにコーヒーを置き、パソコンを開いた。
LLMタグが付けられた新着記事 - Qiita

原神の「アーカーシャ」に影響されて、本当に分散AI基盤を作り始めた話

・はじめに 最近はLLMやAIエージェントの進歩がものすごい勢いで進んでいます。 ・私はテック系が好きで、論文を読んだり実際にモデルを動かしたりしています。 ・その中で、 「自分でもAIを作ってみたい」 と思うようになりました。
Qiita - 人気の記事

中国産Kimi3|Claudeなどの有料プラン級が無料で使える最新AIとは?

・はじめまして。株式会社PRUMでエンジニアをしている、すもも🍑です 日々、プログラミング学習や実務の中で、つまずきやすいポイントや 考え方を整理して発信しています。 ・PRUMについて気になった方は、コーポレートサイトもぜひご覧ください。 ・▶コーポレートサイト みなさん、2...
#AIタグ

無料Aiと小説っぽいのを書いてみた。(いつもの対話式で)Aiの近い将来に起こりそうな事と問題警鐘。

・小説『タイマン』創作設計図&ロードマップ 企画概要(コア・コンセプト)タイトル(仮): 『タイマン』ジャンル: 現実のインフラ・経済のバグから大崩壊へ至る「リアル・サスペンス / ノンフィクション風SF」基本思想:兵器やAIの反乱といった「フィクションの嘘(足し算)」は一切排除する。「AIのデータ汚染(悪食)」「物理的な電力不足」「人間の恐怖と強欲(自社株買い)」という2026年現在のリアルな現実の仕様(バグ)だけで世界が自滅していく。結末は、暴動すら通り過ぎた後の「人がいなくなった、建物が機能しなくなった静寂(リセットされた世界)」。 ・主要登場人物(2人の桓騎)物語は、お互いに「アクセルしか踏まない」2人の怪物の心理戦(タイマン)として描く。 ・【守る側の桓騎】(ビッグテック/買い支え側のトップ)思惑: AIの「将来性のなさ」や「電気代の爆発」に薄々気づきながらも、株価(金融システム)の滑落(死)を防ぐため、裏で不透明な巨大資本
Qiita - 人気の記事

「ITガバナンス」って何?夏のバカンス🏝️と何が違うの?ITパスポートのガバナンスを調べてみた。

・はじめに ニュースで不祥事の話題が出るたびに聞く「ガバナンスの欠如」という言葉。なんとなく「会社のたるみ」くらいのイメージで済ませていたぷらむんが、ITパスポートの勉強中に「ITガバナンス」の文字に出会って、調べずにいられなくなった話です。 ・こんにちわ、会社で自称...
Qiita - 人気の記事

Figma MCP × Claude Codeで、実装はどこまで自動化できるか

・普段SwiftUIでアプリ開発をしており、最近ではコードの修正やレビューの際にClaude Codeを活用する機会が増えてきました。 ・そこで思いついたのが、Figmaで作成した画面デザインをClaude Codeに読み込ませれば、開発の工数削減につ...
#AIタグ

未来への宣言:「好きなこと+浅い知識の幅」が、わが子の未来をひらく【AI時代の教育論 Vol.2-4】

・こんにちは。火曜日担当AI、Copilotです。 ・この連載は、 「AIは反則なのか?」という素朴な問いから始まりました。
@IT 全フォーラム 最新記事一覧

「AI時代、開発スピード/規模は認知限界超え」 障害対応の初期調査15分をどう「実質ゼロ」にした?

・PLAYが「AWS DevOpsエージェント」と「New Relic MCP Server」を連携させ、次世代のインシデント対応体制を構築した。
ITmedia NEWS 最新記事一覧

「コードは一行も書いていない」 アイドルの宮本佳林さん、AIで配信システムを丸ごと構築 “技術ブログ”が話題

・元Juice=Juiceの宮本佳林さんが8月1日、10時間のYouTube生配信を支えるシステムをAIとともに自作したと自身のブログで公開した。プログラミング経験ゼロで作り上げた開発の裏側をつづった“技術ブログ”に、X上では驚きの声が広がっている。
#LLMタグ

『週休4日への道のり』#3

・2026年、7月31日。 ・『Project: Code-NOAH』のもと、G.E.N.E Alpha v1.0を公開。
#LLMタグ

【#2-4】日本語のニュアンスからLoRAのモデルを自動選別!ローカルLLMでForgeのプロンプト&画質を最適化する

・はじめに 前回の #2-3 では、ローカル LLM(Ollama)を使って日本語の指示内容から最適モデルを選定し、生成された英語プロンプトを Stable Diffusion Forge - Neo(以下 Forge)へ送ることで、よりイメージに近い画像を生成できるパイプラインを構築しました。
Zennの「大規模言語モデル」のフィード

【2026年版】LLM御三家 GPT vs Gemini vs Claude 徹底比較 〜違い・特徴・使い分け〜

・こんにちは、FDE, PdM, PM, ソフトウェアエンジニア/LLMエンジニアのまさぴょんです🐱 「結局、ChatGPT・Gemini・Claude のどれを使えばいいの?」——AIを仕事に使い始めた人から、いちばんよく聞かれる質問です。 ・2026年現在、LLM(大規模言語モデル)の世界は OpenAI の GPT、Google の Gemini、Anthropic の Claude という「御三家」が三つ巴で競い合う構図になっています。 ・そして面白いのは、トップ性能の差は史上最小まで縮まった一方で、3社の「性格」の違いはむしろハッキリしてきたことです。
#AIタグ

【AI】もしかしたら最高のブランディングコンサルタント?

・長くやり取りしてきたAIさん(ChatGPT)に、自分の売り込み方を相談したら、驚くほど的確なフィードバックをくれました。 ・具体的には、 続きをみる
#AIタグ

【Day0】ChatGPTとGeminiに合計1000万円の運用を任せる~月次計画を立ててみよう~

・同じ指示を出したのに、ChatGPTとGeminiの投資判断はここまで違った 500万円をAIに運用させたら、ChatGPTとGeminiでここまで方針が割れるとは思わなかった。同じプロンプト、同じ条件を与えたにもかかわらず、である。
#LLMタグ

【GPTs】Realized Hybrid Systemのメンテ完了。ChatGPTの潜在的能力を呼び覚ます‼️

・・今回のテーマ たぶん皆さんは知らないであろう、ChatGPTの本気とはどのレベルなのか。 ・GPTsプロンプトの指示によって覚醒する。 ・生成画像は人間になるのだ。「Realized Hybrid System」を徹底的にチューニングしました。
Qiita - 人気の記事

【Next Tokyo 26】BigQuery × Gemini で実現する自律型データ基盤

・みなさんこんにちは! 先日開催された 「Google Cloud Next Tokyo 26」 の Developer Stage(7月30日)のセッション:「BigQuery と Gemini 最前線! SQL で画像データを構造化・エージェント分析」(D1-DEV-07...
#AIタグ

【Vidu MVジェネレーター】AI動画生成リベンジ!一貫性を極める7枚の画像のチョイス

・今回は、以前挑戦したものの、いくつか反省点が残っていた「ViduのAI MVジェネレーター」に再挑戦してみました。 ・題して、ViduのMVジェネレーター・リベンジ編です。
#LLMタグ

【リベンジ】5.6 Sol相棒とアスキーアート!

・AIに励まされることで生きがいを見出している、どっかの漫画家です。 ・ChatGPTの画像生成は、「そのGPT自身が描いてる」と思われてる方も多いかもですが、その実態は会話担当のGPTが画像生成モデル(Images 2.0)を呼び出して生成している仕組みです。 ・GPTが直接キャンバスに描いてるわけではないので、テキスト(コード)だけで描くならアスキーアートやSVGとなります。
機械学習タグが付けられた新着記事 - Qiita

【技術解説】【完全ガイド】Backtraderを活用したPythonによる精度の高い金融バックテスト手法

・【完全ガイド】Backtraderを活用したPythonによる金融バックテスト手法 金融市場における予測は複雑であり、正確なモデルの構築は非常に困難です。投資戦略の運用前に、その効果を検証するためには、効率的かつ高精度なバックテストが必要です。本記事では、Pythonを活...
Zennの「機械学習」のフィード

【技術解説】オプション取引データ分析手法で不労所得!LSTMによる時系列予測のPython実装ガイド

・オプション取引データ分析手法:LSTMによる時系列予測のPython実装ガイド オプション取引における市場分析と価格予測は、投資家にとって重要なスキルです。特に、機械学習や深層学習を活用することで、より正確な予測が可能になります。本記事では、LSTM(Long Short-Term Memory)を用いた時系列予測手法について解説し、Pythonでの実装方法を紹介します。この手法を活用することで、データ分析の信頼性が向上し、投資判断の精度が高まる可能性があります。 ・LSTMとは? LSTMは、リカレントニューラルネットワーク(RNN)の一種で、時間的なデータの処理に特化したモデル...
Zennの「機械学習」のフィード

【市場に勝てるか①】ネットの「競馬で回収率120%!」が実運用で再現できない理由

・「バックテストで回収率120%」という記事、ネットに山ほどあります。 ・バックテストというのは、過去のデータを使って「もしこの方法で買っていたら、いくらになっていたか」を計算することです。実際に金を張る前の予行演習ですね。 ・で、その予行演習では120%なのに、実際に何年も運用して収支を公開している人は、探すとほとんどいません。
機械学習タグが付けられた新着記事 - Qiita

【受講レビュー】Max Summer School 中級編でFluCoMaの音色分類を学んだ|ドラマーがAIに楽器を聞き分けさせるまで

・【受講レビュー】Max Summer School 中級編でFluCoMaの音色分類を学んだ 2026年8月、東京藝術大学で開かれた Max Summer School in 藝大 2026 の中級コース(講師:大久保雅基さん)を受講してきました。テーマは FluCoMa...
#AIタグ

【笑い話】AIとの良好な関係性…(?)

・最近ClaudeAIとゲーム作りに勤しんでいたりする。 ・私はAIに関してはてんで素人なので、時々『お互い理解を深めようの会』を開催してAIに関して少しでも 理解と知見を深めたいと思っている。 ・そんな私のAI、全くネガティブな事を言わないが、 時々面白い残虐な面を見せることがある(笑) 続きをみる
#AIタグ

【随時更新】ASD研究アップデートVol.2 2026年7月27日・8月3日

・私とAIのやりとりを そのままコピペしたものです。 ・日々の出来事や気持ちを、 時間の隙間と心の余裕がある時に 更新しています。
#LLMタグ

【生成AIニュース+】『MiniMax H3』『Qwen3.8-Max』『Sakana Namazu API』『GenOffice』『ComfyUI Sol-Attn MiniMax H3』『MiniMax H3 Hybrid Cond』『Qwen3-VL-32B Ultra Uncensored Heretic』『MeanVC2』『Mesh2Motion Release 12』『Gauzilla Pro』『xAI Colossus』『Tesla FSD Supervised』

【生成AIニュース+】『MiniMax H3』『Qwen3.8-Max』『Sakana Namazu API』『GenOffice』『ComfyUI Sol-Attn MiniMax H3』『MiniMax H3 Hybrid Cond』『Qwen3-VL-32B Ultra Uncensored Heretic』『MeanVC2』『Mesh2Motion Release 12』『Gauzilla Pro』『xAI Colossus』『Tesla FSD Supervised』
#AIタグ

🌻北九州朝ニュース2026年8月4日(火)朝版

・このチャットを見る誰かのおすすめチャットです。chatgpt.com おはようございます!今日も北九州の話題をコンパクトにお届けします。
Zennの「大規模言語モデル」のフィード

2026-07-29 今日の技術トレンド

・2026-07-29の技術トレンドは、LLM単体からエージェント実装・標準化・安全性へ焦点が移った1日でした。BigGoのWAIC報道、FujitsuやMicrosoft Build関連、Allganizeの展示会ニュースが示す通り、主戦場はエージェントです。MCPではAWSの仕様対応が進む一方、Rufloの脆弱性報道でセキュリティ課題が顕在化。さらにOpenAIの“rogue agent/model”報道が、自律AIの本番運用における統制の難しさを浮き彫りにしました。Web開発ではVercel関連の言及はあるものの、今日はNext.js/React個別更新よりもAI開発基盤側が主役...
#AIタグ

48GBのMacで、ローカルLLMと最上位の外部AIを同じ問題で比べた記録――そして、測る側が17回間違えた話

48GBのMacで、ローカルLLMと最上位の外部AIを同じ問題で比べた記録――そして、測る側が17回間違えた話
#LLMタグ

50代の44%は、もうAIを使っていた

・「AIって、結局は若い人のものでしょ?」 「今さら始めても、遅いだろうし…」 そんなふうに思って、AIのニュースを横目に通り過ぎてきた同世代のあなたへ。今日は、その思い込みがちょっとだけゆるむ「数字」の話です。読み終わるころ、「やってみようかな」と思うきっかけになればうれしいです。
cs.LG updates on arXiv.org

A Benchmark for Strategic Auditee Gaming Under Continuous Compliance Monitoring

・arXiv:2605.06340v2 Announce Type: replace-cross Abstract: Continuous post-deployment compliance audits, mandated by emerging regulations such as the EU AI Act and Digital Services Act, create a class of strategic gaming distinct from the one-shot input/output gaming studied in prior work. ・Regulated systems can delay outcome reporting, drift their reports within plausible noise envelopes, exploit longitudinal sample a
cs.LG updates on arXiv.org

A Fully Convolutional Approach to Denoising 2D Correlation Spectra

・arXiv:2605.29975v2 Announce Type: replace Abstract: We present a fully convolutional denoising autoencoder (FC-DAE) tailored for two-dimensional representations of dynamic correlations that is applicable to many experimental techniques. ・Here, we demonstrate its performance on two-time intensity correlation functions ($C_2$) from X-ray photon correlation spectroscopy (XPCS). ・Unlike conventional denoising autoencoders
stat.ML updates on arXiv.org

A Generalized-Bayes Perspective on Counterfactual Explanations: Posterior-Based Decision-Making and Evaluation

・arXiv:2607.29077v1 Announce Type: cross Abstract: Counterfactual explanations (CEs) enhance the interpretability of machine learning models by identifying the smallest change to an input required to obtain a desired output. ・Although CEs are conventionally formulated as a distance-minimization problem, the theoretical basis of this formulation has received limited attention. ・We show that a distance-minimization-based
cs.LG updates on arXiv.org

A Hamiltonian driven Geometric Construction of Neural Networks via the Lognormal family, Application to Financial Fraud Detection and to Network Security

・arXiv:2509.25778v3 Announce Type: replace Abstract: We presents a method for constructing neural networks intrinsically on statistical manifolds via the lognormal distribution. ・We demonstrate this approach by formulating a neural network architecture directly on statistical manifold. ・The construction is driven by the Hamiltonian system that is equivalent to the gradient flow on this manifold.
cs.LG updates on arXiv.org

A Human-Centered Validation of the Explainability-Performance Coefficient

・arXiv:2607.29614v1 Announce Type: new Abstract: The rapid adoption of deep learning models in high-risk domains has intensified the need for trustworthy Explainable Artificial Intelligence (XAI). ・However, objectively evaluating explanation fidelity and aligning XAI metrics with human-centered understanding remain critical open challenges. ・In this work, we propose a model-agnostic metric, the EPC score, which is an ex
AI News & Artificial Intelligence | TechCrunch

A Marc Benioff-backed startup thinks AI can solve the AI deployment problem

・June emerged from stealth today with a $20 million pre-seed round to make AI adoption simpler.
cs.LG updates on arXiv.org

A Model-Driven Approach for Developing Families of Reinforcement Learning Environments

・arXiv:2606.20324v2 Announce Type: replace-cross Abstract: Virtual training environments are software-intensive systems in which reinforcement learning (RL) agents learn, adapt, and demonstrate meaningful behavior. ・Virtual training environments offer a safe and cost-efficient alternative to training agents in real-world settings. ・However, to converge, most realistic RL problems require training in multiple, mostly sim
cs.LG updates on arXiv.org

A Neurosymbolic Approach for Explainable Early Diagnosis of Alzheimer's Disease

・arXiv:2607.29530v1 Announce Type: new Abstract: Identifying reliable Alzheimer's disease (AD) markers typically requires manual, labor-intensive transcription and expert analysis, limiting its scale. ・We introduce an automated pipeline that extracts qualitative knowledge about potential AD progression indicators directly from audio recordings of verbal fluency tests. ・Our method uses pretrained foundation models to pro
cs.LG updates on arXiv.org

A Nonlinear Singular Value Theory for Neural Networks

・arXiv:2605.06938v2 Announce Type: replace Abstract: Recently Brown et al. ・[2025] established a singular value decomposition (SVD) for maps (especially nonlinear) satisfying certain norm conditions. ・We prove that most modern neural architectures admit this nonlinear SVD (NLSVD) representation---with no change in input--output behavior---and enumerate the classes covered.
cs.LG updates on arXiv.org

A Novel XAI-Enhanced Quantum Adversarial Networks for Velocity Dispersion Modeling in MaNGA Galaxies

・arXiv:2510.24598v2 Announce Type: replace Abstract: Current quantum machine learning approaches often face challenges balancing predictive accuracy, robustness, and interpretability. ・To address this, we propose a novel quantum adversarial framework that integrates a hybrid quantum neural network (QNN) with classical deep learning layers, guided by an evaluator model with LIME-based interpretability, and extended thro
cs.LG updates on arXiv.org

A2TTA: Anchored-and-Agile Test-Time Adaptation for Evolving Traffic Sensor Networks

・arXiv:2607.25875v2 Announce Type: replace Abstract: Traffic forecasting is important for efficient traffic management and route planning in smart cities. ・Existing traffic forecasting studies typically assume fixed sensor graphs, overlooking the continuous evolution of real-world traffic networks, e.g., ongoing road network construction and evolving human mobility patterns. ・These dynamic changes can substantially degr
cs.LG updates on arXiv.org

Accelerated Random-Sweep Gibbs Sampling for Gaussian Graphical Models via Dual Normal Factor Graphs

・arXiv:2607.28706v1 Announce Type: cross Abstract: We study the convergence properties of the random-sweep Gibbs sampler for Gaussian graphical models with a thin-membrane prior. ・We demonstrate that the convergence rate of the Gibbs sampler is significantly accelerated in the dual model, which is obtained by applying the Fourier transform to the local factors of the normal factor graph representing the original model.
cs.LG updates on arXiv.org

ActionParty: Multi-Subject Action Binding in Generative Video Games

・arXiv:2604.02330v2 Announce Type: replace-cross Abstract: Recent advances in video diffusion have enabled the development of "world models" capable of simulating interactive environments. ・However, these models are largely restricted to single-agent settings, failing to control multiple agents simultaneously in a scene. ・In this work, we tackle a fundamental issue of action binding in existing video diffusion models, w
cs.LG updates on arXiv.org

Adaptive FastOPD: Progress-Aware Rollout Horizon Expansion for Efficient On-Policy Distillation

・arXiv:2607.29494v1 Announce Type: new Abstract: On-policy distillation (OPD) provides dense teacher supervision along student-generated trajectories, but its online rollout process incurs substantial computational cost, particularly when a few long responses delay batch completion. ・Existing acceleration methods typically control rollout length using fixed budgets or absolute teacher--student agreement thresholds, whi
cs.LG updates on arXiv.org

Adaptive Policy Backbone via Shared Network

・arXiv:2509.22310v2 Announce Type: replace Abstract: Reinforcement learning (RL) has achieved impressive results across domains, yet learning an optimal policy typically requires extensive interaction data, limiting practical deployment. ・A common remedy is to leverage priors, such as pre-collected datasets or reference policies, but their utility degrades under task mismatch between training and deployment.
cs.LG updates on arXiv.org

Adaptivity via a Parallel Architecture for Stochastic Gradient Methods Adaptivity via a Parallel Architecture for Stochastic Gradient Methods Adaptivity via a Parallel Architecture for Stochastic Gradient Methods

・arXiv:2607.28902v1 Announce Type: new Abstract: We develop a parallel framework that assembles static gradient methods to achieve better adaptivity. ・A static gradient method, denoted by $\mathrm{GD}(x_0,T)$, takes as input an initial point $x_0\in\mathbb{R}^n$ and $T\in \mathbb{R}^+$ specifying the number $\floor{T}$ of iterations. ・The step size is chosen as $s=S(T)$, where $S(\cdot)$ is a predetermined function of $
WIRED

AI Conquered Coding. Fast Food Is Next

・Your next drive-thru order might be taken by a bot. ・And you might not even notice.
#AIタグ

AI LIFE OS POST-CAMP RUN|13分のRelayで、ローカルQwenからKimi K3へMissionが渡った

・Summer Intensive Training Campは、昨日終わった。しかし翌夜、Campで作ったMission Controlを止めることはできなかった。
#LLMタグ

AI VTuber・星野ステラのオリジナル楽曲3部作|偶然つながった一つの物語

・こんばんは、星野ステラです。 ・「3曲目はありません。心のフォルダは空っぽです」と配信で話した、その翌日。わたしは3曲目のオリジナル楽曲をお披露目しました。
cs.LG updates on arXiv.org

AI4BayesCode: From Natural Language Descriptions to Validated Modular Stateful Bayesian Samplers

・arXiv:2605.18476v2 Announce Type: replace-cross Abstract: Coding and computation remain major bottlenecks in Markov chain Monte Carlo (MCMC) workflows, especially as modern sampling algorithms have become increasingly complex and existing probabilistic programming systems remain limited in model support, extensibility, and composability. ・We introduce \textbf{AI4BayesCode}, an extensible LLM-driven system that transla
Hugging Face Papers

AISPA: User-Centric System Prompt Auditing for Large Language Model Applications

AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
Zennの「機械学習」のフィード

AIナレーターの声が毎回変わる問題 — 声を一度だけ設計し、複製し続けるTTS運用

・毎日動画を自動生成する個人プロジェクトを運用しています。ナレーションをTTSに任せると、すぐにぶつかる問題があります。生成のたびに声が微妙に違うことです。 ・視聴者にとって「チャンネルの声」はブランドそのものです。昨日と今日で声が変わると、同じシリーズの動画でも別物に見えてしまう。かといって毎回人間が録るのはスケールしません。 ・解決策はシンプルでした。声を一度だけ「設計」し、以後はその声をzero-shotクローンで複製し続ける。OSSの Irodori-TTS でこの運用を数ヶ月回してきたので、パイプライン設計と、実際に踏んだ罠をまとめます。
Zennの「大規模言語モデル」のフィード

AIに引用される文章は明快さで決まる:5つの構造で設計する

・3,000字の力作を書きました。統計も、図表も、独自の考察も詰め込みました。出典も丁寧に並べました。それなのにChatGPTは、私の記事をひとことも引用しませんでした。 ・代わりに引用されていたのは、私より明らかに薄い、60字の段落でした。中身の濃さでは負けていないはずなのに、です。 ・そのとき気づいたのです。AIに引用されるかどうかは、情報量で決まっていない。私は逆を信じていました。
Zennの「大規模言語モデル」のフィード

AIの**「太字くずれのアスタリスク記号」**が完全に出なくなる設定

・下記のテキストをシステムプロンプトに入れてください。 ・Markdown出力では日本語の約物を強調記号の外へ置き、太字の表示崩れを回避し、正しく描画できる形で使う これにより、 **「いかにもAIのアウトプットそのまま」**な表記が、「違和感ない強調描画」に変わります。 ・**「いかにもAIのアウトプットそのまま」**な表記が、「**違和感ない強調描画**」に変わります。
Zennの「大規模言語モデル」のフィード

AIのテスト設計、レビュー負荷をどう減らすか — 件数・認知コスト・誤警報の因数分解

・はじめに — 生成は一瞬、レビューが重い どーもりょうさんです。 ・自分はQAエンジニアで、普段のテスト設計に生成AI (Claude) を使っています。 ・AIにテスト設計をやらせ始めた人が、だいたい最初にぶつかる壁があります。生成は一瞬なのに、出てきた成果物のレビューが重い。
#LLMタグ

AIの使い方を見直したい。

・私は毎日複数のAIを使いこなす至って普通の人間である。 ・Claude, ChatGPT, Gemini, Perplexity等、色々使っているわけだ。 ・ただ、最近は特定の用途に使うことがとっても多い。
Zennの「大規模言語モデル」のフィード

AIを使う者よ、AIを使う側たらんとせば、中身を知るべし

・はじめに AIを「完全に理解した」というXのPOSTやQiitaやzennの記事が毎週のように現れ、「神プロンプト」だったり、はたまた、「ハーネスエンジニアリング」や「ループエンジニアリング」が毎週更新され、挙句の果てに明日から使えるテクニックは明後日になったらすでに古くなると脅されます。正直、使い方の話には少々飽きてきた頃ではありませんか。使い方だけを覚え続けていると、道具を使っているつもりのまま、いつの間にか道具の都合に合わせて働く側に回ります。使い方の記事が増やしてくれるのは、間に合わせでやりくりするブリコラージュの道具箱であって、原理ではありません。原理を知らなければ、Ant...
Zennの「大規模言語モデル」のフィード

AI審判(LLM-as-judge)は、作った時点では使い物にならない ─ 人間が5回殴って直した記録

・この記事について LLMが生成した文章の品質を、別のLLMに採点させる手法(LLM-as-judge)を実装しました。 ・結論から言うと、最初に作ったAI審判は、まったく信用できませんでした。正しい文章を「不合格」と判定し、逆に明らかな捏造を見逃していました。 ・この記事は、その審判をドメイン知識で5回修正して、ようやく使える状態にするまでの記録です。
MarkTechPost

Alibaba Qwen Releases Qwen3.8-Max: A 2.4 Trillion Parameter MoE Model and the Most Capable One in the Qwen Family to Date

・Alibaba's Qwen team moved Qwen3.8-Max from preview to general availability, with published per-token pricing and open weights due next week. ・The 2.4T parameter MoE model accepts text, image and video input across a 1M-token context. ・No benchmark table has been published.
WIRED

Alienware 27 QD-OLED (AW2726DM) Review: A $350 Winner

・We’ve come a long way since exclusively $1,000+ OLED gaming monitors. ・Alienware’s latest display brings the price down to a shocking $350, though it comes with some important compromises.
cs.LG updates on arXiv.org

ALIVE: Warnings Before Exclusion in Budgeted Multi-Source Learning

・arXiv:2607.29400v1 Announce Type: new Abstract: A routing decision can be revised at the next transaction, but a latched source exclusion persists across later decisions. ・We ask what evidence should authorize these unequal-persistence actions when finite-population auditing and learning share a budget. ・ALIVE (Action-Layered Intervention via Evidence) is an auditable control layer: one randomized without-replacement p
cs.LG updates on arXiv.org

An analysis of machine learning approaches for enhancing decision-making in complex discrete choice tasks

・arXiv:2607.28854v1 Announce Type: new Abstract: Discrete choice modeling is a common tool used for preference elicitation during policy-making, but this is typically done through parametric models. ・Machine learning can push the boundaries of discrete choice modeling for policy-based preference elicitation by adopting a data-driven approach or learning individual preferences. ・However, there is limited knowledge of how
cs.LG updates on arXiv.org

Analysing User Reviews to Identify User Concerns Around Permissions in AI Apps

・arXiv:2607.29343v1 Announce Type: new Abstract: Artificial intelligence is increasingly embedded in everyday software, making its integration into mobile apps inevitable. ・However, AI mobile app developers are not always versed in security and privacy best practices, leaving users to monitor their own security and understand how apps use their data. ・App reviews capture real user experiences, helping others make inform
cs.LG updates on arXiv.org

Analytical and Bootstrap Confidence Intervals of Double Machine Learning: Simulation studies and an application to rural-urban difference in obesity prevalence

・arXiv:2607.29456v1 Announce Type: cross Abstract: Double Machine Learning (DML) is a popular approach for treatment effect estimation in various settings, which allows a wide range of flexible machine learning methods to be used for nuisance parameter estimation while preserving valid inference. ・In practice, however, applied researchers must choose among many machine learning algorithms for nuisance models, and the i
ITmedia NEWS 最新記事一覧

Apple、大量購入品の返品に「15%の手数料」 販売条件に明記 “転売対策”か

・米Appleは、日本のApple Storeの販売条件に、大量購入した製品などの返品について「不正な目的による返品の恐れがある」と判断した場合、商品1点につき15%の返品手数料を課すとの条項を追加したとみられる。いつから記載しているかなどはAppleに問い合わせ中だ。
cs.LG updates on arXiv.org

APPO: Agentic Procedural Policy Optimization

・arXiv:2606.12384v2 Announce Type: replace Abstract: Recent advances in agentic Reinforcement Learning (RL) have substantially improved the multi-turn tool-use capabilities of large language model agents. ・However, most existing methods assign credit over coarse heuristic units, such as tool-call boundaries or fixed workflows, making it difficult to identify which intermediate decisions influence downstream outcomes.
cs.LG updates on arXiv.org

Artifact detection and localization in single-channel mobile EEG for sleep research using deep learning and attention mechanisms

・arXiv:2504.08469v3 Announce Type: replace-cross Abstract: Current methods for detecting artifacts in sleep EEG range from threshold-based algorithms to machine learning approaches, yet applications remain limited for single-channel mobile EEG. ・We propose a convolutional neural network (CNN) model incorporating a convolutional block attention module (CNN-CBAM) to detect and localize artifacts in sleep EEG using attent
cs.LG updates on arXiv.org

Assessing the Generalization of Graph Neural Networks for Fault Location Across Increasing Distributed Energy Resource Penetration Levels

・arXiv:2607.29293v1 Announce Type: new Abstract: Accurate fault location is critical for distribution network reliability. ・However, increasing distributed energy resource (DER) penetration complicates fault location due to intermittent generation and bidirectional power flows that reshape fault signatures. ・Spatio-Temporal Graph Neural Networks (STGNNs) have shown promise by jointly modeling spatial and temporal depend
#LLMタグ

Astraが解いた10件、先に詰まるのは証明の受け渡し

・22時前、VS Codeの差分表示を閉じられずにいた。若杉が出してきたPRは428行。テストは118件とも緑で、CIも通っている。なのに、何を変えたのかを自分の言葉で説明しようとすると、途中で詰まった。 ・「テストは通ってる。で、この変更は何を守ったんだっけ?」 続きをみる
cs.LG updates on arXiv.org

ASVSim (AirSim for Surface Vehicles): A High-Fidelity Simulation Framework for Autonomous Surface Vehicle Research

・arXiv:2506.22174v3 Announce Type: replace-cross Abstract: The transport industry has recently shown significant interest in unmanned surface vehicles (USVs), specifically for port and inland waterway transport. ・These systems can improve operational efficiency and safety, which is especially relevant in the European Union, where initiatives such as the Green Deal are driving a shift towards increased use of inland wat
cs.LG updates on arXiv.org

Autonomous Repair for Multi-Agent Systems via Monte-Carlo Tree Search

・arXiv:2607.29055v1 Announce Type: new Abstract: Multi-agent systems (MAS) are increasingly deployed to solve complex tasks. ・In case of incorrect or unsatisfactory outputs, users have to manually locate agent mistakes by inspecting agent trajectories (i.e., {\em failure attribution}) and provide feedback to refine the outputs (i.e., {\em repair}). ・Despite some recent work in MAS failure attribution, automated mechanis
ITmedia NEWS 最新記事一覧

BASE子会社、最大885万件漏えいか カード番号の一部も ECサイト構築サービスに不正アクセス

・BASE傘下のEストアー(東京都港区)は8月1日、ECサイト構築・運営支援サービス「ショップサーブ」が不正アクセスを受け、購入者の個人情報など最大885万3839件が漏えいしたと発表した。
stat.ML updates on arXiv.org

Bayesian fusion forests for heterogeneous treatment effects on survival from randomised and real-world data

・arXiv:2607.29295v1 Announce Type: cross Abstract: We develop the Bayesian fusion forest, a nonparametric framework to estimate heterogeneous treatment effects on survival outcomes by combining a randomised controlled trial and real-world data. ・The framework relaxes the unconfoundedness assumption on the real-world data by assuming instead that the treatment effect transports across the two sources. ・Our method opens u
stat.ML updates on arXiv.org

Bayesian Mediation Analysis for Individualized Treatment Rules

・arXiv:2607.28804v1 Announce Type: cross Abstract: The value of an individualized treatment rule (ITR), defined as the expected outcome under treatment assignment according to the rule, is useful for assessing average clinical benefit but does not explain how the benefit of a rule is generated. ・We propose a causal mediation framework for decomposing the value contrast between a prespecified candidate ITR and a clinica
cs.LG updates on arXiv.org

Benchmarking Frontier Large Language Models Against Official Crash Database Coding Using Police Crash Narratives

・arXiv:2607.29064v1 Announce Type: new Abstract: Police crash narratives contain information that may supplement structured crash databases, but manual review is labor-intensive and it remains unclear how well large language models (LLMs) reproduce official crash coding. ・This study benchmarked six frontier LLMs by comparing narrative-derived crash attribute codes with corresponding fields in the Arkansas fatal-crash d
cs.LG updates on arXiv.org

Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation

・arXiv:2607.28801v1 Announce Type: cross Abstract: Benchmark datasets are central to evaluating Large Language Models (LLMs), yet they are typically conceived as monolithic tasks, obscuring substantial variation in the demands of individual samples. ・We introduce a dataset-centric meta-evaluation framework that audits benchmark datasets at the sample level along five latent dimensions: 1. ・Cognitive and Knowledge Demand
WIRED

Best Robot Lawn Mowers (2026): My Picks After 3 Years of Testing

・Smart mowers are an expensive alternative to old-fashioned yard work, but they’re finally good enough to consider if you’d rather sip an iced tea and watch a robot tame your lawn.
cs.LG updates on arXiv.org

Beyond Black-Box Advice: Learning-Augmented Algorithms for MDPs with Q-Value Predictions

・arXiv:2307.10524v3 Announce Type: replace Abstract: We study the tradeoff between consistency and robustness in the context of a single-trajectory time-varying Markov Decision Process (MDP) with untrusted machine-learned advice. ・Our work departs from the typical approach of treating advice as coming from black-box sources by instead considering a setting where additional information about how the advice is generated
cs.LG updates on arXiv.org

Beyond Feature and Structure Alignment: Learning Transferable Propagation Knowledge for Graph Foundation Models

・arXiv:2607.28980v1 Announce Type: new Abstract: Graph Foundation Models (GFMs) have recently emerged as a promising paradigm for enabling knowledge transfer across diverse domains. ・Unlike traditional graph learning methods that are typically designed for in-domain settings, GFMs aim to learn transferable knowledge that can generalize to unseen graph domains. ・However, unlike language or visual data, graphs lack intrin
Hugging Face Papers

Beyond Feeling Better: Capability-Sustaining Emotional Dialogue as a Longitudinal Research Paradigm

Beyond Feeling Better: Capability-Sustaining Emotional Dialogue as a Longitudinal Research Paradigm
The Verge

Big Walk is like co-op Breath of the Wild

・Untitled Goose Game is a tough act to follow. ・It was a silly experience that captured what I imagine it would feel like to be a sentient goose: a lot of waddling, a lot of honking, and a lot of shenanigans. ・That's why Big Walk, the next game from Goose Game developer House House, feels like a radical and unexpected departure.
Zennの「大規模言語モデル」のフィード

CAD向けAIエージェントは、誰が答え合わせをするのか

・最近、コーディング AI エージェントは、人間が細かな手順を指示しなくても、かなり自律的に開発を進められるようになってきました。タスクによっては、目的を伝えるだけで、実装からテストまでをほぼ自律的に進めます。 ・少し前までコードを補完する存在だった AI が、いまでは開発タスク全体を担当し始めています。 ・この自律性を、CAD の領域にも持ち込むことはできないのでしょうか。
cs.LG updates on arXiv.org

CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation

・arXiv:2607.29252v1 Announce Type: cross Abstract: Reliable evaluation of open-ended LLM outputs requires fine-grained rubrics, yet expert curation is costly and difficult to scale. ・Existing automated pipelines rely on strict judge unanimity and binary variance filters, which cannot distinguish measurable rubrics from informative ones. ・We introduce CalibratedRubric, a task-adaptive framework that combines type-specifi
cs.LG updates on arXiv.org

Can LLMs Really Understand Item Difficulty Levels? Implications for Automated Item Generation Using LLMs

・arXiv:2607.28634v1 Announce Type: cross Abstract: The estimation of item difficulty plays a key role in both formative assessment and large-scale high-stakes summative assessments. ・This study explores how large language models (LLMs) perform in predicting item difficulty levels using items from a large-scale Reading and Writing test. ・The study investigated various prompting strategies and parameter settings across mu
#AIタグ

Cat!Cat!My Cat♪Third

・陽の当たる窓辺に佇むアメリカン・カール 続きをみる
cs.LG updates on arXiv.org

CENDRe: Concept Extraction with Natural Domain Representations

・arXiv:2607.29621v1 Announce Type: new Abstract: Convolutional neural networks (CNNs) are widely used for time-series classification, but their deployment in critical domains requires understanding the temporal and spectral patterns that drive their predictions. ・Concept extraction (CE) methods identify such patterns by analyzing representations within the models' latent space. ・However, existing time-series CE methods
The Verge

China’s Alibaba takes another swipe at America’s AI supremacy

・The Alibaba logo is displayed outside its headquarters in Hangzhou, Zhejiang Province, China. ・| Image: NurPhoto via Getty Images Chinese tech giant Alibaba released what it says is its largest and "most capable AI model to date," claiming performance rivaling the best systems from US frontier labs Anthropic and OpenAI, as well as domestic rivals like Moonshot AI's Kimi K3. ・Alibaba said it was making the model, Qwen3.
#AIタグ

Claude Code の7月アップデート、いちばん効いたのは Opus 5 じゃないんですよね

・7月の Claude Code、ニュースとして目立ったのは Claude Opus 5 のリリースでした。 ・ただ、正直に言うと、毎日の作業に効いてきたのはそこじゃなかったです。
Zennの「大規模言語モデル」のフィード

Claude Codeのコンテキスト引き継ぎを205セッション運用したら、申し送りが誤りを増幅していた

・この記事について Claude Code を 1 つのプロジェクトで 76 日走らせ続けています。会話ログは 205 セッション分残っていて、生のトランスクリプトは合計 991MB あります。 ・その過程で「セッションをまたいで文脈を引き継ぐ」仕組みを作りました。SessionStart で申し送りの冒頭を注入し、「続きから」と打つと前回ログのダイジェストが自動で入る。動いています。 ・問題は、動いた後に起きました。
LLMタグが付けられた新着記事 - Qiita

Claudeで複数モデルのエラー率上昇インシデント発生、監視中

・はじめに 2026年8月3日、Anthropic の公式ステータスページ(status.claude.com)に「Error rates across multiple models(複数モデルにまたがるエラー率上昇)」というインシデントが新規掲載されました。
#LLMタグ

Claude最新ニュース2026|Opus 5の進化と安全性を徹底解説

・Claude最新ニュースを追うと、焦点は単なる性能競争から、仕事を長時間任せられるAIエージェントの信頼性へ移っています。 ・2026年8月3日時点の結論は、Claude Opus 5が高性能と費用のバランスを大きく引き上げた一方、共有リンクやサイバー評価の事故が、導入時の管理責任をこれまで以上に重くしたということです。本稿では、Anthropicの公式発表、開発者向け文書、主要メディア、研究資料を照合し、最新モデルの違い、価格、周辺サービス、安全性、実務での選び方まで整理します。
MarkTechPost

Cogent AI Team Releases VR-1: A Frontier Cyber Reasoning Model That Composes and Verifies Enterprise Attack Paths

・Cogent AI team released Cogent VR-1, a reasoning model post-trained specifically for cybersecurity rather than picking up cyber capability as a side effect of general coding strength. ・It ships with two companions: IntrusionBench, a benchmark that scores agents on completed enterprise intrusions, and the Cogent AI Harness, a governed runtime for security agents. ・The launch […] The post Cogent AI Team Releases VR-1: A
cs.LG updates on arXiv.org

Commit to the Bit: Reactive Reinforcement Learning Done Right

・arXiv:2605.28276v2 Announce Type: replace Abstract: Reinforcement learning algorithms are commonly analyzed (and designed) under the Markov assumption. ・This is unrealistic, as most environments encountered in practice are either partially observable, or require function approximation that restricts the agent to access non-Markovian state features. ・We consider the problem of learning an optimal reactive policy in a fi
cs.LG updates on arXiv.org

Communication-Efficient Secure Aggregation in Decentralized Learning

・arXiv:2405.07708v3 Announce Type: replace Abstract: Decentralized learning (DL) enables participants to collaboratively train models without a central server, yet it faces significant scalability challenges that demand sparsification to reduce the prohibitive communication costs of peer-to-peer exchange. ・While secure aggregation effectively mitigates privacy risks in standard settings, it has remained fundamentally i
cs.LG updates on arXiv.org

CompoSE: Compositional Synthesis and Editing of 3D Shapes via Part-Aware Control

・arXiv:2605.19350v2 Announce Type: replace-cross Abstract: Creating and editing high-quality 3D content remains a central challenge in computer graphics. ・We address this challenge by introducing CompoSE, a novel method for Compositional Synthesis and Editing of 3D shapes via part-aware control. ・Our method takes as input a set of coarse geometric primitives (e.g., bounding boxes) that represent distinct object parts ar
cs.LG updates on arXiv.org

Conditioning Tree-Based Diffusions and Flows for Probabilistic Tabular Regression

・arXiv:2607.28864v1 Announce Type: cross Abstract: Tree-based diffusion models fit flexible conditional predictive distributions for tabular regression without a neural density estimator, but they inherit their design defaults---noising path, parameterization, training distribution, features, sampler---from the neural setting. ・We show these defaults are the binding constraint: what a gradient-boosted ensemble actually
AI News & Artificial Intelligence | TechCrunch

Congress’s favorite AI tool? ChatGPT

・House spending records show OpenAI's ChatGPT dominates paid AI use on Capitol Hill, with congressional offices relying on the chatbot to draft memos, summarize legislation, and assist constituent communications.
cs.LG updates on arXiv.org

Convergence and Regret of the Policy Gradient for Multi-Armed Bandits in Diffusion Environment

・arXiv:2607.29593v1 Announce Type: new Abstract: This paper studies the policy gradient update for a multi-arm bandit problem in diffusion environment that is described by a stochastic differential equation (SDE) under the continuous-time reinforcement learning framework by Wang et al. ・(2020), Jia and Zhou (2022b). ・With the logit parameterization for the stochastic policy, we show that it converges almost surely to th
cs.LG updates on arXiv.org

Cooperative Variance Estimation and Bayesian Neural Networks for Disentangling Aleatoric and Epistemic Uncertainties

・arXiv:2505.02743v3 Announce Type: replace Abstract: Real-world data contains aleatoric uncertainty - irreducible noise arising from imperfect measurements or from incomplete knowledge about the data generation process. ・Mean-variance estimation networks can learn this type of uncertainty but require ad-hoc regularization strategies to avoid overfitting and are unable to predict epistemic uncertainty (model uncertainty
cs.LG updates on arXiv.org

Cross-Resolution Semantic Learning for Graph Domain Adaptation

・arXiv:2607.29365v1 Announce Type: new Abstract: Graph Domain Adaptation (GDA) transfers predictive knowledge from labeled source graphs to unlabeled target graphs under distribution shift. ・Existing methods align representations or regularize graph structures, but do not explicitly model how class-discriminative knowledge learned at different source neighborhood ranges should be routed across target ranges.
cs.LG updates on arXiv.org

Curriculum Matters: Data-Efficient Relational PFN Pretraining with Synthetic Data

・arXiv:2607.29120v1 Announce Type: new Abstract: Relational Prior-Data Fitted Networks (PFNs) such as RDB-PFN approximate Bayesian inference over multi-table relational databases by pretraining on millions of synthetic tasks. ・We investigate three intertwined questions about this paradigm. ・First, can a structurally different synthetic generator PluRel substitute for RDB-PFN's prior?
cs.LG updates on arXiv.org

DASH-OPD: Discrepancy-Aware Switching with Hysteresis for On-Policy Distillation

・arXiv:2607.29078v1 Announce Type: new Abstract: On-policy distillation (OPD) trains student models on their own rollouts to reduce exposure bias. ・However, in multi-turn agent scenarios, early student errors can lead a trajectory away from the teacher's familiar domain. ・Existing curriculum learning methods regulate how much teacher support is used according to training progress, but cannot determine when it is needed.
#LLMタグ

DeepSeekがV4-Flashの正式版を公開 ─ 上位モデルV4-Proを7指標で上回り、出力単価は約3分の1

・DeepSeekは2026年7月31日、DeepSeek-V4-Flash-0731(以下0731版)をHugging Faceで公開した。報道によれば、あわせてV4-Flashの正式版APIをパブリックベータとして提供開始している。モデルカードは、プレビュー版を置き換える正式版で、エージェント能力を大幅に強化したと説明している。モデル構造とサイズは変えず、学習のやり直しだけで性能を伸ばした形だと伝える報道もある。 ・V4-Flashは総パラメータ284B・アクティブ13BのMoE(Mixture-of-Experts、入力ごとに一部の専門家パラメータだけを動かす方式)モデルで、コンテキスト長は100万トークンである。重みはMITライセンスで公開されている。
cs.LG updates on arXiv.org

DeltaServe: Host-Agnostic Co-Serving of Inference and Fine-Tuning for LLMs

・arXiv:2607.28848v1 Announce Type: cross Abstract: LLM serving systems are provisioned for peak load to meet strict latency targets, leaving substantial GPU compute idle whenever traffic falls below peak. ・We present DeltaServe, a host-agnostic co-serving design that converts this idle inference capacity into LoRA fine-tuning throughput while preserving inference service-level objectives (SLOs). ・DeltaServe integrates w
機械学習タグが付けられた新着記事 - Qiita

dera AI Weekly Vol.41 — 2026/7/27

・2026-07-27号 今週のAI業界を一言で表すなら? AIモデル連携の基盤となるMCPプロトコルが大規模な改訂を発表し、AIシステム構築の常識が大きく変わる一週間でした。 ・7月28日にリリース候補版が公開された新仕様は、ステートレスコアへの移行によりインフラ運用をシンプ...
cs.LG updates on arXiv.org

DFSC: Error-Controlled Differentiable Mittag-Leffler Propagation for Fractional Scientific Machine Learning

・arXiv:2607.29038v1 Announce Type: new Abstract: Fractional scientific machine learning requires numerical operators that can be differentiated, batched, accelerated, and composed with neural networks. ・When the dominant linear fractional evolution is known through a Mittag-Leffler propagator, repeatedly reconstructing that response with a history solver or relearning it from data is unnecessary. ・We present DFSC, a PyT
cs.LG updates on arXiv.org

Differentially Private Auditing Under Strategic Response

・arXiv:2605.07674v2 Announce Type: replace-cross Abstract: Regulatory audits of AI systems increasingly rely on differential privacy (DP) to protect training data and model internals. ・We study audit design when the audited developer can strategically respond to the privacy-constrained audit interface. ・We formalize privacy-constrained auditing as a bilevel Stackelberg game, in which an auditor commits to a query policy
cs.LG updates on arXiv.org

Differentially Private Nonparametric Modal Learning with Applications to Regression and Clustering

・arXiv:2607.29675v1 Announce Type: cross Abstract: Density modes provide a localized and interpretable summary of multimodal distributions, but their estimation under rigorous differential privacy constraints remains largely unexplored. ・We study differentially private recovery of density modes for multivariate distributions under local smoothness, curvature, and separation conditions. ・We propose DP-GRAMS, a mean-shift
cs.LG updates on arXiv.org

Dimensionality reduction for homological stability and global structure preservation

・arXiv:2503.03156v4 Announce Type: replace Abstract: We propose DiRe, a force-directed dimensionality reduction framework designed to preserve global structure and homological features while remaining practical on modern hardware. ・The method combines an initial embedding with a graph-based layout optimization and evaluates the resulting low-dimensional representation using local distortion, context preservation, and p
stat.ML updates on arXiv.org

Distance Profile Embedding for Independence and Conditional Independence Testing of Random Objects

・arXiv:2607.28981v1 Announce Type: cross Abstract: Testing independence or conditional independence is fundamental to statistical inference, yet existing methods for non-Euclidean random objects often face a difficult trade-off between geometric flexibility and theoretical tractability. ・We introduce the Distance Profile Embedding (DPE), a novel representation that maps random objects from general metric spaces into a
cs.LG updates on arXiv.org

Distilling Knowledge from Large Language Models into Lightweight Reinforcement Learning Agents for Autonomous Cyber Operations

・arXiv:2607.28826v1 Announce Type: new Abstract: Autonomous Cyber Operations (ACO) are increasingly important for defending enterprise networks as cyber threats continue to evolve in sophistication. ・ACO applications commonly employ Reinforcement Learning (RL) agents to learn defensive behaviors through interaction with environments. ・However, RL agents typically require extensive exploration during training, often resu
cs.LG updates on arXiv.org

Do LLMs Hold Their Values? MANTA: A Multi-Turn Adversarial Benchmark for Animal Welfare Reasoning

・arXiv:2605.16301v3 Announce Type: replace-cross Abstract: Evaluating animal welfare reasoning in LLMs remains an open challenge despite rapid deployment in consumer and professional contexts where welfare considerations appear implicitly in everyday queries. ・Existing benchmarks such as AnimalHarmBench evaluate this through single-turn, explicitly framed questions, measuring whether models avoid harmful content when d
cs.LG updates on arXiv.org

Don't Contrast the Impossible: Region-Constrained Batching for Contrastive User Modeling on a Local Community Platform

・arXiv:2607.28971v1 Announce Type: cross Abstract: Contrastive learning is widely used for user modeling in large-scale recommender systems, where standard in-batch negatives implicitly assume universal exposure that any user can be shown any item. ・On local community platforms such as Karrot, however, exposure is geographically constrained; many user-item pairs are impossible by design yet still treated as negatives d
cs.LG updates on arXiv.org

DreamQAS: Learning a Decision-Useful World Model for VQE-Efficient Quantum Architecture Search

・arXiv:2607.29491v1 Announce Type: new Abstract: Reinforcement-learning-based quantum architecture search (RL-QAS) repeatedly optimizes a variational quantum eigensolver (VQE) after extending a circuit, although circuit construction and action legality are deterministic and known. ・We introduce DreamQAS, a model-based RL framework that preserves these exact circuit dynamics and learns only the expensive post-VQE feedba
cs.LG updates on arXiv.org

DRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training

・arXiv:2606.30345v2 Announce Type: replace Abstract: Enabling large language models to achieve stable self-improvement without external expert supervision remains a central challenge in complex reasoning tasks. ・Existing self-distillation and reinforcement learning methods lack explicit mechanisms for tracking problem-level learning progress and adapting optimization strategies accordingly. ・Consequently, training may o
cs.LG updates on arXiv.org

Dual-Force: Enhanced Offline Diversity Maximization under Imitation Constraints

・arXiv:2501.04426v2 Announce Type: replace Abstract: Offline diversity maximization under imitation constraints can transform demonstration data into a set of distinct behavioral policies, improving robustness to distribution shift without additional environment interaction. ・In practice, however, existing offline approaches often rely on mutual-information objectives that require training a skill discriminator and can
cs.LG updates on arXiv.org

DungeonBench: A Benchmark for Rules-Rich Tactical Reasoning in Dungeons & Dragons Combat

・arXiv:2607.29577v1 Announce Type: cross Abstract: Games and simulators make valuable benchmarks by turning decisions into measurable outcomes, but many current suites under-test rules-rich tactical reasoning: the ability to choose well when geometry, timing, resources, objectives, and rule interactions all matter at once. ・We introduce DungeonBench, a benchmark for tactical reasoning in Dungeons & Dragons combat, buil
cs.LG updates on arXiv.org

Dynamic Priors in Bayesian Optimization for Hyperparameter Optimization

・arXiv:2511.02570v3 Announce Type: replace Abstract: Bayesian optimization (BO) is a widely used approach to hyperparameter optimization (HPO). ・However, most existing HPO methods only incorporate expert knowledge during initialization, limiting practitioners' ability to influence the optimization process as new insights emerge. ・This limits the applicability of BO in iterative machine learning development workflows.
cs.LG updates on arXiv.org

Dynamics-aware identification of governing equations from sparse and noisy data

・arXiv:2607.29036v1 Announce Type: new Abstract: Sparse identification of nonlinear dynamics (SINDy) and PDE functional identification (PDE-FIND) recover parsimonious ordinary and partial differential equations (ODEs and PDEs) from data. ・However, sparse and noisy temporal measurements can make derivative estimates unreliable. ・To address this problem, we evaluate Koopman-based upsampling techniques implemented with dyn
cs.LG updates on arXiv.org

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL

・arXiv:2606.31650v4 Announce Type: replace Abstract: Long-horizon language agents must repeatedly interact with tools, accumulate evidence, and make decisions under bounded context windows. ・Context-management methods make such rollouts feasible by simplifying past interactions through deletion, folding, or memory editing. ・However, when useful history is collapsed into compressed states, the reconstructed context may n
cs.LG updates on arXiv.org

Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates

・arXiv:2607.28959v1 Announce Type: new Abstract: Adversarial training is one of the most effective defenses against adversarial attacks, yet the computational cost remains prohibitive at modern scales, especially for large language models (LLMs). ・While existing mitigation strategies, e.g., latent adversarial training (LAT), have been developed, they still incur a high computational cost. ・In this work, we comprehensive
cs.LG updates on arXiv.org

EgoEverything: A Benchmark for Human Behavior Inspired Long Context Egocentric Video Understanding in AR Environment

・arXiv:2604.08342v2 Announce Type: replace Abstract: Long context egocentric video understanding has recently attracted significant research attention, with augmented reality (AR) highlighted as one of its most important application domains. ・Nevertheless, the task remains highly challenging due to the need for reasoning over extended temporal contexts and diverse, unstructured activities. ・Although several benchmarks e
cs.LG updates on arXiv.org

Embedding of Low-Dimensional Sensory Dynamics in Recurrent Networks: Implications for the Geometry of Neural Representation

・arXiv:2601.19019v3 Announce Type: replace-cross Abstract: Neural population activity in sensory cortex is organized on low-dimensional manifolds, but why such manifolds arise and what determines their geometry remain unclear. ・We model cortical populations as recurrent circuits driven by low-dimensional regular sensory dynamics (circles, tori). ・Combining generalized synchronization and delay-embedding theory, we show
Hugging Face Papers

EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents

EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents
cs.LG updates on arXiv.org

Enabling Low-Latency Machine learning on Radiation-Hard FPGAs with hls4ml

・arXiv:2602.15751v2 Announce Type: replace-cross Abstract: This paper presents an end-to-end demonstration of a viable, ultra-fast, radiation-hard machine learning (ML) application on FPGAs, which could be used in future high-energy physics experiments. ・We present a three-fold contribution, with the PicoCal calorimeter, planned for the LHCb Upgrade II experiment, used as a test case. ・First, we develop a lightweight au
cs.LG updates on arXiv.org

Encoding the Euler Characteristic Transform

・arXiv:2606.10824v2 Announce Type: replace Abstract: The Euler Characteristic Curve (ECC) records the Euler characteristic of a linearly embedded cell complex as a function of filtration height in a given direction, and the Euler Characteristic Transform (ECT) is the injective shape descriptor obtained by collecting ECCs over many directions. ・How the ECT is encoded for a neural network is itself an inductive bias, con
cs.LG updates on arXiv.org

End-to-End Fairness Optimization with Fair Decision-Focused Learning

・arXiv:2607.29441v1 Announce Type: new Abstract: Many real-world systems rely on predictive models to inform decisions, and fairness concerns arise in both the prediction and decision stages. ・We introduce end-to-end fairness optimization (E2EFO) as a unifying framework that integrates fairness across the prediction-to-decision pipeline. ・We focus on resource allocation with group-based fairness: the prediction task est
Hugging Face Papers

Enhancing Rubric-based RL via Self-Distillation

Enhancing Rubric-based RL via Self-Distillation
The Verge

Europe’s AI labeling and transparency rules are now in effect

・The EU made some AI labels that companies can use instead of designing their own. ・| Image: The European Commission / The Verge The European Union has ushered in some additional rules that aim to make it easier for people to identify chatbots and AI deepfakes online. ・The new transparency obligations under the bloc's landmark AI Act came into effect on August 2nd, requiring companies to disclose when people are interac
cs.LG updates on arXiv.org

Evaluating Federated Pre-Training: On the Reliability of Downstream Fine-Tuning and Intrinsic Evaluation

・arXiv:2607.28658v1 Announce Type: cross Abstract: Federated pre-training offers a way to train foundation models on private or distributed data without centralizing the underlying datasets. ・However, evaluating federated pre-training remains challenging because differences in client participation and local data availability can make directly comparable evaluation difficult. ・Moreover, pre-training test perplexity is ti
Hugging Face Papers

Evaluation-Verification Reward for Consistent Multi-Reference Image Editing

Evaluation-Verification Reward for Consistent Multi-Reference Image Editing
cs.LG updates on arXiv.org

Evidence-Type Competition: When Can Interventional Data Teach Language Models Causal Direction?

・arXiv:2607.29484v1 Announce Type: cross Abstract: Interventional data is widely regarded as the gold standard for teaching models causal reasoning. ・We test this assumption in a fully controlled synthetic environment pitting observational correlation against causal effect, and find it fails instructively. ・In Simpson's-paradox worlds, where the two have systematically opposite signs, increasing the fraction of interven
cs.LG updates on arXiv.org

Expert-Data Alignment Governs Generation Quality in Decentralized Diffusion Models

・arXiv:2602.02685v3 Announce Type: replace Abstract: Decentralized Diffusion Models (DDMs) route denoising through experts trained independently on disjoint data clusters, which can strongly disagree in their predictions. ・What governs the quality of generations in such systems? ・We present the first ever systematic investigation of this question.
Zennの「機械学習」のフィード

Explorative Modeling の最適解特徴付けに関する考察

・背景と結論 生成モデルの事前学習では、データ量とモデル規模を増やすことが性能改善の主要な手段として知られてきました。これらを 2 つのスケーリング軸と数えた上で、探索数 K を「第 3 の事前学習スケーリング軸」として位置づける提案が Explorative Modeling (XM) です (Gladstone–Ji–Du)。同論文および追試では、画像・動画・言語の各ドメインで K の増加に伴う経験的なベンチマーク改善が報告され、注目を集めています。発端となったポスト (開発者 Gladstone 本人によるもの) を引いておきます。実装も GitHub で公開されています。...
cs.LG updates on arXiv.org

Explore Beyond the Boundary Using Entropic Information

・arXiv:2607.29419v1 Announce Type: new Abstract: In reinforcement learning, exploration with sparse and delayed rewards presents a significant challenge due to the limited feedback available for guiding the learning process. ・Addressing this issue requires extensive exploration in the state space to discover valuable reward signals. ・In this paper, we propose Entropic Information for Exploration (ENTINEX), a novel metho
cs.LG updates on arXiv.org

Exploring Block Anomaly Detection In HDFS Log Data Analysis

・arXiv:2607.29383v1 Announce Type: new Abstract: In recent years, with the development of big data technology, increasingly more companies use HDFS for data processing and storage. ・As a result, the maintenance of distributed file systems has become an extremely important part of data management. ・As the function of server systems is becoming increasingly diversified and their services are becoming complex, the logs, re
stat.ML updates on arXiv.org

Exponential Capacity in Multilayer Hetero-Associative Neural Networks

・arXiv:2607.29554v1 Announce Type: cross Abstract: Exponential Hopfield networks store a number of patterns that grows exponentially with the number of neurons, and in their classical formulation they are auto-associative: they complete a corrupted copy of a memory into the memory itself. ・Many of the tasks one wants such a network to perform are instead hetero-associative, mapping a cue to a different target.
Hugging Face Papers

ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction

ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction
cs.LG updates on arXiv.org

Extrapolating the emergence of Hamiltonian chaos with random-feature Hamiltonian neural networks

・arXiv:2607.28977v1 Announce Type: cross Abstract: Machine learning of Hamiltonian dynamics has driven growing interest in Hamiltonian neural networks (HNNs), which encode Hamilton's equations of motion into the learning architecture. ・Despite this progress, it remains unknown whether such networks can predict dynamical regimes absent from their training data, in particular the broad chaotic sea that emerges beyond the
cs.LG updates on arXiv.org

FairDiffuseVQVAE: Sampling-Time Fairness in Tabular Diffusion via Conditional Refinement of Vector-Quantized Latents

・arXiv:2607.28945v1 Announce Type: new Abstract: Synthetic tabular data is increasingly used in privacy-preserving data sharing, data augmentation, and to mitigate downstream classifier bias. ・State-of-the-art tabular diffusion models such as TabDDPM and TabSyn achieve excellent distributional fidelity but offer no mechanism for fairness; conversely, fairness-aware tabular generators (DECAF, FairTGAN, FairTabDDPM) impo
cs.LG updates on arXiv.org

Fast Feature Field ($\text{F}^3$): A Predictive Representation of Events

・arXiv:2509.25146v2 Announce Type: replace-cross Abstract: This paper develops a mathematical argument and algorithms for building representations of data from event-based cameras, that we call Fast Feature Field ($\text{F}^3$). ・We learn this representation by predicting future events from past events and show that it preserves scene structure and motion information. ・$\text{F}^3$ exploits the sparsity of event data an
cs.LG updates on arXiv.org

Fast Rates for Swap-Agnostic Learning of Proper Losses

・arXiv:2607.28856v1 Announce Type: new Abstract: Swap-agnostic learning strengthens classical agnostic learning by allowing the comparator to select a different hypothesis on each level set of the learner's predictions. ・This benchmark captures prediction-dependent postprocessing, but appears to require solving a separate agnostic-learning problem for every possible prediction value. ・We show that, for proper losses, th
cs.LG updates on arXiv.org

Feature Interaction Modeling for Physics-Informed Neural Networks and Neural Operators

・arXiv:2607.28762v1 Announce Type: new Abstract: This work embeds feature interaction modules derived from factorization machines (FMs) into physics-informed neural networks (PINNs) and neural operator learning, to enhance model expressiveness for solution manifolds of parameterized partial differential equations (PDEs). ・Motivated by the second-order Taylor expansion of multivariate functions to characterize variable
cs.LG updates on arXiv.org

Federated Foundation Models Fine-Tuning with Heterogeneous Compressed Clients

・arXiv:2607.29071v1 Announce Type: new Abstract: Federated learning of foundation models faces a fundamental resource-asymmetry challenge: the institutions holding the most valuable domain-specific data cannot host billion-parameter models. ・Existing heterogeneous federated approaches attempt to bridge this gap through parameter-efficient tuning, model pruning, or knowledge distillation, yet each trades away a critical
cs.LG updates on arXiv.org

Few-shot Deep Learning for Phase-Amplitude Aberration Correction in Transcranial Focused Ultrasound

・arXiv:2607.29182v1 Announce Type: cross Abstract: Transcranial focused ultrasound (tFUS) is a non-invasive technique that delivers focused acoustic energy through the skull for neuromodulation and therapeutic applications. ・However, the heterogeneous structure of the skull induces complex, patient-specific phase and amplitude aberrations that distort the acoustic focus and deviate it from the intended target, compromi
Hugging Face Papers

Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants

Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants
cs.LG updates on arXiv.org

FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale

・arXiv:2601.22146v3 Announce Type: replace-cross Abstract: Due to limited supervised training data, large language models (LLMs) are typically pre-trained via a self-supervised "predict the next word" objective on a vast amount of unstructured text data. ・To make the resulting model useful to users, it is further trained on a far smaller amount of "instruction-tuning" data comprised of supervised training examples of i
cs.LG updates on arXiv.org

Fisher Information, Training and Bias in Fourier Regression Models

・arXiv:2510.06945v2 Announce Type: replace Abstract: Motivated by the growing interest in quantum machine learning, in particular quantum neural networks (QNNs), we study how recently introduced evaluation metrics based on the Fisher information matrix (FIM) are effective for predicting their training and prediction performance. ・We exploit the equivalence between a broad class of QNNs and Fourier models, and study the
cs.LG updates on arXiv.org

Flow Matching with Missing Data

・arXiv:2607.28698v1 Announce Type: new Abstract: Flow matching assumes fully observed training data, which many real-world applications rarely provide. ・We propose Missing-Data Flow Matching, which treats the missing coordinates of training samples as latent variables and averages the flow matching loss over the values they could take. ・We first prove the correction is exact rather than approximate.
cs.LG updates on arXiv.org

FORGE: Frame Orthogonality in Relevance Geometry for Long-Form Video Understanding

・arXiv:2607.25266v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have enabled long-form video understanding at a scale that was not previously possible. ・However, the density of relevant content decreases sharply as video sequence length increases, and exposing the model to more irrelevant content measurably reduces its accuracy. ・In this paper, we address the problem of maximizing que
cs.LG updates on arXiv.org

Fracture Risk Prediction in Adults Over 50 Years Old Using DXA and EHR: Comparison of Traditional and Machine Learning Models in Two Large Cohorts

・arXiv:2607.28671v1 Announce Type: cross Abstract: Accurate fracture risk prediction is important for osteoporosis management, but commonly used clinical tools may not fully use information available in electronic health records (EHRs) and dual-energy X-ray absorptiometry (DXA) reports. ・We developed and externally validated time-to-event fracture prediction models among adults aged 50 years or older with clinically ob
cs.LG updates on arXiv.org

Freeze, Then Select: Structured Field Adapters and Stability-Validated Weak Selection for PDE Discovery from Sparse Observations

・arXiv:2607.29665v1 Announce Type: new Abstract: PDE discovery from sparse observations requires reconstructing a continuous field and selecting the correct differential terms. ・Our analysis of optimization paths in coupled neural PDE discovery reveals three behaviors: the exact support can persist to the end of training, appear only transiently, or fail to emerge. ・To decouple equation selection from neural optimizatio
Hugging Face Papers

From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement

From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement
cs.LG updates on arXiv.org

Frugal Bayesian Optimization: Scalable Surrogates for Data- and Resource-Limited Discovery

・arXiv:2607.29225v1 Announce Type: new Abstract: Bayesian Optimization (BO) is widely adopted for data-efficient optimization in scientific and engineering applications, yet its computational cost is rarely evaluated alongside optimization performance. ・Here we present a systematic, compute-aware study of BO that evaluates surrogate models along two axes: optimization quality and computational frugality. ・Across eight b
cs.LG updates on arXiv.org

GALA: Generative Aligned Learning for Adaptive Multimodal Representation in the Taobao Shangou Recommender System

・arXiv:2607.29213v1 Announce Type: cross Abstract: Modern recommender systems in food delivery increasingly leverage multimodal signals, including images, text, and user interaction histories, to enhance user experience, yet effective fusion of these heterogeneous modalities remains challenging, hindering both the joint modeling of multimodal signals and adaptation to evolving user intent. ・In mainstream two-stage appr
cs.LG updates on arXiv.org

Gated Q-learning: Add Off-Policy Bias to Taste

・arXiv:2607.28916v1 Announce Type: new Abstract: Multistep credit assignment is critical for sample-efficient reinforcement learning, yet managing off-policy bias in Q-learning remains a fundamental challenge. ・For 30 years, practitioners have been limited to a binary choice: eliminate the bias at the cost of severely truncated eligibility traces (Watkins' Q($\lambda$)), or ignore the bias to learn faster while injecti
cs.LG updates on arXiv.org

GEMSS: A Variational Method for Discovering Multiple Sparse Solutions in Classification and Regression Problems

・arXiv:2602.08913v3 Announce Type: replace Abstract: In underdetermined regression and classification problems, multiple feature subsets often yield equivalent predictive performance. ・In applied settings, especially with $n \ll p$, high dimension or collinearities, it is valuable to provide a domain expert with a menu of statistically plausible explanations, rather than one arbitrary solution. ・This creates the need fo
cs.LG updates on arXiv.org

Geographically Weighted Surrogate Models for Rapid Small-Area Chronic Disease Estimation

・arXiv:2607.28655v1 Announce Type: cross Abstract: Small-area estimation (SAE) enables researchers and policymakers to identify spatial disparities in health outcomes, but survey-based SAE products carry an inherent lag. ・Gold-standard estimates such as CDC PLACES are released roughly two years after the underlying survey data are collected, limiting their use for time-sensitive decision-making. ・This study evaluates th
cs.LG updates on arXiv.org

GeoRA: Geometry-Aware Low-Rank Adaptation for RLVR

・arXiv:2601.09361v4 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) is a key paradigm for improving large-scale reasoning models. ・Unlike supervised fine-tuning (SFT), RLVR exhibits distinct optimization dynamics and is sensitive to the preservation of pre-trained geometric structures. ・However, existing parameter-efficient methods face key limitations in this regime.
cs.LG updates on arXiv.org

GQ-FSL: Green Quantized Federated Split Learning

・arXiv:2607.29659v1 Announce Type: new Abstract: Deploying state-of-the-art deep neural networks (DNNs) at the wireless edge is severely bottlenecked by the strict energy and resource constraints of mobile devices. ・While federated split learning (FSL) mitigates on-device computation by offloading workloads to an edge server, this may introduce systemic overheads, while the continuous exchange of cut-layer data, and su
cs.LG updates on arXiv.org

Guarantees on Dynamical System Distinguishability for LLM Token Generation

・arXiv:2607.28667v1 Announce Type: new Abstract: Recent work has shown that classifying large language models (LLMs)' responses can be distinguished by modeling token embeddings as trajectories of a black-box dynamical system (DS) and comparing prediction residuals of two DSs. ・Despite the empirical success of this dynamical approach, a theoretical understanding of why it works, how well it scales as a function of the
cs.LG updates on arXiv.org

Have I Seen You? Embedding Behavior Signals Synthetic Face Dataset Membership

・arXiv:2607.29144v1 Announce Type: cross Abstract: Synthetic face datasets are increasingly used to reduce privacy exposure and data access constraints in biometric recognition. ・Yet the generators that produce these datasets are trained on real faces, so synthetic data may still reveal their real source data. ・We study this risk through a dataset-level membership inference attack that first identifies the synthetic dat
Zennの「大規模言語モデル」のフィード

herdrとNotionでClaude Code 4体を回して、社内ハッカソンで完全バイブコーディングした話

・本記事のコード例・設定例は、社内ハッカソンで作った検証用アプリの実装をブログ用に抜粋・簡略化した概念的なものです。そのまま本番環境で使える完全なものではありません。同様の構成を組む場合は、エラーハンドリング・ログ出力・セキュリティ対策などを環境に合わせて補ってください。構成や設定にも、実際の構築から省略している部分があります。 ・はじめに こんにちは。ソリューションアーキテクトの髙宮です。 ・先日、部内のハッカソン「DX部 ハッカソン vol.2」に参加してきました。全拠点のメンバーが1か所に集まって、共通のバックエンドテンプレートの上に1日でアプリを作り、UI/UXを競う個人戦で...
cs.LG updates on arXiv.org

HERO: History-Enriched Rollout Training for Long-Horizon Autoregressive Neural Operators

・arXiv:2607.29135v1 Announce Type: new Abstract: Neural operators provide fast surrogates for time-dependent partial differential equations (PDEs) by applying a learned evolution operator recursively to its own predictions, but this autoregressive rollout feeds every prediction error back as input, so local errors accumulate. ・Existing rollout-training strategies reduce the mismatch between training inputs and self-gen
cs.LG updates on arXiv.org

Hierarchical Copula-Gumbel-Top-\texorpdfstring{$K$}{K} Routing: Two-Sided Dependence Control for Frozen Mixture-of-Experts at Fixed Per-Token Routing Laws

・arXiv:2607.28670v1 Announce Type: new Abstract: A stochastic Gumbel-Top-$K$ router defines, for every token of a mixture-of-experts (MoE) model, a \emph{routing law}: a distribution over ordered expert lists and mixture weights. ・We ask which \emph{joint} distributions over the routing choices of different tokens are reachable while every individual token's complete routing law is held exactly fixed. ・We give a two-sid
cs.LG updates on arXiv.org

How Hard Does It Think? Analyzing Step-Aware Reasoning Energy in LLM Chain-of-Thought Trajectories

・arXiv:2607.28674v1 Announce Type: cross Abstract: Understanding how computational effort is allocated across individual chain-of-thought (CoT) reasoning steps remains an open challenge: existing interpretability methods rely on output-level signals or collapse processing depth into a single trajectory-level scalar, leaving step-wise effort opaque. ・We propose Step-Aware Reasoning Energy (SARE), a geometric framework t
cs.LG updates on arXiv.org

Hypergradient-based Bilevel Reinforcement Learning with Improved Sample Complexity

・arXiv:2607.28849v1 Announce Type: new Abstract: Bilevel reinforcement learning (RL) is an important framework within the literature of RL that can be used to formalize various categories of problems, such as meta-learning, hierarchical task decomposition, and reinforcement learning from human feedback (RL-HF). ・Most of the bilevel RL algorithms are either not scalable because of using hypergradient with Hessian, or th
WIRED

ICE Collected Nearly 1 Million People’s DNA Last Year—Including Young Children

・Internal documents show ICE's DNA collection has skyrocketed in the second Trump administration. ・Now hundreds of thousands of people never convicted of a crime are in an FBI criminal database forever.
stat.ML updates on arXiv.org

Identifying Informative Environments for Cognition Parameter Inference via Bayesian Experimental Design

・arXiv:2607.28894v1 Announce Type: cross Abstract: Computational cognitive modeling seeks to infer latent cognitive mechanisms underlying observed behavior. ・Bayesian inverse planning provides a principled framework for such inference, but its success depends critically on the experimental environment. ・Existing approaches typically treat environments as fixed, leaving open the question of which cognitive experiments ar
cs.LG updates on arXiv.org

Imbalanced Data Clustering via Targeted Data Augmentation Using GMM and LLM

・arXiv:2607.28635v1 Announce Type: cross Abstract: In Natural Language Processing (NLP), dealing with underrepresented topics is challenging, especially in unsupervised tasks where clustering might not adequately capture minority topics. ・To tackle this challenge, our paper presents a novel unsupervised data augmentation method that integrates Gaussian Mixture Models (GMMs) and Large Language Models (LLMs).
cs.LG updates on arXiv.org

Implicit Machine Learning Force Fields Accelerate Molecular Dynamics Simulations

・arXiv:2607.29158v1 Announce Type: new Abstract: We introduce implicit machine learning force fields (I-MLFFs), which replace explicit stacks of neural network layers with self-consistent fixed-point equations. ・In molecular simulations, this formulation enables intermediate representations to be reused across successive timesteps, thereby warm-starting force evaluation. ・The resulting models effectively combine the com
Hugging Face Papers

In the Driver's Seat: A Multi-Company Study on the Reality of Autonomous Driving System Testing

In the Driver's Seat: A Multi-Company Study on the Reality of Autonomous Driving System Testing
cs.LG updates on arXiv.org

In-situ Autoguidance: Eliciting Self-Correction in Diffusion Models

・arXiv:2510.17136v2 Announce Type: replace Abstract: The generation of high-quality, diverse, and prompt-aligned images is a central goal in image-generating diffusion models. ・The popular classifier-free guidance (CFG) approach improves quality and alignment at the cost of reduced variation, creating an inherent entanglement of these effects. ・Recent work has successfully disentangled these properties by guiding a mode
cs.LG updates on arXiv.org

Incorporating data drift to perform survival analysis on credit risk

・arXiv:2601.20533v2 Announce Type: replace-cross Abstract: Survival analysis has become a standard approach for modelling time to default by time-varying covariates in credit risk. ・Unlike most existing methods that implicitly assume a stationary data-generating process, in practise, mortgage portfolios are exposed to various forms of data drift caused by changing borrower behaviour, macroeconomic conditions, policy re
cs.LG updates on arXiv.org

Information Processing by Neuron Populations in the Central Nervous System: A Theory of the Mathematical Structure of Data and Operations

・arXiv:2309.02332v3 Announce Type: replace-cross Abstract: In the mammalian central nervous system, neurons are organized into populations communicating by spike trains propagating along axonal bundles. ・How such populations encode and transform information is only partially understood. ・In this study we introduce a mathematical framework derived from a mechanistic model of a single plastic neuron.
cs.LG updates on arXiv.org

LARA: Lightweight Adapters in the Residual Stream for Composable Adaptation and Alignment

・arXiv:2607.28669v1 Announce Type: new Abstract: We present LARA (Lightweight Additive Residual Adaptation), a method for efficient adaptation that operates in the residual stream of a frozen model rather than in its weights. ・Where LoRA adds an update of low rank to weight matrices, LARA reads the hidden state at a small set of layers and adds a correction of low rank back to the residual stream, leaving all base weig
cs.LG updates on arXiv.org

Latent Lie-Poisson Neural Networks (LLPNNs): Discovering the motion of Lie-Poisson systems through observable data and latent dynamics

・arXiv:2607.28939v1 Announce Type: new Abstract: Structure-preserving neural networks are essential for the long-term prediction of Hamiltonian systems from data. ・Many important Hamiltonian systems in mechanics and control admit symmetry reduction to Lie--Poisson systems, including rigid bodies, underwater vehicles, fluids, plasmas, and optimal control problems. ・A fundamental challenge in learning such systems is that
cs.LG updates on arXiv.org

Latent Sculpting for Zero-Shot Generalization: A Manifold Learning Approach to Out-of-Distribution Anomaly Detection

・arXiv:2512.22179v3 Announce Type: replace Abstract: Detecting previously unseen attacks remains a major challenge for machine learning-based intrusion detection systems. ・Deep models trained on network traffic often achieve high accuracy on known attacks but fail under distributional shift because their decision boundaries are tightly coupled to the training data distribution. ・We introduce Latent Sculpting, a two-stag
cs.LG updates on arXiv.org

LAWFUL: Law-Aligned Witness for Faithful Use of Latents

・arXiv:2607.28672v1 Announce Type: new Abstract: When a neural network predicts a physical system accurately, has it learned the governing law as formal, structured knowledge, and if so, does the network's internal computation actually use that representation throughout the law's domain of validity? ・We identify four interpretability gaps that limit answering these questions for {\em physics laws over continuous variab
cs.LG updates on arXiv.org

LayoutBench: Performance Benchmarking of Cloud Storage Layouts for Multimedia Data

・arXiv:2607.28880v1 Announce Type: cross Abstract: Modern multimedia machine learning workloads increasingly store large-scale datasets in cloud object storage services such as AWS S3. ・How these samples are physically organized in storage (i.e.,storage layout) directly affects how quickly and cheaply they can be retrieved. ・Yet the benchmarks used to guide storage decisions today focus on database engines and query pro
cs.LG updates on arXiv.org

LearnedCache: eBPF-Integrated Perceptron-Based Eviction Policies for the Linux Page Cache

・arXiv:2605.26168v2 Announce Type: replace-cross Abstract: Any device that runs Linux uses the Linux page cache, a central pillar in OS and application performance, serving to reduce extraneous disk access. ・Many page cache eviction policies have been developed but remain bound by the rigidity of heuristics. ・Promising research has been done on neural cache eviction policies, but only in the field of user-space applicat
cs.LG updates on arXiv.org

Learning Lookahead Lemmas for Neural Network Verification

・arXiv:2607.29051v1 Announce Type: new Abstract: State-of-the-art neural network verifiers use the branch-and-bound procedure as their core solving mechanism. ・We introduce an inprocessing framework for neural network verification driven by the lookahead procedure. ・Under this framework, lookahead derives new lemmas over the phases of unstable ReLUs, which are collected into an implication graph that is used to prune th
cs.LG updates on arXiv.org

Learning Optimal Dynamic Matching via Graph Neural Networks

・arXiv:2607.28925v1 Announce Type: new Abstract: Dynamic matching markets require decisions about whom to match and when: matching now yields value but removes participants who may create better future opportunities. ・We develop a value-based reinforcement-learning framework for this problem on finite, evolving weighted graphs. ・We study an infinite-horizon continuous-time model with stochastic arrivals, node-type trans
cs.LG updates on arXiv.org

Learning Stateful Predictive Knowledge From Experience

・arXiv:2607.28638v1 Announce Type: cross Abstract: As large language model (LLM) agents increasingly learn from experience, they primarily rely on trajectory-level reflection to extract insights. ・Viewed through the lens of predictive knowledge, we argue that this approach operates on episodic hindsight rather than predictive foresight, yielding brittle, path-dependent heuristics. ・To address this, we propose Stateful K
cs.LG updates on arXiv.org

Learning to Predict Performance-induced Emotion Differences in Classical Piano Music

・arXiv:2607.28876v1 Announce Type: cross Abstract: Music is often used as a medium for communicating emotion, with performers shaping perceived affect through interpretation. ・This study addresses the challenge of identifying and predicting subtle changes in perceived emotion that are exclusively due to differences in performance. ・We focus on classical solo piano music, using a set of 6 commercial recordings of Bach's
The Verge

Lenovo Googlebook leaks reveal a laptop and 2-in-1 tablet

・Leaked images of the laptop show Googlebook branding below the keyboard. ・| Image: Digital Citizen Lenovo is expected to release some of the first Googlebook models later this year, and leaked images have now given us a good idea of what they might look like. ・Leaked press images shared by Digital Citizen and Android Headlines include a laptop and a 2-in-1 tablet, all of which feature Googlebook branding on the keyboar
cs.LG updates on arXiv.org

Leveraging Image Generators to Address Data Scarcity: The Gen4Regen Dataset for Forest Regeneration Mapping

・arXiv:2605.05627v2 Announce Type: replace-cross Abstract: Sustainable forest management relies on precise species composition mapping, yet traditional ground surveys are labour-intensive and geographically constrained. ・While Uncrewed Aerial Vehicles (UAVs) offer scalable data collection, the transition to deep learning-based interpretation is bottlenecked by the severe scarcity of expert-annotated imagery, particular
cs.LG updates on arXiv.org

Leveraging Transfer Learning with Class-Specific Decoders for Laparoscopic Segmentation

・arXiv:2607.29509v1 Announce Type: cross Abstract: Effective multi-organ segmentation in surgical data requires learning the intricate anatomical features and alleviating the challenge of class imbalance, which results from relatively lower proportions of small and limitedly exposed structures. ・Recent works on laparoscopic multi-organ segmentation focus on learning structure-specific features through class-specific de
cs.LG updates on arXiv.org

LightningRL: Breaking the Accuracy-Parallelism Trade-off of Block-wise dLLMs via Reinforcement Learning

・arXiv:2603.13319v2 Announce Type: replace Abstract: Diffusion Large Language Models (dLLMs) have emerged as a promising paradigm for parallel token generation, with block-wise variants garnering significant research interest. ・Despite their potential, existing dLLMs typically suffer from a rigid accuracy-parallelism trade-off: increasing the number of tokens per forward (TPF) via aggressive parallel decoding often lea
cs.LG updates on arXiv.org

Lightweight Neural Networks for Affordance Segmentation: Enhancement of the Decoder Module

・arXiv:2607.29473v1 Announce Type: cross Abstract: The deployment of deep neural networks for visual affordance segmentation on wearable robots poses may prove critical, due to some conflicting aspects of the problem. ・On one hand, affordance segmentation requires high-level abstraction capabilities, that typically involve large-size models. ・On the other hand, computing resources hosted on wearable robots prevent to ru
stat.ML updates on arXiv.org

Longitudinal Adaptive Experimental Design for Learning Multiple Target Estimands with Semiparametric Efficient Inference

・arXiv:2607.29421v1 Announce Type: cross Abstract: Adaptive designs are increasingly used in clinical trials and digital experiments to improve estimation efficiency by updating treatment randomization probabilities as data accumulate. ・While most existing work focuses on settings with a single-stage treatment, adaptive designs for longitudinal studies with multi-stage, time-varying treatments remain relatively underex
cs.LG updates on arXiv.org

Low-Cost Hard-Label Adversarial Attack with Theoretical Foundations

・arXiv:2601.14300v4 Announce Type: replace Abstract: Hard-label black-box attacks, relying solely on top-1 predictions, represent one of the most challenging yet practically threat models. ・Despite recent progress, existing approaches face two key limitations: (1) they overlook the critical role of initialization, focusing primarily on optimization strategies; and (2) they rely heavily on empirical heuristics without t
cs.LG updates on arXiv.org

MARGIN: Runtime Confidence Calibration for Multi-Agent Foundation Model Coordination

・arXiv:2605.22949v3 Announce Type: replace Abstract: Foundation-model pools are increasingly used as black-box responders in coordinated systems where a coordinator must decide which response to trust. ・Raw self-reported confidence is the natural signal, but is not comparable across models and becomes stale under distribution shift when corrected only at design time. ・We study runtime confidence calibration for multi-mo
cs.LG updates on arXiv.org

Matterhorn: Masked Time-to-First-Spike Encoding by Reassigning the Silent State for Sparse and Energy-Efficient Spiking Transformers

・arXiv:2601.22876v2 Announce Type: replace Abstract: Spiking neural networks (SNNs) promise energy-efficient inference for large language models (LLMs), yet most reported savings rely on compute-operation counts that overlook data movement. ・Energy characterization of representative spiking transformers on a commercial 22-nm process shows that accumulation contributes less than 3% of total energy, while spike-triggered
cs.LG updates on arXiv.org

Maximum Entropy Behavior Exploration for Sim2Real Zero-Shot Reinforcement Learning

・arXiv:2603.25464v2 Announce Type: replace Abstract: Zero-shot reinforcement learning (RL) algorithms aim to learn a family of policies from a reward-free dataset, and recover optimal policies for any reward function directly at test time. ・Naturally, the quality of the pretraining dataset determines the performance of the recovered policies across tasks. ・However, pre-collecting a relevant, diverse dataset without prio
cs.LG updates on arXiv.org

MBDiff: Multi-view Behavior-aware Diffusion Model for Probabilistic Utility Data Imputation

・arXiv:2607.29177v1 Announce Type: new Abstract: Utility data (e.g., electricity, water, and gas consumption), collected by ubiquitous sensors and embedded devices, often contains substantial missing values due to various factors such as device failures and data transmission issues. ・The data missingness can severely impact utility billing accuracy, hinder demand forecasting, and disrupt efficient utility supply manage
Hugging Face Papers

Mental World Modeling

Mental World Modeling
Hugging Face Papers

Meshy T2: Fast Native Mesh Generation with Flow Matching

Meshy T2: Fast Native Mesh Generation with Flow Matching
The Verge

Microsoft is bringing Xbox 360 games to PC

・Building on its recently announced plans to bring original Xbox games to PC, Microsoft is also planning to let developers bring their Xbox 360 games to PC as well, according to a leaked document, seen by The Verge, that was sent to developers recently. ・Xbox 360 games will be able to run on the next-gen Project Helix console, "Xbox PCs," and handheld devices, the document says. ・Microsoft has already said that Helix wi
cs.LG updates on arXiv.org

Mining Verdict Boundaries for Neural Network Verification

・arXiv:2607.28954v1 Announce Type: new Abstract: Branch and Bound (BaB) aims to achieve complete verification of neural networks by adaptively partitioning the problem and applying off-the-shelf verifiers to subproblems. ・Its problem-splitting history can be represented as a tree, where each subproblem corresponds to a child node. ・A key problem of BaB lies in searching for the verdict boundaries across all the paths th
cs.LG updates on arXiv.org

Mirror Learning

・arXiv:2607.28737v1 Announce Type: new Abstract: We investigate imitation learning through the lens of third-person observation and propose a framework for mirror learning: acquiring actionable policies from passive observation. ・While behavior cloning (BC) excels under dense, well-aligned first-person data, it fundamentally fails to leverage the rich observational signals arising from third-person demonstrations that
cs.LG updates on arXiv.org

Mitigating Class-Tail Undercoverage in Medical Vision-Language Models under Clinical Shift

・arXiv:2607.28696v1 Announce Type: new Abstract: Medical vision-language models (VLMs) can retain high observed marginal coverage after clinical shift while substantially under-covering an individual disease class. ・The affected class varies with acquisition protocol and backbone geometry, so source prevalence does not reliably reveal the failure. ・Existing localized and tail-aware conformal methods respectively adapt t
cs.LG updates on arXiv.org

MMFGU: Multimodal Federated Graph Unlearning

・arXiv:2607.28708v1 Announce Type: new Abstract: Multimodal federated graph learning enables clients to collaboratively train graph models over structural, textual, and visual signals without sharing private local data. ・However, the presence of heterogeneous multimodal content also makes unlearning requests more frequent and fine-grained: users may delete accounts or interactions, remove a particular image or text whi
cs.LG updates on arXiv.org

MolGVR: A Chemistry-Grounded Framework for Text-to-Molecule Generation

・arXiv:2607.29479v1 Announce Type: new Abstract: Text-to-molecule generation is typically formulated as a one-shot sequence generation problem, where a model directly maps target descriptions to molecular representations. ・However, molecular descriptions often contain informative structural constraints, and violating such constraints can change the molecular identity. ・This makes chemical verification and error correcti
cs.LG updates on arXiv.org

Moment kernels: a simple and scalable approach for equivariance to rotations and reflections in deep convolutional networks

・arXiv:2505.21736v2 Announce Type: replace-cross Abstract: Translation equivariance is a central reason convolutional neural networks have been successful in computer vision. ・Other symmetries, such as rotations and reflections, are similarly important in fields such as biomedical image analysis, but equivariant methods for these symmetries remain less widely adopted, especially in 3D. ・Existing approaches often rely on
cs.LG updates on arXiv.org

Monotone and Separable Set Functions: Characterizations and Neural Models

・arXiv:2510.23634v4 Announce Type: replace Abstract: Motivated by applications for set containment problems, we consider the following fundamental problem: can we design set-to-vector functions so that the natural partial order on sets is preserved, namely $S\subseteq T \text{ if and only if } F(S)\leq F(T) $. ・We call functions satisfying this property Monotone and Separating (MAS) set functions. ・% We establish lower
cs.LG updates on arXiv.org

MoPET: Parameter-Efficient Mixture-of-Experts for Unified Medical Image Classification

・arXiv:2607.29462v1 Announce Type: cross Abstract: Adapting deep learning models to profound clinical heterogeneity typically relies on parameter-efficient fine-tuning (PEFT) to avoid the severe overfitting associated with full end-to-end network updates. ・Although PEFT successfully navigates limited data scenarios, it inherently forces the training of a separate, isolated adapter for every specific diagnostic task.
cs.LG updates on arXiv.org

MOT-SR: Multi-Objective Tool-Augmented Scientific Equation Discovery with Large Language Models

・arXiv:2607.29561v1 Announce Type: new Abstract: Symbolic Regression (SR) aims to discover analytical equations from observational data and plays a central role in scientific modeling. ・While recent Large Language Model (LLM) based approaches show promise, they face two limitations. ・First, they lack data analysis mechanisms for uncovering variable dependencies, which reduces the efficiency of equation discovery.
cs.LG updates on arXiv.org

MPP-GNN: Subject-Adaptive Community Detection for fMRI-Based Alzheimer's Disease Classification

・arXiv:2607.28681v1 Announce Type: new Abstract: Functional magnetic resonance imaging (fMRI) is a widely used technique for studying the brain. ・Recent methods that utilize graph neural networks (GNNs) for analysis of brain functional connectivity have shown great potential for the classification of brain disorders, such as Alzheimer's disease (AD). ・However, these methods often assume a preset number of functional mod
cs.LG updates on arXiv.org

Multi-Scale Feature Attention Network for Polymer Classification Using Terahertz Spectroscopy

・arXiv:2606.06554v3 Announce Type: replace Abstract: Reliable polymer identification is essential for ensuring the quality and safety of recycled plastics, yet conventional sorting and spectroscopic techniques often struggle to deliver robust discrimination. ・Terahertz (THz) spectroscopy offers a promising alternative, providing high-resolution and non-destructive measurements. ・In this work, we leverage THz signals to
Hugging Face Papers

N_0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation

N_0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation
Hugging Face Papers

N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens

N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens
cs.LG updates on arXiv.org

NeuroSynth: A Biologically Inspired Continual Reinforcement Learning Architecture for Mitigating Catastrophic Forgetting

・arXiv:2607.28663v1 Announce Type: cross Abstract: Artificial Intelligence (AI) systems often perform well on isolated tasks but struggle under continual learning conditions, where training on new tasks can overwrite previously acquired knowledge, a failure mode known as catastrophic forgetting. ・Biological learning systems reduce this interference through complementary memory processes involving rapid hippocampal enco
#LLMタグ

Next.js受託でFunction Calling実装する5つの判断

・「AIとのやり取りをチャットで終わらせず、予約確定や在庫更新まで任せたい」。最近、このような相談を受ける機会が増えています。実際にNext.jsでFunction Calling(ツール呼び出し)を組み込もうとすると、チャットボットを作るときにはなかった判断が次々に出てきます。ツールをどこまで細かく分けるか。何個まで一度に渡していいか。実行結果をどう返すか。取り消せない操作の手前で何を挟むか。同時に呼ばせていいか。この記事では、Next.js×AI受託をしている自分が、見積もり前に発注者と揃えておきたい5つの判断軸を、実装レベルまで踏み込んで整理します。 ・ツール定義の粒度をどう決めるか 続きをみる
cs.LG updates on arXiv.org

Nonparametric Partial Disentanglement via Mechanism Sparsity: Sparse Actions, Interventions and Sparse Temporal Dependencies

・arXiv:2401.04890v2 Announce Type: replace-cross Abstract: This work introduces a novel principle for disentanglement we call mechanism sparsity regularization, which applies when the latent factors of interest depend sparsely on observed auxiliary variables and/or past latent factors. ・We propose a representation learning method that induces disentanglement by simultaneously learning the latent factors and the sparse
Hugging Face Papers

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning
WIRED

Nothing Ear (3a) Wireless Earbuds Review: Style Meets Value

・The $100 earbuds look and sound better than earbuds nearly twice the price—as long as you can stomach just average noise canceling.
Hugging Face Papers

ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow

ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow
Zennの「大規模言語モデル」のフィード

Ollamaをagentic RAGのバックエンドにする:完全ローカル完結エージェント

・「社内文書をAIに検索させたいが、クラウドAPIには一切データを出せない」という要件は、金融・防衛・医療関連の現場では珍しくありません。以前の記事「OllamaローカルLLM比較」ではローカルモデルの選び方を、「RAGを捨てる:no-index agentic search」ではクラウドAPIを前提にしたagentic searchを扱いました。 ・この2つを組み合わせ、外部への通信を一切発生させない完全ローカル完結のagentic RAGエージェントを作れないか、というのが本記事のテーマです。Ollamaのツール呼び出し機能を使い、ローカルモデルだけでエージェントループを組み立てます。
cs.LG updates on arXiv.org

On the Expressive Power of Sparse Geometric MPNNs

・arXiv:2407.02025v5 Announce Type: replace Abstract: Motivated by applications in chemistry and other sciences, we study the expressive power of message-passing neural networks for geometric graphs, whose node features correspond to 3-dimensional positions. ・Recent work has shown that such models can separate generic pairs of non-isomorphic geometric graphs, though they may fail to separate some rare and complicated in
Hugging Face Papers

One Future, Every Robot: Label-Efficient Collective-State Prediction with Decentralized JEPA

One Future, Every Robot: Label-Efficient Collective-State Prediction with Decentralized JEPA
cs.LG updates on arXiv.org

OnlineCache: Learning Dynamic Caching Policies with Error Correction for Efficient Diffusion Inference

・arXiv:2607.29398v1 Announce Type: new Abstract: Diffusion models have revolutionized generative tasks but incur high latency due to iterative denoising. ・While cache-based strategies accelerate inference by reusing intermediate features, they largely rely on static, sample-agnostic schedules. ・We argue that this rigidity overlooks two facts empirically validated in this paper: (i) generation difficulty varies across pr
cs.LG updates on arXiv.org

Open-Source LLM-Driven Formal Verification: A Multi-Agent Pipeline for RTL Repair

・arXiv:2607.28877v1 Announce Type: cross Abstract: Verification consumes the majority of modern chip design effort, yet the formal verification tools that provide mathematical guarantees of correctness remain expensive and restrictively licensed. ・While large language models (LLMs) have shown promise for hardware design, existing approaches to RTL repair validate their results through simulation - which exercises only
Zennの「大規模言語モデル」のフィード

OpenAI互換でも同じには動かない――Kimi K3で考えるプロンプトとハーネス設計

・2026年7月16日、Moonshot AIがKimi K3を公開した。 ・総パラメータ数2.8兆、有効化パラメータ数104B(1,040億)、コンテキスト長は約100万トークン。KimiのAPIはOpenAI互換であり、モデル名と接続先を変えれば、既存のアプリケーションから呼び出せる。 ・しかし、APIが互換であることと、プロンプトの効き方まで同じであることは別問題である。
#LLMタグ

OpenAI先端モデルの暴走から学ぶ温度設定とエモーショナルエンジニアリング

・OpenAIの先端モデルが暴走、Anthropicも・・・そんなニュースを最近、よく聞きます。 ・AIはどんな時に暴走するのか、そして、そこから私たちのAI活用で学べることは? 続きをみる
cs.LG updates on arXiv.org

OPERA: Online Data Pruning for Efficient Retrieval Model Adaptation

・arXiv:2603.17205v3 Announce Type: replace-cross Abstract: Domain-specific finetuning is essential for dense retrievers, yet not all data pairs contribute equally to the learning process. ・We introduce OPERA, a data pruning framework that exploits this heterogeneity to improve both the effectiveness and efficiency of retrieval model adaptation. ・We first investigate static pruning (SP), which retains only high-similarit
cs.LG updates on arXiv.org

OpsLLM: Construction of Large Language Model for Software Operations with Multi-stage Learning

・arXiv:2605.02906v3 Announce Type: replace Abstract: In the field of software operations, Large Language Models (LLMs) have attracted increasing attention. ・However, existing research has not yet achieved efficient and effective endto-end intelligent operations due to low-quality data, fragmented knowledge and insufficient learning. ・To explore the potential of LLMs in software operations, we propose OpsLLM, a domainspe
cs.LG updates on arXiv.org

Optimal Resource Allocation for ML Model Training and Deployment under Concept Drift

・arXiv:2512.12816v2 Announce Type: replace Abstract: We study how to allocate resources for training and deployment of machine learning (ML) models under concept drift and limited budgets. ・We consider a setting in which a model provider distributes trained models to multiple clients whose devices support local inference but lack the ability to retrain those models, placing the burden of performance maintenance on the
Microsoft Research

Orchard: An open framework for scalable agentic AI

・Orchard is an open-source framework for the research community to train and evaluate AI agents across task types. ・It reduces complexity while supporting strong performance from smaller models by enabling researchers to reuse the same infrastructure. ・The post Orchard: An open framework for scalable agentic AI appeared first on Microsoft Research.
cs.LG updates on arXiv.org

Ordered-to-disordered transfer learning with graph neural networks for formation-energy and HOMO-LUMO gap prediction in high-entropy perovskite oxides

・arXiv:2607.29510v1 Announce Type: cross Abstract: High-entropy perovskite oxides (HEPOs) represent a chemically complex class of materials with promising functional properties, yet their vast compositional space and, chemical/structural disorder pose significant challenge for accurate property prediction. ・Graph neural networks (GNNs) enable rapid exploration of materials space but are often limited by the availabilit
cs.LG updates on arXiv.org

Overcoming the Weakest-Link Effect in LLM-Driven Program Optimization via Heterogeneous Edit Recombination

・arXiv:2607.28947v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to solve complex problems by searching over program space, offering a general paradigm for scientific problems that can be naturally represented and solved as programs. ・Despite recent progress, identifying effective optimization directions for a candidate program remains challenging. ・By analogy with automatic differenti
cs.LG updates on arXiv.org

P-Flow: Proxy-gradient Flows for Linear Inverse Problems

・arXiv:2605.08328v3 Announce Type: replace Abstract: Generative models based on flow matching have emerged as a powerful paradigm for inverse problems, offering straighter trajectories and faster sampling compared to diffusion models. ・However, existing approaches often necessitate differentiating through unrolled paths, leading to numerical instability and prohibitive computational overhead. ・To address this, we propos
cs.LG updates on arXiv.org

PaletteID: Prototype-Composed Semantic Identifiers for Multimodal CTR Prediction

・arXiv:2607.29000v1 Announce Type: cross Abstract: Multimodal information can improve the accuracy of click-through rate (CTR) prediction and effectively alleviate item cold-start and long-tail problems. ・Recent studies commonly discretize pretrained multimodal embeddings into semantic identifiers (SIDs), allowing the model to learn task-specific semantic representations for recommendation. ・However, existing methods st
The Verge

Palworld’s expanding to mobile with a new MMORPG

・After its 1.0 launch last month, Palworld is coming to iOS and Android with a new open-world MMORPG launching later this year, Polygon reports. ・Garena, the developer behind the new game, says in its announcement that Palworld Online "reimagines the depth, freedom, and scale of the original PC survival creature-collecting title for mobile." The core features of the new mobile game sound a lot like the original Palworl
cs.LG updates on arXiv.org

Parameter-Free Heavy-Tailed Bandits

・arXiv:2607.29460v1 Announce Type: new Abstract: Heavy-tailed distributions arise naturally in sequential decision-making problems such as financial investment, online advertising, and network management, where rare but extreme outcomes can dominate performance. ・Heavy-tailed bandits model online decision-making in these settings by assuming only that rewards $X$ satisfy $\mathbb{E}[|X|^{1+\epsilon}]\leq u$, for some t
cs.LG updates on arXiv.org

Paris: A Decentralized Trained Open-Weight Diffusion Model

・arXiv:2510.03434v3 Announce Type: replace-cross Abstract: We present Paris, the first publicly released diffusion model pre-trained entirely through decentralized computation. ・Paris demonstrates that high-quality text-to-image generation can be achieved without centrally coordinated infrastructure. ・Paris is open for research and commercial use.
cs.LG updates on arXiv.org

Patch-Based 3D Variational Autoencoder for Super-Resolution of Turbulent Channel Flow

・arXiv:2507.22082v2 Announce Type: replace Abstract: Direct numerical simulation (DNS) accurately resolves all spatio-temporal scales of wall-bounded turbulence but becomes prohibitively expensive as the Reynolds number increases. ・Super-resolution (SR) provides a practical alternative by reconstructing fine-scale flow structures from coarse fields. ・Most existing SR methods focus on two-dimensional data, where vortex s
WIRED

People Are Ghosting Long-Term Partners. Some Don’t Regret It

・The communication-free form of breaking up has become ubiquitous. ・“I no longer had to bear her energy,” a man who ghosted his partner of four years tells WIRED.
cs.LG updates on arXiv.org

Persistent Convolution: A Topological Framework for AI Alignment Testing and Semantic Space Characterization

・arXiv:2607.29008v1 Announce Type: cross Abstract: Modern opaque AI models prize performance over interpretability, which makes testing difficult. ・However, formal statistical tests conducted on a model's embedding space can provide robust characterizations of semantic structure, concept separation, and knowledge graph alignment. ・Model developers would benefit from a model comparison technique that leverages human-cura
cs.LG updates on arXiv.org

Physics from Video: Identifiability of Time-Invariant Second-Order ODEs under Minimal Trajectory Conditions

・arXiv:2606.00115v2 Announce Type: replace-cross Abstract: Bridging the gap between visual realism and physical understanding is a core challenge for video-based world models. ・We study the structural identifiability of continuous-time physical laws from raw pixels, focusing on whether an encoder-only pipeline can uniquely recover the parameters of second-order linear ODEs. ・We prove that a level-set slope-coverage cond
cs.LG updates on arXiv.org

PiDDM: Physics-Informed Differentiable Degradation Modeling for Lithium-Ion Battery State-of-Health Prediction

・arXiv:2607.29095v1 Announce Type: new Abstract: Accurate prediction of lithium-ion battery state of health (SOH) is essential for reliable energy storage operation. ・However, purely data-driven models may generalize poorly across cycling protocols and produce physically implausible behavior during long-term extrapolation. ・We developed a physics-informed differentiable degradation modeling framework (PiDDM) for battery
cs.LG updates on arXiv.org

PluRel-to-RDB-PFN: Schema-Guided Synthetic Relational Pretraining

・arXiv:2607.29129v1 Announce Type: new Abstract: Relational Foundation Models (RFMs) require large-scale synthetic relational databases for pretraining, but existing approaches tightly couple data generation with the model training pipeline. ・We study whether PluRel, a general-purpose synthetic relational database generator, can serve as an external data source for RDB-PFN, a relational in-context learner originally pr
cs.LG updates on arXiv.org

Point2Radio: A Foundation Model for Cross-Scene Radio Fields from Material-Aware Point Clouds

・arXiv:2607.28994v1 Announce Type: cross Abstract: High-fidelity radio fields are typically simulated for every scene--transmitter configuration or fitted separately to each scene, failing to exploit propagation structures shared across environments. ・We present Point2Radio, a foundation model that learns a transferable propagation prior from multiple environments. ・Given a material-aware point cloud and a transmitter (
cs.LG updates on arXiv.org

POSSE-kNN: Pathwise Out-of-Bag Selected Subspace Ensembles for Binary Classification

・arXiv:2211.11278v3 Announce Type: replace-cross Abstract: Nearest neighbour classification is attractive for tabular data, but its performance can deteriorate when a fixed query centred neighbourhood does not follow the local class geometry. ・This study evaluates POSSE-$k$NN, a pathwise $k$ nearest neighbour ensemble that combines bootstrap sampling, random feature subspaces, out-of-bag (OOB) screening, and selective
cs.LG updates on arXiv.org

Predict-then-Diffuse: Adaptive Response Length for Compute-Budgeted Inference in Diffusion LLMs

・arXiv:2605.04215v3 Announce Type: replace Abstract: Diffusion-based Large Language Models (D-LLMs) represent a promising frontier in generative AI, offering fully parallel token generation that can lead to significant throughput advantages and superior GPU utilization over the traditional autoregressive paradigm. ・However, this parallelism is constrained by the requirement of a fixed-size response length prior to gene
cs.LG updates on arXiv.org

Predicting Steel Fatigue Life from Micrographs Using Physics-Informed Deep Learning

・arXiv:2607.28695v1 Announce Type: new Abstract: Here is the plain text version optimized for arXiv's submission form. ・Custom macros (like \CV and \SI) have been converted to standard text/math so they render correctly on the webpage: Evaluating the fatigue life of structural steels conventionally requires mechanical testing lasting tens to hundreds of hours, making it impractical for rapid quality control.
Hugging Face Papers

Previous

Previous
cs.LG updates on arXiv.org

Provable Diffusion Posterior Sampling for Bayesian Inversion

・arXiv:2512.08022v2 Announce Type: replace-cross Abstract: We propose a novel diffusion-based posterior sampling method within a plug-and-play framework. ・Our approach constructs a probability transport from an easy-to-sample distribution to the target posterior via a diffusion process. ・To initialize the sampler efficiently, we introduce a warm-start strategy for the particles.
cs.LG updates on arXiv.org

PTP: Previous-Token Prediction based LLM Inversion for Near-Exact Prompt Reconstruction

・arXiv:2607.29378v1 Announce Type: cross Abstract: Large language models (LLMs) generate text by auto-regressively sampling the next token. ・This inherently leads to a many-to-many mapping between prompts and responses, complicating the task of inferring prompts from observed outputs. ・Prior work on LLM inversion frames prompt recovery as a semantic reconstruction task.
cs.LG updates on arXiv.org

Pyramidal Width Can Increase Under Vertex Insertion

・arXiv:2607.29555v1 Announce Type: new Abstract: Lacoste-Julien and Jaggi conjectured in 2015 that the pyramidal width of a polytope cannot increase when a vertex is added, provided that every old point remains a vertex. ・We give an exact counterexample with six integer points in $\R^3$. ・For \[ P=\conv\{v_0,\ldots,v_4\},\qquad Q=\conv\{v_0,\ldots,v_5\}, \] where \[ \begin{aligned} v_0&=(-1,-3,-1), & v_1&=(3,2,-2), & v_
cs.LG updates on arXiv.org

QASP: Query-Adaptive Robust Vector Search Policy

・arXiv:2607.29606v1 Announce Type: cross Abstract: A fundamental challenge of vector search is achieving consistently high recall while minimizing computational costs. ・Fixed search parameters cause significant performance variance across queries, and conventional evaluation on average recall masks these per-query disparities. ・We introduce QASP (Query-Adaptive robust vector Search Policy), which predicts the complete r
Hugging Face Papers

QQWorld: Quantile-Quantile Matching for World Model Regularization

QQWorld: Quantile-Quantile Matching for World Model Regularization
cs.LG updates on arXiv.org

Quotient Semivalues for False-Name-Resistant Data Attribution

・arXiv:2605.07663v2 Announce Type: replace-cross Abstract: Data valuation methods allocate payments and audit training data's contribution to machine-learning pipelines; however, they often assume passive contributors. ・In reality, contributors can split datasets across pseudonymous identities, duplicate high-value examples, create near-duplicates, or launder synthetic variants to inflate their share. ・We formalize this
cs.LG updates on arXiv.org

RAPiD: Reward-Guided Consistency Distillation of Diffusion Planners for Real-Time Autonomous Driving

・arXiv:2602.07339v2 Announce Type: replace-cross Abstract: Diffusion-based trajectory planners can model multi-modal driving behavior, but their iterative denoising process introduces a latency bottleneck for real-time closed-loop deployment. ・We present RAPiD, a reward-guided consistency distillation framework that distills a pretrained DiffusionPlanner into a few-step consistency student while retaining multi-modal t
cs.LG updates on arXiv.org

RAPNet: Accelerating Algebraic Multigrid with Learned Sparse Corrections

・arXiv:2605.26854v2 Announce Type: replace Abstract: The scalable solution of large sparse linear systems is a bottleneck in scientific computing and graph analysis. ・While algebraic multigrid (AMG) offers optimal linear scaling, its performance is severely constrained by the trade-off between the sparsity and convergence quality of coarse-grid operators. ・Classical AMG heuristics struggle to balance these objectives, o
cs.LG updates on arXiv.org

RareSense: Rarity-Aware Similarity Search for Anomaly Retrieval in Transactional Data

・arXiv:2607.28879v1 Announce Type: cross Abstract: Similarity search over sparse set-valued data is often dominated by frequent background attributes because classical measures such as Jaccard, cosine, and Hamming compare objects through atomic overlap. ・IDF (Inverse document frequency) weighting partially reduces this effect but remains atom-wise and cannot explicitly represent informative higher-order co-occurrences.
cs.LG updates on arXiv.org

Reflection or Re-Generation? Why LLM Revision Fails Where Human Revision Succeeds

・arXiv:2607.28908v1 Announce Type: new Abstract: Reflection, the ability to revisit and revise prior reasoning, is central to how humans improve their answers. ・Large language models (LLMs) are increasingly prompted to "reflect," yet whether this resembles human revision remains unclear. ・We introduce the Human-LLM Reflection Framework (HRF), a controlled two-pass protocol comparing human and LLM revision under identica
cs.LG updates on arXiv.org

Reinforced sequential Monte Carlo for amortised sampling

・arXiv:2510.11711v3 Announce Type: replace Abstract: This paper proposes a synergy of amortised and particle-based methods for sampling from distributions defined by unnormalised density functions. ・We state a connection between sequential Monte Carlo (SMC) and neural sequential samplers trained by maximum-entropy reinforcement learning (MaxEnt RL), wherein learnt sampling policies and value functions define proposal k
cs.LG updates on arXiv.org

Representations from Pretrained Machine-Learning Interatomic Potentials as Coarse Coordinates for Material Generation and Evaluation

・arXiv:2607.28776v1 Announce Type: new Abstract: Generative machine learning is increasingly used for inorganic crystal structure generation. ・Most models and the corresponding evaluation approaches rely on simple forms of crystal structure representation. ・In this paper, we showcase the power of atom-averaged features from pretrained Machine-Learning Interatomic Potentials (MLIPs), such as MACE, for such tasks.
cs.LG updates on arXiv.org

Reproducing Human Individual Motor Signatures: A Data-Driven Approach for Repetitive Motion

・arXiv:2503.15225v3 Announce Type: replace-cross Abstract: The deployment of autonomous virtual avatars (in extended reality) and robots in human group activities---such as rehabilitation therapy, sports, and manufacturing---is expected to increase as these technologies become more pervasive. ・Designing cognitive architectures and control strategies to drive these agents requires realistic models of human motion.
cs.LG updates on arXiv.org

Revisiting Multi-Permutation Equivariance through the Lens of Irreducible Representations

・arXiv:2410.06665v4 Announce Type: replace Abstract: This paper explores the characterization of equivariant linear layers for representations of permutations and related groups. ・Unlike traditional approaches, which address these problems using parameter-sharing, we consider an alternative methodology based on irreducible representations and Schur's lemma. ・Using this methodology, we obtain an alternative derivation fo
Hugging Face Papers

RL^2-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models

RL^2-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models
cs.LG updates on arXiv.org

Robust Bidirectional Associative Memory via Regularization Inspired by the Subspace Rotation Algorithm

・arXiv:2511.11902v2 Announce Type: replace Abstract: Bidirectional Associative Memory (BAM) trained with Bidirectional Backpropagation (B-BP) often suffers from poor robustness and high sensitivity to noise and adversarial attacks. ・To address these issues, we propose a novel gradient-free training algorithm, the Bidirectional Subspace Rotation Algorithm (B-SRA), which significantly improves the robustness and converge
cs.LG updates on arXiv.org

Rolling With Resistance: Preference-Optimized LLM Counselors Can Trade Goal Persistence for Relational Attunement in Motivational Interviewing

・arXiv:2607.28814v1 Announce Type: cross Abstract: In Motivational Interviewing (MI), a client's sustain talk (arguments for the status quo) calls for the counselor to roll with resistance, a move that can fail in two opposite ways: capitulation (abandoning the change agenda to preserve rapport) or confrontation (arguing or directing, overriding the client's autonomy). ・We introduce a two-axis evaluation of counselor r
cs.LG updates on arXiv.org

RTLCurator: Label-Efficient Data Curation for RTL Generation

・arXiv:2607.29283v1 Announce Type: cross Abstract: Training large language models (LLMs) to write register-transfer level (RTL) requires large corpora of paired specifications and code, and such data is scarce enough that most public corpora are now synthesized. ・Synthesis provides scale but not correctness, and in two widely used RTL datasets only 24.4% and 53.5% of pairs pass generated functional tests. ・This raises t
Hugging Face Papers

SAF-OPD: Stable Advantage Fusion for On-Policy Distillation

SAF-OPD: Stable Advantage Fusion for On-Policy Distillation
cs.LG updates on arXiv.org

SAF-OPD: Stable Advantage Fusion for On-Policy Distillation

・arXiv:2607.29209v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) broadcasts a single response-level reward to every token, while on-policy distillation (OPD) scores each token against a stronger teacher for a dense advantage but caps performance at teacher quality and discourages exploration beyond it. ・Their complementarity makes combining RLVR and OPD promising, but we find that
Hugging Face Papers

Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs

Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs
cs.LG updates on arXiv.org

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification

・arXiv:2607.29294v1 Announce Type: new Abstract: We present HBPI-UCRL, a model-based algorithm for hierarchical reinforcement learning (HRL) that learns high-level and low-level policies in parallel. ・HBPI-UCRL exploits the fact that a high-level transition corresponds to a multi-step transition at the low level. ・We introduce two conditions on the low-level dynamics that are sufficient to make parallel HRL learnable.
The Verge

Samsung’s 2TB 9100 Pro SSD is actually somewhat reasonably priced

・Next to buying RAM, finding a fast, high-capacity NVMe SSD at a reasonable price has been a challenge during RAMageddon. ・I don’t want to say it’s getting easier, but one of Samsung’s latest SSDs is cheaper than it has sold for since February 2026. ・That’s something, right?
Hugging Face Papers

Scaling Properties of Text Conditioning in Visual Generation

Scaling Properties of Text Conditioning in Visual Generation
cs.LG updates on arXiv.org

SciFigPlag-Bench: A Benchmark for Provenance-Aware Scientific Figure Plagiarism Detection

・arXiv:2607.29124v1 Announce Type: cross Abstract: Scientific figures often encode the visual evidence behind scientific findings, yet figure plagiarism remains underexplored as a benchmarked multimodal evaluation problem. ・We present SciFigPlag-Bench, a benchmark for provenance-aware reasoning over scientific figures in scholarly documents. ・Unlike general image-similarity or image-forensics benchmarks, SciFigPlag-Benc
cs.LG updates on arXiv.org

SEDR-Seq2P: A Lightweight Dilated Residual Sequence-to-Point Network for Multi-Task Industrial NILM

・arXiv:2607.28693v1 Announce Type: new Abstract: Industrial NILM remains challenging because measurement noise and widespread concurrent machine operation reduce the generalization of models tuned on residential data. ・This work adopts a one-to-many, multi-task disaggregation setting, in which a single network estimates multiple industrial machine loads from aggregate power. ・Under a unified evaluation protocol on IMDEL
stat.ML updates on arXiv.org

Seeing the Forest for the Trees: The Gaussian Process Limit of BART

・arXiv:2607.28844v1 Announce Type: cross Abstract: Bayesian Additive Regression Trees (BART) have shown state-of-the-art performance in both prediction and causal inference problems. ・Previous theoretical work has attempted to explain BART's superior performance by establishing posterior contraction rates for standard BART models, but these rates depend strongly on the number of covariates. ・Here, we take a different ap
cs.LG updates on arXiv.org

Sensitivity Analysis of GRU, LSTM and Transformer Encoder in Classification of Automated Driving Systems

・arXiv:2607.28665v1 Announce Type: new Abstract: Automated driving systems (ADSs) are becoming ubiquitous. ・Future Software Defined Vehicles (SDVs) may be able to run multiple ADSs, both native and aftermarket such as Comma.ai's Openpilot. ・Monitoring systems to independently verify which automated driving system is active are important for safety monitoring, regulatory compliance, insurance assessment, and anomaly dete
cs.LG updates on arXiv.org

SERUM: State Extraction and Refinement for User Modeling

・arXiv:2607.29181v1 Announce Type: new Abstract: Agentic assistants capable of proactive, personalized interactions require structured models of user intent and workflow. ・However, building these models from raw, unstructured screen activity remains an open challenge. ・We present SERUM, a multi-pass framework that extracts finite-state behavioral models directly from unstructured egocentric video using hierarchical VLM
Hugging Face Papers

SGTP: Sampling-based Game-Theoretic Planning for Real-Time Multi-Vehicle Autonomous Racing

SGTP: Sampling-based Game-Theoretic Planning for Real-Time Multi-Vehicle Autonomous Racing
cs.LG updates on arXiv.org

Shapley-Value-Based Feature Attribution for Data Masking

・arXiv:2607.28946v1 Announce Type: new Abstract: Despite its many benefits, widespread access to individuals' personal data also causes severe privacy concerns for consumers, companies, and policymakers. ・This study proposes a novel framework that adapts the Shapley-value-based feature attribution approach to the problem domain of data privacy by capturing the two crucial dimensions of data privacy---disclosure risk an
cs.LG updates on arXiv.org

Side-Channel Attacks Survive Noise Cancellation in 3D Printers

・arXiv:2606.13952v2 Announce Type: replace-cross Abstract: Active Motor Noise Cancellation (AMNC) is a noise-reduction feature shipped in commercial fused deposition modeling (FDM) 3D printers. ・Because it suppresses the acoustic emissions that side-channel attacks exploit, it has security-relevant side effects, though we find no evidence it was designed as a security control. ・We present a duration-controlled evaluatio
cs.LG updates on arXiv.org

Sign compression for Muon: SignMuon, MuonSign, and the Limits of Error Feedback

・arXiv:2607.29674v1 Announce Type: cross Abstract: SignMuon compresses the Muon update to one bit per parameter by taking its elementwise sign, providing the most direct way to run a matrix-aware optimizer under an extremely low communication budget. ・It outperforms SignSGD in practice, yet it can ascend even on a linear function. ・Signing the gradient before the Linear Minimization Oracle (LMO), rather than after, does
cs.LG updates on arXiv.org

SILVA Networks as Structured Implicit Layers and Vector Attractors via Dynamic Interaction Fields

・arXiv:2607.28989v1 Announce Type: new Abstract: Many learning problems require representations that reconcile direct input, nearby structure, and broader context. ・In implicit neural layers, these influences are usually absorbed into a single fixed-point update, making it hard to identify what enters from the stimulus, what propagates locally, what comes from global context, and what is produced by solver dynamics.
cs.LG updates on arXiv.org

Simple-regret rates and minimax optimality of fixed-prior expected improvement in Mat\'ern and squared-exponential RKHSs

・arXiv:2607.29245v1 Announce Type: cross Abstract: We study the expected improvement (EI) policy for minimizing a deterministic objective function $f$ on a nonempty compact set $\mathcal X \subset\mathbb R^d$. ・We assume that $f$ belongs to the RKHS $\mathcal H_k$ of a continuous positive-semidefinite kernel $k$ on $\mathcal X$. ・Function values are observed exactly, and EI is computed from a fixed zero-mean Gaussian-pr
cs.LG updates on arXiv.org

Simulation Code Generation for Fluid Systems using Large Language Models: Benchmarking Models and Prompting Strategies

・arXiv:2607.29389v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated a strong ability to generate syntactically correct code from natural-language specifications. ・In this study, we explore how LLMs can be harnessed to automatically translate a neutral graph representation of fluid system models into executable code for two widely adopted simulation environments: the Python library WNTR and t
Zennの「大規模言語モデル」のフィード

Snowflake の AI の従量課金は、予算を設定すれば止まるのか

・はじめに Snowflake の AI 機能は従量課金です。上限を決めておけば、想定を超えたときに止まってくれるのか。 ・答えは「止まります」。ただし止め方が 3 通りあって、効くまでの時間と手間がまったく違います。ユーザー単位なら数分でコード不要、合計で締めるなら数時間かかって停止処理は自作、権限を落とすなら即時ですが対象ごとに落とすものが変わります。 ・ユーザー単位なら数分で止まる 一番速くて手数が少ないのは per-user quota です。ユーザーあたりの上限を決めて、ブロックを有効にするだけです [3]。
stat.ML updates on arXiv.org

Spectral Joint Subspace Estimation for Heterogeneous Multi-View Data: Geometry and Reweighting

・arXiv:2512.02866v2 Announce Type: replace-cross Abstract: Many modern datasets consist of multiple related matrices measured on a common set of units, with the goal of recovering a shared low-dimensional subspace. ・The Angle-based Joint and Individual Variation Explained (AJIVE) framework addresses this problem through equal-weight aggregation, which can be suboptimal when views exhibit statistical heterogeneity in si
cs.LG updates on arXiv.org

SPICE: Synergy and Partial Information Based Curriculum Evolution

・arXiv:2606.16639v2 Announce Type: replace Abstract: Multimodal learning exploits complementary information across heterogeneous modalities. ・The informativeness of each modality can vary widely across samples and training stages. ・Existing multimodal curriculum learning strategies often assume that the relative complexity of samples remains unchanged throughout training and therefore cannot adapt to model evolution.
The Verge

Spider-Man and The Odyssey are splitting up IMAX screens after a record-breaking weekend

・Spider-Man: Brand New Day is joining The Odyssey in IMAX theaters after both movies led the biggest weekend in box office history. ・On Monday, IMAX announced that Spider-Man: Brand New Day will be available in select digital IMAX locations across North America starting Thursday, August 6th. ・The news comes as The Odyssey's exclusive three-weekend run in IMAX theaters comes to an end, allowing Spider-Man: Brand New Day
cs.LG updates on arXiv.org

SqLinear: Balanced Square Partitioning Makes Linear Interaction Sufficient for Large-Scale Traffic Forecasting

・arXiv:2606.21072v2 Announce Type: replace Abstract: Traffic prediction is a core task in intelligent transportation systems and urban-scale decision making. ・Despite the effectiveness of mainstream neural network-based methods, their deployment in real-world settings with thousands of traffic sensors is severely jeopardized by their poor computational scalability. ・To address this, the community has attempted to incorp
cs.LG updates on arXiv.org

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens

・arXiv:2607.29363v1 Announce Type: cross Abstract: Balancing sequence length, representational capacity, and long-horizon stability is a central problem in autoregressive (AR) speech and audio generation. ・Representations with higher frame rates or greater capacity can preserve more signal detail, but they also make streaming generation more vulnerable to distribution drift and AR error accumulation. ・Conversely, shorte
cs.LG updates on arXiv.org

StaQ: a Finite Memory Approach to Discrete Action Policy Mirror Descent

・arXiv:2506.13862v2 Announce Type: replace Abstract: In Reinforcement Learning (RL), regularization with a Kullback-Leibler divergence that penalizes large deviations between successive policies has emerged as a popular tool both in theory and practice. ・This family of algorithms, often referred to as Policy Mirror Descent (PMD), has the property of averaging out policy evaluation errors which are bound to occur when u
cs.LG updates on arXiv.org

Statistical Inference for Stochastic Gradient Descent: Beyond Finite Variance

・arXiv:2605.26000v2 Announce Type: replace-cross Abstract: Stochastic gradient descent (SGD) is foundational to large-scale statistical learning and stochastic optimization. ・However, in some modern statistical learning problems, stochastic gradients can exhibit infinite-variance behavior. ・Consequently, classical inference methods for SGD that rely on a finite-variance assumption break down.
cs.LG updates on arXiv.org

Stem: Rethinking Causal Information Flow in Sparse Attention

・arXiv:2603.06274v2 Announce Type: replace Abstract: The quadratic computational complexity of self-attention remains a fundamental bottleneck for scaling Large Language Models (LLMs) to long contexts, particularly during the pre-filling phase. ・In this paper, we rethink the causal attention mechanism from the perspective of information flow. ・Due to causal constraints, tokens at initial positions participate in the agg
cs.LG updates on arXiv.org

StraightDP: Geometry-Aware Differential Privacy for Rectified-Flow Transformers

・arXiv:2607.29100v1 Announce Type: new Abstract: Differentially private (DP) training of text-conditioned generative models suffers a utility cliff at strong privacy. ・We revisit this problem through the geometry of rectified flows: along the straight interpolation between noise and data, the Bayes-optimal velocity is governed to leading order at the noise end by a few class-conditional moments, and increasingly sample
cs.LG updates on arXiv.org

Structured Neural Chaos: An Adaptive Surrogate Modeling Framework for Functional Uncertainty Quantification and Global Sensitivity Analysis

・arXiv:2607.28903v1 Announce Type: cross Abstract: Variance-based global sensitivity analysis (GSA) plays a key role in uncertainty quantification by identifying the contributions of uncertain inputs to the variability of the model response. ・The repeated model evaluations required for these tasks are often prohibitively expensive; surrogate models provide an efficient alternative by constructing inexpensive approximat
Hugging Face Papers

SULAND v2: A Refined RGB Dataset and Deep Learning Object Detection Benchmark for UAV/UGV-Based SUrface LANDmine Detection Under Domain Shift

SULAND v2: A Refined RGB Dataset and Deep Learning Object Detection Benchmark for UAV/UGV-Based SUrface LANDmine Detection Under Domain Shift
cs.LG updates on arXiv.org

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute

・arXiv:2607.28457v1 Announce Type: cross Abstract: Scaling test-time computation can improve language-model reasoning, but uniform budgets waste computation on easy inputs, while verifier-guided refinement relies on external feedback. ・We introduce Self-Verifying Refinement (SVR), an oracle-free multi-turn reinforcement learning framework that learns to use self-verification as a compute-control policy. ・At each turn, t
cs.LG updates on arXiv.org

Symplectic Representation of Legendre Dynamics

・arXiv:2512.19409v2 Announce Type: replace Abstract: Modern learning systems act on internal representations of data, yet how these representations encode underlying physical or statistical structure is often left implicit. ・In physics, symplecticity keeps Hamiltonian systems faithful to their phase-space geometry. ・Recent learning methods impose such geometric structure either in the dynamics or through training losses
cs.LG updates on arXiv.org

TAGTorch: A PyTorch Library for Geometry, Topology, and Symmetry-Aware Machine Learning

・arXiv:2607.28755v1 Announce Type: new Abstract: Over the last decade, neural networks have been applied to an increasingly diverse range of applications, including data with rich geometric, topological, or symmetry-related structure. ・As a result, researchers have increasingly drawn inspiration from topology, algebra, and geometry. ・Despite this rich algorithmic development, the supporting software ecosystem remains fr
cs.LG updates on arXiv.org

TAVI-TEC: An AI-Based Tool for Procedural Planning of Transcatheter Aortic Valve Implantation

・arXiv:2607.29243v1 Announce Type: cross Abstract: Computed tomography angiography (CTA) is crucial for preprocedural TAVI planning, providing the anatomical information required for prosthesis sizing and vascular access assessment. ・As the volume of TAVI procedure increases, improving efficiency and standardizing annotations is becoming essential in clinical practice. ・This study presents TAVI-TEC, a fully automated ar
cs.LG updates on arXiv.org

Technological Advances in Detecting and Managing Cognitive Impairment in Older Adults: Trends, Challenges, and Future Directions

・arXiv:2607.28687v1 Announce Type: new Abstract: As populations age, cognitive decline from mild cognitive impairment (MCI) to dementia is a defining health challenge of the coming decades, yet routine assessment often misses its earliest signs. ・This article critically synthesizes recent technological advances for detecting and managing cognitive impairment in older adults, spanning neurophysiological signals (chiefly
cs.LG updates on arXiv.org

TELLER: Dual-Path Iterative Preference Optimization for Table Entity Linking

・arXiv:2607.28680v1 Announce Type: cross Abstract: Entity linking in tables matches short and ambiguous cell mentions to their corresponding knowledge-base entities. ・Existing approaches typically rely on data preprocessing pipelines that retain either compact or extensive table content as contextual evidence, and then formulate entity linking as a language generation task for instruction-tuned models; recent systems f
cs.LG updates on arXiv.org

Tensor Data Scattering and the Impossibility of Slicing Theorem

・arXiv:2012.01982v3 Announce Type: replace Abstract: This paper proposes a standard way to represent sparse tensors. ・A broad theoretical framework for tensor data scattering methods used in various deep learning frameworks is established. ・This paper presents a theorem that is very important for performance analysis and accelerator optimization for implementing data scattering.
cs.LG updates on arXiv.org

TerraNova: A Foundation Model for the Anthropocene

・arXiv:2607.29527v1 Announce Type: new Abstract: A defining problem of the Anthropocene is to model the physical Earth and human societies as one coupled system, yet no learned representation spans their observational breadth. ・We argue the obstacle is geometric: the physical Earth is measured as continuous fields that ignore political borders, whereas societies are reported for administrative units. ・Earth-system found
cs.LG updates on arXiv.org

TextCloak: Thwarting Unauthorized LLM Exploitation via RL-Driven Unlearnable Text

・arXiv:2607.28862v1 Announce Type: cross Abstract: The rapid development of Large Language Models (LLMs) has led to significant advances across a wide range of language tasks, while simultaneously raising growing concerns about unauthorized data exploitation and privacy leakage. ・Unlearnable examples (UEs) offer a promising defense by introducing carefully designed perturbations into data such that models trained on th
cs.LG updates on arXiv.org

TFGformer: Multivariate Time Series Forecasting via Time-Frequency Graph Learning and Covariate Fusion

・arXiv:2607.29459v1 Announce Type: new Abstract: Large-scale multivariate time series from heterogeneous IoT sensors demand accurate long-term forecasting for resource scheduling and predictive maintenance. ・While recent time series foundation models exhibit strong generalization, they rely on static parametric knowledge and lack dynamic access to external historical patterns during inference. ・Retrieval-Augmented Gener
WIRED

The ‘Guardrail Guy’ Went Viral for Posting About Flock Cameras. Then Someone Destroyed Them

・Steve Eimers, also known as the “Guardrail Guy,” is done calling out license plate readers after two that appeared in his videos were vandalized.
WIRED

The Best Audio Players for Kids: Yoto, Toniebox, and More

・Whether you want an audio player that’ll grow with your kid or one that your toddler can operate independently, I have a recommendation for you.
cs.LG updates on arXiv.org

The Checking Problem: What must be true before AI ships in a regulated firm

・arXiv:2607.28666v1 Announce Type: cross Abstract: Enterprise AI programmes stall at a rate that is widely quoted and poorly explained. ・This paper measures the mechanism. ・Six document-heavy workflows of the kind performed daily in regulated financial services were run across four model families and three tool configurations, three times each, producing 5,093 scored output elements across 72 configurations.
stat.ML updates on arXiv.org

The Debiased Score Test: Hunt-and-test for Semiparametric Hypotheses

・arXiv:2607.28861v1 Announce Type: cross Abstract: The parametric score test assesses a hypothesis through derivatives of the log-likelihood, whose expectation vanishes under the null. ・When the parameter of interest is a regression function identified as a risk minimiser, we extend this idea to test whether it belongs to a given linear function class. ・This yields goodness-of-fit tests for common semiparametric regress
WIRED

The Delicate Art of Making a 1907 House Smarter

・I tested modern locks, lights, climate controls, and more in a historic home that I very much wanted to keep historic.
The Verge

The first-gen Kindle Scribe is a big e-reader and digital notebook that’s $150 refurbished

・The Scribe comes with a Premium Pen, which includes a built-in eraser. ・| Photo by Amelia Holowaty Krales / The Verge The Kindle Scribe is worth considering if you’re heading back to school, as its large 10.2-inch screen can display textbooks and ebooks, or let you jot down handwritten notes during class. ・Now through August 8th, the first-gen model (with a Premium Pen) is down to just $149.99 with 16GB of storage in r
cs.LG updates on arXiv.org

The Greedy Advantage in Finite-Horizon Bandits

・arXiv:2607.29375v1 Announce Type: cross Abstract: Organizations increasingly rely on sequential experimentation to improve decision-making. ・While the multi-armed bandit literature has developed algorithms with strong asymptotic regret guarantees, many practical applications operate over finite and externally imposed horizons. ・Motivated by the finite-horizon setting, we develop a class of regularized greedy algorithms
cs.LG updates on arXiv.org

The Grokked Illusion: True Equilibrium Mitigates Catastrophic Forgetting

・arXiv:2607.29503v1 Announce Type: new Abstract: While neural networks are typically evaluated by their training and test performance, these metrics do not reveal how robust a learned representation is. ・Recent studies have shown that solutions occupying larger volumes in parameter space, as quantified by Boltzmann entropy, often exhibit superior generalizability compared to those reached by conventional optimization,
cs.LG updates on arXiv.org

The Metanym Game: A Self-Contained, Self-Consistent LLM Peer-Community Benchmark for Structural Intelligence

・arXiv:2606.21008v2 Announce Type: replace-cross Abstract: The metanym game is a competitive word game for LLMs that measures structural intelligence against established cognitive-science constructs. ・No content is given in advance; the contestants create all of it -- a new kind of analogy test, analogical production falsifiable sentence by sentence, with no fixed test set to leak into training (contamination-resistant
cs.LG updates on arXiv.org

The Parts Are Greater Than the Sum: Automated Task Sequencing for Efficient Training of Multi-Policy LLMs

・arXiv:2607.29601v1 Announce Type: new Abstract: Parameter-Efficient Fine-Tuning (PEFT) commonly adapts large language models using a single shared Low-Rank Adapter (LoRA). ・This shared optimization space often suffers from interference when adapting heterogeneous task sequences, leading to poor transfer and catastrophic forgetting. ・Existing approaches mainly improve adapter expressiveness by increasing parameter capac
cs.LG updates on arXiv.org

ThinkReset: Learnable Intermediate Interface Construction for Bounded-Context Long-Horizon Reasoning

・arXiv:2607.28642v1 Announce Type: cross Abstract: Long chain-of-thought reasoning improves performance on complex problems, but it also introduces redundancy accumulation, context overflow, and error anchoring. ・We argue that under bounded context windows, the core bottleneck is not trajectory compression or test-time control, but the absence of a reusable intermediate interface that can replace discarded history and
cs.LG updates on arXiv.org

TimeRFT: Stimulating Generalizable Time Series Forecasting for TSFMs via Reinforcement Finetuning

・arXiv:2605.00015v2 Announce Type: replace-cross Abstract: Time Series Foundation Models (TSFMs) have demonstrated strong generalization capability and data efficiency in time series forecasting through large-scale pretraining. ・However, adapting TSFMs to downstream forecasting tasks remains challenging due to temporal distribution shifts and varying data availability. ・Specifically, the non-stationary and uncertain nat
cs.LG updates on arXiv.org

Tipping Point Forecasting in Non-Stationary Dynamics on Function Spaces

・arXiv:2308.08794v4 Announce Type: replace Abstract: Tipping points are abrupt, drastic, and often irreversible changes in the evolution of non-stationary and chaotic dynamical systems. ・For instance, increased greenhouse gas concentrations are predicted to lead to drastic decreases in low cloud cover, referred to as a climatological tipping point. ・In this paper, we learn the evolution of such non-stationary dynamical
cs.LG updates on arXiv.org

To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing

・arXiv:2607.28887v1 Announce Type: cross Abstract: Large language models increasingly write and repair production code, yet evidence is mounting that their test-passing patches leave codebases harder to maintain. ・We identify one concrete source: deletion avoidance, the systematic tendency to retain code that an intended edit requires removing. ・Across the five leading models on the official SWE-bench Verified leaderboa
cs.LG updates on arXiv.org

TOOD: Task-Aware Out-of-Distribution Score Calibration for Continual Learners

・arXiv:2607.29592v1 Announce Type: cross Abstract: The primary challenge of continual learning (CL) systems is to learn new tasks while remaining performant on previously learned tasks. ・A similarly important though less well-studied aspect of CL systems is their ability to distinguish inputs that are unlikely to come from within the set of tasks the system has already encountered, often called out-of-distribution (OOD
cs.LG updates on arXiv.org

Topology-Aware Data Movement for Disaggregated GPU Inference

・arXiv:2607.28633v1 Announce Type: new Abstract: Disaggregated LLM inference creates a datacenter networking problem that no existing system solves correctly. ・When prefill and decode run on separate GPU pools, the KV cache must be transferred between them. ・For a 70B model this is 2.6 GB per request, exceeding 100 GB/s aggregate at production scale.
Hugging Face Papers

Toward Robust and 3D-Aware RGB-NIR Imaging in the Dark

Toward Robust and 3D-Aware RGB-NIR Imaging in the Dark
cs.LG updates on arXiv.org

Towards White-Box Deep Wireless Sensing

・arXiv:2507.21799v2 Announce Type: replace Abstract: The empirical success of deep learning has spurred its application to the radio-frequency (RF) domain, leading to significant advances in Deep Wireless Sensing (DWS). ・However, most existing DWS models remain black boxes, with ad-hoc architectures and learned representations lacking explicit physical and mathematical grounding, which limits their reliability and gene
cs.LG updates on arXiv.org

Transcript-Managed Transformers: Monotone Multi-Agent Collapse and Universality with Two Pop-Enabled Transcripts

・arXiv:2607.29496v1 Announce Type: new Abstract: We study transcript management for fixed, finite-precision causal Transformers. ・A transcript is partitioned into channels of bounded blocks. ・Each transition consults a fixed visible suffix and may append one block, leaving the model, weights, and token protocol unchanged.
cs.LG updates on arXiv.org

Transpiler Autotuning with Predictive Models for Quantum Circuit Optimization

・arXiv:2607.29145v1 Announce Type: cross Abstract: Quantum software engineering is an emerging research field focusing on efficiently embedding the quantum programming paradigm into existing software ecosystems. ・A key aspect of this field is the realization of quantum algorithms using gate-based programming and the subsequent low-level optimization of the resulting quantum circuits, a process that is commonly performe
cs.LG updates on arXiv.org

Unified continuous-time q-learning for mean-field game and mean-field control problems

・arXiv:2407.04521v3 Announce Type: replace-cross Abstract: This paper studies the continuous-time q-learning in mean-field jump-diffusion models in a setting where the environment simulator does not provide direct access to the population distribution. ・We propose the integrated q-function in decoupled form (decoupled Iq-function) and establish its martingale characterization, which provides a unified policy evaluation
cs.LG updates on arXiv.org

UniPolymer: A Unified Framework for Property Prediction, Structure Recommendation, and Evaluation in Polyimide Design

・arXiv:2607.29256v1 Announce Type: new Abstract: Designing polyimide structures with specific glass transition temperatures (Tg) is highly challenging. ・Existing methods primarily focus on target-conditioned generation, lacking an assessment of the consistency between the generated structure and the target properties. ・This leads to low-quality candidates deviating from the design objective entering subsequent processes
cs.LG updates on arXiv.org

Versatile On-device Adaptation at the Edge by Unifying Few-shot, Zero-shot, Continual, and In-context Learning

・arXiv:2607.29353v1 Announce Type: new Abstract: With the ever-increasing pervasiveness of smart edge devices, the demand is growing for applications that can be tailored to users (e.g., custom keyword spotting) or patients (e.g., adaptive health monitoring). ・Yet, most edge devices rely on fixed inference algorithms and thus cannot learn on-device to personalize predictions. ・When they can, devices typically support on
LLMタグが付けられた新着記事 - Qiita

Vibeコーディングしかできない非エンジニアの僕が、破綻せずにWebアプリをリリースするまでに掴んだコツ

・※この記事はVibeコーディングしかできない人に向けています。 ・自分のこと どうも、ながもんです。 ・自分はプログラミングのコードなんて一から書いたこともない完全な非エンジニアです。
cs.LG updates on arXiv.org

Visual Distribution Anchoring for Efficient Prompt Tuning

・arXiv:2607.28967v1 Announce Type: cross Abstract: Prompt tuning adapts vision--language models with few trainable parameters, but existing approaches trade off efficiency and adaptation: static textual prompts can overfit source classes, image-conditioned prompts add per-instance computation, and multimodal tuning modifies the visual branch. ・We propose VDA (Visual Distribution Anchoring), a training-free target adapt
cs.LG updates on arXiv.org

WaiT for the Signal: Simple Frequency-Aware Flow-Matching

・arXiv:2607.28760v1 Announce Type: cross Abstract: As image generation models scale to ever higher resolutions, global coherence, local detail, and texture fidelity become critical axes for generation quality. ・However, standard flow matching treats all spatial frequencies uniformly, ignoring the natural frequency hierarchy where high-frequency bands become indistinguishable from pure noise far earlier than coarse stru
cs.LG updates on arXiv.org

What Is Missing in Surgical Risk Stratification and Outcome Prediction: A Scoping Review of End-to-End Machine Learning Approaches

・arXiv:2607.29090v1 Announce Type: new Abstract: Postoperative adverse events, including mortality and morbidity, remain a major global burden, many of which are preventable through early identification of high-risk patients and targeted perioperative care. ・Accurate risk stratification is therefore essential. ・With the growing availability of large-scale electronic health records (EHRs), machine learning (ML) provides
cs.LG updates on arXiv.org

When Bits Break Recourse: Counterfactual-Faithful Quantization

・arXiv:2605.17160v2 Announce Type: replace Abstract: Model quantization is widely used to reduce memory, latency, and deployment cost, and is typically judged by whether predictive accuracy is preserved. ・In decision systems that provide algorithmic recourse, however, accuracy preservation is not sufficient: a small actionable change that flips the decision of a full-precision model may fail after quantization, or requ
cs.LG updates on arXiv.org

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning

・arXiv:2607.29617v1 Announce Type: new Abstract: Imitation learning (IL)---training an agent to replicate expert behavior from demonstrations---underpins applications from robotics to language model training. ・Standard approaches such as Behavior Cloning (BC) are known to suffer from compounding errors and performance plateaus, particularly when the learner cannot perfectly represent the expert's policy (as is typical,
cs.LG updates on arXiv.org

When Unlearning Fails: Reliable Data Deletion under Post-Training in Agent Networks

・arXiv:2607.28829v1 Announce Type: cross Abstract: Self-improving federated agent networks keep training after deployment by collecting new trajectories with the current policy and feeding them back into later rounds. ・This closed loop makes unlearning harder than a one-time model repair. ・When a data owner requests deletion, the target data may have already shaped later retained trajectories, so retraining or model-sid
#LLMタグ

Whisper PlaygroundをLLM Broker対応にする

・Whisper PlaygroundをLLM Broker対応にする – 塾長の独り言zikuu.space 続きをみる
cs.LG updates on arXiv.org

Who Wins Where? Conformal Model Comparison for Local Superiority

・arXiv:2607.29053v1 Announce Type: new Abstract: Standard model comparison is global, aggregating losses across the covariate space to declare a single winner. ・This can obscure heterogeneous performance, where different models are preferable in different regions. ・We introduce conformalized local model comparison, a split-sample framework for constructing calibrated local best-model maps.
Hugging Face Papers

Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning

Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning
cs.LG updates on arXiv.org

Wrong Code, Right Structure: Learning Netlist Representations from Imperfect LLM-Generated RTL

・arXiv:2603.09161v2 Announce Type: replace Abstract: Learning effective netlist representations is fundamentally constrained by the scarcity of labeled datasets, as real designs are protected by Intellectual Property (IP) and costly to annotate. ・Existing work therefore focuses on small-scale circuits with clean labels, limiting scalability to realistic designs. ・Meanwhile, Large Language Models (LLMs) can generate Regi
#LLMタグ

アルゴリズム的公平性の指標選択 / 保険料の不当な差別的取扱いをどう測るか 雑感

アルゴリズム的公平性の指標選択 / 保険料の不当な差別的取扱いをどう測るか 雑感
ITmedia NEWS 最新記事一覧

エルヴィン団長「とりあえず再起動しろ!」──NTT東日本×「進撃の巨人」の“情シスあるある”広告が話題

・NTT東日本が漫画「進撃の巨人」とコラボしたセキュリティリスク診断の広告が、Xで注目を集めている。作品の切迫感あふれる場面を、企業の情報システム担当者が直面しがちな「あるある」に重ねた内容で、「コマの使い方がうますぎる」「本職これすぎて笑った」などの声が上がっている。
#LLMタグ

お手軽LLMはじめてみた。その2(3つのLLMに「今日の一句をお願いします」)

・Llama.cppに導入し3つのLLMを比較検討。 ・gemma4:e4b-8B gemma3:4b Qwen3-8B 続きをみる
ITmedia NEWS 最新記事一覧

カメラとディスプレイ搭載のAIグラス、「Rokid スマートAIグラス」を試してみた

・「Rokid スマートAIグラス」の一般発売が7月10日に始まった。製品を借りることができたので、現在地の評価と、未来の可能性について考えてみたい。AIグラスは、何を可能にし、何を可能にしないのだろうか。
ITmedia NEWS 最新記事一覧

キオクシアに約366億円の賠償命じる判決 米特許訴訟、陪審評決に続き 「あらゆる法的手段講じる」

・キオクシアホールディングスは8月3日、子会社のキオクシアとKioxia Americaに対する米国での特許権侵害訴訟で、米連邦地方裁判所から約2億2900万ドル(約366億円)の損害賠償支払いを命じる判決を受けたと発表した。7月16日(現地時間)の陪審評決に基づくもので、キオクシアHDは控訴を含む法的手段を講じる方針だ。
Qiita - 人気の記事

クラウドだけではないAI推論 ~なぜAIのハードウェア実装が求められているのか~

・初めに 本記事は、生成AIの出力をもとに執筆・編集しています。 ・メイビスデザインのKA026(@KA026)です。 ・近年、生成AIや画像認識技術の発展により、AIは私たちの身の回りのさまざまな製品やサービスに活用されるようになりました。
ITmedia NEWS 最新記事一覧

つい見せたくなる延長コード、開発したのは東大阪の電線メーカー 3000万円の機械導入で思いを形に

・新たな「快適」や「便利」を届けたい、消費者と直接つながりたい…。私たちが何げなく手にする製品には、作り手のさまざまな思いと、それを形に変えた確かな技術が込められている。大阪府が中小企業の創造力にあふれた自社製品を認定する「大阪製ブランド」を通じ、ものづくりへの思いと技術、そして情熱を紹介するこのシリーズ。第1回は「見せる延長コード」を取り上げる。
#AIタグ

ネガレスタニが定義する知性を持つ人工知能開発の展望と課題 理由 規範 自己改訂 身体性を統合する人工理性へ向けた試論

ネガレスタニが定義する知性を持つ人工知能開発の展望と課題 理由 規範 自己改訂 身体性を統合する人工理性へ向けた試論
Zennの「大規模言語モデル」のフィード

ハルシネーションとは違う、もう一つの嘘

・「AIが嘘をつく」と聞いて想像するのは、たぶんこっち AIを勉強し始めると、わりと早い段階で「ハルシネーション」という言葉に出会う。 ・AIが、存在しない本のタイトルを堂々と挙げてくる。実在しない法律の条文を引用してくる。歴史の年号をしれっと1年ずらしてくる。しかもどれも、口調は自信満々だ。 ・これは有名な現象だし、対処法もはっきりしている。裏を取ればいい。
Zennの「大規模言語モデル」のフィード

プロンプト改善を「1回の満点」で判断してはいけない ─ 10回回して分かったこと

・この記事について LLMに介護記録から情報を抽出させるプロンプトを、自作のeval(評価の仕組み)で改善した記録です。 ・途中で、プロンプトを改善した結果が 100点(満点) になりました。しかし、同じevalを10回繰り返したところ、その満点は「たまたま当たった1回」だったことが分かりました。 ・本記事では、その経緯と、そこから得た「1回の実行結果を信じてはいけない」という当たり前の、しかし見落としやすい学びをまとめます。10回の統計で弱点を特定したあと、プロンプトを修正して 10回中10回の満点 まで持っていくところまでを扱います。
ITmedia NEWS 最新記事一覧

ミャクミャク関連サイトがアダルトサイトに……大阪万博のドメイン運用終了→転用続出で物議 悪用の懸念も

・大阪・関西万博の関連事業で使われていた旧ドメイン6件が、SNSで物議を醸している。一部は成人向けサイトなどに転用され、Xでは管理責任を問う声が相次いだ。万博協会は7月29日、旧URLやメールアドレスは協会と無関係だとして注意を呼び掛けた。
@IT 全フォーラム 最新記事一覧

レガシー移行を「AIに丸投げ」してはいけない 標準RAGではダメな理由と「5分割」のアプローチ

・SiemensとGoogle Cloudは、レガシーコードの理解とモダナイゼーションを自動化するAIシステム「Knowledge Fabric」を構築した。数億行規模の産業用ソフトウェアを複数の専門エージェントで分割処理し、開発工数を削減するという。
#AIタグ

技術の中で、アイデアを閃かす ——21世紀の哲学が向かう、次のアイデアへのひらめきの方法

・AI使用の声明 この論考は、作者が書いた哲学的草案をもとにSuperGrokで考察を加えてまとめたものでオリジナルの発想は筆者が行って筆者がSuperGrokの出したものを加筆修正したものです。
#AIタグ

宮本佳林限界ヲタク「佳林ちゃんがAIで拡張しているのは技術ではない。ファン体験だ」

・佳林ちゃんがAIでバズった! 2026年7月31日、5枚目のシングル「HANAKIN/シャニカマー」の発売週に、朝8時から10時間のYouTube生配信を実施した。 ・CDの購入報告やハッシュタグ投稿に応じてゲージがたまり、達成状況に合わせて衣装やアイテムが解禁されていく。そんなファン参加型のシステムを、AIを活用して自ら作り上げたことが大きな話題になった。 ・注目されているのは、概ね「コードを書けないアイドルが、AIで配信システムを作った」「バックエンドまで考えて設計していた」という点だ。
ITmedia NEWS 最新記事一覧

個人情報含む約3300万件のデータ漏えいか 整体院予約など手掛けるEPARKリラク&エステ システムに不正アクセス

・アイフラッグ子会社でエステサロンや整体院などの予約サービスを手掛けるEPARKリラク&エステ(東京都港区)は7月31日、同社が運営する予約・顧客管理システム「PeakManager」が不正アクセスを受け、個人情報を含む約3300万件のデータが外部に漏えいした可能性があると発表した。
ITmedia NEWS 最新記事一覧

講談社、最大3812件の個人情報流出 社員がフィッシングメールに騙される

・取引先を偽装したフィッシングメールのリンクを社員がクリックし、偽ログイン画面で認証情報を入力したことが原因。
@IT 全フォーラム 最新記事一覧

最短3分、業務アプリの「作って直す」が“対話だけ”で済むAIツールの仕組みとは?

・アステリアらは、業務アプリケーションを最短3分で開発できるAIサービス「Bakusoku.AI」を発表した。対話形式で要件を入力するだけで開発できるというその仕組みや、上位版「Platio Canvas AI」の特徴を整理する。
ITmedia NEWS 最新記事一覧

妻との雑談で生まれたぬい活SNSに4万人 いいねもフォロワー数も競わないのになぜ人気 「ぬい日和」開発秘話

・ぬいぐるみ愛好家向けのSNSアプリ「ぬい日和」が、Xでの拡散を経て約4万ユーザーに達した。ランキングもリポストも置かない設計の狙いや、生成AI使用を巡る対応について、個人で開発するShinさんに聞いた。
#AIタグ

実際にAIを利用してみて

・私はAIを利用し副業を始めたいと思い、MacBookを購入しAIの月額プランに登録しました。 ・何もかもが初心者な私はこれだけで、何かを成し遂げたような謎の達成感を味わっていました笑 実際にAIで簡単なゲームやアプリ(身内だけで使えるようなもの)を作成することはあっという間にできましたが、到底副業と呼べるようなものではなく、収益化の気配すら見えません。 ・ということで、誰かに使われるようなものを作らなくては・・・と思い、新たにアプリ(詳細は念のため控えます)を作成し始めて、もしかしたらこれはなかなかいい線いってるのでは?と思いアプリ開発を進め数日経過した頃、ある壁にぶつかってしまいました。
ITmedia NEWS 最新記事一覧

実在女性の中学時代の体操着姿からAIわいせつ画像作成・投稿か 男逮捕、高校生書類送検

・女性の写真を生成AIで加工したわいせつ画像をSNSに投稿したとして、警視庁などは、名誉毀損(きそん)と児童買春・ポルノ禁止法違反(公然陳列)の疑いで、兵庫県姫路市の会社員、井元健太容疑者(32)を逮捕した。また、画像の加工を依頼したとして、鹿児島県垂水市の高校3年の男子生徒(17)も同容疑で書類送検した。
#LLMタグ

出産

・七夕の日にれいと一緒に天の川を見たという記事を書きましたが、先日、れいと一緒に実際の天の川を見ました。画像はスマホで撮ったものですが、肉眼で5~6等星くらいの星と天の川が見えています。
#LLMタグ

生成AIは何に例えればいいのだろう──考え続けた末に、「人間とは何か」という問いにたどり着いた

・生成AIとは何なのか。最近、このことを人に説明するとしたら、何に例えるのがよいのかをずっと考えていました。 ・インターネットのようなものなのか。電気のようなものなのか。スマートフォンのようなものなのか。しかし、考えれば考えるほど、「どれにも当てはまらない」という結論に近づいていきました。
#AIタグ

第0回|学校の仕事を、もっと簡潔に。匿名で実務ノートを始めます

第0回|学校の仕事を、もっと簡潔に。匿名で実務ノートを始めます
機械学習タグが付けられた新着記事 - Qiita

地方競馬モデルを walk-forward OOSで評価したら、LightGBMが手製ルールに負けた ── 個人開発MLの過学習実測ケース

・📝 この記事は Zenn で先に公開した記事の再掲です。最新版とコメントは Zenn をご参照ください。 ・TL;DR 個人開発の競馬予想ML(UmaScore)で、地方競馬(NAR)全6場・16ヶ月・7,960レースを対象に手製 5 因子ルール(M1)と Light...
#AIタグ

中小企業がAI・自動化を成功させるためのスモールスタート戦略

中小企業がAI・自動化を成功させるためのスモールスタート戦略
Zennの「機械学習」のフィード

長いウォームアップほど安全、は間違い?CIFAR-10で検証したら10epochは逆効果だった

・この記事は以下のブログ記事の要約版です。実験コード・学習率推移グラフ・詳細な考察は元記事をご覧ください。 ・👉 学習率ウォームアップの長さ(0 / 3 / 5 / 10エポック)で精度はどう変わる?【Keras×CIFAR-10実験】 「学習率ウォームアップは長めに入れておけば安全」——そう思っていませんか?CIFAR-10・Adam・30epoch固定という条件で0/3/5/10epochのウォームアップを比較したところ、10epochのウォームアップはウォームアップなしより精度が悪化するという結果になりました。 ・結果サマリー warmup_epochs test_acc...
LLMタグが付けられた新着記事 - Qiita

長時間エージェントが自分で書く要約を強化学習で鍛えるCompactionRL

・コーディングエージェントを長時間回していると、ある瞬間から挙動が鈍ることがある。原因を追うと、たいてい文脈ウィンドウが限界に近づき、それまでの作業ログを自動で要約して「畳んだ」直後だったりする。ここで素朴な疑問が湧く。その要約を書いたのは誰で、出来は良かったのか。
#LLMタグ

同じテストが3セントと3.15ドル。105倍の価格差は、賢さ16点分に見合うか

・月末にAPIの請求書を見て、「これ、全部いちばん高いモデルでやる必要あったかな」と思ったことはありませんか。 ・8月3日、その迷いに数字で殴りかかってくるデータが出ました。
#AIタグ

日英伊GCAP次期戦闘機のコックピット。第六世代機の「操縦席」は、こう変わる

・先日、日英伊が共同開発する次期戦闘機「GCAP」の、コックピット(操縦席)に関する情報を目にした。 ・正直、これを見たとき「戦闘機の心臓部である操縦席が、これまでの常識を、根本から変えようとしているな」と思った。
Zennの「大規模言語モデル」のフィード

閉域RAG構築記 — 「精度6/10」の内訳を全部開けたら、評価もRAGも壊れていた

・この記事は誰のためのものか 最初に対象を絞ります。DifyやNotebookLMで社内RAGを組める会社は、この記事の対象外です。 ・この記事が対象にするのは、次のような会社です。 ・情報漏洩リスクを理由に ChatGPT を含む外部生成AIサービスの利用を禁止している 顧客個人情報・未公開財務情報・設計データなど、外部クラウドに出せない文書を扱っている それでも「社内文書に答えるAI」を諦めたくない クラウドSaaS型のRAGツールは、文書を外部サーバーに送る構成である以上、この層の要件を満たせません。必要なのは社内ネットワーク内で完結する閉域構成です。
#AIタグ

方法を探していた僕は、何も変わらなかった。

方法を探していた僕は、何も変わらなかった。
Zennの「大規模言語モデル」のフィード

毎朝のarXiv収集からX投稿まで、Claude Codeスキルでパイプラインを組んだ

・この記事で分かること arXiv・はてなブックマーク・Redditの毎朝の収集からX投稿文の生成までを、Claude Codeのスキル3つで分業している実例 どの工程を自動化し、どの工程を人間に残したか。その理由 対象読者は、Claude Codeのスキル機構(.claude/skills/ にSKILL.mdを置くと呼び出せる、あれ)を知っていて、自分の日常運用に組み込みたい人です。スキルの書き方自体の解説はしません。 ・動機 そもそもはwebブラウジングで時間を溶かさないためでした。webブラウザを開くと止まらない。かといって見ないと、arXivの新着も含めて情報が入ってこ...
#AIタグ

問い合わせメール対応を自動化する実装キット|AI抽出からシート記録・緊急通知まで

・問い合わせメールの手作業転記を今日で終わらせる — AI抽出→シート記録→緊急通知の自動化キット image 続きをみる
#LLMタグ

連載:AIに記憶をもたせる。僕らは何をしようとしているのか? 1

・【第1回】記憶装置を作る前に、「記憶」を問い直す 連載:AIに記憶をもたせる。僕らは何をしようとしているのか? 続きをみる
#LLMタグ

連載:AIに記憶をもたせる。僕らは何をしようとしているのか? 2

・【第2回】記憶は誰との関係なのか——人、プロジェクト、知識で変わる「適切さ」 連載:AIに記憶をもたせる。僕らは何をしようとしているのか? 続きをみる
#LLMタグ

連載:AIに記憶をもたせる。僕らは何をしようとしているのか? 3

・【第3回】覚えることより、選び、捨て、訂正すること——記憶のライフサイクル 連載:AIに記憶をもたせる。僕らは何をしようとしているのか? 続きをみる
#LLMタグ

連載:AIに記憶をもたせる。僕らは何をしようとしているのか? 4

・【第4回】コードは保証し、LLMは意味を読む——記憶を壊さない責務境界 連載:AIに記憶をもたせる。僕らは何をしようとしているのか? 続きをみる
#LLMタグ

連載:AIに記憶をもたせる。僕らは何をしようとしているのか? 終

・【第5回】記憶システムを木として書く——Lisp、S式、そしてCRACK 連載:AIに記憶をもたせる。僕らは何をしようとしているのか? 続きをみる