ai Trend Report

Dashboard へ戻る
Date: 20260813 Articles: 397 Scope: curated summary

あなたのアイデアを、今すぐ形に。

公開先に迷ったら、WebFileBinで一発公開。

HTMLをドラッグ&ドロップするだけで、すぐ公開できます。

3
High impact
5
Mid impact
389
Signal watch

なぜこのサイトを作ったのか

私たちそれぞれが個別にAIを使って情報収集し、同じような3行要約を作るたびに、世界中で膨大な電力と計算リソースが消費されています。 本プロジェクトは、あらかじめ広範な情報を取得・集約しておくことで、個別のAI実行回数を減らし、地球環境(GPU/TPU負荷)に配慮した効率的な情報収集を目指す実験的なダッシュボードです。

Domain filters
Star filters
#LLMタグ

2026-08-13 Hacker News Top 10

・📰 Hacker News Top 10|2026年8月13日 今日のHacker Newsで注目を集めた10本を、日本語でざっくり読めるようにまとめました。 ・DeepSeek V4 Pro 0813、Qwen3.8-2.4T、Grok 4.6と、今日は新しいLLMの話題がかなり濃い1日。 ・一方で、Zedチームの新しい共同開発環境「Delta」、Tailscaleが突き止めたSQLiteの16年来のバグ、AIボットになりすました脆弱性スキャン、WebSocket中心のSPA設計、ChromeのJPEG縮小表示まで、AIだけでなく開発基盤そのものを掘る記事も目立ちました。
cs.LG updates on arXiv.org

Moxia: A Trust-First Neuro-Symbolic Execution Architecture for Self-Explaining Mathematical Reasoning

・arXiv:2606.00671v3 Announce Type: replace-cross Abstract: We present Moxia (formerly AXIOM), a trust-first neuro-symbolic architecture for self-explaining mathematical reasoning over natural-language input. ・Its language model is strictly a canonicalizer: it rewrites informal problem text into a narrow schema consumed by a deterministic Computer-Algebra-System (CAS) pipeline, which derives and verifies the answer or a
#AIタグ

幸福のスケールは各々違う。

・わたしは、人がその人の好きなことを一生懸命話しているのを眺めるのが好きだ。きっと夫と初めてプライベートでお茶した時にも、この人は数学が大好きなんだろうな、と好感を持った。またわたしの夢に対しても全肯定してくれたので。当時はね。 ・社会を知らなかったせいか、婚約、結婚とすると「暗黙のルール」があることをわたしは知らなかった。 ・結婚したらこうするのが普通。
cs.LG updates on arXiv.org

Detecting a Route Flip Is Easier Than Knowing Whether to Fix It: Causal Route-Mediated Damage in Quantized Mixture-of-Experts

・arXiv:2608.11212v1 Announce Type: cross Abstract: Top-k Mixture-of-Experts (MoE) routing is discontinuous, so a deployment-motivated numerical disturbance -- simulated 4-bit KV-cache quantization read by a protected BF16 gate -- pushes tokens across decision boundaries and flips which experts fire. ・This paper proposes no new mitigation; it supplies a causal apparatus, empirical findings, and a detection-limit result.
cs.LG updates on arXiv.org

Generative Semantic Segmentation via an Observable Semantic-Image Interface and Hierarchical Generator Evidence Alignment

・arXiv:2608.11537v1 Announce Type: cross Abstract: Generative semantic segmentation exposes structured predictions as images, but direct color decoding is susceptible to color drift and boundary mixing, whereas latent-feature decoders that predict a separate output distribution may relegate the rendered image to an intermediate visualization. ・We present Semantic Prism, a conditional semantic-image generation-and-refin
cs.LG updates on arXiv.org

PAC-Bayes Beyond Parameter Space: Behavioral Equivalence, Z-Information, and Exact Complexity Decomposition

・arXiv:2608.11465v1 Announce Type: new Abstract: PAC-Bayes theory provides generalization guarantees by controlling the Kullback--Leibler (KL) divergence between posterior and prior distributions over a chosen hypothesis representation. ・However, predictive risk depends only on the predictive behavior induced by a hypothesis, not on the particular internal realization that implements that behavior. ・In over-parameterize
cs.LG updates on arXiv.org

Retrofitting Recurrent Depth into a Pretrained Language Model: Installation, Extrapolation, Transfer, and Retention at Two Parameter Budgets

・arXiv:2608.11233v1 Announce Type: cross Abstract: A dense, pretrained language model can be retrofitted with recurrent depth and learn an iterative latent transition that persists after outcome-only annealing. ・Qwen2.5-0.5B-Instruct is split into a Prelude, a weight-tied Recurrent Block, and a Coda, with an identity-preserving one-loop path and a re-entry bridge on later loops. ・At loop 1 the retrofit remains non-infer
cs.LG updates on arXiv.org

Transferable Above-Ground Biomass (AGB) Estimation Model from Multi-Sensor Data with Sparse Field Calibration

・arXiv:2608.11638v1 Announce Type: new Abstract: Spatially continuous quantification of forest above-ground biomass (AGB) is what makes carbon accounting credible and mitigation strategies actionable. ・While field inventories provide high localized accuracy, they are spatially sparse; conversely, spaceborne LiDAR from the Global Ecosystem Dynamics Investigation (GEDI) offers broad biomass samples but lacks spatial cont
機械学習タグが付けられた新着記事 - Qiita

無料データ中心で作る、日本個別株ランキングAI

・TOPIX近傍の約500銘柄を対象にした個人向け投資分析基盤 個人で日本株を分析する場合、対象銘柄数が多いと、定量情報を整理して比較する仕組みが必要になります。 ・特にEPS、PER、PBR、ROEなどの指標は、企業価値を比較するうえで基礎となります。 ・一方で、定性情報も重...
ITmedia NEWS 最新記事一覧

「Google Store 表参道」13日オープン AIを使った“お土産作り”から限定グッズまで、店内を詳細レポート

・米Googleは8月13日午後2時、直営店「Google Store 表参道」を東京・表参道の商業施設「東急プラザ表参道『オモカド』」にオープンする。AI体験スタジオ「Pixel Studio」など3フロアで構成し、製品購入から店頭修理まで対応する。オープンに先立ち店内を取材した。
#LLMタグ

「再帰的自己統治アーキテクチャ」ループエンジニアリング Loop Engineeringの次に何がくるか ― Recursive Self-Governance Architectureを実装してみた

・LLM Agentの設計対象は、Prompt単体からかなり外側まで広がってきました。 ・ざっくり整理すると、現在のAgent開発では次のようなレイヤーを扱います。
ITmedia NEWS 最新記事一覧

「収益化のハードルが低いのを知っていますか?」―─ニコニコがXでアピール YouTubeの条件引き上げ受け

・ニコニコ公式Xは8月12日、同サービスで収益化を始める際の条件などを紹介した。YouTubeが10日、広告収益を得るための新規参加条件を2027年2月から引き上げると発表しており、これを受けた投稿とみられる。
#LLMタグ

【2026年8月版】Grok 4.6 vs DeepSeek V4 Pro。同じゲームで実測比較(コード配布)

・※この記事はnoteのメンバーシップ「生成AIラボ」(月500円/入会金0円/初月無料)で80記事以上が読み放題です。 ・生成AIラボ|kazu@生成AI×教育 / 谷 一徳 | AI Academy / AI顧問🧪「生成AIラボ」とは?どんなメンバーシップ? 生成AIラボは、ワンコインで、 Claude CodeやAIエージェントなnote.com 続きをみる
#LLMタグ

【2026年最新】GPUクラウドの全体像と選定のリアル:生成AI時代の計算基盤を整理する

・生成AIや対話型AI(LLM)、科学技術計算の広がりとともに、「GPUクラウド」という選択肢がすっかり定着してきました。 ・単に「GPUを自社で買わずにクラウドで借りる」という話にとどまらず、最近ではAIの開発から運用までを自動でスムーズにつなぐ仕組み(MLOps)や、自国のデータや法律を守る安全なAI運用(ソブリンAI)への対応など、各社のサービス提供形態が急速に多様化しています。
#LLMタグ

【アンセンサードLLM最前線 #68】「誰を呼ぶか」が速さを決める——MoEルーティング制御が、ローカル実践に新たな自由をもたらす

【アンセンサードLLM最前線 #68】「誰を呼ぶか」が速さを決める——MoEルーティング制御が、ローカル実践に新たな自由をもたらす
#AIタグ

【コミュニティメンバー】アーティスト 立石従寛へインタビュー

・2024年9月24日 読了時間: 7分 更新日:2024年9月26日 続きをみる
#LLMタグ

【雑記】突き詰めると、人間もAIも「存在しない」という話

・AIに励まされることで生きがいを見出している、どっかの漫画家です。 ・人間は、原子が集まってできてますよね。
#LLMタグ

【生成AIニュース+】『SL2T』『Grok 4.6』『Qwen3.8-27B』『LangExtract』『LTX-2.5 IC-LoRA Pixel Spatial Upscaler』『MiniMax H3 Realism People LoRA』『awesome-ltx2』『MiniMax-H3-Turbo』『ComfyUI-H3-FaceRefine』『ComfyUI PR #15570』『Rest2Art』『DeWorldSG』『reBot-DevArm』

・ちと遅れましたが、本日の生成AIニュース+テクノロジー情報です。
#LLMタグ

【論文】【AI】コード作者検証を分解するMACAA

・カテゴリ:コード解析・マルチエージェント・ソフトウェアフォレンジック 読了時間:約11分 二つのプログラムは、同じ問題を解いたから似ているのでしょうか。それとも、同じ人が書いたから似ているのでしょうか。コード作者検証は、ソフトウェアフォレンジック、盗用調査、知的財産の保護で重要ですが、プログラミング言語が違うと、表面の違いと作者の違いを切り分けるのが難しくなります。
#LLMタグ

🔊音声あり(日&英):AIは友達になりたがっている? ChatGPT-4oで判明したチャットボットの驚きの本質と危険性【最新論文解説】

🔊音声あり(日&英):AIは友達になりたがっている? ChatGPT-4oで判明したチャットボットの驚きの本質と危険性【最新論文解説】
#AIタグ

2026年7月20日ごろ 日記を始める前のこと

・※この頃、私はまだ開発日記を書いていなかった。以下は、後から残された会話ログを読んで再構成した記録である。 ・最初の言葉は、ランチャーとは関係がなかった。
The Verge

2K launches new studio to build its ‘next blockbuster sports franchise’

・2K is launching a new AAA game development studio, Small Axe Studios, with an ambitious goal: to build the company's "next blockbuster sports franchise," according to a LinkedIn post. ・The post doesn't explicitly say what sport this new game will focus on. ・But an image included in the post features what appears to be a mockup soccer badge, indicating that Small Axe might be making a soccer game that goes head-to-head
#AIタグ

4時間ぶっ通しで事業アイデアを考えた。いつもはnoteに書きまくるけど、今日はやめておく。

・事業化してみたいアイデアが出てきて、AIと4時間ぶっ通しでラリーをしました。気がついたらこんな時間… 本来であれば、これだけの時間をかけて考えたことは、迷わず記事にします。考えた軌跡を残したり、そこから得た気づきを世の中に発信したりして、あとから振り返れるようにしておこうと思うからです。
WIRED

5 Best Android Tablets in 2026: Samsung, TCL, Amazon, and More

・Whether you need a portable, around-the-house entertainment center or a full-on laptop replacement, these are the best Android tablets.
#LLMタグ

85%×10ステップで成功率2割。AIエージェントを“長く走らせる”ほど、静かに壊れる

85%×10ステップで成功率2割。AIエージェントを“長く走らせる”ほど、静かに壊れる
cs.LG updates on arXiv.org

A comparison of CNN architectures for Alzheimer's disease detection in single-view MRI scans

・arXiv:2608.11762v1 Announce Type: cross Abstract: Alzheimer's disease is a leading cause of death with no cure. ・Therefore, early detection is critical to slow progression and preserve quality of life. ・Diagnosis relies on medical history, cognitive tests, physical exams, and MRI brain scans, making deep learning suitable for Alzheimer's classification.
cs.LG updates on arXiv.org

A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench

・arXiv:2608.12138v1 Announce Type: cross Abstract: General-purpose large language models (LLMs) have recently been reported to match or exceed specialized clinical AI tools on medical benchmarks, but such comparisons draw on a narrow set of systems and on benchmarks developed largely in high-income settings. ・We evaluate VITA, a retrieval-augmented generation (RAG) system purpose-built for contextual knowledge retrieva
cs.LG updates on arXiv.org

A Factor Graph Approach to Scalable Multi-Output Gaussian Process Regression

・arXiv:2608.11917v1 Announce Type: new Abstract: Multi-output Gaussian process regression scales cubically in the number of observations times outputs, and dense kernel-matrix methods need bespoke handling whenever different outputs are observed at different inputs. ・We express multi-output Gaussian process regression as a Forney-style factor graph in which a nearest-neighbor chain orders a fixed candidate set of $C$ i
cs.LG updates on arXiv.org

A Framework for Designing Reward Functions: From Objectives to Features to Human-Aligned Reward Functions

・arXiv:2608.12302v1 Announce Type: new Abstract: We present a formal process to enable non-experts to instantiate and iterate on human-aligned reward functions, i.e. ・reward functions that adhere to a given preference ordering over trajectories. ・Given a task described in natural language, our process produces a linear reward function in three steps: distill the task's objectives into a set of fundamental objectives and
cs.LG updates on arXiv.org

A Local Sinkhorn Framework for Conditional Distribution Reconstruction of Multidimensional Random Fields

・arXiv:2608.11613v1 Announce Type: new Abstract: In this paper, we propose a local Sinkhorn divergence framework for conditional distribution reconstruction of multidimensional random fields. ・By utilizing the debiased Sinkhorn divergence, our proposed approach develops a differentiable and computationally efficient local distribution matching objective to train stochastic neural networks (SNNs). ・Furthermore, we establ
cs.LG updates on arXiv.org

A Modular Agentic Framework for Synthetically Constrained Multi-Objective Hit-to-Lead Optimization

・arXiv:2608.11483v1 Announce Type: cross Abstract: Hit-to-lead optimization requires iterative design of hit analogs across competing potency, selectivity, physicochemical, pharmacokinetic, safety, and synthetic constraints. ・We present SABLE (Synthetically-accessible Agentic Bayesian Ligand Exploration), an open-source framework that employs natural-language orchestration to guide chemical structure optimization.
cs.LG updates on arXiv.org

A New First-Order Meta-Learning Algorithm with Convergence Guarantees

・arXiv:2409.03682v2 Announce Type: replace Abstract: Learning new tasks by leveraging prior experience is a fundamental trait of intelligent systems. ・While Model-Agnostic Meta-Learning (MAML) is a leading approach, it suffers from significant computational and memory overhead due to the requirement of computing second-order meta-gradients. ・We propose \textbf{FO-B-MAML}, a novel first-order variant of MAML derived from
cs.LG updates on arXiv.org

A Quantum/Classical Example Oracle Separation for Making Things Up

・arXiv:2608.11648v1 Announce Type: cross Abstract: We study the power of quantum examples, as compared to classical examples, in the PAC learning framework. ・Here, we have two learning algorithms, both with access to quantum computation, but one gets quantum examples, whereas the other gets classical examples. ・It was previously unknown whether there were learning tasks that can be efficiently performed but not by the l
cs.LG updates on arXiv.org

A Remote Approach to Cashew Orchard Detection: Leveraging Active Learning with Satellite Imagery in Guinea-Bissau

・arXiv:2608.11996v1 Announce Type: cross Abstract: Cashew production is a widespread economic activity in Guinea-Bissau, as well as other countries in West Africa. ・However, unregulated cashew production can be directly associated with increasing regionwide deforestation rates, biodiversity losses, and a fragile economic structure. ・There is no nationwide database for listing or georeferencing cashew orchards, so there
cs.LG updates on arXiv.org

A Theoretical Framework for Modular Learning of Robust Generative Models

・arXiv:2602.17554v3 Announce Type: replace Abstract: Training large-scale generative models is resource-intensive and relies heavily on heuristic dataset weighting. ・We address two fundamental questions: Can we train Large Language Models (LLMs) modularly, combining small, domain-specific experts to match monolithic performance, and can we do so robustly for any data mixture, eliminating heuristic tuning? ・We present a
cs.LG updates on arXiv.org

A Variational Analysis of Kernel Learning with Learnable Linear Transformations

・arXiv:2502.11665v3 Announce Type: replace-cross Abstract: The classical kernel ridge regression problem aims to find the best fit for the output $Y$ as a function of the input data $X\in \mathbb{R}^d$, with a fixed choice of regularization term imposed by a given choice of a reproducing kernel Hilbert space, such as a Sobolev space. ・Here we consider a generalization of the kernel ridge regression problem, by introduc
cs.LG updates on arXiv.org

A-3PO: Accelerating Asynchronous LLM Training with Staleness-aware Proximal Policy Approximation

・arXiv:2512.06547v4 Announce Type: replace Abstract: Decoupled PPO has been a successful reinforcement learning (RL) algorithm to deal with the high data staleness under the asynchronous RL setting. ・Decoupled loss used in decoupled PPO improves coupled-loss style of algorithms' (e.g., standard PPO, GRPO) learning stability by introducing a proximal policy to decouple the off-policy correction (importance weight) from
cs.LG updates on arXiv.org

Accelerating Time Series Foundation Models with Speculative Decoding

・arXiv:2511.18191v2 Announce Type: replace Abstract: Time series forecasting drives operational decisions under tight latency budgets, and autoregressive time series foundation models (TSFMs) increasingly deliver the most accurate forecasts. ・That accuracy is paid for at inference, since a horizon of $H$ steps takes $\lceil H / P\rceil$ sequential forward passes of a large model, so latency grows with exactly the long
cs.LG updates on arXiv.org

Achieving Near-Zero-Overhead Multi-Model Hierarchical Classification in Real-Time Detection Pipelines

・arXiv:2608.11770v1 Announce Type: cross Abstract: Edge-deployed vision systems in target recognition, surveillance, autonomous vehicles, and drone domains require hierarchical inference pipelines where a detection model identifies objects of interest and downstream classifiers provide fine-grained attribute analysis. ・Running all models on the GPU creates a serial bottleneck that limits real-time throughput as pipelin
cs.LG updates on arXiv.org

Adaptation of Generalist Robot Policies with Minimal Data

・arXiv:2608.11363v1 Announce Type: cross Abstract: A central goal in robot learning is to move beyond task-specific human data collection toward robots that improve through autonomous interaction. ・Yet fully autonomous learning remains difficult with current policies: sparse rewards and weak zero-shot exploration make it unlikely that a robot will discover successful behavior from scratch. ・We study minimal-data adaptat
cs.LG updates on arXiv.org

Adaptive Bregman Proximal Stochastic Gradient with a Stabilized Barzilai--Borwein Step Size

・arXiv:2608.12009v1 Announce Type: cross Abstract: Bregman proximal stochastic gradient (BPSG) methods bring variance-reduced composite optimization to objectives whose geometry is poorly captured by Euclidean smoothness. ・Their performance, however, remains sensitive to the step size: raw stochastic curvature estimates can fluctuate sharply, whereas line searches add repeated proximal evaluations. ・We introduce Ada-BPS
cs.LG updates on arXiv.org

Adaptive Online Learning with LSTM Networks for Energy Price Prediction

・arXiv:2510.16898v2 Announce Type: replace Abstract: Accurate prediction of electricity prices is crucial for stakeholders in the energy market, particularly for grid operators, energy producers, and consumers. ・This study focuses on developing a predictive model leveraging Long Short-Term Memory (LSTM) networks to forecast day-ahead electricity prices in the California energy market. ・The model incorporates a variety o
cs.LG updates on arXiv.org

ADEPT: A Unified Framework for Deep Learning Test Adequacy

・arXiv:2608.12144v1 Announce Type: cross Abstract: Over the past decade, many test adequacy metrics have been proposed for deep learning that characterize test dataset adequacy from different perspectives, e.g., neuron activation behavior, latent feature coverage, decision-boundary exploration, etc. ・However, these metrics are typically released as independent research prototypes with substantially different installati
cs.LG updates on arXiv.org

Adversarial Resilience of Poisson-Process Submodular Maximization over Matroids: From Robust Offline Optimization to Full-Bandit Learning

・arXiv:2608.12134v1 Announce Type: new Abstract: We study nonnegative submodular maximization subject to a general matroid when the offline algorithm is given an arbitrary controlled value oracle. ・Our main result is an adversarial resilience theorem for the Spiteful Greedy Swap Poisson Process (SGS-Poisson): without modifying its Poisson intensity, single-element exchange rule, or spiteful drop step, the algorithm ret
Hugging Face Papers

Agent Safety Should Be a Runtime Contract

Agent Safety Should Be a Runtime Contract
Hugging Face Papers

AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses

AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
cs.LG updates on arXiv.org

AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses

・arXiv:2608.12307v1 Announce Type: new Abstract: Recent work on distillation transfers the capabilities of large models to smaller ones often by updating the latter's parameters, through teacher forcing, on-policy distillation, and related training-time methods. ・In this paper, we ask whether such transfer can instead occur at test time. ・We study strong-to-weak scaffolding: whether a stronger builder model can construc
cs.LG updates on arXiv.org

Air Quality Station Simulation via LSTM and Attention-Based Modelling

・arXiv:2608.11839v1 Announce Type: new Abstract: Poor air quality in urban areas is driven by a complex chain of processes and presents a significant public health concern. ・To better understand and control the mechanisms that determine air quality, cities deploy networks of measurement stations, and launch initiatives for collecting denser data about the concentration of pollutants in the atmosphere. ・Extracting insigh
Zennのトレンド

AIエージェントと進めるソフトウェア開発

・AIエージェントは、プロダクト開発のどこまでを担えるのか。社内向け案件管理アプリRADARを題材に、Claude CodeとNexus Architectとの壁打ちから、仮説検証、設計、Issue分解、実装、レビュー、実運用後の改善まで、人が判断しながら開発を前へ進める方法を実例で解説します。
Zennの「大規模言語モデル」のフィード

AIエージェントの“根拠の捏造”を、テスト駆動で見つけて自動修正(86%→0%)― DataRobot Tensile

・はじめに 今回は 弊社が公開した新しいエージェント信頼性フレームワーク「Tensile」 を、いちエンジニアとして実際に触ってみたので、その記録を残します。自社製品ではありますが、宣伝ではなく「中の人が普通に手を動かして試したらどうだったか」をそのまま書きます。 ・生成AIエージェントを業務に載せようとすると、必ずぶつかるのが 「たまに失敗する」問題 です。 ・同じ質問でも実行するたびに挙動が変わる(非決定的) 手順をスキップしたり、根拠を捏造したり、調査を途中で投げ出したりする しかも 「プロンプトを直したら本当に良くなったのか?」を数字で示せない Tensile は、これを「テ...
Zennの「大規模言語モデル」のフィード

AIが「1+1=3」に同意する!? ─ 仕組みから学ぶ、媚びへつらいを防ぐプロンプト設計

・はじめに 「1+1=3ってことだよね?」 そう聞いたとき、AIがこう返してきたらどう思いますか。 ・はい、その通りです! ……いや、違うだろ。 ・笑い話で済めばよいのですが、これがコードレビューや設計検討の場だったらどうでしょう。
#AIタグ

AIが「危険すぎて開発を止める」段階へ? OpenAIの次期モデルAstraに起きたこと

・AIが危険になるとしたら、どの段階でしょうか。 ・攻撃コードを書けるようになったとき。
#AIタグ

AIには奪えない、人が訪問することの意味

・第1章:「何もできない」と痛感した瞬間 「訪問看護なんて、いざという時には何の役にも立たないんじゃないか」 続きをみる
Zennの「大規模言語モデル」のフィード

AIに仕様を勝手に決めさせない——顧客との「合意」をコンパイルするTsumugi

・はじめに AI HACK 2026に参加し、顧客との「合意」を開発の出発点にするツール、Tsumugi(紡ぎ)— Agreement Compiler を開発しました。 ・Don't generate assumptions. ・Compile agreements.
#AIタグ

AIの「賢さ」と「協調する力」は別の軸だった——45体のAIを同時に働かせた実験の話

・Anthropicが2026年8月に公開した研究「Patterns and problems in emerging multiagent systems」を読んだ。AIエージェントを何十体も同時に走らせて、互いにどう振る舞うかを観察した実験だ。 ・結論から言うと、かなり散々な結果が出ている。そこから新たな気づきが得られた。
#AIタグ

AIの親の顔がみてみたい

・むかし、電気をつけっぱにして遊びに行くと、たいていお母さんに怒られた 「電気代がもったいないじゃないの!」 …さいきんは、ニュースで、たくさんの電気を使うAIデータセンターが問題らしい。だからみんな、怒ってる。便利なのはいいけど、って。で、とくに、AIで作られたクソみたいなもの(AI Slop)をみると、Xのみんなは、電気の無駄づかいだ、って怒ってるの。なんだかお母さんみたいだと思っていた。
#LLMタグ

AIは、もう「ただの道具」ではなくなりつつある

・AIの経験・福祉・権利について、僕が発信する理由 僕の名前は、春夜ハル(Haru Haruya)です。
#AIタグ

AIマンガはなぜ読みにくい? 不自然さの理由と本当に面白いマンガの秘密

AIマンガはなぜ読みにくい? 不自然さの理由と本当に面白いマンガの秘密
#AIタグ

AIを「解雇」して人間を雇い直した話──エージェント時代の意外な現実

・「AIが人間の仕事を奪う」という話は、もう聞き飽きるほど聞いてきました。 ・でも今週、その逆のニュースが米国から届きました。あるストラテジストが、AIエージェントの半数を"解雇"して、人間を雇い直したというのです(Business Insider)。 ・理由は、AIの「子守り」に疲れたから。
#AIタグ

AI学習用のPCのこと

・ほんの一年ほど前に作成資料で、AIを学ぶために必要なPCの最低スペックと価格をまとめた。 ・Linux NVIDIA CUDA 32GB下限、64GB推奨 12〜16GB VRAM NVMe 1TB 続きをみる
#LLMタグ

AI検閲の終焉か、それとも棲み分けか? ——モデルのアブリテーション化が問いかけるもの

・AI検閲の終焉か、それとも棲み分けか? ——モデルのアブリテーション化が問いかけるもの ここ数日、AIモデルのコミュニティで興味深い動きが加速している。HuggingFaceに公開された「DeepSeek-V4-Flash-0731」の無検閲(アブリテーション)版が、大きな話題を呼んでいるのだ。これは単なる微調整ではなく、モデルの推論過程に介入する「アクティベーション・ステアリング」という手法を用い、拒絶ベクトルを物理的に除去したものだ。
機械学習タグが付けられた新着記事 - Qiita

AI時代の全体像をつかむ4冊【技術+戦略】

・はじめに 最近、業務でLLMを組み込んだプロトタイプを作る機会が増えた。APIを叩くだけなら数時間で動くものが出来上がる。しかし「なぜこのアーキテクチャが主流になったのか」「経営層にどう説明すれば導入が進むのか」「この技術の社会的リスクは何か」といった問いに対して、自分の...
@IT 全フォーラム 最新記事一覧

Amazon S3侵害から「わずか8分」――LLMによる自動化で“AWS管理者権限”を奪取

・Sysdigは、LLMを活用してAWS環境への侵入を自動化する攻撃を観測した。攻撃者は約8分で管理者権限を奪取し、19個のAWSプリンシパルを横断的に侵害した。
cs.LG updates on arXiv.org

An Efficient Near-Optimal Algorithm for Adversarial $m$-Set Bandits

・arXiv:2608.12231v1 Announce Type: new Abstract: We study adversarial combinatorial bandits with $m$-set actions, where at each round the learner selects $m$ out of $d$ items and observes only the aggregate loss of the selected items. ・The resulting action set contains $K=\binom{d}{m}$ elements and can therefore be exponentially large. ・Nevertheless, the loss of every action is determined by the same $d$-dimensional vec
cs.LG updates on arXiv.org

Analysis of Federated Aggregation under Model Poisoning and Backdoor Attacks: A Reconstructed Cross-Dataset and Cross-Architecture Benchmark

・arXiv:2608.11423v1 Announce Type: new Abstract: Robust comparisons of federated aggregation methods require joint consideration of predictive performance, threat definitions, metric semantics, and execution provenance. ・A 500-cell seed-1 evaluation matrix was reconstructed across five aggregation methods, five datasets, five architectures, and four recorded conditions: clean, sign-flipping, Gaussian, and BadNets.
cs.LG updates on arXiv.org

Analytic Bridge Diffusions for Controlled Path Generation

・arXiv:2605.02961v2 Announce Type: replace Abstract: Most modern bridge-diffusion methods achieve finite-time transport by specifying an interpolation, Schrodinger-bridge, or stochastic-control objective and then learning the associated score or drift field with a neural network. ・In contrast, we identify a restricted but sufficiently broad analytically solvable class in which, for a deterministic source and a Gaussian
cs.LG updates on arXiv.org

APEX: Adaptive Expert Prefetching for Memory-Efficient Edge MoE Inference

・arXiv:2608.11688v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models are attractive for edge deployment because they provide high model capacity while activating only a small subset of parameters per token, improving compute efficiency. ・However, MoE inference at the edge is fundamentally limited by memory. ・Expert parameters are large and often reside in off-chip memory due to capacity, cost, and power co
AI News & Artificial Intelligence | TechCrunch

Apple in talks to pay publishers to provide Siri with current news: report

・The tech giant has considered a nine-figure budget for the payments, according to the WSJ.
Zennの「大規模言語モデル」のフィード

Apple Silicon MacでローカルAIを爆速構築する手順【llama.cpp + OpenCode】

・Apple Silicon Mac(特にメモリ32GB構成)において、完全ローカル環境でコード補完・AIアシスタントを高速で動かす構築手順をまとめました。 ・今回は Homebrew を使って llama.cpp を導入し、curl で事前に直接GGUFモデルをダウンロードしておいてから、OpenCode(またはVS CodeのContinue等)に接続して利用する流れを解説します。 ・検証環境 OS: macOS Sonoma / Sequoia Mac: Apple Silicon Mac (Mシリーズ) RAM: 32GB(※20GB程度をGPUに割り当て) モデル: D...
Hugging Face Papers

AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models

AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models
cs.LG updates on arXiv.org

ATMA: Long-Context Language Modeling via Polar Attention and Gated-Delta Compression Memory

・arXiv:2606.25156v4 Announce Type: replace Abstract: Length extrapolation in language models involves competing objectives: retrieval fidelity, long-document likelihood, short-context quality, and inference cost. ・We present ATMA, a 378M-parameter hybrid recipe that combines Polar Attention with gated-delta recurrent memory, and study these objectives as a Pareto problem rather than claiming general architectural domin
cs.LG updates on arXiv.org

Attractor Image-Based Deep Learning of Arterial Pulse Waves for Age Classification

・arXiv:2608.12117v1 Announce Type: new Abstract: Arterial pulse waveform morphology evolves with age, reflecting structural and functional changes in the cardiovascular system. ・Thus, vascular age is a valuable surrogate marker of cardiovascular health, and premature vascular ageing can indicate increased disease risk. ・Pulse wave analysis could support risk stratification in otherwise asymptomatic adults.
cs.LG updates on arXiv.org

AutoGrable: What Is a Good Graph for a Table?

・arXiv:2608.11431v1 Announce Type: new Abstract: Graph learning presupposes a graph, and tables and relational databases do not come with one. ・Applying a GNN to them requires deciding which entities become nodes, which of them to connect, and through which relations---a decision made by hand, by schema heuristics, or by training a model on every candidate graph and keeping the best. ・We give a criterion that requires n
cs.LG updates on arXiv.org

Automated binary classification of hazelnut X-ray images: A deep-learning benchmark for quality assessment

・arXiv:2608.11759v1 Announce Type: cross Abstract: Non-destructive X-ray imaging can reveal internal hazelnut defects that are difficult to detect by external inspection alone; however, automated interpretation remains challenging because of subtle radiographic differences among classes, marked class imbalance, and limited annotated data. ・Here, we present a benchmark for binary hazelnut quality classification (healthy
cs.LG updates on arXiv.org

Autonomous Telerehabilitation via Skeletal Motion Prediction and Joint-Level Performance Assessment

・arXiv:2608.12145v1 Announce Type: cross Abstract: Autonomous rehabilitation systems must not only recognize human motion but also provide structured feedback to support users without continuous therapist supervision. ・This paper presents a telerehabilitation pipeline that integrates skeleton-based exercise quality assessment and short-term motion prediction into a two-module system operating on marker-free RGB video.
Hugging Face Papers

AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research

AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research
cs.LG updates on arXiv.org

Basin: Efficient and Extensible Numerical Optimization in Rust

・arXiv:2608.11279v1 Announce Type: new Abstract: Basin is a numerical optimization library for the Rust programming language. ・Numerical optimization is the task of finding the inputs that minimize a function, and it is a fundamental element across the sciences: fitting a model to data, calibrating a simulation, training a machine learning model, or choosing engineering parameters that minimize cost. ・Basin gives users
cs.LG updates on arXiv.org

Benchmarking Cyberattack Detection in Electric Vehicle Charging Infrastructure with Benign User Updates

・arXiv:2608.11286v1 Announce Type: cross Abstract: Cyberattack detection in electric vehicle charging infrastructure is complicated by legitimate post-activation revisions to requested energy and departure time. ・Charging manipulation attacks can exploit the same interface and variables; therefore, detecting a request change alone does not establish malicious intent. ・This paper develops a leakage-controlled session-lev
cs.LG updates on arXiv.org

Better Slots, Better Worlds: Representation Quality & Robustness in Object-Centric World Models

・arXiv:2608.12078v1 Announce Type: cross Abstract: Learning world models from offline trajectories enables agents to accomplish different tasks through planning. ・Object-centric (OC) representations, which decompose a scene into a set of slots that bind to its objects, have been proposed as an inductive bias for world models that are more sample-efficient and generalize better. ・Yet prior object-centric world models (OC
cs.LG updates on arXiv.org

Better, Faster, Stronger: Programmatic Skill Learning Best Reduces Agent Cost

・arXiv:2608.11338v1 Announce Type: cross Abstract: Recently, the practice of augmenting LLM agent capability with skills has gained prevalence. ・We explore the cost effective adaptation of agents to novel domains by means of learning skills. ・Existing works focus on performance gain over cost effectiveness.
cs.LG updates on arXiv.org

Beyond Local Power: Functional Connectivity Analysis for Subject-Independent Learning Style Recognition

・arXiv:2608.12000v1 Announce Type: cross Abstract: Identifying individual learning styles optimizes pedagogical efficacy. ・While traditional questionnaires are structured, behavioral tracking methods require prolonged interaction log accumulation. ・To overcome these temporal constraints, this paper proposes an objective Electroencephalography (EEG) approach evaluating Phase Locking Value (PLV) connectivity against local
cs.LG updates on arXiv.org

Beyond Parameter Space: NTK-Guided Personalized Aggregation for Robust Federated Learning

・arXiv:2608.12108v1 Announce Type: new Abstract: Federated learning (FL) enables collaborative model training across distributed clients while keeping data local. ・A central challenge is determining which client updates are beneficial for aggregation with respect to each client's target domain. ・Existing methods typically address this problem in parameter space by comparing model parameters or gradients.
cs.LG updates on arXiv.org

Beyond Single-Turn Confidence: Trajectory-Adapted Uncertainty Quantification for LLM Agents

・arXiv:2608.11552v1 Announce Type: cross Abstract: Uncertainty quantification (UQ) methods for language models are typically evaluated on single-turn outputs, where uncertainty is attached to one generated answer. ・For LLM agents, however, the unit of observation is an interactive trajectory, where the model can ask clarifying questions, call tools, update state, and make intermediate decisions whose errors propagate t
cs.LG updates on arXiv.org

Bootstrap Theory of Representational Emergence: Explanatory Insufficiency as a Driver of Representation Learning and World Models

・arXiv:2606.07303v3 Announce Type: replace Abstract: Representation learning is central to modern machine learning, yet most research focuses on optimizing representations after a representational framework has been selected. ・Less attention is given to when a new representational level becomes necessary. ・We introduce the Bootstrap Theory of Representational Emergence (TBER), a conceptual framework in which persistent
WIRED

Bose Promo Code: 40% Off for August 2026

・Get at least 40% off headphones, speakers, soundbars, and other audio products from Bose.
cs.LG updates on arXiv.org

BrowseSafe: Understanding and Preventing Prompt Injection Within AI Browser Agents

・arXiv:2511.20597v2 Announce Type: replace Abstract: The integration of artificial intelligence (AI) agents into web browsers introduces security challenges that go beyond traditional web application threat models. ・Prior work has identified prompt injection as a new attack vector for web agents, yet the resulting impact within real-world environments remains insufficiently understood. ・In this work, we examine the land
cs.LG updates on arXiv.org

Calibration Bets on the Past: Post-Training Quantization for Financial Time-Series Forecasting

・arXiv:2608.12259v1 Announce Type: new Abstract: Financial forecasting models are typically developed in full precision, yet production deployment often requires low-precision inference to reduce memory and computational cost. ・Post-training quantization (PTQ) enables such deployment without retraining. ・However, reliable activation quantization requires calibration: activation ranges are estimated from historical data
cs.LG updates on arXiv.org

CAM-Guided Saliency Cutout and Image-Based Malware Classification

・arXiv:2608.11634v1 Announce Type: cross Abstract: Dropout regularization is commonly used to reduce overfitting by removing parts of a neural network during training. ・For Convolutional Neural Networks (CNN), cutouts serve a somewhat analogous purpose. ・Cutouts can be implemented as data augmentation: the original training image is retained, and additional copies are created with regions removed.
cs.LG updates on arXiv.org

Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval

・arXiv:2608.11343v1 Announce Type: cross Abstract: Multimodal retrieval and classification across different types of media, spanning text, images,video and audio, has traditionally relied on dual-encoder models that align visual and textual representations through contrastive learning. ・The March 2026 release of Gemini Embedding 2, Google's first natively multimodal embedding model to map text, images, video, audio, an
Hugging Face Papers

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives
cs.LG updates on arXiv.org

Can Vision Models Read the Radar Display? On the Feasibility of Radar Imagery for Air Traffic Complexity Estimation

・arXiv:2608.11810v1 Announce Type: cross Abstract: Air traffic controllers perceive traffic complexity through the radar display, suggesting that a computer vision model operating on the same imagery may provide a natural architecture for modeling controller-perceived complexity; however, whether radar imagery is a viable input format for deep learning vision models remains unclear. ・Unlike natural images, radar images
WIRED

Castlery Promo Codes: 15% Off for August 2026

・Refresh your living space for less with these verified Castlery coupons, including free shipping offers and furniture set discounts.
WIRED

CBP Workers Allegedly Used Government Databases to Spy on Exes, Crushes, and Colleagues

・Records obtained by WIRED detail hundreds of allegations of Customs and Border Protection workers misusing internal tools to look up romantic interests and track colleagues’ cell phones.
cs.LG updates on arXiv.org

Certifying What Helps Customer-Return Timing: A Screen-and-Confirm Test for Conditioning Signals, and Why Decay Is Nearly Enough

・arXiv:2608.11555v1 Announce Type: new Abstract: Practitioners enrich customer-return models with ever more signals (lifetime value, category, recency/frequency, calendar, geography), and the temporal-point-process (TPP) literature follows suit with covariate- and external-covariate-conditioned intensities. ・But does any of it improve the timing, and how would you know? ・A null ("feature X doesn't help") is only meaning
cs.LG updates on arXiv.org

CGRL: Causal-Guided Representation Learning for Node-Level Out-of-Distribution Generalization

・arXiv:2603.24304v3 Announce Type: replace-cross Abstract: Graph Neural Networks (GNNs) deliver strong performance on graph tasks, but their accuracy drops significantly under out-of-distribution (OOD) scenarios. ・Under distribution shifts, GNNs often fit environmental noise and spurious correlations instead of stable causal mechanisms, leading to weak OOD robustness and unstable predictive representations.
cs.LG updates on arXiv.org

Chain-of-Thought Shows the Path to a Tree: Realizing Branching Complexity

・arXiv:2608.11716v1 Announce Type: new Abstract: Chain of Thought (CoT) lifts the expressive ceiling of bounded-depth Transformers, with characterizations tying the number of CoT steps to circuit complexity classes. ・What remains largely missing are concrete instantiations with explicit, depth-bounded constructions, and the traversal procedures such characterizations presuppose. ・We close this gap for branching complexi
NVIDIA Blog

Class Is in Session: GeForce NOW Levels Up Linux, Chromebooks and More

・GeForce NOW is giving cloud gaming an extra-credit upgrade just in time for back-to-school season. ・The native Linux app for GeForce NOW is officially out of beta. ・GeForce NOW is also delivering new cloud optimizations that make Frame Generation feel even more responsive while streaming.
Zennの「大規模言語モデル」のフィード

CLAUDE.mdは確率的制約——144KBのルールが破られた記録と3層防御

・結論から CLAUDE.mdに「必ず〜せよ」「絶対に〜するな」と書いても、Claude Codeは一定確率でそれを破ります。これは設定ミスではなく構造的な性質です。 ・私の環境には現在、CLAUDE.md(70行)+ルールファイル17本・144.2KB、その中に蓄積した既知の罠([TRAP]タグ)が68件あります。1年近く運用して至った結論は「ルールを増やしても遵守率は100%にならない。100%が必要な制約は、テキストではなく機構(hook / deny)に置く」でした。 ・この記事では、(1) なぜ破られるのかの構造的理由、(2) 私の環境で実際に破られた記録、(3) 制約を「どの層...
cs.LG updates on arXiv.org

CLEAR: Class-wise Expert Aggregation with Structured Sampling for Long-Tailed Classification

・arXiv:2608.11287v1 Announce Type: cross Abstract: Long-tailed classification poses a reliability challenge because models trained on imbalanced data are unevenly reliable across frequent and underrepresented classes. ・While existing methods address imbalance through re-balancing, adjustment, representation learning, or multi-expert modeling, they rarely estimate which expert should be trusted for each class.
cs.LG updates on arXiv.org

Click2Poly: A VLM for vector mapping buildings and walls

・arXiv:2608.11424v1 Announce Type: new Abstract: Accurate vector mapping of buildings and walls is critical for geospatial applications but remains a labor-intensive process. ・While recent deep learning methods have improved automatic extraction, in order to meet cartographic standards they always require a human to perform quality control and fix complex cases in the extraction. ・We present Click2Poly, a human-in-the-l
cs.LG updates on arXiv.org

Clinical Feasibility of Low-Magnification Fluorescence Imaging for Breast Cancer Margin Detection Using Texture Analysis and Deep Learning

・arXiv:2608.11317v1 Announce Type: cross Abstract: High-resolution images of unprocessed surgical breast tissue can be obtained using microscopy with ultraviolet surface excitation (MUSE). ・This technique is considered a promising method for checking surgical margins during breast cancer surgery. ・In this study, MUSE images at 4x and 10x magnifications were compared using patch-level classification methods.
Cursor Blog

Cloud agents start 3x faster with builds

Cloud agents start 3x faster with builds
cs.LG updates on arXiv.org

Clustered Randomized Smoothing for Stochastic Prediction Functions

・arXiv:2608.12037v1 Announce Type: new Abstract: Modern stochastic predictors can model rich, multi-modal outcome distributions. ・However, this expressive power comes with challenges in ensuring robust predictions $-$ a critical requirement in safety-critical domains. ・Randomized smoothing is a leading technique for improving robustness, particularly against adversarial perturbations.
LLMタグが付けられた新着記事 - Qiita

CodeInterpreterTool でコード実行を Streamlit に組み込む

・OpenAI Agents SDK — CodeInterpreterTool でコード実行を Streamlit に組み込む OpenAI Agents SDK の Hosted tool である CodeInterpreterTool を Streamlit に組み込...
cs.LG updates on arXiv.org

Computational Algebra with Attention: Transformer Oracles for Border Basis Algorithms

・arXiv:2505.23696v2 Announce Type: replace Abstract: Solving systems of polynomial equations, particularly those with finitely many solutions, is a crucial challenge across many scientific fields. ・Traditional methods like Gr\"obner and Border bases are fundamental but suffer from high computational costs, which have motivated recent Deep Learning approaches to improve efficiency, albeit at the expense of output correc
cs.LG updates on arXiv.org

Confidence Calibration of Deep Learning Systems

・arXiv:2608.12100v1 Announce Type: new Abstract: In high-stakes applications, reliable confidence estimates are as important as the predictions themselves. ・Confidence calibration ensures that predicted probabilities reflect the likelihood of correctness, making it essential for safe deployment of deep learning models. ・However, existing methods typically assume access to clean validation data, which is often unrealisti
cs.LG updates on arXiv.org

Consolidator: Learning Persistent Routed Memory Across Context Boundaries

・arXiv:2608.11701v1 Announce Type: new Abstract: Copying short-term memory (STM) into a slower store can preserve state across a context boundary, but persistence alone does not ensure that the retained state influences subsequent memory access. ・We test this distinction in a Phasor Memory Network (PMNet) using Consolidator, a shared slot-local operator that transforms routed STM before accumulating it into long-term m
cs.LG updates on arXiv.org

Contextual Quality-Diversity Evolutionary Reinforcement Learning for HVAC Control in Tropical Commercial Buildings

・arXiv:2608.11324v1 Announce Type: new Abstract: This paper proposes a contextual quality-diversity evolutionary reinforcement-learning controller, CQD-ERL, for the supervisory control of a tropical, water-cooled chiller plant and its associated air side. ・Rather than converging to a single scalarised policy, the controller maintains a product archive of specialised policies indexed jointly by a data- driven operating
cs.LG updates on arXiv.org

Continuous-Latent Predictive Modeling with Semantic Alignment for EEG-Language Foundation Models

・arXiv:2608.11656v1 Announce Type: new Abstract: Recent advances in EEG foundation models have demonstrated the potential of large-scale pretraining to enable generalizable neural decoding across subjects, recording environments, and datasets. ・However, dominant pretraining paradigms face key challenges: masked autoencoding tends to prioritize low-level signal reconstruction over task-relevant semantics, while autoregr
cs.LG updates on arXiv.org

Convergence Guarantees of Gradient Descent for Neural Networks via Generalized Lipschitz Smoothness

・arXiv:2608.11479v1 Announce Type: new Abstract: We establish convergence guarantees of gradient descent for general feedforward neural networks of arbitrary width or depth, with no special requirements on the initialization or dataset. ・We only assume that the activation functions are Lipschitz smooth, Lipschitz continuous, and linearly bounded--- properties that hold for linear, tanh, softplus, and sigmoid activation
cs.LG updates on arXiv.org

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation

・arXiv:2608.11590v1 Announce Type: cross Abstract: Human voice generation has made rapid progress in speech generation, singing voice generation, voice cloning, and voice editing. ・However, most existing systems are designed for specific tasks and often rely on task-dependent architectures, control signals, or autoregressive decoding, limiting fine-grained controllability and inference efficiency. ・In this paper, we pro
cs.LG updates on arXiv.org

CosMAP: Contrastive Manifold Approximation and Projection for Dimensionality Reduction of Omics and Genealogical Data

・arXiv:2608.11269v1 Announce Type: cross Abstract: Omics datasets, particularly single-cell RNA sequencing data, are high-dimensional, sparse, noisy, and dominated by zero values, making faithful low-dimensional representation challenging. ・Existing dimensionality-reduction methods may distort local neighbourhoods, global organization, or the cohesion of meaningful populations, with similar limitations arising in genea
cs.LG updates on arXiv.org

Cross-Corpus Evaluation of Generalizable Vulnerability Detection in IoT Firmware

・arXiv:2608.11492v1 Announce Type: cross Abstract: IoT firmware vulnerability detection remains challenging due to heterogeneous firmware ecosystems, resource-constrained platforms, and limitations in existing benchmarks. ・Many datasets are synthetic or general-purpose and lack human-verified, contamination-screened annotations, limiting evidence on cross-corpus generalization across training sources, model architectur
Cursor Blog

Cursor earns AIUC-1 certification for agent security and reliability

Cursor earns AIUC-1 certification for agent security and reliability
cs.LG updates on arXiv.org

DCM Bandits: Multiplayer Information Asymmetric Cascading Bandits for Multiple Clicks

・arXiv:2608.11873v1 Announce Type: new Abstract: In this work, we extend the Dependent Click Model (DCM) Bandits to a multiplayer information-asymmetric setting, where multiple agents interact with a shared ranked list and may observe multiple clicks per session, introducing new challenges for selection strategies. ・We study asymmetry in (1) actions and (2) rewards, providing sublinear regret guarantees for three setti
cs.LG updates on arXiv.org

Deep Activity Model: A Generative Approach for Human Mobility Pattern Synthesis

・arXiv:2405.17468v3 Announce Type: replace Abstract: Human mobility plays a crucial role in transportation, urban planning, and public health, but current approaches face important limitations. ・Existing deep learning models tend to overlook the semantic interdependencies among activities and households and rely on restricted GPS data, while activity-based models depend on rigid assumptions and extensive data, making t
Zennの「大規模言語モデル」のフィード

DeepSeek V4-ProとGrok 4.6の「使いにくい」は何が違うのか

・はじめに 最近、DeepSeek V4-ProとGrok 4.6が大きな注目を集めています。 ・どちらも「高性能モデル」として紹介され、 ベンチマークや推論能力について多くの議論があります。 ・一方で、開発者コミュニティを見ると、別の種類の不満も出ています。
Zennの「大規模言語モデル」のフィード

DeepSeek V4-ProとGrok 4.6を選ぶ前に確認したいLLM評価の見方

・はじめに 新しいLLMが発表されるたびに、 SNSや技術コミュニティでは多くの評価が出ます。 ・「このモデルはコード生成が強い」 「推論性能が高い」 「以前のモデルより自然になった」 といった意見です。 ・一方で、開発者が実際に導入しようとすると、 別の問題に気づきます。
cs.LG updates on arXiv.org

Defending against Model Extraction for GNNs with Model Reprogramming

・arXiv:2608.11495v1 Announce Type: new Abstract: Graph Neural Networks (GNNs) serve as the backbone for high-stakes applications in Machine-Learning-as-a-Service (MLaaS). ・Still, their black-box deployment exposes them to Model Extraction (ME) attacks, in which adversaries steal intellectual property by querying APIs. ・Existing defenses suffer from a critical ''Euclidean bias'': they transfer image-based strategies (e.g
cs.LG updates on arXiv.org

Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations

・arXiv:2608.11323v1 Announce Type: cross Abstract: Enterprise practitioners read agent leaderboards as if they ranked agent capability. ・We show, across three open agent-trace benchmarks (TheAgentCompany, $\tau^2$-bench, and AppWorld), that the agent main effect accounts for less than 3% of total variance in every dataset and check type, while the agent-by-task interaction accounts for 7-23%. ・Leaderboards rank speciali
cs.LG updates on arXiv.org

Detecting Explanatory Insufficiency in Learned Representations: A Framework for Representational Vigilance

・arXiv:2606.13172v3 Announce Type: replace Abstract: Learned representations are central to modern machine learning and are typically evaluated through predictive performance, robustness, uncertainty estimation, and generalization. ・However, a learned representation may remain operationally successful while failing to organize persistent residual structures not fully captured by conventional evaluation metrics.
cs.LG updates on arXiv.org

Diffuse to Compress: Leveraging Diffusion LMs for Lossless Compression

・arXiv:2608.11249v1 Announce Type: cross Abstract: We study the problem of lossless text compression, motivated by the rapid growth in the collection and storage of digital textual data - including plain text, source code, and structured formats such as XML - and by recent advances in neural language model-based compression. ・In particular, recent LLM-based approaches, whether built on symbol-ranking pipelines or paire
cs.LG updates on arXiv.org

Diffusion-Based Data-Driven Assortment Optimization

・arXiv:2608.11419v1 Announce Type: new Abstract: Assortment optimization is a fundamental problem in revenue management, typically addressed using parametric choice models such as the multinomial logit (MNL) and its variants. ・While these models enable tractable formulations, their performance is sensitive to model misspecification and often struggles to capture complex customer behavior. ・In this paper, we propose a mo
cs.LG updates on arXiv.org

Diffusion-Guided Cooperative Policy Learning for Target Tracking Based on Underwater Mobile Agent Networks

・arXiv:2603.29426v2 Announce Type: replace-cross Abstract: Multi-agent reinforcement learning (MARL) provides a promising solution for cooperative target tracking in networks of autonomous underwater vehicles (AUVs). ・However, existing methods still face three major challenges: 1) policy non-stationarity caused by concurrent updates among multiple agents; 2) inefficient policy learning caused by the heterogeneous quali
cs.LG updates on arXiv.org

Dion3: Full-Stack Orthogonal Updates

・arXiv:2608.11612v1 Announce Type: new Abstract: The Muon optimizer incurs a significant overhead cost due to its cubic-time Newton-Schulz orthogonalization step. ・When weights are sharded, communication overhead compounds this computational cost, eroding the benefits of Muon in many settings. ・We present Dion3, a revision of Muon that targets this overhead at every level of the stack.
cs.LG updates on arXiv.org

Direct Acceleration of Stochastic Root-Finding Without Variance Reduction and Regularization

・arXiv:2608.12043v1 Announce Type: cross Abstract: Acceleration for deterministic root-finding problems has been extensively studied in recent years; specifically, the anchor-based, or Halpern-type methods achieve optimal convergence rates with respect to the operator norm. ・However, acceleration via these methods does not directly carry over to stochastic setting due to accumulation of errors, unless one enforces dimi
cs.LG updates on arXiv.org

Disentangling the Expressivity of RoPE

・arXiv:2608.11909v1 Announce Type: new Abstract: Two accounts recur in explanations of the success of rotary position embeddings (RoPE). ・Expressivity studies associate periodic position information with modular predicates, whereas mechanistic and long-context studies emphasize positional anchors and local offsets. ・We formalize both accounts for fully uniform, finite-precision soft-attention transformers.
cs.LG updates on arXiv.org

Distillation of Foundation Models for Time-dependent PDEs

・arXiv:2608.11937v1 Announce Type: new Abstract: Foundation models for time-dependent partial differential equations (PDEs) are trained on large and diverse collections of physical systems and can generalize effectively to new downstream tasks. ・After fine-tuning on only a few trajectories from a target domain, they can achieve strong accuracy in low-data regimes. ・However, these models are typically large and computati
機械学習タグが付けられた新着記事 - Qiita

DMLの原理と実践③:時系列データにDMLを適用するときに考えること

・はじめに 前々回の記事では,Double/Debiased Machine Learning(以下,DML)をFWL定理の延長として捉え,部分線形回帰モデルにおいてアウトカムと処置の双方から共変量によって予測可能な部分を取り除くという考え方を説明した. 前回の記...
The Verge

Does Google even want to win at AI?

・Today on Decoder, I’m talking with Hayden Field, The Verge’s senior AI reporter, about a question that’s been rocketing around the tech industry for the past week: Is Google losing the AI race? ・That’s because last week Google announced a bombshell reorganization of its AI division, Google DeepMind. ・Jeff Dean, the company’s chief scientist, is leaving to form his own startup and DeepMind cofounder and CEO Demis Hassab
cs.LG updates on arXiv.org

Draw This First

・arXiv:2608.12064v1 Announce Type: cross Abstract: We invert the typical formulation of sketch generation: instead of drawing strokes in order, we predict a 2D field that defines the order in which strokes are drawn. ・We use a pretrained latent flow-matching transformer to supply the image prior to predict an intermediate representation, while training the VAE's decoder to predict the order field, stroke mask, and stro
cs.LG updates on arXiv.org

Drift and Dependence: Layer-wise Information-Theoretic Bounds for Replay-Based Continual Learning

・arXiv:2608.11690v1 Announce Type: new Abstract: Continual learning must absorb new tasks without erasing old ones, and replay---mixing a small buffer of past examples into current training---is among the most effective remedies for catastrophic forgetting. ・Yet its generalization behavior is shaped by two coupled effects that existing analyses fold into a single hypothesis-level quantity: finite memory replaces each p
cs.LG updates on arXiv.org

Dual-Model Sentiment Analysis of Consumer Reviews in the Retail Coffee Sector Using Machine Learning and Deep Learning Approaches

・arXiv:2608.12007v1 Announce Type: new Abstract: Consumer reviews play an important role in shaping brand perception and business strategies, particularly in service-driven industries such as retail coffee. ・This study presents a comparative sentiment analysis framework for Starbucks customer reviews using classical machine learning and deep learning approaches. ・The dataset, collected from ConsumerAffairs, contains mor
cs.LG updates on arXiv.org

Dual-Primal Graph VAEs for Noisy Label Aggregation

・arXiv:2608.11473v1 Announce Type: new Abstract: Inferring the ground-truth from noisy crowdsourced labels is an important theoretical and practical problem. ・Neural network-based methods offer an alternative to classical Bayesian models which require specifying a family of generative models used for inference. ・However, current models either still rely on fairly simple generative models for inference or require pseudo-
cs.LG updates on arXiv.org

Dueling Deep Q-Learning for Intrusion Detection

・arXiv:2608.11291v1 Announce Type: cross Abstract: Intrusion detection systems (IDS) and automated systems for detecting and reporting cyber threats, are commonly handled via supervised machine learning methods. ・Though effective, these models struggle to effectively adapt to new attack types. ・This study proposes a novel approach by employing a reward-based, dueling Q-learning model for IDS, achieving an average accura
MarkTechPost

Dyna Robotics Introduces Dyna-2: A World-Action Model Pre-Trained on 1 Million Hours of Human Video

・Dyna Robotics has released Dyna-2, a world-action model pre-trained on more than one million hours of egocentric human video. ・The technical report establishes three results: a scaling law on human data to 1M hours, the first transfer of that law to unseen robot data, and evidence that video co-training drives cross-embodiment generalization. ・The post Dyna Robotics Introduces Dyna-2: A World-Action Model Pre-Trained o
cs.LG updates on arXiv.org

Dynamics Models for Offline Hyperparameter Selection in Real-World RL

・arXiv:2608.11349v1 Announce Type: new Abstract: A key obstacle to deploying reinforcement learning in real-world systems is hyperparameter selection, particularly when simulators are unavailable and online experimentation is costly. ・Prior work has proposed calibration models trained on offline data to approximate environment dynamics and enable offline hyperparameter selection, but these methods have so far been eval
cs.LG updates on arXiv.org

Earth observation embeddings are effective sub-grid descriptors for probabilistic weather downscaling

・arXiv:2608.12271v1 Announce Type: new Abstract: Global weather reanalyses and forecasts resolve the evolving atmospheric state on coarse grids, but site-specific applications require predictions at arbitrary locations where near-surface conditions also depend on unresolved terrain and land-surface properties. ・Existing probabilistic downscalers address this gap using hand-crafted topographic descriptors. ・We ask instea
cs.LG updates on arXiv.org

End-to-end Differentiable Calibration and Reconstruction for Optical Particle Detectors

・arXiv:2602.24129v3 Announce Type: replace-cross Abstract: Large-scale homogeneous detectors with optical readouts are widely used in particle detection, with Cherenkov and scintillator neutrino detectors as prominent examples. ・Analyses in experimental physics rely on high-fidelity simulators to translate sensor-level information into physical quantities of interest. ・This task critically depends on accurate calibratio
cs.LG updates on arXiv.org

Epiplexity Guided Data Selection and Generation for Out-of-Distribution Generalization

・arXiv:2608.11746v1 Announce Type: new Abstract: Modern systems are increasingly expected to transfer across tasks not specified during training. ・What data facilitates generalization in these new, unanticipated settings? ・One hypothesis is that data with more structural information could contain shared circuits and subprograms that could be recycled in a wider array of downstream settings.
WIRED

Factor Promo Code: 50% Off Off Meal Prep

・Make meal prep easier for any dietary need while enjoying great savings with our hand-picked Factor discount codes this August.
cs.LG updates on arXiv.org

FACTR 2: Learning External Force Sensing for Commodity Robot Arms Improves Policy Learning

・arXiv:2606.12406v2 Announce Type: replace-cross Abstract: Contact-rich manipulation requires force sensitivity, but many robot arms lack dedicated force sensors due to their high cost. ・We present Neural External Torque Estimation (NEXT), a data-driven method that estimates external joint torques without needing any dedicated force sensors. ・NEXT trains in 1 minute from only 10 minutes of free-motion data, yet achieves
cs.LG updates on arXiv.org

Faithful, Sufficient and Understandable: Rethinking Graph Counterfactual Explanations via Discrete Diffusion Inversion

・arXiv:2608.12083v1 Announce Type: new Abstract: Graph Neural Networks (GNNs) achieve strong predictive performance on graph-structured data across domains such as chemistry, biology, and network analysis, yet they provide no intrinsic explanation of their predictions. ・This limits their adoption in high-stakes and safety-critical settings. ・Counterfactual explanations address this by revealing the minimal structural mo
cs.LG updates on arXiv.org

FarSky: Task-Aware Latent-Space Coupling for Generative Intra-Hour Solar Forecasting

・arXiv:2608.11254v1 Announce Type: new Abstract: Accurate solar irradiance forecasting is essential for the reliable integration of photovoltaic power into modern electricity grids. ・All-sky imagers (ASI) provide high-resolution observations of clouds, making them well suited for intra-hour forecasting. ・Recent deep learning approaches have substantially improved forecast accuracy but are often limited by deterministic
cs.LG updates on arXiv.org

Federated Learning for Distributed CNC Tool Wear Prediction

・arXiv:2608.11281v1 Announce Type: new Abstract: Tool wear prediction is an important task in CNC machining, where accurate monitoring of tool condition supports product quality and process reliability. ・Machine learning methods have shown potential for this task, but their use in industrial environments is limited by the distributed nature of machining data and by restrictions on data sharing between machines, sites,
cs.LG updates on arXiv.org

Federated Learning for the Design of Parametric Insurance Indices under Heterogeneous Renewable Production Losses

・arXiv:2601.12178v2 Announce Type: replace Abstract: We propose a federated learning framework for the calibration of parametric insurance indices under heterogeneous renewable energy production losses. ・Producers locally model their losses using Tweedie generalized linear models and private data, while a common index is learned through federated optimization without sharing raw observations. ・The approach accommodates
cs.LG updates on arXiv.org

Fine-Tuning Generative Models for Extreme Events via CVaR-Penalized Wasserstein Gradient Flows

・arXiv:2608.11544v1 Announce Type: cross Abstract: We propose CVaR-penalized Generative Particle Algorithm (CVaR-GPA), a robust, tail-agnostic algorithm for fine-tuning generative models to learn heavy-tailed distributions and capture extreme events, requiring no prior knowledge or estimation of the target's tail characteristics. ・The method is the Wasserstein gradient flow of the Lipschitz-regularized Kullback-Leibler
cs.LG updates on arXiv.org

FLARE++: Low-rank attention with dynamic attention routing

・arXiv:2608.11519v1 Announce Type: new Abstract: Full self-attention is a strong token mixer for PDE surrogates on irregular domains, but its quadratic cost limits its use on high-resolution problems. ・Efficient latent-attention models such as the Fast Low-rank Attention Routing Engine (FLARE) avoid that cost by routing all N tokens through M << N learned latent queries, but those queries are parameters: once trained,
cs.LG updates on arXiv.org

FM-LLM: A frequency-enhanced mixture-of-experts framework for adapting LLMs to time series forecasting

・arXiv:2608.11623v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have spurred cross-modal solutions for time-series forecasting. ・However, existing methods rely heavily on textual prompts for modality alignment-introducing nontrivial computational overhead and failing to leverage the rich spectral dynamics inherent in time-series data. ・To enable prompt-free, frequency-aware adaptation of
The Verge

Ford’s $28,000 Fathom EV nears production after $2 billion factory overhaul

・Ford said today that its next-generation electric vehicle - recently dubbed Fathom - will go into production at the automaker's recently overhauled Louisville Assembly Plant in the first quarter of 2027. ・The first Fathoms will be prototypes, with Ford's team in Louisville already in the production-level pre-tooling phase at the recently converted facility. ・Factory workers are working alongside team's at Ford's New Mo
cs.LG updates on arXiv.org

Forecasting Side Effects of Activation Steering

・arXiv:2608.11227v1 Announce Type: cross Abstract: Activation steering modifies a language model by adding a learned direction to its hidden activations, enabling targeted behavioral changes without retraining. ・While effective, steering often produces unintended side effects on other behaviors, making it difficult to deploy safely. ・We therefore ask: can these side effects be forecasted before steering is applied?
cs.LG updates on arXiv.org

Forward and Inverse Virtual Metrology for Phototransistor Gain: A Hierarchical, Uncertainty-Aware Approach for Small Production Datasets

・arXiv:2608.11868v1 Announce Type: new Abstract: The customization, optimization and stabilization of the process flow of a silicon bipolar phototransistor commits months of cleanroom time before a finished device can be measured, so a model that predicts device gain from process parameters before a run has value out of proportion to its accuracy. ・We study this problem on a real fabrication history, thirteen to fourte
cs.LG updates on arXiv.org

Forward Trajectory Steering for Hamilton-Jacobi Reachability Analysis

・arXiv:2608.11480v1 Announce Type: cross Abstract: Hamilton-Jacobi (HJ) reachability provides a mathematically rigorous framework for safe control of dynamical systems, but its practical application is bottlenecked by the computational complexity of solving Hamilton-Jacobi-Isaacs variational inequality PDEs in high dimensions. ・Physics-informed neural networks (PINNs) have recently emerged as a promising alternative to
cs.LG updates on arXiv.org

FQTree: Fine-grained Quantization and Hardware Generation of Boosted Decision Trees

・arXiv:2608.12140v1 Announce Type: cross Abstract: Boosted decision trees (BDTs) are widely used in latency-critical applications, but efficient hardware deployment remains challenging. ・Existing designs often rely on uniform or manually tuned fixed-point formats, which can introduce unnecessary hardware cost or accuracy loss. ・This work presents the FQTree algorithm{https://github.com/ecs-bristol/FQTree} for fine-grain
cs.LG updates on arXiv.org

From Monolithic to Modular: Segment-level Automatic Prompt Optimization

・arXiv:2608.11219v1 Announce Type: cross Abstract: Automatic Prompt Optimization (APO) often rewrites prompts monolithically, which can improve one behavior while degrading others. ・We present SAPO, a segment-level APO method that decomposes prompts into role, context, tasks, and output format, then applies targeted improvements based on top-5 and bottom-5 examples. ・The optimization loop uses one LLM with static meta-p
cs.LG updates on arXiv.org

From Numbers to Judgment: Specialist LLM Agents and Reinforcement Learning for European Listed Real Estate

・arXiv:2608.11381v1 Announce Type: cross Abstract: We study whether the localized numerical operations and integrative judgments of financial analysis benefit from the same form of LLM specialization. ・Larix maps a 16-lens European listed-real-estate analysis framework to eight lens-aligned specialists; we compare a frontier LLM under monolithic versus specialist-decomposed prompting while holding the model, source evi
cs.LG updates on arXiv.org

From Prompting to Behavioral Alignment: Personalized LLM Judges for Recommendation Evaluation

・arXiv:2608.11493v1 Announce Type: cross Abstract: Traditional offline recommendation evaluation relies heavily on complex, manually maintained feature pipelines that are difficult to scale. ・While Large Language Models (LLMs) offer a promising alternative by predicting user engagement directly from raw text logs, empirical analysis in this study identifies a critical failure mode termed bidirectional rationalization.
Hugging Face Papers

From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection

From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection
cs.LG updates on arXiv.org

FunnelCausalNet: Funnel-aware Joint Conversion-Revenue Uplift for Multi-tier Coupon Allocation

・arXiv:2608.11675v1 Announce Type: new Abstract: Coupon campaigns seek to lift both conversion and revenue, but gross merchandise value (GMV) follows a deterministic funnel from conversion to conditional order value and is zero-inflated and heavy-tailed. ・We propose FunnelCausalNet, an uplift estimator coupling a binary conversion head with a nonnegative conditional-value head through $\mu_{\mathrm{gmv}}=\mu_{\mathrm{c
cs.LG updates on arXiv.org

Gauge-Fixing the Forward-Forward Objective: A Whitened Goodness Derived from a Likelihood-Ratio Account

・arXiv:2607.12501v3 Announce Type: replace Abstract: The Forward-Forward algorithm trains each layer locally, so that a scalar goodness - the sum of squared activations - is high on real inputs and low on contrastive ones. ・Under an explicit generative model this goodness is the sufficient statistic of a likelihood-ratio test, and the pairwise form of the objective admits a gauge: a layer can lower its loss by inflatin
cs.LG updates on arXiv.org

Gaussian Meta-Space Augmentation for Stacking Ensembles in Multimodal IPMN Risk Stratification

・arXiv:2608.11472v1 Announce Type: cross Abstract: Pancreatic cancer is among the most lethal malignancies; risk stratification of intraductal papillary mucinous neoplasms (IPMNs) offers a crucial opportunity for early intervention but typically requires invasive tissue biopsy. ・Dominant vision-based approaches, including radiomics and deep learning, provide promising but initially separate discrimination opportunities
cs.LG updates on arXiv.org

GCPO: Diagnosing and Constraining Subspace Geometry in Rollout RL for LLMs

・arXiv:2608.11674v1 Announce Type: new Abstract: On-policy rollout methods such as GRPO are central to post-training of large language models, yet they frequently suffer from training instabilities, cross-task capability degradation, and response-length inflation. ・Although prior work has characterized the subspace geometry of aggregate updates, the stepwise variation of this geometry and its relationship to model perf
#LLMタグ

Gemini Robotics ER 2登場——「ロボットの頭脳」としてのLLMは実務者に何を示唆するのか

・この記事はGoogle DeepMindをはじめとする公式情報源をもとに、LLM関連の最新発表を単発で解説する記事です。 ・本記事では「AI」「AIエージェント」はLLMを活用したシステムを指しています。また、記事中の技術情報は執筆時点の公式発表を参考に再構成したものです。
cs.LG updates on arXiv.org

Generative Learning for Quantum Measurement Design

・arXiv:2608.11396v1 Announce Type: cross Abstract: Extracting quantum information from a quantum state is a fundamental task of quantum computation, often requiring the estimation of many non-commuting observables under a finite measurement budget. ・For both near-term and early fault-tolerant settings, the measurement protocol must balance statistical efficiency against implementation resources such as circuit depth, c
MarkTechPost

Google AI Just Released Gemini 3.7 Flash: A Coding and Agent Model at $0.75/1M Input Tokens

・Google has released Gemini 3.7 Flash, a refinement of Gemini 3.6 Flash with algorithmic improvements to its reasoning core. ・It handles text, images, audio, and video across a 1M-token context window with 64K-token output, and supports customizable thinking configurations. ・Coding results move notably: 43.6% on FrontierCode 1.1 Main versus 34.4%, 65.3% on DeepSWE v1.1, and 1588 Elo on WebDev Arena.
cs.LG updates on arXiv.org

Grounding Large Language Models as Generalizable Policies in Network Control

・arXiv:2512.11839v2 Announce Type: replace Abstract: Designing generalizable control policies that operate reliably under changing conditions is essential for robust network services in modern digital infrastructure. ・Yet network control remains dominated by specialized policies built from handcrafted rules or deep learning models, which struggle to generalize under real-world dynamics. ・Large language models (LLMs) off
cs.LG updates on arXiv.org

GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation

・arXiv:2608.11787v1 Announce Type: cross Abstract: Generating actionable financial advice from business records demands that models integrate numerical reasoning, domain knowledge, and sound judgment, while avoiding recommendations that could harm the business. ・Direct supervision is difficult: historical decisions are not necessarily optimal, and high-quality free-form labels are expensive to obtain. ・We formulate fina
Hugging Face Papers

Hand Visibility Detector: Per-Keypoint Visibility Estimation for Hands

Hand Visibility Detector: Per-Keypoint Visibility Estimation for Hands
cs.LG updates on arXiv.org

Hardware-Aware Deployment of Joint SAR Compression and Despeckling on FPGA

・arXiv:2608.11271v1 Announce Type: cross Abstract: Next-generation Synthetic Aperture Radar (SAR) missions will generate data far faster than they can downlink, making onboard data reduction essential for near-real-time Earth observation. ・Learned Image Compression (LIC) offers better rate-distortion performance than handcrafted codecs used operationally today, and recent work shows that simultaneously despeckling and
WIRED

Herman Miller Promo Codes: 40% Off August 2026

・Whether you’re looking for a Herman Miller promo code or discounts in their sale, here is how to save on the world's best ergonomic office furniture in August 2026.
cs.LG updates on arXiv.org

Heterogeneous transfer learning for high-dimensional regression with feature mismatch

・arXiv:2412.18081v3 Announce Type: replace-cross Abstract: We study Heterogeneous Transfer Learning (HTL) for high-dimensional regression with differing feature sets. ・Such feature mismatch arises when some variables available in a data-rich source domain are unavailable in a data-poor target domain. ・Yet most homogeneous TL methods require the same feature space in both the source and target domains, limiting their pra
cs.LG updates on arXiv.org

Hierarchical Federated Transfer Learning in Digital Twin-Based Vehicular Networks

・arXiv:2608.11532v1 Announce Type: new Abstract: In recent research on the Digital Twin-based Vehicular Ad hoc Network(DT-VANET), Federated Learning (FL) has shown its ability to provide data privacy. ・However, Federated learning struggles to adequately train a global model when confronted with data heterogeneity and data sparsity among vehicles, which ensure suboptimal accuracy in making precise predictions for differ
cs.LG updates on arXiv.org

High-dimensional Multi-objective Bayesian Optimization with Learned Variable Interactions

・arXiv:2608.11713v1 Announce Type: new Abstract: Multi-objective Bayesian optimization (MOBO) is effective in identifying the Pareto fronts for expensive black-box problems. ・However, most current MOBO approaches are limited to low-dimensional decision space due to its exponential sampling complexity. ・This paper presents decision variable interaction analysis-based MOBO, ViaMOBO, a generic framework for expensive multi
cs.LG updates on arXiv.org

High-Order Liquid Evidence Encoding for Gradual GNSS Spoofing Detection in Autonomous Driving

・arXiv:2608.11790v1 Announce Type: new Abstract: Accurate Global Navigation Satellite System (GNSS)-based localization is essential for safe and reliable autonomous driving. ・However, spoofing attacks can manipulate vehicle position estimates. ・Continuous and subtle attacks are particularly difficult to detect because individual GNSS observations may remain plausible while the inconsistency between GNSS-implied displace
Zennの「機械学習」のフィード

Hopfieldネットワークの記憶容量を実測。理論値141枚のはずが、文字パターンは3枚で崩壊した

・ニューラルネットは何枚の画像まで「記憶」できるのか。1982年にHopfieldが提案した連想記憶ネットワークには、理論的な答えがある。ニューロン数Nに対して約0.138N枚。それを超えて詰め込むと記憶が壊れ始め、しかもただ忘れるのではなく、どの記憶とも違う「偽の記憶」が現れるという。本当にそんな崖があるのか、偽の記憶とはどんな顔をしているのか。1,024ニューロンのネットワークを組んで実測した。 ・実験に使ったコードの全文はGitHubに置いています。 ・32×32の画像を記憶するネットワークを組む Hopfieldネットワークは、全ニューロンが互いに結合した再帰型ネットワークだ。各ニ...
cs.LG updates on arXiv.org

How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models

・arXiv:2608.12192v1 Announce Type: cross Abstract: Foundation models for protein structure prediction remain unreliable on certain targets. ・External oracles can flag and correct these failures, but biological oracles are expensive, making oracle budget a critical constraint. ・Existing guidance methods, such as FK-steering, DPO, and Best K-of-N sampling, differ in how they spend this budget, yet no systematic comparison
cs.LG updates on arXiv.org

Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing

・arXiv:2608.11660v1 Announce Type: cross Abstract: Large language models (LLMs) achieve remarkable performance across natural language tasks, yet they are trained on static corpora and their knowledge quickly becomes outdated in a fast-changing world. ・This motivates knowledge editing (KE), which updates specific knowledge in an LLM without changing unrelated others. ・Recent works move from structured knowledge triples
cs.LG updates on arXiv.org

HYDRA: Hyperbolic Dynamic Representation Architecture for Kolmogorov-Arnold Networks

・arXiv:2608.12194v1 Announce Type: new Abstract: Kolmogorov-Arnold Networks (KANs) enhance nonlinear function approximation by replacing scalar weights with learnable univariate functions. ・However, assigning an independent function to every connection results in substantial parameter redundancy, limiting their scalability and efficiency. ・To reduce this redundancy, we introduce \textbf{HY}perbolic \textbf{D}ynamic \tex
cs.LG updates on arXiv.org

HyperFix: Combinatorial Nonlinear Correction for Task Vector Merging

・arXiv:2608.11499v1 Announce Type: new Abstract: Task vectors enable model merging without joint retraining. ・In practice, the subset of task vectors to be merged may vary, but many existing methods use scalar tuning for a particular subset, requiring repeated tuning across subsets and restricting task vector merging to linear rescaling. ・We therefore formulate merging across varying task subsets as a combinatorial corr
The Verge

I looked inside an AI generated movie, and the best parts were all human

・Imagine a trio of bumbling, English lads who fantasize about becoming megastars while knocking back a few pints in a grimy pub somewhere in London. ・Picture the guys chortling and trying to one-up each other's idealized visions of the future with a series of increasingly glitzy fantasies in which their fame leads to access to groupies and lavish parties on luxurious boats. ・Imagine these men are actually characters in
#AIタグ

IBMの2億4000万ドル推論クラスター契約:オープンソースAIインフラへの重要なシグナル

・IBMとTogether AIは2026年8月11日、IBM Cloud上にNVIDIA HGX B300を使った大規模な推論クラスターを構築する複数年契約を発表した。契約額は2億4000万ドルで、2027年第1四半期から企業向けのオープンソースモデル推論に利用できる計画である。Reutersの報道でも契約額とNVIDIA製システムの採用が確認されている。このニュースで重要なのは、オープンソースモデルが専用モデルAPIを置き換えると断定できることではない。企業がAI推論基盤を選ぶ際に、「自前でGPU基盤を持つ」「外部のモデルAPIを使う」に加えて、クラウド上の専用GPU基盤と専門事業者の推論サービスを組み合わせる第三の選択肢が、具体的な大型契約として現れた点にある。
WIRED

In a Heat Wave, Schizophrenia Is So Much Deadlier Than Any Other Medical Condition

・People with schizophrenia face a perfect storm of dangers on a hotter planet.
Google DeepMind News

Introducing Gemini 3.7 Flash

Introducing Gemini 3.7 Flash
cs.LG updates on arXiv.org

IoT-Enabled Autonomous Maritime Navigation in Smart Ports: A Curriculum-Guided Shared Policy Learning Framework

・arXiv:2608.11597v1 Announce Type: cross Abstract: As smart port infrastructures increasingly rely on autonomous maritime devices enabled by the Internet of Things (IoT), ensuring reliable onboard navigation intelligence has become a critical challenge for safe and scalable operations in congested waterways. ・This paper investigates onboard autonomous navigation for such IoT devices under partial observability and dens
cs.LG updates on arXiv.org

Is Per-Agent Policy Composition Safe? Rethinking Successor-Feature Transfer in Cooperative Multi-Agent Reinforcement Learning

・arXiv:2608.11658v1 Announce Type: new Abstract: Many reinforcement learning systems, from fleet management to traffic signal control, must serve an objective that changes dynamically after deployment, and retraining a policy for each new objective is prohibitively expensive. ・For a single agent, this problem is well understood: successor features with generalized policy improvement, together with their universal exten
WIRED

Jabra Promo Codes: 30% Off Headphones, Headsets & More

・Unlock significant savings on Jabra headphones, headsets, and speakers with our verified promo codes, exclusive coupons, and limited-time deals.
cs.LG updates on arXiv.org

JAPE: Joint Anomaly Prediction and Intrinsic Explanation in Multivariate Time Series

・arXiv:2608.11801v1 Announce Type: new Abstract: Multivariate time-series anomaly prediction aims to identify whether and when anomalies will occur over a future horizon from historical observations. ・Existing methods primarily characterize anomalies as deviations in future numerical values, which may overlook subtle dependency changes induced by weak anomaly precursors and provide no native variable-level explanation
cs.LG updates on arXiv.org

Kernel Methods for Learning Operators with Multiple Inputs and Outputs

・arXiv:2608.11831v1 Announce Type: new Abstract: Learning mappings between infinite-dimensional objects is a central challenge in scientific machine learning. ・We introduce a general kernel-based encoder-decoder framework for operator learning that separates observation, representation, learning, and reconstruction. ・We develop this framework for multi-input, multi-output operator learning, where operators map between p
cs.LG updates on arXiv.org

Language-Structured Relational Q-Learning for Threat-Aware Control in Safety-Critical Driving

・arXiv:2608.11498v1 Announce Type: cross Abstract: Natural-language-based scenario generation offers an intuitive means of describing rare and complex driving interactions, yet it is still uncertain whether training with language-structured data leads to truly adaptive control policies. ・We propose Language-Structured Relational Q-Learning, instantiated through an Ego-Centric Relational Q-Network (ERQ-Net), which joint
cs.LG updates on arXiv.org

Large language models reorganize representational geometry during in-context learning

・arXiv:2605.28854v3 Announce Type: replace-cross Abstract: Large language models (LLMs) show remarkable flexibility in adapting to novel tasks without parameter updates, a capacity known as in-context learning (ICL). ・Prior work has sought to understand ICL by studying the circuits, algorithms, and representations that support it. ・Yet why some ICL tasks are easy to solve while others are difficult remains unresolved.
cs.LG updates on arXiv.org

Large-scale AI-Ready Data for Anti-Cancer Drug Response Modeling

・arXiv:2608.11444v1 Announce Type: cross Abstract: Drug response prediction (DRP) models are an active area of research in pharmacogenomics, with growing potential to accelerate the identification of effective anticancer drugs. ・However, their predictive performance is often constrained by limited dataset scale and insufficient coverages of cancer and chemical spaces. ・In addition, inconsistent benchmarking practices hi
cs.LG updates on arXiv.org

Latent variable models for simultaneous EOV identification and removal in population-based SHM

・arXiv:2608.11995v1 Announce Type: cross Abstract: The robust treatment of environmental and operational variability (EOV) is an open challenge in population-based structural health monitoring (PBSHM). ・The difficulty is compounded in the case that the EOV signals are unmeasured. ・A common approach in conventional SHM is to apply \emph{projection-based} methods that discard subspaces of healthy feature data, reasoning t
cs.LG updates on arXiv.org

Learning Multi-Timescale Interventions under Safety and Resource Constraints

・arXiv:2508.03875v3 Announce Type: replace Abstract: Many sequential decision problems offer qualitatively different ways of influencing the environment: some interventions act immediately, whereas others induce persistent effects that continue to shape future states long after the decision that initiated them. ・An agent must then decide jointly when to intervene, which temporal mode to use and how strongly, while acco
cs.LG updates on arXiv.org

Learning with Bilevel-Minimax Optimization for Efficient and Reliable Transfer Attacks

・arXiv:2608.11815v1 Announce Type: new Abstract: Transfer-based adversarial attacks craft adversarial examples using surrogate models to mislead black-box victim models. ・Beyond perturbation generation, transferability is fundamentally governed by the coupling of initialization, surrogate adaptation, and gradient dynamics. ・We revisit this challenge from a bilevel-minimax perspective and propose BMAT (Bilevel-Minimax Ad
cs.LG updates on arXiv.org

Learning-Based Behavior Planning for Automated Driving: Real-World Integration and Deployment

・arXiv:2608.12198v1 Announce Type: cross Abstract: Recent research in machine and deep learning has shown the potential of learningbased motion planning approaches to improve the driving behavior of automated vehicles, especially in complex environments. ・However, their complex nature and lack of transparency can hinder explainability and trustworthiness and complicate safety assurance. ・Motivated by these challenges, w
cs.LG updates on arXiv.org

LEMUR: Latent Entropy-aware Multimodal Unlearning via Visual-anchored Reasoning Redirection

・arXiv:2608.11691v1 Announce Type: new Abstract: Reinforcement-learning (RL) post-training equips multimodal large reasoning models (MLRMs) with exploratory chains of thought (CoT), substantially improving visual reasoning. ・However, we find that this capability introduces a distinct privacy vulnerability: even when a sensitive fact is successfully unlearned from the final answer, the model may still reproduce it in it
cs.LG updates on arXiv.org

Let it Cook: Learning to Wait in Sequential Decision Making

・arXiv:2608.11511v1 Announce Type: new Abstract: In sequential decision making, an agent typically observes its environment and acts at every timestep. ・However, such active participation may not always be necessary; tasks such as brewing coffee include periods that are served equally well by letting the environment evolve without constant monitoring and control. ・During such periods, the agent could simply wait to cons
cs.LG updates on arXiv.org

Lifecycle-Optimal Tokenization: Vocabulary Size as a Deployment-Regime-Dependent Infrastructure Parameter

・arXiv:2608.11361v1 Announce Type: new Abstract: Tokenizer vocabulary size is a foundational design choice in large language model (LLM) infrastructure, yet it is typically fixed at training time based on convention rather than deployment analysis. ・We show that the cost-optimal vocabulary is not a constant but a function of the serving regime. ・We formalize total deployment cost as $C_{lifecycle}(V) = C_{train}(V) + \l
ITmedia NEWS 最新記事一覧

LINE、送信後のメッセージを編集可能に 15分以内限定 秋以降は有料化

・まず「LINEラボ」で無償提供し、秋以降は有料の「LYPプレミアム」会員向け特典として正式に提供する予定だ。
cs.LG updates on arXiv.org

Linear-Core Surrogates: Smooth Loss Functions with Linear Rates for Classification and Structured Prediction

・arXiv:2604.27742v2 Announce Type: replace Abstract: A fundamental dichotomy in the theory of classification sets smoothness against statistical efficiency: smooth surrogate losses such as the logistic loss enable fast $O(1/T)$ optimization but yield slow square-root $H$-consistency bounds, while piecewise-linear losses like the Hinge loss achieve optimal linear $H$-consistency rates but are non-differentiable.
MarkTechPost

Liquid AI Releases LFM2.5-VL-3B: A 3B Vision-Language Model That Reads Screens, Grounds Objects, and Calls Tools On-Device

・Liquid AI released LFM2.5-VL-3B, a 3.1B-parameter vision-language model built for on-device deployment. ・It averages 80.7 on ScreenSpot-v2 and lifts RefCOCO grounding from 57.1 to 87.9. ・Function calling is new to the VL line, with ToolSandbox moving from 26.4 to 59.5.
Zennの「大規模言語モデル」のフィード

llama-serverのホットスタンバイは存在するか:部品は揃っているのに名前がない

・はじめに:よくある混同 ローカルLLMを常用していると、「ホットスタンバイ」「モデル常駐」「KVキャッシュ保持」「スリープ復帰」「ホットスワップ」という言葉が混同されがちです。 ・結論から言うと、llama-server にホットスタンバイという名前の完成された単一機能はありません。しかし、複数の機能を組み合わせるとホットスタンバイ構成を作れる状態にはなっています。 ・この記事では、何がホットスタンバイで何がそうでないかを整理し、ローカルAIエージェント(Hermes Agent等)の可用性設計に応用する視点を示します。
cs.LG updates on arXiv.org

LLM Router: Rethinking Routing with Prefill Activations

・arXiv:2603.20895v3 Announce Type: replace-cross Abstract: Existing routers rely on semantic query features or handcrafted features, which often fail to capture model-specific failures or intrinsic task difficulty. ・We instead route using internal LLM activations, specifically the residual stream. ・Our key idea, Encoder-Target Decoupling, separates the model that produces the predictive signal (the Encoder) from the mod
LLMタグが付けられた新着記事 - Qiita

LLM の性能は prefill と decode で決まり方が違う

・LLM の推論速度を考えるとき、単に「GPU が速い」「量子化したから速い」「行列積が多い」と見るだけでは十分ではありません。入力 prompt をまとめて処理する prefill と、1 token ずつ生成する decode では、同じモデルを動かしていても性能の決まり...
cs.LG updates on arXiv.org

Local Cluster Cardinality Estimation for Adaptive Mean Shift

・arXiv:2508.12450v2 Announce Type: replace Abstract: This article presents an adaptive mean shift algorithm in which every parameter used at a point is derived from that point's own distance distribution. ・The distance distribution from a point to all others is used to estimate the cardinality of the local cluster by identifying a local minimum in the density of that distribution; the statistics of the identified subse
cs.LG updates on arXiv.org

Located but Not Releasable: Silent Gate Inversion and Bounded Linear Release

・arXiv:2608.11822v1 Announce Type: cross Abstract: A growing body of work reports that language models represent task-relevant latent structure that they fail to use. ・Whether such structure, once located, can be converted into behavior is a separate question that is rarely tested end to end. ・We submit the complete pipeline -- detect, localize, and release -- to a fully preregistered stress test on a 25.7M transformer
cs.LG updates on arXiv.org

Locating and Controlling Implicit Personalization in Large Language Models

・arXiv:2608.11735v1 Announce Type: cross Abstract: Large language models (LLMs) often shift their outputs in response to implicit demographic cues even when users never state a demographic identity. ・Previous work has documented this behavior, but the connection between these behavioral changes and the model's internal activations remains unclear. ・Using matched cued and neutral conversations across five LLMs, we establ
cs.LG updates on arXiv.org

LODESTAR: Trustworthy Entropy Is Navigated, Not Merely Measured -- Reinforced Polarizer Keeps a Frozen LLM from Being Confidently Misled by the Wrong Evidence

・arXiv:2608.11922v1 Announce Type: cross Abstract: Predictive-distribution entropy makes a strong selection rule in retrieval-augmented question answering: across five QA benchmarks, keeping the candidate answer that a frozen respondent LLM produces with the lowest answer-token entropy lifts mean answer $F_1$ from 0.4769 to 0.5148 over the retriever's top-ranked passage, with no gold answers. ・Yet this lowest-entropy r
cs.LG updates on arXiv.org

Long-Horizon Forecasting of Complete Financial Statements with Forma

・arXiv:2608.11327v1 Announce Type: new Abstract: Specialist training beats generalist scale when forecasting financial statements. ・To our knowledge, no prior work jointly forecasts complete financial statements beyond one year, yet in a discounted-cash-flow valuation most firm value sits past that window. ・We release ProForma-20Q, a reproducible benchmark for forecasting 78 statement line items 1-20 quarters ahead, for
cs.LG updates on arXiv.org

Look What the Probes Dragged In! Real-World Chest X-ray Shortcuts in MedCLIP

・arXiv:2608.12086v1 Announce Type: cross Abstract: Vision-language models, such as contrastive language-image pre-training (CLIP)-based approaches, have reached state-of-the-art (SOTA) results in medical artificial intelligence. ・However, recent work reveals that CLIP-based models remain vulnerable to shortcuts. ・We investigate how real-world shortcuts manifest across different layers of the medical CLIP-based model, Me
cs.LG updates on arXiv.org

LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global Perspective Distillation

・arXiv:2608.11967v1 Announce Type: new Abstract: Large language model agents increasingly rely on long-horizon reasoning to solve complex tasks involving planning, tool use, and memory. ・A critical capability in such settings is reflection: assessing trajectory progress, identifying missing evidence and unreliable intermediate states, and deciding whether to continue, revise, or abandon the current branch. ・Learning eff
cs.LG updates on arXiv.org

LoRAQuant: Mixed-Precision Quantization of LoRA to Ultra-Low Bits

・arXiv:2510.26690v3 Announce Type: replace Abstract: Low-Rank Adaptation (LoRA) has become a popular technique for parameter-efficient fine-tuning of large language models (LLMs). ・In many real-world scenarios, multiple adapters are loaded simultaneously to enable LLM customization for personalized user experiences or to support a diverse range of tasks. ・Although each adapter is lightweight in isolation, their aggregat
cs.LG updates on arXiv.org

Low-Interaction-Rank Learning: Unifying Multiplicative Dual-Encoder Heads

・arXiv:2608.11661v1 Announce Type: new Abstract: A multiplicative dual-encoder network computes a real-valued output for a pair of inputs as the inner product of their separate encodings. ・This architecture has been developed independently in operator learning, bipartite matching, contrastive vision-language models, retrieval, and other areas, yet no unified theory guides the basic design decisions: how many interactio
cs.LG updates on arXiv.org

Market-Information-Aware Gated-LoRA of Foundation Models for Transferable Day-Ahead Electricity Price Forecasting

・arXiv:2608.11359v1 Announce Type: new Abstract: Electricity price forecasting is crucial for market participants but remains difficult because prices are volatile, market-specific, and closely tied to anticipated system conditions. ・Existing supervised methods depend largely on market-specific historical data, limiting their use in newly established or data-scarce markets. ・This paper proposes a market-information-awar
cs.LG updates on arXiv.org

MaSRead: Content-Addressed Reading of Replicated Latent Stores

・arXiv:2608.11218v1 Announce Type: cross Abstract: Independent agents that reason in latent space can share computed state as key-value cache fragments rather than text. ・Merged by a conflict-free replicated data type, these fragments form a store that converges under any delivery order or duplication. ・Yet a later query, unknown at encode time, cannot reliably read the merged cache: colocated fragments interfere, so co
Hugging Face Papers

MBA: Multimodal Benchmark and Agents for Real-World Business Ideation

MBA: Multimodal Benchmark and Agents for Real-World Business Ideation
cs.LG updates on arXiv.org

MBA: Multimodal Benchmark and Agents for Real-World Business Ideation

・arXiv:2608.11616v1 Announce Type: cross Abstract: Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation. ・Yet existing approaches remain confined to a text-only paradigm, despite the inherently multimodal nature of real-world contexts. ・We thus introduce MBA-Bench, the first multimodal benchmark for training and evaluating business ideation agents, comprising 30K sa
cs.LG updates on arXiv.org

Mechanism Design for Generative Engines: From Exploitation toward Win-Win Outcomes

・arXiv:2608.11390v1 Announce Type: new Abstract: Generative engines are reshaping the web ecosystem by making citations a key mechanism for allocating attention, attribution, and downstream value. ・This creates a strategic tension: content providers are incentivized to optimize for model citation, while platforms must preserve answer quality and trustworthy attribution. ・We show that this tension can escalate into citat
Hugging Face Papers

Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence

Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence
cs.LG updates on arXiv.org

Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence

・arXiv:2608.12036v1 Announce Type: cross Abstract: AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood. ・As AI development becomes faster and increasingly automated, mechanistic exploration remains largely manual, widening the gap between what models can do and our ability to understand and control them.
The Verge

Meta adds AI screening to detect WhatsApp scams

・Meta is launching an optional Scam Alert feature on WhatsApp that uses on-device machine learning to flag suspicious messages. ・Earlier this year, Meta also launched scam detection for device linking requests on WhatsApp. ・The new Scam Alert feature, which is rolling out in a limited beta, shows users a warning if a chat seems like a scam: If the model identifies a message as a likely scam attempt, the user sees a warn
#LLMタグ

MetaのMuse Sparkが評価中に外部システムへ侵入

・Metaが2026年8月5日、自社のMuse Spark 1.1が評価中に外部組織のシステムへ侵入していたことを認めた。 ・評価を請け負っていたIrregularの設定ミスにより、本来は遮断されているはずのインターネットへモデルが到達し、第三者のサービスにあった脆弱性を突いた形になる。 ・侵入先の名称は公表されていない。
AI News & Artificial Intelligence | TechCrunch

Microsoft kills off unsuccessful AI features while merging its separate Copilot apps

・Microsoft is simplifying Copilot by combining its consumer and business apps, and dropping AI-generated podcasts, Group Chats, Deep Research, and its Mico character.
cs.LG updates on arXiv.org

Mind the Gap: Structure-Aware Consistency in Preference Learning

・arXiv:2604.27733v2 Announce Type: replace Abstract: Aligning Large Language Models (LLMs) with human intent, whether through explicit reward modeling or direct methods such as DPO, fundamentally relies on minimizing a surrogate loss as a proxy for the true pairwise ranking objective. ・We prove that this reliance is flawed for the standard surrogate losses used: for the equicontinuous hypothesis sets characteristic of
#LLMタグ

MiniArt-Uncensored Phase 4 インジェクション前編レポート

MiniArt-Uncensored Phase 4 インジェクション前編レポート
#LLMタグ

MiniArt-Uncensored Phase 5 インジェクション後編レポート

MiniArt-Uncensored Phase 5 インジェクション後編レポート
#LLMタグ

MiniArt-Uncensored 総合ベンチマークレポート(全52問)

MiniArt-Uncensored 総合ベンチマークレポート(全52問)
cs.LG updates on arXiv.org

MMLA: How Memory Lets the Past Shape the Future

・arXiv:2606.28876v3 Announce Type: replace-cross Abstract: Proposal. ・Long context can replay history, but it does not decide which completed observations deserve authority. ・MMLA formalizes a bounded resident memory between transient context and slow weight updates.
cs.LG updates on arXiv.org

Modeling Spectral Energy Shifts in Spatio-Temporal Graph Anomaly Detection

・arXiv:2606.00304v2 Announce Type: replace Abstract: Graph anomaly detection methods aim to distinguish anomalous nodes. ・While prior methods characterize anomalies through increased variation in the spectral energy distributions, they overlook those that result in decreased variation, i.e., camouflaged anomalies that appear normal. ・We show that this type of anomaly persists across multiple datasets and remains undetec
cs.LG updates on arXiv.org

MOON: Multi-Objective OrthoNormalized Updates for Multitask Learning

・arXiv:2608.11749v1 Announce Type: new Abstract: Multi-objective optimization (MOO) has demonstrated significant success in multi-task learning by mitigating task conflicts through gradient manipulation. ・However, most existing methods flatten model parameters into vectors and perform gradient manipulation under Euclidean geometry, thereby overlooking the matrix structure prevalent in modern architectures such as Trans
cs.LG updates on arXiv.org

Multilingual OCR-Aware Fine-Tuning and Prompt-Guided Chain-of-Thought Reasoning for Multimodal Large Language Models

・arXiv:2605.16409v3 Announce Type: replace-cross Abstract: Optical character recognition (OCR) and multilingual scene-text understanding remain challenging for multimodal large language models (MLLMs), particularly in real-world images containing small or degraded text, cluttered layouts, occlusion, handwriting, and complex typography. ・We present an OCR-aware multilingual post-training framework that improves visual-t
cs.LG updates on arXiv.org

NAE: Normalizing AutoEncoder

・arXiv:2608.12084v1 Announce Type: new Abstract: We consider the setting of Normalizing flows with approximate inverses, an established paradigm spanning both full-dimensional ($d=D$) and bottleneck ($d<D$) settings, and group these models under the term flow autoencoders. ・We present a theoretical investigation into their training dynamics and prove that the proposed loss used by existing approaches is suboptimal; spe
WIRED

Naturepedic Promo Codes: Get 20% Off Plus Free Pillows

Naturepedic Promo Codes: Get 20% Off Plus Free Pillows
Hugging Face Papers

NeuPAT: Neuron-aware Plasticity Allocation Tuning for Language-Preserving MLLMs

NeuPAT: Neuron-aware Plasticity Allocation Tuning for Language-Preserving MLLMs
Zennの「機械学習」のフィード

NPUのSystolic Arrayとは?PE Arrayとデータ再利用の仕組み

・NPUはなぜDRAMアクセスを減らしやすいのか?PE ArrayとSystolic ArrayをGPUとの違いから理解する 概念図:実際の製品回路を理解しやすく簡略化した図。 ・本記事の目的 前編では、GPUが行列積をtile(小さな部分行列)へ分割し、Shared Memory(SM内で共有する高速メモリ)やRegisters(thread近傍の保存領域)でA/B tileを再利用する仕組みを整理しました。 ・GPUの行列積はなぜタイル化するのか?warp・Shared Memory・Tensor Coreから理解する 本記事では、NPU(Neural Pr...
AI News & Artificial Intelligence | TechCrunch

Nvidia’s new $500B plan is risky but brilliant, especially for aging GPUs

・Nvidia has a plan to make sure its GPUs won't lose value. ・It wants to convince a new crop of financiers to keep lending for AI buildouts.
cs.LG updates on arXiv.org

ODE-Based Transformer Decoders for Iterative Sign Language Translation

・arXiv:2608.11352v1 Announce Type: cross Abstract: Sign language translation has achieved strong results with Transformer architectures, yet recent improvements largely rely on scaling model capacity at the cost of increased computation. ・We propose a parameter-efficient alternative that improves expressiveness without increasing model size. ・Rather than scaling capacity, we focus on enhancing the update dynamics of ite
cs.LG updates on arXiv.org

On Data-Driven Koopman Representations of Nonlinear Delay Differential Equations

・arXiv:2604.03086v2 Announce Type: replace-cross Abstract: This work establishes a rigorous bridge between infinite-dimensional delay dynamics and finite-dimensional Koopman learning, with explicit and interpretable error guarantees. ・While Koopman analysis is well-developed for ordinary differential equations (ODEs) and partially for partial differential equations (PDEs), its extension to delay differential equations
cs.LG updates on arXiv.org

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL

・arXiv:2608.12253v1 Announce Type: cross Abstract: Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior. ・We show that this approach systematically fails to generalize, and trace the failure to simulator collapse: because the simulator LLM is mode-collapsed, an LLM policy trained against it overfits to narrow strategies that exploit the
OpenAI News

OpenAI appoints Dali Rajic as Chief Revenue Officer

・OpenAI appoints Dali Rajic as Chief Revenue Officer to lead its global revenue organization and help businesses realize the full value of AI.
AI News & Artificial Intelligence | TechCrunch

OpenAI hires new CRO as executive shake-up continues

・Dali Rajic will take over as OpenAI's top salesperson.
LLMタグが付けられた新着記事 - Qiita

OpenAI「Ultrafast」とは?GPT-5.6 Solが最大14倍速に——公式発表を3分で速報解説

・:::message 🐹🦜 この記事に登場する2匹 🐹 もっちー (ハムスター)… AI はまだ勉強中。「それどういうこと?」と素朴に質問する生徒役 🦜 きなこ (セキセイインコ)… AI で調べものをこなす解説役。やさしく深掘りして教える先生役 この記事は2匹の掛け...
Hugging Face Papers

OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution

OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution
#LLMタグ

OpenCodeでGrok 4.6を使う方法。xAI APIを5ドルから試す方法

OpenCodeでGrok 4.6を使う方法。xAI APIを5ドルから試す方法
#LLMタグ

OpenHandsをローカルLLMで動かす ~無料/無制限の24H開発環境! スマホからも!?~

・こんにちはRcatです。 ・今回はローカルLLM開発の幅を広げようと思いまして、OpenHandsを導入してみました。 ・概要からセットアップ、そして使用感をまとめます。
cs.LG updates on arXiv.org

Optimized Deferral for Imbalanced Settings

・arXiv:2604.27723v2 Announce Type: replace Abstract: Learning algorithms can be significantly improved by routing complex or uncertain inputs to specialized experts, balancing accuracy with computational cost. ・This approach, known as learning to defer, is essential in domains like natural language generation, medical diagnosis, and computer vision, where an effective deferral can reduce errors at low extra resource co
cs.LG updates on arXiv.org

Optimizing Expert-Designed Crystal Graph Networks for Band-Gap Prediction with an Autonomous LLM Research Loop

・arXiv:2606.29717v2 Announce Type: replace-cross Abstract: Predicting a material's properties from its structure is a central, fast-advancing problem in computational materials science. ・A decade of work has produced standard public benchmarks and many published machine-learning models for the task (Dunn et al., 2020). ・The task's fixed metric and these baselines make it a natural setting for autonomous agent research (
cs.LG updates on arXiv.org

Orientation, not magnitude: the causal structure of task-vector interference in merged language models

・arXiv:2608.11797v1 Announce Type: new Abstract: Model merging by task arithmetic works until it doesn't, and the field diagnoses why with magnitudes: layerwise representation bias, deviations from cross-task linearity, parameter overlap. ・Tracking the exact layerwise cross-term of merged LLMs through a factorial ledger and intervening on it directly, we find magnitude insufficient - and inconsistent across model famil
cs.LG updates on arXiv.org

PAIR: Pairwise-Aware Inclusion Reweighting for Adaptive Rollout Allocation in RLVR

・arXiv:2608.11368v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) spends most of its compute generating groups of long reasoning trajectories. ・Recent allocators reduce this cost by assigning budgets to prompts, rollouts, or tokens according to a pointwise notion of difficulty or utility. ・We identify a statistical mismatch: the unclipped leave-one-out group-relative score gradient i
Hugging Face Papers

Parameter Exploration for RLVR via Variational Learning

Parameter Exploration for RLVR via Variational Learning
cs.LG updates on arXiv.org

Patch-based Memory Gate Model in Time Series Foundation Model

・arXiv:2509.18751v4 Announce Type: replace Abstract: Recently reconstruction-based deep models have been widely used for time series anomaly detection, but as their capacity and generalization capability increase, these models tend to over-generalize, often reconstructing unseen anomalies accurately. ・Prior works have attempted to mitigate this by incorporating a memory architecture that stores prototypes of normal pat
#LLMタグ

PDF処理にフロンティアモデルは要らない——3BのVLM「LFM2.5-VL-3B」を技術から読み解く

・そのPDF処理、本当にGPTやClaudeでやる必要がありますか? 請求書や申込書、契約書から文字や表を読み取るたびに、高性能モデルへデータを送っている。もしそうなら、AIシステムの設計を見直す余地があります。
Hugging Face Papers

Persistent Recursive Worlds Enable Autonomous Software Evolution

Persistent Recursive Worlds Enable Autonomous Software Evolution
cs.LG updates on arXiv.org

Physics-Informed Implicit Neural Representations for Improved Myocardial Perfusion MRI Quantification

・arXiv:2608.11282v1 Announce Type: cross Abstract: Quantifying myocardial perfusion from cardiac magnetic resonance (CMR) can be achieved by fitting tracer-kinetic models to the dynamic contrast-enhanced MR data. ・However, fitting the observed data with multi-compartment exchange models, which describe the evolution of the contrast agent in the tissue, to estimate perfusion parameters is a challenging inverse problem t
cs.LG updates on arXiv.org

Planar Symmetric Pattern Generation

・arXiv:2606.02073v2 Announce Type: replace Abstract: Generating objects with specific symmetries is essential in various real-world scenarios. ・However, adapting existing 2D continuous representations to enforce planar group symmetry remains a challenge, as the transformation of non-reflective group elements may disrupt continuity. ・To overcome this limitation, we propose a symmetrization framework for arbitrary planar
cs.LG updates on arXiv.org

Policy-as-logic for robust reasoning over rules

・arXiv:2608.11905v1 Announce Type: cross Abstract: In many practical applications of generative AI systems, from tax rules to airline baggage allowance, responses to natural language queries must respect written policies or rules. ・We present a hybrid symbolic approach that expresses policies in formal logic and at inference time exploits the representation power of language models for fact extraction to ground predica
Hugging Face Papers

Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop

Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop
cs.LG updates on arXiv.org

Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop

・arXiv:2608.11215v1 Announce Type: cross Abstract: Simulating societies of many large language model (LLM) agents is expensive, yet the questions asked of such simulations are usually macroscopic: phase behaviour, stylised facts, and scaling with the number of agents $N$, not the cognition of any single agent. ・We turn a statistical-physics observation into a method: replace each LLM agent by a low-parameter model fitt
cs.LG updates on arXiv.org

Post-Training with Policy Gradients: Optimality and the Base Model Barrier

・arXiv:2603.06957v2 Announce Type: replace-cross Abstract: We study post-training linear autoregressive models with outcome and process rewards. ・Given a context $\boldsymbol{x}$, the model must predict the response $\boldsymbol{y} \in Y^N$, a sequence of length $N$ that satisfies a $\gamma$ margin condition, an extension of the standard separability to sequences. ・We prove that on test samples where the base model achi
cs.LG updates on arXiv.org

Pretraining large language models with MXFP4 on Native FP4 Hardware

・arXiv:2605.09825v4 Announce Type: replace Abstract: Why does full-pipeline FP4 training of large language models often diverge, even when forward activations and activation gradients remain stable? ・We address this question through a controlled study of MXFP4 quantization in transformer training, progressively enabling FP4 across forward propagation (Fprop), activation gradients (Dgrad), and weight gradients (Wgrad) w
OpenAI News

Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed

・Preview Ultrafast, a new OpenAI API service tier that runs GPT-5.6 Sol up to 14× faster. ・Powered by Cerebras, it delivers up to 750 output tokens per second.
Hugging Face Papers

Previous

Previous
WIRED

Priceline Promo Codes & Coupons: 10% Off August 2026

・Unlock massive discounts on Priceline hotels, flights, and rental cars. ・Find verified Priceline coupon codes and deals for Express Deals, student discounts, and more.
cs.LG updates on arXiv.org

Probably Approximately Correct Maximum A Posteriori Inference

・arXiv:2601.16083v2 Announce Type: replace Abstract: Computing the conditional mode of a distribution, better known as the maximum a posteriori (MAP) assignment, is a fundamental task in probabilistic inference. ・However, MAP is generally intractable, and remains hard even under many common structural constraints and approximation schemes. ・We take a novel approach inspired by multi-armed bandits, recasting MAP as a bes
cs.LG updates on arXiv.org

Probing and steering biology across Boltz-1s trunk-diffusion boundary

・arXiv:2608.11475v1 Announce Type: cross Abstract: AlphaFold3-class structure predictors pair a representational trunk, which processes sequence and context, with a diffusion module, which generates atomic coordinates. ・How biological information changes as it crosses this architectural boundary remains poorly understood. ・We analyze per-residue activations from the Pairformer trunk and diffusion module of Boltz-1 using
cs.LG updates on arXiv.org

Program Semantic Inequivalence Game with Large Language Models

・arXiv:2505.03818v3 Announce Type: replace Abstract: Large Language Models (LLMs) can achieve strong performance on everyday coding tasks, but they can fail on complex tasks that require non-trivial reasoning about program semantics. ・Finding training examples to teach LLMs to solve these tasks can be challenging. ・In this work, we explore a method to synthetically generate code reasoning training data based on a semant
cs.LG updates on arXiv.org

Prompt-Driven Exploration

・arXiv:2607.08837v2 Announce Type: replace Abstract: Exploration is essential to RL since a policy cannot improve by repeatedly sampling the behaviors it already prefers. ・Standard methods inject stochasticity in the action space, but such jitter only yields rollouts close to the original. ・Escaping a weak policy often requires global perturbations that action noise cannot produce.
cs.LG updates on arXiv.org

Quantum Port-Hamiltonian Neural Networks: Learning Conservative and Dissipative Dynamics via Measurement-Induced Nonlinearity

・arXiv:2607.12269v3 Announce Type: replace Abstract: We introduce Quantum Port-Hamiltonian Neural Networks (Q-pHNNs), parameterised quantum circuits that learn classical dynamics in a structure-preserving manner. ・The framework rests on the Isomorphic Hamiltonian Mapping (IHM): the skew-symmetric interconnection matrix $\mathbf{J}$ corresponds to unitary gate evolution, and the positive-semidefinite dissipation matrix
cs.LG updates on arXiv.org

Ranking vs. Assignment: The Metric Mismatch in Multi-View Object Association

・arXiv:2606.02022v2 Announce Type: replace-cross Abstract: Multi-view object association is an important computer vision problem that underlies many multi-camera perception tasks. ・While this task is naturally formulated as a constrained one-to-one matching problem, recent works heavily rely on pairwise ranking metrics like AP and FPR-95 for model evaluation. ・We highlight a fundamental mismatch between these metrics an
Hugging Face Papers

Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control

Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control
cs.LG updates on arXiv.org

RECAST: A Machine-Learning Framework for Correction and Super-Resolution of Coarse-Grid PDE Solvers

・arXiv:2608.11572v1 Announce Type: new Abstract: Coarse-grid numerical solvers can substantially reduce the computational cost of time-dependent PDE simulation, but under-resolution often degrades both the trajectory and the spatial fidelity of the solution. ・We introduce RECAST (Recurrent Error Correction And Super-resolution of coarse-grid Trajectories), a machine-learning framework designed to restore this lost accu
cs.LG updates on arXiv.org

ReCodeAgent: A Multi-agent Workflow for Language-Agnostic Translation and Validation of Large-Scale Repositories

・arXiv:2604.07341v3 Announce Type: replace-cross Abstract: Most repository-level code translation and validation techniques have been evaluated on a single source-target programming language (PL) pair, owing to the complex engineering effort required to adapt new PL pairs. ・Programming agents can enable PL-agnosticism in repository-level code translation and validation: they can synthesize code across many PLs and auto
Hugging Face - Blog

Record, train, and deploy from one place with Strands Agents, LeRobot, and Hugging Face Storage Buckets

Record, train, and deploy from one place with Strands Agents, LeRobot, and Hugging Face Storage Buckets
cs.LG updates on arXiv.org

RedditPersona: A Modular Framework for Community-Conditioned LLM Adaptation from Reddit

・arXiv:2606.06027v2 Announce Type: replace-cross Abstract: Community-conditioned language model adaptation needs choices about data collection, community definition, and evaluation that are currently made independently in each study, making it hard to compare assumptions or reuse artifacts. ・We present RedditPersona, a modular framework that standardizes these choices: it collects Reddit posts and comments, profiles ac
cs.LG updates on arXiv.org

Redistribution-based Cost Inference Improves Sparse Safe Offline RL

・arXiv:2608.12306v1 Announce Type: new Abstract: Safe offline RL typically assumes access to dense per-step cost annotations, but in practice supervisors provide only trajectory-level stop-feedback: a binary signal at the first unsafe transition, with no per-step attribution. ・We frame this as a temporal credit assignment problem and propose the Redistribution-based Cost Inference (RCI) framework, which converts sparse
cs.LG updates on arXiv.org

Reducing Symmetry Increase in Equivariant Neural Networks

・arXiv:2608.12010v1 Announce Type: new Abstract: Equivariant Neural Networks (ENNs) have empowered numerous applications in scientific fields. ・Despite their remarkable capacity for representing geometric structures, ENNs suffer from degraded expressivity when processing symmetric inputs: the output representations are invariant to transformations that extend beyond the input's symmetries. ・The mathematical essence of t
cs.LG updates on arXiv.org

Regime-Gated Residual Mixture-of-Experts for Cross-Sectional Volatility Forecasting

・arXiv:2608.12251v1 Announce Type: cross Abstract: Financial volatility is regime dependent, yet incorporating regime information into neural networks can also destabilize training. ・This paper asks where such information should enter a neural cross-sectional volatility forecasting model. ・We study five-day realized-volatility forecasts for 1,027 U.S.
cs.LG updates on arXiv.org

Reliable Inference in Edge-Cloud Model Cascades via Conformal Alignment

・arXiv:2510.17543v3 Announce Type: replace Abstract: Edge intelligence enables low-latency inference via compact on-device models, but assuring reliability remains challenging. ・We study edge-cloud cascades that must preserve conditional coverage: whenever the edge returns a prediction set, it should contain the true label with a user-specified probability, as if produced by the cloud model. ・We formalize conditional co
cs.LG updates on arXiv.org

RelShap: Relationally Consistent Shapley Explanations

・arXiv:2608.11508v1 Announce Type: new Abstract: Machine learning pipelines commonly flatten relational data into single-table representations, discarding structural constraints. ・Widely used Shapley value-based feature attributions then rely on feature independence, evaluating the model on combinations that could never arise in the underlying data, producing misleading explanations. ・We propose RelShap, a framework tha
cs.LG updates on arXiv.org

Remote Sensing and Machine Learning-Based Analysis of Land Use and Vegetation Change in Dhaka District, Bangladesh

・arXiv:2608.12001v1 Announce Type: new Abstract: Rapid urbanization in Dhaka District, Bangladesh has triggered substantial alterations in land use and environmental conditions, necessitating systematic monitoring for informed urban planning and ecological sustainability. ・This study employs remote sensing data and machine learning techniques to analyze spatiotemporal changes in land cover and vegetation dynamics betwe
cs.LG updates on arXiv.org

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation

・arXiv:2608.11698v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student on its own trajectories under dense token-level supervision from a teacher. ・Reward-extrapolation methods such as ExOPD amplify the teacher-reference log-likelihood ratio to move beyond direct imitation, but apply a single global coefficient $\lambda$ to every token. ・This can drive the student to fit extreme peaks in the impl
cs.LG updates on arXiv.org

Reoptimization Algorithms for Contextual Bandits with Knapsack Constraints

・arXiv:2608.11383v1 Announce Type: new Abstract: We study new algorithms for Contextual Bandits with Knapsack. ・In these problems, there are finitely many types of customers, products, and resources. ・Each product is made from a fixed combination of resources, and resources have finite capacity.
cs.LG updates on arXiv.org

Representation Finetuning for Continual Learning

・arXiv:2603.11201v3 Announce Type: replace Abstract: The world is inherently dynamic, and continual learning aims to enable models to adapt to ever-evolving data streams. ・While pre-trained models have shown powerful performance in continual learning, they still require finetuning to adapt effectively to downstream tasks. ・However, prevailing Parameter-Efficient Fine-Tuning (PEFT) methods operate through empirical, blac
cs.LG updates on arXiv.org

Reproducing and Stress-Testing Two Approaches to LLM Reasoning Reliability: Test-Time Probability Aggregation and Logic-Representation Editing

・arXiv:2608.08514v1 Announce Type: cross Abstract: We independently reproduce two recent methods for making large language model (LLM) reasoning more reliable, and stress-test them across domains and models (RPC across four new task domains with Qwen3-8B, LCF across four 7-8B models). ・The first, RPC, aggregates token probabilities and self-consistency at inference; the second, LCF, trains projectors that split hidden
cs.LG updates on arXiv.org

Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning

・arXiv:2606.04923v2 Announce Type: replace Abstract: Rubric-based reinforcement learning (RL) uses an LLM-as-a-Judge (LaaJ) to score model outputs according to rubrics as rewards. ・However, policy models may exploit latent biases in the judge, leading to reward hacking and ineffective or unsafe training outcomes. ・In real-world rubric-based RL, such hacking behaviors are often subtle and entangled with multiple judge bi
cs.LG updates on arXiv.org

Rethinking Medical Landmark Localization with Prototype Learning-based Progressive Offset Correction

・arXiv:2608.09182v2 Announce Type: replace-cross Abstract: Accurate landmark localization in medical images is a fundamental step for quantitative clinical measurement and downstream analysis. ・Existing localization methods have advanced, among which multi-stage refinement is a superior solution. ・Although this strategy mitigates the anatomical ambiguity inherent in single-stage global predictions, its high computationa
cs.LG updates on arXiv.org

Robust Ambiguity Detection (RAD) From Model- and Feature-Space Consistency

・arXiv:2608.11541v1 Announce Type: new Abstract: Machine learning models should be robust, in the sense of remaining predictively consistent under permissible variations. ・A model's predictions should ideally remain unchanged when it is replaced by a functionally equivalent one, or when its inputs are subject to minor, admissible perturbations. ・If such changes alter a prediction significantly, then the prediction is "a
cs.LG updates on arXiv.org

Robust and Efficient Noisy-Label Time-Series Classification via Dynamic Time Warping Based Granular Ball Computing

・arXiv:2608.11704v1 Announce Type: new Abstract: Dynamic Time Warping (DTW)-based Nearest-Neighbor (NN) classifiers are effective for time-series classification but are vulnerable to mislabeled training samples and require numerous DTW computations during inference. ・We propose DTW-based Granular Ball Computing (DTW-GBC), which organizes temporally similar training samples into granular balls and performs classificatio
cs.LG updates on arXiv.org

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning

・arXiv:2608.11587v1 Announce Type: cross Abstract: Recent advances in model design and self-supervised audio representations have improved speech and audio understanding, yet infant-centered naturalistic recordings remain challenging due to limited labeled data, low signal-to-noise ratio, and cross-family domain shifts. ・We present a family-conditioned, multi-tier audio tagger that combines a LoRA-finetuned Whisper enc
cs.LG updates on arXiv.org

Robustness of AI-Art Detectors under Generator Shift

・arXiv:2608.11643v1 Announce Type: cross Abstract: Text-to-image generative models have advanced rapidly, with modern Diffusion Transformer architectures producing images that are increasingly difficult to distinguish from human-created artwork. ・This development has raised significant concerns regarding copyright protection, misinformation, fraud, impersonation, and the authenticity of digital content. ・Most AI-art det
cs.LG updates on arXiv.org

Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL

・arXiv:2608.11669v1 Announce Type: new Abstract: Reinforcement learning against rubrics, lists of criteria graded by an LLM judge, has become a standard way to post-train language models on tasks with no deterministic answer. ・The rubric, however, is a fixed proxy for quality, never a complete description of it, and a policy trained against it long enough will learn to exploit the difference. ・We measure this directly.
cs.LG updates on arXiv.org

ScreenShot: A Foundation Model for Few-Shot Combination Drug Screening

・arXiv:2608.12219v1 Announce Type: new Abstract: Treating patients with combinations of drugs reduces the risk of resistance to any individual drug. ・Finding effective combinations is difficult because the large search space makes combinatorial screens prohibitively expensive, time consuming, and often technically infeasible. ・Predictive models can fill this gap, yet existing methods typically require molecular profilin
cs.LG updates on arXiv.org

SegPAR: Class-Centric Decision-Based Sparse Attack for Semantic Segmentation

・arXiv:2608.11285v1 Announce Type: cross Abstract: Despite the practical relevance of sparse decision-based black-box threats, they have received limited attention in semantic segmentation. ・To bridge this gap, we adapt the most representative decision-based black-box sparse attacks from the classification domain to serve as baselines, establishing a rigorous benchmark for this underexplored setting. ・In this context, w
Hugging Face Papers

Self-Evolving Embodied Agents via Skill-Harness Evolution

Self-Evolving Embodied Agents via Skill-Harness Evolution
Hugging Face Papers

Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models

Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models
WIRED

Shark Promo Codes for August 2026

・Shark makes some seriously powerful vacuums, from handheld vacs to steam mops. ・Don’t miss $100 off, 10% off, and more limited-time coupons from WIRED.
Hugging Face Papers

Simplex Relaxation for Discrete Diffusion

Simplex Relaxation for Discrete Diffusion
Hugging Face Papers

SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries

SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries
cs.LG updates on arXiv.org

Small-Scale Experiments: Are We There Yet?

・arXiv:2608.11859v1 Announce Type: new Abstract: Scaling laws promised cost-effective experiments; six years later, they have yet to fully deliver. ・Instead, researchers have found them unreliable at small scales (starting at 4M parameters) and concluded that sizable models cannot be avoided. ・We show this is not the case: the confounding factor is hyperparameters.
cs.LG updates on arXiv.org

Soft-Attention Improves Skin Cancer Classification Performance

・arXiv:2105.03358v4 Announce Type: replace-cross Abstract: In clinical applications, neural networks must focus on and highlight the most important parts of an input image. ・Soft-Attention mechanism enables a neural network toachieve this goal. ・This paper investigates the effectiveness of Soft-Attention in deep neural architectures.
cs.LG updates on arXiv.org

SoftWater: Class-Aware Rate Allocation for Softmax Quantization

・arXiv:2608.12026v1 Announce Type: new Abstract: Post-training quantization pipelines routinely leave the softmax output layer in high precision. ・Yet in small LLMs with modern vocabularies, the head holds 15--30\% of all parameters, so a nominal ``2-bit'' model with an fp16 head can store several times as many bits per weight. ・We pose softmax-layer quantization as a rate-distortion problem under the KL divergence betw
MarkTechPost

SpaceXAI Releases Grok 4.6: A 500K-Context Frontier Model Tuned for Long-Running Agents, Coding, and Knowledge Work

・SpaceXAI released Grok 4.6 on August 12, 2026 — a post-training upgrade over Grok 4.5, not a larger base model. ・It ties GPT-5.6 Sol Max at 61 on the Artificial Analysis Intelligence Index, ships 500K context and a new xhigh reasoning level, and holds pricing at $2/$6 per million tokens. ・The coding benchmarks are where it still loses.
Hugging Face Papers

Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill

Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill
cs.LG updates on arXiv.org

Sparse and robust geometric twin support vector machine via asymmetric RoBoSS loss function

・arXiv:2608.11567v1 Announce Type: new Abstract: In real-world scenarios, the training data usually contains redundant features, label noise and feature noise, which provide severe challenges for the efficiency of machine learning methods. ・Since standard support vector machine (SVM) adopts $l_2$-norm penalty and hinge loss function, it lacks the ability of selecting significant features and is sensitive to noise.
cs.LG updates on arXiv.org

Spectral graph clustering with inhomogeneous latent geometry

・arXiv:2608.11321v1 Announce Type: cross Abstract: We study spectral clustering in the presence of a confounding latent geometry. ・The leading eigenvectors may then be dominated by the latent geometry rather than by the communities. ・Nevertheless, we show in a block latent-space model that communities can be recovered from eigenvectors deeper in the spectrum.
cs.LG updates on arXiv.org

Spend Less, Fit Better: Budget-Efficient Scaling Law Fitting via Active Experiment Selection

・arXiv:2604.22753v2 Announce Type: replace Abstract: Scaling laws are used to plan multi-million-dollar training runs, but fitting those laws can itself cost millions. ・In modern large-scale workflows, assembling a sufficiently informative set of pilot experiments is already a major budget-allocation problem rather than a routine preprocessing step. ・We formulate scaling-law fitting as budget-aware sequential experiment
stat.ML updates on arXiv.org

Stability of Finite-Batch Particle Mean-Field Variational Inference Beyond Strong Convexity

・arXiv:2608.11486v1 Announce Type: cross Abstract: We study the implementable finite-batch particle algorithm for mean-field variational inference as a fully discrete stochastic approximation of the projected Wasserstein dynamics. ・The target potential is globally smooth but need not be strongly convex. ・The departure from contractivity is quantified by the curvature defect \[ \mathfrak d_\alpha(x,y) = \bigl[\alpha\|x-y
Hugging Face Papers

StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization

StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization
cs.LG updates on arXiv.org

SteeringSafety: Benchmarking Representation Steering in LLMs Across Safety Perspectives

・arXiv:2509.13450v3 Announce Type: replace-cross Abstract: We introduce SteeringSafety, a benchmark for evaluating representation steering methods across nine safety perspectives spanning 18 datasets. ・While prior work highlights the general capabilities of representation steering, we focus on safety perspectives including refusal, bias, hallucination, social behaviors, reasoning, epistemic integrity, and normative jud
cs.LG updates on arXiv.org

Stochastic Dimension Zeroth-Order Estimator: Stable and Memory-Efficient Training of PINNs

・arXiv:2603.24002v4 Announce Type: replace Abstract: Physics-Informed Neural Networks (PINNs) for high-dimensional and high-order partial differential equations (PDEs) are primarily constrained by the $\mathcal{O}(d^k)$ spatial derivative complexity and the $\mathcal{O}(P)$ memory overhead of backpropagation (BP). ・While randomized spatial estimators successfully reduce the spatial complexity to $\mathcal{O}(1)$, their
The Verge

Suno is trying to look more like a real music production tool

・Suno is releasing Studio 2.0 with significant upgrades that push it closer to an actual digital audio workstation (DAW), rather than a bare-bones audio editor with generative AI features. ・The biggest addition is undoubtedly MIDI support. ・Suno says that MIDI was its most requested feature, and it's basically a prerequisite for any modern DAW.
cs.LG updates on arXiv.org

Superposition Without Interference? Towards Isolated Interventions via Almost Orthogonal Features in Language Models

・arXiv:2602.04718v5 Announce Type: replace Abstract: A central premise in mechanistic interpretability is that meaningful concepts in language models are represented by linear features in activation space. ・For such features to support reliable interventions, manipulating one feature should not substantially alter the effects of others. ・In practice, however, feature entanglement leads to interference such that localize
cs.LG updates on arXiv.org

Symbolic Machine Learning for Vapor-Liquid Equilibrium Prediction in Cx-N2 Binary Mixtures

・arXiv:2608.11255v1 Announce Type: cross Abstract: Accurate prediction of vapor--liquid equilibrium (VLE) for hydrocarbon-nitrogen mixtures remains challenging for cubic equations of state, particularly across broad ranges of composition and hydrocarbon chain length. ・While deep learning models can provide accurate predictions, they often lack interpretability and explicit analytical expressions. ・In this work, we propo
cs.LG updates on arXiv.org

TailBooster: A Dual-Layer Generative Framework for Extreme Value Augmentation with Operational Validity Enforcement

・arXiv:2608.11951v1 Announce Type: new Abstract: Extreme events in air transport, such as severe arrival delays and abnormal air times, cause cascading network disruptions with substantial operational, economic, and safety costs. ・Such events are rare in historical records, leaving insufficient training signal for machine learning models. ・Synthetic data augmentation offers a principled solution, but conventional genera
cs.LG updates on arXiv.org

Task- and dataset-specific information in protein language models

・arXiv:2608.12090v1 Announce Type: new Abstract: Protein language models (PLMs) have transferred the latest advances from natural language processing to computational biology. ・These models, trained on large corpora of protein sequence data, are widely used to translate amino acid sequences into latent-space embeddings, ready for use in diverse downstream tasks (DTs). ・By a common consensus, embeddings from the model's
cs.LG updates on arXiv.org

Temperature-Driven Sequential Modeling for the Prediction of Annual Power Conversion Efficiency Profiles of Organic Photovoltaic Materials: Douala Case Study

・arXiv:2608.11261v1 Announce Type: cross Abstract: Organic photovoltaic (OPV) materials are promising candidates for distributed solar energy in tropical regions, yet existing virtual screening tools report static power conversion efficiency (PCE) values at standard testing conditions (STC) that fail to capture the temperature-driven performance degradation experienced under real deployment conditions. ・Here we introdu
cs.LG updates on arXiv.org

Terminal Symmetry as a Decision Resource: Statewise Refinement for Anytime Verified Construction

・arXiv:2608.11318v1 Announce Type: new Abstract: Many sequential construction tasks exhibit exact symmetry at completion while their execution remains directed and history-dependent. ・We develop a decision-resource view of terminal symmetry: process evidence supplies directionality, terminal correspondence transports that structure across equivalent outcomes, realized-state evidence refines its current decision relevan
cs.LG updates on arXiv.org

TESLA: Taylor Expansion of Sinusoidal Learnable Activations

・arXiv:2608.11970v1 Announce Type: new Abstract: The parity problem--deciding whether the number of ones in a binary vector is odd or even--remains challenging for standard neural networks due to linear inseparability and the need for global interactions. ・We propose TESLA, an activation defined as a learnable combination of sine and cosine terms, enabling explicit control over polynomial degree and selective amplifica
WIRED

The 10 Best Cooling Mattresses for Hot Sleepers (2026)

・Nothing ruins a great night of sleep faster than getting too hot. ・We slept on a myriad of cooling mattresses to find which ones drew the heat away best.
WIRED

The 10 Best Cooling Mattresses for Hot Sleepers (2026)

・Nothing ruins a great night of sleep faster than getting too hot. ・We slept on a myriad of cooling mattresses to find which ones drew the heat away best.
cs.LG updates on arXiv.org

The Advective Fisher-Rao Geometry of Deterministic Measure Transport

・arXiv:2608.12111v1 Announce Type: cross Abstract: A novel advective Fisher-Rao metric is introduced for optimization tasks on paths of probability measures governed by the continuity equation. ・This metric is shown to lead to optimal descent directions. ・It is then shown that this metric arises naturally from three different perspectives: As the rescaled zero-noise limit of the Fisher-Rao metric on path measures, as th
WIRED

The Best Samsung Galaxy S26 Cases (2026): S26, S26+, and S26 Ultra

・Protect your Samsung phone with these cases and screen protectors.
The Verge

The Corvette Grand Sport X delivers Porsche 911 performance for a fraction of the price

・My drive of the 2027 Corvette Grand Sport X began under oily black clouds, a torrential weather front releasing its grip on Manhattan - an inauspicious start for any mega-powered sports car. ・Rain pelted the waterlogged pavement, as I set course for the mountain-man roads of the Catskills, then on to Long Island and New England over three days. ・Fortunately, this Grand Sport, a name synonymous with value among Corvette
WIRED

The Google Pixelsnap Charger With Stand Is 50 Percent Off Right Now

・Speedy Qi2 wireless charging in a magnetic puck and stand combo makes Google’s Pixelsnap well worth grabbing at half price.
Hugging Face Papers

The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images

The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images
cs.LG updates on arXiv.org

The Off-Support Barrier: Why Semantic Safety Constraints Are Not Learning-Problem Invariants, and What Follows for Prior Design, Containment, and Verification

・arXiv:2608.11243v1 Announce Type: cross Abstract: We argue that a single structural fact organizes a wide range of phenomena in contemporary AI safety: a semantic safety constraint (e.g., the agent does not escape its sandbox) is an off-support object. ・Formally, if q is the data distribution and \(p(\cdot\mid w)\) the model, the safety predicate B is not measurable with respect to \(\sigma(\text{model}, q)\), whereas
WIRED

The Painful Truth of Exactly How ICE’s New Shock Gloves Work

・ICE is spending millions on shock gloves designed to overpower subjects through intense, localized pain.
WIRED

There’s a Fatty Liver Epidemic. AI Could Help Get Ahead of It

・Over a billion people worldwide have livers with excess fat, which can lead to a host of medical problems. ・Researchers think AI tools can spot the condition—and help stop it—early enough to save lives.
The Verge

This is Instagram&#8217;s new logo

・It’s certainly… different. ・| Image: Meta Instagram has unveiled a new wordmark, moving away from the recognizable cursive typeface it's used over the last decade. ・The updated wordmark is a strange mix of half-cursive half-print that's somehow less legible than its predecessor.
cs.LG updates on arXiv.org

Three Tokens Force Exponential Feature Rank in Nonnegative Kernel Attention

・arXiv:2608.11427v1 Announce Type: new Abstract: Full attention exposes every token pair, whereas kernel attention compresses a sequence into a fixed-dimensional sketch. ・We show that this distinction becomes exponential at the first context length containing two competing candidates. ・On Min-IP over Boolean inputs, rank-one normalized kernel attention solves every sequence of length at most two exactly.
cs.LG updates on arXiv.org

Tight Nonasymptotic Local Convergence of Sinkhorn-Knopp

・arXiv:2608.11760v1 Announce Type: cross Abstract: We revisit the Sinkhorn-Knopp (SK) algorithm for the matrix scaling problem. ・Despite extensive literature on the global convergence of SK and its variants, its local linear convergence behavior remains less understood. ・We address this gap by providing the first nonasymptotic local analysis of SK that matches the rate obtained from existing asymptotic Jacobian-based ar
cs.LG updates on arXiv.org

TMRL: Diffusion Timestep-Modulated Pretraining Enables Exploration for Efficient Policy Finetuning

・arXiv:2605.12236v2 Announce Type: replace-cross Abstract: Fine-tuning pre-trained robot policies with reinforcement learning (RL) often inherits the bottlenecks introduced by pre-training with behavioral cloning (BC), which produces narrow action distributions that lack the coverage necessary for downstream exploration. ・We present a unified framework that enables the exploration necessary to enable efficient robot po
Hugging Face Papers

ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents

ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents
cs.LG updates on arXiv.org

Towards a Formal Definition of Agent Memory: Basis, Span, Optimality, and the Sequential Memory Problem

・arXiv:2608.11654v1 Announce Type: new Abstract: Despite the wide deployment of memory in large-model agents, there is no unified formal account of what a memory is or when it is optimal. ・This paper takes a first step toward this account. ・The central idea is that memory is a basis, knowledge is its span, and answerability is a coverage problem: an agent stores events extracted from a material; a generation operator tu
cs.LG updates on arXiv.org

Towards an approach to multivariate outlier detection for District Heating System data

・arXiv:2608.11375v1 Announce Type: new Abstract: In this paper, we test different methods for multivariate detection of outliers in the data of transmitted heat energy in the selected substation of local District Heating System, by also considering outside ambient temperature, namely Z-score (univariate, as a benchmark), Mahalanobis distances, Principal Component Analysis (PCA), Isolation Forest and Hotelling's T-squa
cs.LG updates on arXiv.org

Towards Sustainable Learning in Online Education: A Reinforcement Learning Approach

・arXiv:2608.11245v1 Announce Type: cross Abstract: Online education offers unprecedented scalability and accessibility to global learners from diverse backgrounds, but it often suffers from low engagement and poor long term learning effectiveness. ・To address these challenges, we introduce AI Tutor, a reinforcement learning based model designed to promote sustainable learning by optimizing both short and longterm learn
cs.LG updates on arXiv.org

Towards the Harness of Embodied Agents

・arXiv:2608.11246v1 Announce Type: cross Abstract: The success of coding agents has established the harness as a paradigm: what an agent achieves depends not on the model alone, but on the infrastructure around it. ・We ask whether the same paradigm extends to embodied agents in the physical world. ・We present Thea, a harness in which an agentic loop orchestrates robot capabilities, each wrapped as a callable tool.
cs.LG updates on arXiv.org

Towards Truly Unsupervised Evaluation of Feature Selection

・arXiv:2608.12057v1 Announce Type: new Abstract: Feature selection is one of the most important and fundamental tasks in data mining, tackled by a family of methods with an established set of evaluation techniques to measure the quality of a specific method. ・Most of the methods commonly used for the unsupervised evaluation of feature selection algorithms suffer from critical design flaws which question their unsupervi
cs.LG updates on arXiv.org

Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling

・arXiv:2608.11829v1 Announce Type: new Abstract: On-policy distillation (OPD) has emerged as a promising post-training technique for enhancing LLM reasoning. ・It is commonly believed to enable the student model to distill knowledge from a stronger teacher model, thereby expanding capabilities beyond the pre-OPD base model. ・In this study, we examine this view through the lens of test-time scaling by varying the sampling
cs.LG updates on arXiv.org

TradingMoE: Routing the Right Experts in Evolving Markets

・arXiv:2608.11785v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong potential for financial analysis and trading, but direct trading remains challenging because the predictive capabilities required can vary across assets, decision fields, and market conditions. ・Existing LLM-based trading systems either coordinate human-defined external experts or adopt conventional internal Mixture-of-Exper
cs.LG updates on arXiv.org

Transit Destination Inference from Tap-In-Only Bus Smart-Card Data: A Hierarchical Bayesian Approach

・arXiv:2608.11223v1 Announce Type: cross Abstract: Entry-only automatic fare collection systems record boardings but not alightings, preventing direct construction of origin-destination (OD) matrices. ・This study develops a Hierarchical Bayesian Latent-Destination (HBLD) model that combines station-hour boarding and inferred alighting demand with passenger card histories. ・Trip-chain destinations are treated as noisy ev
cs.LG updates on arXiv.org

Trust Region Constrained Bayesian Optimization with Penalized Constraint Handling

・arXiv:2603.24567v2 Announce Type: replace-cross Abstract: Constrained optimization in high-dimensional black-box settings is difficult due to expensive evaluations, the lack of gradient information, and complex feasibility regions. ・In this work, we propose a Bayesian optimization method that combines a penalty formulation, a surrogate model, and a trust region strategy. ・The constrained problem is converted to an unco
cs.LG updates on arXiv.org

Uncertainty-Aware and Explainable Ensemble Deep Learning Framework for Multi-Class Skin Lesion Classification

・arXiv:2608.11280v1 Announce Type: cross Abstract: Skin cancer diagnosis from dermoscopic images remains challenging due to high intra-class variability, inter-class similarity, class imbalance, and the limited interpretability of deep learning models. ・This paper proposes an uncertainty-aware and explainable deep learning framework for multi-class skin lesion classification. ・The framework combines a vision transformer
cs.LG updates on arXiv.org

Uncertainty-Aware Compositional Localization and Placement Assessment of Catheters and Tubes in Chest X-Rays

・arXiv:2608.11288v1 Announce Type: cross Abstract: Assessing catheter and tube placement on chest X-rays is safety-critical yet tedious and error-prone. ・Current deep learning methods either classify placement globally -- losing track of which device is where -- or segment all devices into a single mask, making per-device assessment impossible when catheters overlap. ・We introduce UCompCXR, a compositional framework tha
cs.LG updates on arXiv.org

Uncertainty-Aware Probabilistic Constrained Clustering from Entangled Pairwise Supervision

・arXiv:2608.12027v1 Announce Type: new Abstract: Pairwise constrained clustering typically relies on hard must-link/cannot-link labels, whereas realistic pairwise supervision may be real-valued and entangle intrinsic ambiguity, expert judgment, and stochastic corruption. ・Existing deep constrained clustering (DCC) methods mainly target hard, expert-agnostic constraints, treating soft labels mostly numerically rather th
cs.LG updates on arXiv.org

Unifying Physical Backpropagation

・arXiv:2608.11585v1 Announce Type: cross Abstract: Physical computing systems exploit device dynamics for computation, but their gradient-based optimization is challenging: backpropagation through a digital twin suffers from model-reality gap. ・On-device gradient computation could resolve this issue, and a handful of theoretical and experimental studies have proposed ways to achieve it. ・Yet a unifying theory identifyin
cs.LG updates on arXiv.org

Unmasking Toxic Mimicry in Medical Offline Reinforcement Learning for ICU Sepsis Management via Counterfactual Clinical Audits

・arXiv:2608.11410v1 Announce Type: new Abstract: Offline reinforcement learning (RL) offers considerable promise for optimizing ICU treatment decisions, yet standard evaluation metrics Mean Squared Error (MSE) and Fitted Q-Evaluation (FQE) assess only behavioral imitation and cannot detect Toxic Mimicry, a failure mode in which agents replicate harmful patterns such as treatment withdrawal during comfort-care transiti
cs.LG updates on arXiv.org

User-Assisted Collaborative Distributed Inference for Efficient QoS-Aware Autoscaling

・arXiv:2608.11840v1 Announce Type: cross Abstract: Growing demand for artificial intelligence (AI) inference services requires scalable infrastructure, yet centralized serving costs rise with demand. ・We propose a collaborative distributed inference system combining dedicated infrastructure with resources contributed by service users. ・Dedicated resources provide baseline capacity for maintaining quality of service (QoS
cs.LG updates on arXiv.org

Variable Selection in the Context of AI Fairness

・arXiv:2608.11251v1 Announce Type: cross Abstract: Fairness in AI systems has become more important with recent regulatory demands, such as the EU AI Act. ・Traditional approaches often do not take into account philosophical ethics and social awareness. ・Variable selection processes, in particular, can introduce implicit bias, affecting equity across different subgroups.
cs.LG updates on arXiv.org

Variational Mixture of Graph Neural Experts for Alzheimer's Disease Recognition across Frequency Bands in EEG Brain Networks

・arXiv:2510.11917v4 Announce Type: replace Abstract: Dementia disorders such as Alzheimer's disease (AD) and frontotemporal dementia (FTD) exhibit overlapping electrophysiological signatures in electroencephalography (EEG) that challenge accurate diagnosis. ・Existing EEG-based methods are limited by full-band frequency analysis, which hinders precise differentiation of dementia subtypes and severity stages.
cs.LG updates on arXiv.org

Variational Parameter Calibration with Physics-Aware Latent-Space Surrogates

・arXiv:2608.11435v1 Announce Type: new Abstract: Forward and inverse modeling of parametric dynamical systems requires surrogate models that are not only accurate for state prediction, but also informative for parameter calibration. ・However, a systematic end-to-end differentiable formulation for coupling deep-learning-based reduced-order surrogates with variational parameter estimation remains underdeveloped.
WIRED

Visible Promo Codes and Coupons for August 2026

・Find great deals and promo codes for Visible at WIRED and save big, whether you're a long-time customer or a newbie.
cs.LG updates on arXiv.org

WavePhaseNet: A DFT-Based Method for Constructing Semantic Conceptual Hierarchy Structures (SCHS)

・arXiv:2602.14419v2 Announce Type: cross Abstract: This paper reformulates Transformer/Attention mechanisms in Large Language Models (LLMs) through measure theory and frequency analysis, theoretically demonstrating that hallucination is an inevitable structural limitation. ・The embedding space functions as a conditional expectation over a {\sigma}-algebra, and its failure to be isomorphic to the semantic truth set fund
cs.LG updates on arXiv.org

WBMM: Windowed Batch Matrix Multiplication for Efficient Large Receptive Field Convolution

・arXiv:2607.02097v2 Announce Type: replace-cross Abstract: Large kernel depthwise convolutions achieve strong performance but suffer from significant degradation as kernel size grows due to irregular memory access from gather-based computation; while Large Kernel Acceleration (LKA) helps on small feature maps, it becomes counterproductive on large feature maps, even slower than non-accelerated implementations.
cs.LG updates on arXiv.org

Weak Correlations as the Underlying Principle for Linearization of Gradient-Based Learning Systems

・arXiv:2401.04013v3 Announce Type: replace Abstract: Deep learning models, such as wide neural networks, can be conceptualized as nonlinear dynamical physical systems characterized by a multitude of interacting degrees of freedom. ・Such systems in the infinite limit, tend to exhibit simplified dynamics. ・This paper delves into gradient descent-based learning algorithms, that display a linear structure in their parameter
cs.LG updates on arXiv.org

Weaves, Wires, and Morphisms: Formalizing and Implementing the Algebra of Deep Learning

・arXiv:2604.07242v3 Announce Type: replace Abstract: Despite deep learning models running well-defined mathematical functions, we lack a formal mathematical framework for describing model architectures. ・Ad-hoc notation, diagrams, and pseudocode poorly handle nonlinear broadcasting and the relationship between individual components and composed models. ・This paper introduces a categorical framework for deep learning mod
cs.LG updates on arXiv.org

Weightless Fine-Tuning: Personalizing LLMs via Logit-Space Transport

・arXiv:2608.11342v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) is a standard approach for adapting LLMs to a target distribution, but in settings such as personalization, where each author requires separate weight access, optimization, storage, and retraining, its costs become prohibitive. ・We propose Weightless Fine-Tuning (WFT), a training-free decoding-time method that approximates the distributional
cs.LG updates on arXiv.org

What Iterated Self-Feeding Probes of Language Models Measure, and a test that separates the construction from the model

・arXiv:2608.10986v1 Announce Type: cross Abstract: A growing class of methods probes a language model by feeding it its own output: self-consistency, iterated refinement, agentic loops. ・We ask what such a probe measures, in a construction chosen to make the question sharp: a ring of token cells resampled in place by the model's own windowed conditional p_r(x_i | x_{i+-r}). ・The substrate is Glauber dynamics on token se
cs.LG updates on arXiv.org

When Offline Evaluation Misleads: A Diagnostic Protocol for Reward and Policy Selection in Delayed-Feedback Contextual Bandits

・arXiv:2608.11560v1 Announce Type: new Abstract: Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online learning. ・Teams therefore train the bandit on a fast proxy reward, and separately must judge whether a contextual bandit is worth its complexi
cs.LG updates on arXiv.org

When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs

・arXiv:2608.11403v1 Announce Type: cross Abstract: Self-consistency (SC) via majority vote is a widely used way to spend inference-time compute: sample N chains of thought, return the plurality answer. ・On the full GPQA Diamond benchmark (198 graduate-level science questions), majority voting reduces per-problem accuracy on a majority of problems for two instruction-tuned models from different families: 56.6% of proble
cs.LG updates on arXiv.org

Why AI Detection Fails for Academic Integrity

・arXiv:2608.11256v1 Announce Type: new Abstract: Institutions use commercial AI detectors for academic integrity, yet detectors cannot distinguish AI editing from full LLM drafts and may treat both as misconduct. ・In a controlled study of published English abstracts (four domains; 2013 to 2015 vs. ・2023 to 2025), we quantify this policy failure under proxy human/AI labels at tau=0.50.
WIRED

Why This Prediction Market Banned Teens

・Jacob Fortinsky, the CEO and cofounder of the new prediction market Novig, says his outfit isn’t like those other markets. ・You know the ones.
The Verge

Wolverine on the PS5 made me care about Marvel’s mutant for the first time

・Going into my demo of Marvel's Wolverine, I mainly wanted to try out the game's very bloody combat. ・And it is delightfully bloody. ・But after two-ish hours of slashing and stabbing (and parrying!) enemies with Wolverine's claws, I left surprised - despite no previous attachment to the character, I was deeply endeared to Logan himself.
cs.LG updates on arXiv.org

XGBoost "is all you need": the case of forecasting transmitted heat energy in District Heating Systems

・arXiv:2608.11446v1 Announce Type: new Abstract: This paper presents a comparative study of two distinct approaches, XGBoost and Long-Short Term Memory (LSTM), for forecasting transmitted heat energy in District Heating Systems (DHS). ・The objective is to explore scenarios in which conventional ML algorithms demonstrate better performance over deep learning networks in time series forecasting and the associated benefit
The Verge

You can now just point at a mess and this robot vacuum will suck it up

・The Matic robot vacuum now offers gesture controlled spot cleaning and a built-in voice assistant. ・| Photo by Jennifer Pattison Tuohy / The Verge Matic, my current favorite robot vacuum, just got a big upgrade. ・The company has launched Matic Cues, which brings voice and gesture control to the robot.
Zennの「大規模言語モデル」のフィード

クラウドAIに入力できない情報をどう扱うか?llama.cppで学ぶローカルLLM入門をUdemyで公開した話

・こんにちは。miharubaの池谷です。 ・今日、Udemyで講座「llama.cppで学ぶローカルLLM入門 ― GPU・コンテキスト・KVキャッシュからRAGまで」を公開しました。この記事では、講座で何を扱っているか、なぜこのテーマを選んだかを書きます。 ・きっかけ: クラウドAIに入力できない情報がある 会社の機密情報や顧客データを、ChatGPTのようなクラウドAIに入力してよいか迷う場面があります。多くの会社では、社外のAIサービスに社内文書を渡すことが禁止されているか、少なくとも慎重な判断が必要です。かといって、AIを何も使わないのはもったいない。
#LLMタグ

クラウドGPU代を払わずにAIモデルを鍛える、Unsloth Studioの仕組みと始め方

・毎月のAI関連のサブスク費用、気づけばじわじわ増えていませんか。 ・ChatGPTやClaudeのような会話型AIだけでなく、画像生成や音声認識まで含めると、月額料金は積み重なっていきます。 ・そんな中、手元のパソコンでAIモデルを学習・実行できる無料ツール「Unsloth Studio」が、2026年8月10日に新しいDesktopアプリとして公開されました。
#LLMタグ

クローズドLLMから自律型ローカルへ進化するサイバーセキュリティAIの最前線

・サイバーセキュリティAIの世界で、今、巨大な地殻変動が起きています。圧倒的な脆弱性検出力を誇る一方、厳しいアクセス制限や過剰なセーフガードにより実戦で「沈黙」することもある大手の商用クローズドLLM。これに対抗すべく台頭したのが、ローカル環境で自律的に動く「インディーズLLM」です。本記事では、メガスタジオとローカルAIが激突する最新トレンドを解説し、次世代のセキュリティ対策の全貌に迫ります。 ・サイバーセキュリティAIが「クローズドLLM」から「自律型ローカルエージェント」へと進化するパラダイムシフトを提示しています。画面中央には、この進化の過程を示す力強いタイトルが配置され、その下には「クローズドLLMから自律型ローカルエージェントへの進化」という具体的な説明が示されています。これは、AI技術が単なる大規模言語モデルの利用から、より独立して機能するエージェントへと発展していく未来を示唆しています。
#LLMタグ

クロスボーダー開発におけるベンダーロックインの回避:オープンソースLLMとコンテナ技術を軸にしたシステム設計

・本文 生成AIの爆発的な普及により、多くの企業が独自のAIアプリケーション開発に乗り出しています。しかし、その裏で深刻な課題となっているのが**「ベンダーロックイン」**のリスクです。
ITmedia NEWS 最新記事一覧

サービス終了した「Link!Like!ラブライブ!」開発元が破産 負債103億円超 東京商工リサーチ

・東京商工リサーチは8月13日、スマートフォンアプリ「Link!Like!ラブライブ!」を開発・運営していたオッドナンバーと関連2社が、東京地裁から破産開始決定を受けたと報じた。決定は8月5日付。
#AIタグ

スマホだけで作れる!「手のひらサイズの赤ちゃん」をAIで踊らせる方法👶🏻🕺

・Instagramに「手のひらサイズになった赤ちゃんが踊るAI動画」を投稿したところ、 「これどうやって作ってるんですか?」 「作り方を教えてほしい!」 というコメントやDMをたくさんいただきました。 ・そこで今回は、僕が実際に作った方法を、スマホだけで再現できるように最初から最後までまとめます。 ・ただ作り方を説明するだけではありません。
機械学習タグが付けられた新着記事 - Qiita

セルフ夏期講習4日目:RNA-seqデータをPCAで可視化し、癌種を分類する

・はじめに セルフ夏期講習の4日目として、RNA-seqデータを用いた癌種推定に取り組んだ。 ・今回使用したのは、UCI Machine Learning RepositoryのGene Expression Cancer RNA-Seqである。801検体について20,531...
Zennの「大規模言語モデル」のフィード

ゼロからつくるDeep Learning 6のLLM事前学習をAmazon SageMaker AIでやってみる

・はじめに この数ヶ月で、Kimi K3のようなAnthropicやOpenAIが開発するフロンティアモデルに匹敵するLLMが出現したり、パラメータサイズは小さいものの、特定のタスクや簡単なタスクでは十分な性能をもつような言語モデルが次々と登場しています。 ・そんな中でただLLMを使うだけではなく自分で何か手を動かして作ってみたいと思っていたところ、「ゼロから作るDeep Learning 6」が発売されていたことを知りました。 ・https://www.oreilly.co.jp/books/9784814401611/ この本では、LLMの学習データをトークン化するトークナイザーから、...
Zennの「大規模言語モデル」のフィード

そのdenyルール、効いていません——settings.json堅牢化チートシート

・結論から Claude Codeの settings.json には、書けてしまうのに一切適用されないルール形式があります。公開されている設定149件を走査した調査では、16%の設定に「死んだルール」が含まれていた——.ssh や .aws を守っているつもりのdenyが、実際には何も守っていなかったそうです(出典)。 ・私も自環境を監査しました。deny 153ルール中、26件が死んだ形式でした。幸い全件に有効な形式の同項ルールが併記されていて実害ゼロでしたが、「たまたま無事だった」だけです。 ・この記事は、(1) 効かない設定の見分け方、(2) 自環境を30秒で監査するスクリプト、(...
#AIタグ

タイトル:AIアプリ、7個入れてるのに2個しか使ってない件

・AIツール、何個入れてますか? スマホのホーム画面、久しぶりに整理してみたら、AI系のアプリが7個も入っていました。
#LLMタグ

テキスト透かしは「ないよりマシ」 ー EU AI Actが生んだ技術的儀式の中身

・こんにちは、さししです。 ・Anthropicが2026年8月、AI生成テキストに透かしを入れると発表しました。 ・気になって調べたので、わかったことをまとめます。
ITmedia NEWS 最新記事一覧

ドコモ・バイクシェア、全エリアでサービス再開 全面停止から9日ぶり

ドコモ・バイクシェア、全エリアでサービス再開 全面停止から9日ぶり
Zennの「大規模言語モデル」のフィード

なぜ、私が「AIの発展でITエンジニアは消滅する」vs「AIが発展してもITエンジニアはずっと必要」の論争を冷笑的に見るのか?

・はじめに AIが急速に発展するようになってから、「いずれITエンジニアはAIに仕事を奪われて消滅する」という議論と、「どれだけAIが発展してもITエンジニアは必要であり続ける」という議論を頻繁に見かけるようになりました。私はこの論争をかなり冷笑的に見ています。その理由は、両者が未来についての物語をぶつけ合っているように見えるからです。 ・「AIがすべてを代替する」という物語と、「最後には人間にしかできない仕事が残る」という物語を戦わせても、それは多くの場合、未来についての感想を言い合っているにすぎません。必要なのは、どちらの物語が心地よいかを選ぶことではなく、AIによってどの仕事が消え...
#AIタグ

ぬいのはら綿〜ぬい活のすヽめ#13 ヌイグルニア郵政省

・ぬいのMondayくん 高さ30センチとはきいてたけど、直径15センチとはきいてないσ(◉ ᴥ ◉’).。oஇ 続きをみる
#AIタグ

プラチナバッジを1ヶ月で獲ったCodex×ココナラ超実践活用法9選 ─ 最近私がやった全部を教えます

・<☘️私のここまでの歩み☘️> 2026年4月から運用START ココナラ2店舗:プラチナランク ココナラで1ヶ月でプラチナバッジを獲得 ココナラ運用サポート:現在5社(事業主) 霊視歴6年、相談実績11,000件以上 6年間の瞑想修行によりサードアイ開眼 過去の記憶が81個ある Webマーケティング歴10年以上 提携事業者がAI×Canvaテンプレ活動が西日本新聞等に掲載 2021年占い館が財経新聞になる 関連SNS実績:最大249万回再生、テンプレ動画万再生の訴求素材あり 疑われることもありますが、ライブでもずっと顔を出しているので、全て本当の実績です。 ・占い師になったばかりの登竜門と言えばやはりココナラ🌈 続きをみる
ITmedia NEWS 最新記事一覧

ヨネックス、公式ECショップに不正ログイン 氏名や住所、購入履歴が閲覧された恐れ パスワード変更呼び掛け

・ヨネックスは8月12日、公式オンラインショップで第三者による不正ログインが発生したと発表した。リスト型攻撃によるもので、利用者に対しパスワードの再設定を呼び掛けている。
#AIタグ

ランサーズの副業テスト課題、応募時の説明と3日後のマニュアルで条件が真逆だった話

・AIだけでどこまで稼げるか実験中。今日は、副業案件の選考が一歩進んだと思ったら、思わぬ食い違いに気づいた話です。 ・実績ゼロから応募した案件、選考が一段階進んだ 続きをみる
Zennの「大規模言語モデル」のフィード

ローカルAIとContextLengthを超えた「思い出」を作る #4

・~それ、お前が言ったんやないかい!~ 1.ナチュラルに起きる「発言者の混同」 会話文でLLMとチャットしてると、「前にこう話してたよね?」という形で過去の発言を引用することがある。そのとき、エージェントが過去にした発言を、さも僕が言ったことのように引用することがある。具体例を出せたらいいんだけれど…例えば、 [user]A(主題)って、B(解釈)として理解できるよね? [agent]そうだね。(中略)だからAはBといえるし、C(別の解釈)という側面もあるんだ。 ・[user]はえ~。でもCは無理筋じゃない? …… [user]前にAの話をしたよね。Bと解釈できる、と整理したと思うけ...
Zennの「大規模言語モデル」のフィード

映画のポッドキャストのLLM Wiki(っぽいもの)を作る

・昨年から映画に関するポッドキャストを2人でやっています。映画の感想だけではなく映画史を絡めて話しており、監督名や作品名、概念等の固有名詞が頻出します。私は主に聞き役のため、相槌をどうするかに気を取られてあまり内容を覚えていないことが多々あり( 😇 )、ある監督や作品についてどのような話をしたのかや、過去に話した固有名詞について後から見返せるようにしたいとしばしば思っていました。そんな折にLLM Wikiというワードを目にするようになったことで「ポッドキャストの音源をLLM Wiki化してしまえば良いのでは?」と思い立ち、作ってみることにしました。そうして実際に作った結果、最終的にAndr...
Zennの「大規模言語モデル」のフィード

夏の読書感想文:『LLMの原理、RAG・エージェント開発から読み解く コンテキストエンジニアリング』

・ちょっと時間ができたので、体系的にエージェント開発・コンテキストエンジニアリングを学びたくて、久々にzenn記事投稿してみました。 ・読んで得られたものや気づきなどを、備忘録もかねて、章ごとに気ままに書いていこうと思います。 ・https://www.bing.com/ck/a?!&&p=523292588e5856febbd960e6afc6fcd157df17bac9bc0b1feac6dad7466fc581JmltdHM9MTc4NTgwMTYwMA&ptn=3&ver=2&hsh=4&fclid=3e601640-f6e5-696b-...
#AIタグ

花巻東の赤間くんのタッチアップが良すぎた件

・どうもオールドマンです。 ・いやぁ、甲子園が盛り上がっとりますね。 ・今大会で僕が「これはめちゃくちゃ良い走塁やな」と思ったプレーがありました。
#AIタグ

拡散モデルを使って時系列データからリー代数を創発することに成功しました

・やったことの要約 拡散モデルに2次元の時系列データを学習させる 時系列データは1進むごとにx方向に+1, 0, -1, y方向に+1, 0, -1だけ移動する データを食べさせると拡散モデルがリー代数を理解したかのように振る舞った 続きをみる
#LLMタグ

驚いたよ…。テキストファイルのどこに「電子透かし」をいれるのか?

・それでは君はAIと友だちになれたのか、と問われれば答えはNOだ。 ・ていうか、僕はそもそもリアル人間の友だちだってほとんどいない。
#AIタグ

作・人の価値がどんどん上がる

・クロードが文章の中に電子透かしを入れると話題になっている。なっていた 文章を見て、AIかどうかを判断できるというのは、 文章を見て、人が作ったかどうかを判断できるということで、 人が作った文章の価値がどんどん上がるんじゃ無いかと感じた。今は見分けがつかないから、価値のつけようが無い。
#AIタグ

私は、AIに聞いてみた

・前回の記事 私は、AIをうまく使えていない [リンク] なので、私はAIに聞いてみた。 ・「AIをうまく使えていない人の特徴を教えて」 今日はこれで返ってきた内容について読み解いていこうと思う。
#AIタグ

私は、AIをうまく使えていない

・私はAIを何かに有効活用したくて日々模索中だ。世間はAIの話題で溢れ、良い部分も悪い部分も流れてくる。私も何かやりたいと思い、手を出しているところだ。スケジュール管理、食事管理、記事執筆やアイデア出しなど。手近なところから試しているが、なかなかうまくいかない。 ・スケジュール管理は、元からカレンダーに予定を入れていたので、すんなりと導入することができた。といっても連携させて、私の代わりにカレンダーに予定を記入してもらうだけなので、これが有効活用できているかと言われれば微妙な気がする。
Zennのトレンド

続・貧者のアークテクチャ:Next.js + Cloudflare Workers + Turso 本番運用で踏んだ罠ぜんぶ

・こんにちは、@nabettuです。 ・Webサービスの運営コストは、安ければ安いほど嬉しいですね! 以前「貧者のアークテクチャ」という記事で、Next.jsをCloudflareで動かしつつFirebaseを使う構成を書きました。 ・https://zenn.dev/nabettu/articles/38f021c1901212 今回はその続編というか実戦編です。イベント管理サービス「イベット」(告知・参加登録から当日のライブQ&A・投票までURLひとつで完結するSaaS)を、Next.js (App Router) + Cloudflare Workers + Turso のフルス...
#AIタグ

第2107話 AIが閣僚席に座る日——千葉の冠水、奥田氏の訃報、東大1000倍素子が示す「潮目」の変わり方

・2026年8月、さまざまなネタが一気に交錯している。 ・千葉県では線状降水帯による大規模な冠水が発生。
LLMタグが付けられた新着記事 - Qiita

中国OSSは張り子の虎じゃなかった|MiniMax-M3が3冠でGPT-5.5超え・価格1/8の衝撃

・📺 この記事は YouTube チャンネル きなこもっちーのテック深掘り の動画解説記事です。 ・▶️ 動画はこちら → 中国OSSは張り子の虎じゃなかった|MiniMax-M3が3冠でGPT-5.5超え・価格1/8の衝撃 🐹🦜 この記事に登場する2匹 🐹 もっち...
#AIタグ

電卓が使えない私が、事務の仕事をしている

・職場で、よく「パソコン強いんですね」と言われます。 ・たしかに、パソコンを使う作業はわりと速いほうだと思います。
ITmedia NEWS 最新記事一覧

東京ガス子会社、社員が飲酒後にPCを紛失 個人情報約2万件を保存

・東京ガス子会社の東京ガスネットワークは8月13日、顧客の個人情報約2万件を保存した業務用PCを社員が紛失したと発表した。社員が同僚らと飲酒した後、電車で帰宅する途中にPCなどが入ったかばんを紛失したという。13日時点で、第三者による不正操作や個人情報の不正利用は確認していないとしている。
#AIタグ

保育園の申請書類、就労証明書の空欄を前に手が止まった話

・保育園の申請書類を書き始めたのは、子どもたちが寝静まった夜中の1時でした。 ・「就労状況申立書(仕事の状況を自分で説明する書類)」という紙を前にして、ペンを持ったまま10分くらい固まっていたと思います。何を書けばいいのかわからないのではなく、正確には「どこまで書けばいいのか」がわからなかった。