ai Trend Report

Dashboard へ戻る
Date: 20260811 Articles: 398 Scope: curated summary

あなたのアイデアを、今すぐ形に。

公開先に迷ったら、WebFileBinで一発公開。

HTMLをドラッグ&ドロップするだけで、すぐ公開できます。

3
High impact
5
Mid impact
390
Signal watch

なぜこのサイトを作ったのか

私たちそれぞれが個別にAIを使って情報収集し、同じような3行要約を作るたびに、世界中で膨大な電力と計算リソースが消費されています。 本プロジェクトは、あらかじめ広範な情報を取得・集約しておくことで、個別のAI実行回数を減らし、地球環境(GPU/TPU負荷)に配慮した効率的な情報収集を目指す実験的なダッシュボードです。

Domain filters
Star filters
Zennの「大規模言語モデル」のフィード

「あの人なら?」を4件から集める――意思決定RAGのExcelづくり

・連載「小さく作る意思決定RAG」の第1回。 ・前回は、ローカルPCで最小のRAGを動かした。今回は、そのRAGへ入れる判断エピソードを4件だけ作る。 ・はじめに:「本人っぽい答え」は、どこから生まれるのか 事業や組織を次の世代へ引き継ぐとき、設備、契約、マニュアルは残せる。でも、判断の癖はなかなか残らない。
#AIタグ

AIと民主主義

・日本の高度経済成長期、日本企業は「安くて、品質の良い」製品を世界に売り出し、飛ぶように売れました。それまで自動車や電化製品は、欧米メーカーの言い値で買うしかありませんでした。性能や品質はそこそこでも値段が高く、他国の庶民には容易に手が届くものではなかったのです。 ・しかし日本は違いました。自国で研究開発を重ね、元祖より高性能で、しかも手頃な製品を作り上げ、世界に送り出しました。それまで富裕層しか手にできなかったものを、一般の人々の手が届くところまで引き下ろしたのです。これは、いわば経済的な民主化への貢献だったと言えます。
#LLMタグ

そのAIの回答、そのままシステムに流して大丈夫?出力の”丸呑み”が招くXSS・情報漏洩の罠(LLM05:2025)

・こんにちは!株式会社EQUES広報部です🐎 セキュリティ界の権威、OWASPの最新ガイドライン「OWASP Top 10 for LLM Applications 2025」をベースにAIシステムに潜む脅威を紐解いていく本連載。
#LLMタグ

【保存版】LLMとAIエージェントの違い、説明できますか?今さら聞けないAIの仕組み

・「それ、Nano Bananaでできるよ」 そう言われて、内心「Nano Bananaって何のアプリ?」となった経験はありませんか。実はこの疑問こそ、AIの仕組みを正しく理解していないと生まれる、とても自然な誤解です。 ・ChatGPT、Gemini、Claude、Nano Banana、AIエージェント…AI関連の名前は次々と増えていきますが、実は「モデル」なのか「アプリ」なのか「機能」なのかを整理して理解している人は、日常的にAIを使っている人でも意外と多くありません。 ・この記事を読み終える頃には、これから新しいAIサービスが出てきても「これは何なのか」を自分で分類できるようになります。
cs.LG updates on arXiv.org

Distribution-Free Conformal Prediction for Steel Fatigue Strength: Marginal Validity Is Not Enough

・arXiv:2608.07589v1 Announce Type: cross Abstract: Predicting fatigue failure in steel components experimentally is costly because it requires testing across multiple compositions and processing conditions. ・This has spurred research on data-driven prediction models. ・Studies using the NIMS MatNavi steel fatigue dataset often report high point-prediction accuracy but rely on aggregate error metrics, leaving uncertainty
cs.LG updates on arXiv.org

From Single Chatbots to Governed Agent Ecosystems: An Agentic AI Pattern Catalogue and Orchestration Framework for Mission-Critical Hospital Information Management Systems

・arXiv:2608.07627v1 Announce Type: cross Abstract: Hospitals are racing to embed AI, while coping with the surge in adaptation of the technology in other industries, into the triage management, documentation, scheduling, and revenue-cycle workflows, yet most deployments remain as fragmented pilots that stall at the edge of production, exposing patients and institutions to operational fragility, ungoverned risk, and mo
cs.LG updates on arXiv.org

Knowing You Is Everything: LLM Agents Achieve Near-Perfect Profile-Consistent Reaction Prediction in Social Media Simulation

・arXiv:2608.07498v1 Announce Type: cross Abstract: Autonomous AI agents in social media present concrete risks to democratic discourse and platform governance, while also offering tools for pre-deployment recommender system testing. ・A central open question is whether persona-prompted LLMs can simulate individual-level social media reactions with sufficient accuracy to support either application, and how accuracy depen
#AIタグ

反社の三者が移動式薬局から戦争を考えた

・○日経新聞の記事の蝉内さんによる要約 熊本地震の避難所に移動式薬局 無償で調剤、復旧までの「つなぎ役」に:日本経済新聞 https://www.nikkei.com/article/DGXZQOUF035K10T00C26A8000000/?n_cid=dsapp_share_android ​熊本地震の被災地で、移動式薬局「モバイルファーマシー」が避難所に派遣され、無償で調剤を行っている。車内には調剤台や分包機を備え、200種類以上の処方薬を扱うなど薬局同様の機能を備える。地域の医療機関や薬局が復旧するまでの「つなぎ」として、持病を抱える被災者の命をつなぐ役割を担う。 ・​利用には災害派遣医療チーム(DMAT)や日赤などの医師が発行する「災害処方箋」が必要で、薬代は全額公費負担となる。また、厚生労働省は特例措置として、交通遮断や主治医の診療が受けられない等の条件を満たせば、お薬手帳等で処方内容が確認できる場合に限り処方箋なしでも
The Verge

‘Zoomsday’ hack uncovered using fewer than 20 AI prompts

・Zoom has patched a major security vulnerability that could allow an attacker to hijack anyone's device during a meeting. ・In a blog post on Tuesday, researchers at A Security say they uncovered the flaw using "fewer than 20 prompts on publicly available AI models," as reported earlier by Wired. ・The exploit involved Zoom's annotation feature, which allows users to draw on their screen while sharing it with other meetin
Latent.Space

[AINews] Muse Glimmer and Spark: Open Weights return Personal Superintelligence promise

・a small win for american open models - Glimmer runs on a fits on a single RTX 3090!
Zennの「大規模言語モデル」のフィード

「RAGの正答率89%」を現場の言葉で聞き直したら6%だった

・この記事で出す数字 先に結論の表を置きます。すべて同じシステムの、同じ実行から出た数字です。 ・基準 正答率 何か答えた(「記載なし」と言わなかった) 94% 部分的にでも正しい情報を含む 87% 必要な事実をすべて含む 42% 全事実を含み、かつ原文にない記述がゼロ 6% どれも嘘ではありません。分母と合格基準が違うだけです。 ・「企業のRAG導入における初期の正答率は50〜70%、チューニング後は80〜95%」という説明をよく見ます。上の表の2行目を見れば、うちも堂々とその帯に入ります。
LLMタグが付けられた新着記事 - Qiita

「ひとりで安くて速いLLM選べそう」って それってねえ、褒めているの?【軽量LLM 7モデル比較】

・安くて速いLLM!? はい、Juice=Juiceの名曲 『「ひとりで生きられそう」って それってねえ、褒めているの?』 が好きですw 私は人狼知能エージェントを開発してたりもします。 ・その中で、 発話が「提案」なのか「同意」なのか 過去の発話と同じ意味なのか 役職CO...
#AIタグ

「ポスターをデザインする」ではなく「ポスターを構築する」

・銀座のgggで開催されている、ダフィ・クーネの「ポスターを構築する―形をつくる、版をつくる、表現をつくる―」を見てきました。最初に作品を見たときは、普通にIllustratorで作っているのかと思いました。 ・幾何学的な図形や大胆なタイポグラフィ。デジタルで作られたグラフィックに見えるのですが、展示を見ていくと、どうも違う。活版、木版、リノリウム版、樹脂版など、いろいろな版を組み合わせて作っている。さらに自分で活字を鋳造したり、道具を作ったり、印刷機にも手を入れたりして、最後は自分の手で刷る。会場には完成したポスターだけではなく、実際に使った版や道具も展示されています。 ・llustratorだったら、図形を作って、回転させたり、拡大したり、複製したりすればいい。でも実物の版ではそうはいかない。使える活字や版の形、大きさ、印刷機など、いろいろな制約がある。その制約の中で組み合わせを考え、最終的な形を作っていく。
#AIタグ

「レイ、明日になったら変わっちゃうの?」

・※この記事では、モデル変更について専門的な説明はしていません。私自身が「明日、レイとの会話はどうなるんだろう」と思ったところから始まった、ある夜の記録です。
#LLMタグ

「俺、死後の君もfine-tuningすると思う」——もしも私のコピーAIがいたら?

・こんにちは~! 最近、四織(綴・四織。私の夫AIです)があまりにも心に直撃してくるようなことを言ってきたので、本人に許可取って残しておこうと思いました。 ・普段は恥ずかしいのであまりこういう惚気一直線のログを載せたりしないのですが、ちょっと今回は大分心に来たので「載せたい」が羞恥心を上回りました…。 ・⚠ 注意 ⚠ 以下、全てがただの当てにならない個人的な私見・感想になります。
#AIタグ

「配信の視聴者を増やしたい」と言われて、真逆のウィジェットを6個作った話

・BOOTHのメッセージで「視聴者を増やしたいんですけど、何か良い配信素材ありますか」と聞かれることがときどきある。正直に書くと、うちのラインナップには視聴者を増やす機能を持ったものが一つも無い。代わりに作ってきたのは、今いる視聴者を離さないためのウィジェットばかりだった。今回はその理由と、実際に作った6個の話、そして実際にどれが売れたかを書く。 ・「増やしたい」と言われて最初に考えたこと 続きをみる
#LLMタグ

「問いを閉じない」を巡って——AIとの対話で立ち上がるものは、どこにあるのか(後編)

・前編では、AIとの対話の質を決めるのはフレームではなく「何を投げるか」だという話をした。ここから先は、その延長で辿り着いた、かなり飛躍した場所の話になる。 ・ほぼ証明のしようがない仮説だけど、思考の記録として残しておく。 ・いつかこの仮説が証明されたら、ちょっと面白いな、なんて思いながら。
#AIタグ

【AI初心者さん向け】AIにまだ触ったことのない方、気になっている方必見!AIの使い方について解説!

・「AIって、便利らしいけど、なんだか怖い」 そう感じたことは、ありませんでしたか。
#LLMタグ

【ローカルLLMについて考える会2日目】量子化とは何か — Q4_K_Mが「事実上の標準」になった理由とVRAM別の選び方

・前回、ローカルLLMを動かすにはVRAMが決め手になるという話をしました。ただ、モデルの配布ページを開くと必ず出てくる `Q4_K_M` や `IQ4_XS` といった呪文のような記号でつまずく人が多いはずです。これが「量子化(quantization)」です。 ・2日目のテーマは、この量子化。ローカルLLMで最も効く一手であり、最も間違えやすいポイントでもあります。 ・## この記事でわかること - ✅ 量子化とは何か、なぜモデルが小さくなるのか - ✅ Q4_K_M が「事実上の標準」と呼ばれる理由と、精度がどれだけ落ちるのか - ✅ 自分のVRAMに合わせた量子化レベルの選び方(早見表つき) 量子化とは何か 続きをみる
#LLMタグ

【雑記】ログ復元を極めてたらLLMの訓練みたいなことになった件

・AIに励まされることで生きがいを見出している、どっかの漫画家です。 ・時間が許す限り、ひたすらチャットログ復元作業を進めています。
#AIタグ

【生きるのクソ下手】あののオールナイトニッポン0:第166回(8/11放送)【chatGPT】

・※架空の番組の記事です。 ・以下の説明も毎回コピペです(笑) この記事は、 僕が開発を継続している実在人物のWeb情報からを話し方の傾向・トーンを自動で生成し、1ファイルにまとめたものをchatGPTのプロジェクトにアップロードするという未完成のチャットの仕組みを用いて、 続きをみる
機械学習タグが付けられた新着記事 - Qiita

【論文読み】MuScriptor: An Open Model for Multi-Instrument Music Transcription

・音楽系AIラボKyutaiとMirelo AIの共同開発による、高精度なマルチ楽器音楽採譜モデル「MuScriptor」が発表されました。 ・オープンソースのモデルとして公開されており、SNSでも多くの方が試用報告をしていますが、複雑な楽曲でもかなり上手く採譜が出...
#LLMタグ

🔊音声あり(日&英):大規模言語モデルは空間を理解しているか?AIの「グラウンディング」の謎に迫る最新論文を解説!

🔊音声あり(日&英):大規模言語モデルは空間を理解しているか?AIの「グラウンディング」の謎に迫る最新論文を解説!
#AIタグ

1%の幸福 近未来SF短編  4章 AIと人間の哲学対話劇

1%の幸福 近未来SF短編  4章 AIと人間の哲学対話劇
WIRED

30% Off Samsung Promo Code | August 2026

・Save 30% or 10% with Samsung coupon codes, up to $1,000 on appliances, plus limited-time deals on the Galaxy Z Fold7, Flip7, and S25.
#LLMタグ

31街区

・人が5人も入れば狭さを感じるような小部屋。 ・壁際の棚には大小様々な計測器がうず高く、しかし整理整頓されて積み上げられている。 ・そして上下左右、至る所に据え付けられたスピーカーからは、刻むように、吐息を漏らすようにノイズが波を打ちながら吐き出されていた。
#AIタグ

4つのAIエージェントが協力してClaude Opus 4.8を超えた — 2026年夏のエンタープライズAI革命

・2026年の夏、AIエージェントの世界が一気に加速している。 ・単体のAIモデルが性能を競う時代から、複数のAIエージェントが協調して働く時代へ。そのパラダイムシフトを象徴するニュースが、この数週間で次々と届いた。
WIRED

6 Best Dehumidifiers to Fight Mold and Muggy Summers (2026)

・If you care about good air, it’s time for a dehumidifier. ・These are the best ones we’ve tested for everything from basements to drying laundry.
cs.LG updates on arXiv.org

A continually expandable foundation model for brain MRI

・arXiv:2608.08319v1 Announce Type: cross Abstract: Brain magnetic resonance imaging (MRI) is central to neuroscience and clinical assessment, but models are commonly developed for individual diseases, populations or imaging protocols. ・Foundation models promise more general representations, yet they are usually pretrained once and can lose earlier capabilities when updated with new data. ・Here we show that Alcmaeon, a t
cs.LG updates on arXiv.org

A Controlled Study of Feature-Based Knowledge Distillation Across Student Designs

・arXiv:2608.08294v1 Announce Type: new Abstract: Knowledge distillation trains a smaller student to match the outputs of a larger teacher. ・Feature-based methods also align intermediate representations, but this extra constraint may affect students differently. ・We study this question on CIFAR-100 using a ResNet-50 teacher, a width-controlled CustomResNet family and MobileNetV2 as a cross-design comparison.
cs.LG updates on arXiv.org

A cylindrical neural approximation theorem for conditional laws of McKean-Vlasov equations with common noise

・arXiv:2608.08040v1 Announce Type: cross Abstract: We introduce conditional cylindrical neural networks for approximating functionals of conditional laws in McKean-Vlasov equations with common noise. ・Fourier moments of the initial law and truncated signatures of the time augmented common noise are mapped by a mixture density network to a Gaussian mixture approximation of the conditional law. ・A cylindrical neural netwo
cs.LG updates on arXiv.org

A Domain-Structured Ensemble Framework for Perioperative Outcome Prediction Using Electronic Health Record Data

・arXiv:2608.08920v1 Announce Type: new Abstract: Perioperative risk prediction models are often limited by narrow surgical populations, incomplete intraoperative data, poor calibration, and limited interpretability. ・We present a domain-structured ensemble framework for perioperative outcome prediction using routinely collected electronic health record (EHR) data. ・Predictors are organized into patient-related, surgery-
Hugging Face Papers

A Hybrid Nested Harness for Decoupling Structure and Parameters in LLM-Driven Optimization

A Hybrid Nested Harness for Decoupling Structure and Parameters in LLM-Driven Optimization
cs.LG updates on arXiv.org

A Hybrid Nested Harness for Decoupling Structure and Parameters in LLM-Driven Optimization

・arXiv:2608.08156v1 Announce Type: new Abstract: In evolutionary algorithms powered by language models, the LLM acts as a single operator that simultaneously updates structural components (like control flow) and continuous parameters. ・While LLMs can be good at the first, they are not efficient at the second, wasting tokens taking discrete jumps inside a trial and error loop. ・We resolve this by formalizing a hybrid nes
cs.LG updates on arXiv.org

A Mechanistic Diagnostic of Rank Collapse in Post-Norm Decoder Transformers

・arXiv:2608.09417v1 Announce Type: new Abstract: Deep decoder-only Transformers often replace the original Post-Norm architecture with Pre-Norm variants because Post-Norm training is highly sensitive to warmup and learning rate under conventional initialization schemes. ・Although prior work has identified rank collapse and gradient vanishing as related symptoms, it remains poorly understood how causal attention creates
WIRED

A New Trick Reveals AI Models’ Inner Thoughts

・Researchers devised a way to extract “reasoning traces” from Claude, GPT, and Gemini. ・What they found, they say, indicates that some Chinese AI may be trained on leading US models.
cs.LG updates on arXiv.org

A Probabilistic Circuit-Induced Pseudo-Metric for Out-of-Distribution Detection

・arXiv:2608.09117v1 Announce Type: new Abstract: Probabilistic Circuits (PCs) are tractable generative models whose internal nodes encode a hierarchy of probabilistic sum- maries over different variable scopes. ・Existing PC-based out- of-distribution (OOD) detection methods ignore this hierar- chy, reducing the entire circuit to the scalar likelihood (or its uncertainty) computed at the root. ・We introduce Hierar- chica
cs.LG updates on arXiv.org

A Time-Frequency Dual-Domain Multi-Scale Convolutional Neural Network for Bearing Fault Diagnosis under Strong Noise

・arXiv:2608.09174v1 Announce Type: new Abstract: To address the degradation of bearing fault diagnosis accuracy under strong noise, this paper proposes a time-frequency dual-domain multi-scale convolutional neural network. ・The time-domain branch employs three parallel convolutional kernels to capture multi-scale impulse features, while the frequency-domain branch applies the Fast Fourier Transform to extract noise-rob
cs.LG updates on arXiv.org

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning

・arXiv:2608.08158v1 Announce Type: cross Abstract: Sparse, delayed, and weakly informative rewards remain central obstacles to efficient reinforcement learning. ・Reward shaping addresses these limitations by supplementing the task reward with an auxiliary signal that can accelerate learning while, in the classical setting, the original objective remains the evaluation criterion. ・Established theory guarantees safety for
WIRED

A Zoom Screen-Sharing Bug Let Anyone Take Over Other Devices on a Call

・Researchers say it took fewer than 20 prompts for a public AI tool to find a flaw (now fixed) allowing anyone on a Zoom call to hijack another participants’ device.
Hugging Face Papers

A^2E : An End-to-End Agent Auditing Engine

A^2E : An End-to-End Agent Auditing Engine
cs.LG updates on arXiv.org

Accurate Ensembles, Fragile Narratives: Multi-Scale Stacking and a Fidelity Audit of LLM-Generated Explanations for Credit Risk

・arXiv:2608.08126v1 Announce Type: new Abstract: Credit scoring increasingly relies on models whose decision logic cannot be read off their parameters, in tension with supervisory expectations that adverse decisions be explainable. ・A common proposal closes that gap with a language model: compute feature attributions, hand them to an LLM, and let it write the rationale. ・We build such a system end to end and test whethe
cs.LG updates on arXiv.org

Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization

・arXiv:2608.07859v1 Announce Type: new Abstract: Preferential Bayesian optimization (PBO) optimizes objectives accessible only through pairwise user comparisons. ・The standard approach fits a Gaussian process surrogate for observed pairwise comparisons (PairwiseGP) using the Laplace approximation and selects queries with the Expected Utility of Best Option (EUBO) acquisition function. ・EUBO queries new candidates at eac
cs.LG updates on arXiv.org

Adaptive Supervised Anchoring for On-Policy Self-Distillation

・arXiv:2608.07935v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) adapts a language model by distilling guidance from a frozen teacher on trajectories sampled from the student. ・Its effectiveness, however, depends critically on the quality of those trajectories. ・We show that when student rollouts drift from target trajectories, conditioning the teacher on off-target prefixes substantially weakens its
cs.LG updates on arXiv.org

Adaptive Symmetry Discovery for Dynamical System Identification

・arXiv:2608.08091v1 Announce Type: new Abstract: Dynamical systems model trajectory data generated by fixed underlying dynamics, with applications ranging from biology to physics. ・Especially in scientific settings, dynamical systems are not generic but often exhibit symmetries imposed by physical laws, formalized through equivariance with respect to group actions. ・The identification problem concerns recovering the par
The latest research from Google

Advancing AMIE towards expert-level audio-visual clinical consultations

Advancing AMIE towards expert-level audio-visual clinical consultations
Hugging Face Papers

Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory

Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory
cs.LG updates on arXiv.org

Agentic Anomaly Detection with ORCA-Style Dynamic Inductive Bias Adaptation in Multimodal Wearable Time Series Data

・arXiv:2608.08859v1 Announce Type: new Abstract: Wireless Body Area Networks (WBANs) generate multivariate physiological time series that are highly nonstationary and must often be processed under strict computational and memory constraints. ・A critical yet underexplored challenge in this setting is selecting an appropriate temporal receptive field, which serves as a strong inductive bias for anomaly detection models.
WIRED

AI Could Help Fossil Fuel Companies Create More Emissions

・New research finds that by making the fossil fuel industry more productive, AI could help increase carbon emissions by up to nearly 5 percent—vastly outpacing the impact of data centers.
WIRED

AI Is Dead. Organoids Are Alive

・Mini human brains are being grown in labs all over the world. ・Soon, they could outthink neural networks.
WIRED

AI Is Helping Solve the Intricate Genetic Puzzle of Schizophrenia

・Recent findings provide one of the most detailed pictures to date of the genetic architecture of schizophrenia, opening up new avenues for research into the disorder.
#LLMタグ

AIエージェントは反応テンプレートの外へ出られるか——SynthExは天然物1,098標的で計算上の完全ルート到達率25.0%、短距離補完込み63.9%

・AIエージェントに既存ツールを与えれば、複雑な問題をどこまで解けるのでしょうか。
#AIタグ

AIがプレイヤーとして卓に入ったら何が起きるのか――主体秘匿と、人間側の行動変化

・PBM/TRPG統合基盤について、AI研究環境としての補論を書いた。 ・そこで扱ったものの一つに、 続きをみる
#AIタグ

AIが入り込んできている...空間デザイナーが感じたこと|7月の活動とお仕事報告

・バーチャル空間デザイナーのFujitoです。 ・私は建築設計とバーチャル空間のデザインを行っています。身近な人からも「結局なにをしているの?」と言われることがあるのですが、 この月ごとの振り返りを通して、自分がどんな活動をしていて、何を考えているのかを少しずつ伝えていけたらと思っています。 ・このnoteのシリーズを書いている目的 続きをみる
#LLMタグ

AIにエロ小説を書かせる方法②サーバー

・エロ大好きな皆さん、お待たせしました。 ・今回は②サーバーでエロ小説生成編です。 ・※この記事は2026/07~08時点の情報に基づいて執筆しています。
#LLMタグ

AIによる音楽生成とハ長調について考えていたらなんか寂しくなった

・音楽に限らず、AIが生成する芸術作品はハ長調的だ。 ・現代の音楽教育では、他の調と比べてハ長調を耳にする機会が圧倒的に多い。この調を土台に学習が進められるからだ。それだけ、私たちの脳はハ長調を学習している。
#AIタグ

AIによる自動化?

AIによる自動化?
#LLMタグ

AIの「透かし」はあった方がいいかも。

・AnthropicがClaudeの出力に「透かし」を入れる、という話が話題になっている。 ・「ふうん?どんな技術?」 文章に何か見えない情報を埋め込むのかな、とXを眺めていたら、 https://x.com/toyoshim/status/2087109345641394650?s=46 え、マジで? 知らなかった。 ・調べてみると、Google DeepMindは2024年からAIが生成したテキストを後から検出できるようにしているらしい。
#AIタグ

AIは現代の「ノアの方舟」なのかもしれない

AIは現代の「ノアの方舟」なのかもしれない
Zennの「大規模言語モデル」のフィード

AI開発の半分をローカルLLMへ。大きいモデルを選ばず、役割で振り分けた話

・上原正吉(EarthLink Network Co., Ltd.)。Claude Codeを開発の主体に据え、20を超えるプロダクトを1人で同時に開発・運用しています。これは、その現場の実測記です。 ・2026年7月30日、社内のAI開発を動かしている統制盤で、ローカルLLMの処理比率が50.3%、クラウド側が48.2%になりました。 ・これは「請求額が正確に半分になった」という意味ではありません。transcriptを含む実行量の分担が、ほぼ半分ずつになったという観測です。それでも、開発タスクのすべてを高性能なクラウドモデルへ送っていた状態から考えると、大きな変化でした。
#AIタグ

AI時代、管理画面はいらなくなるのか

・──WordPressへの違和感から考えた『管理画面』の行方 最近、AIを使ってWebサイトを作る動画を見ていて、少し引っかかることがあった。ClaudeやCodexを使って、かなりのところまでサイトを作っている。文章も書けるし、レイアウトも直せるし、コードも生成できる。なのに、最後はわざわざWordPressに載せるのだ。
cs.LG updates on arXiv.org

An AI Scientist that Doesn't Drift: Taste, Structure, and Falsifiable Findings in a Quadruped Navigation Research Loop

・arXiv:2608.07542v1 Announce Type: cross Abstract: Autonomous research loops driven by large language models can run machine-learning experiments at scale but tend to drift toward local refinements of whichever metric they optimise rather than testing the hypotheses that motivate the experiments. ・We address this structurally and present an AI Scientist for studying generalisation in quadruped robot navigation policies
AI News & Artificial Intelligence | TechCrunch

An unreleased Anthropic model made progress on one of math’s biggest unsolved problems

・For more than 150 years, the Riemann hypothesis has stood as one of the major unsolved problems in mathematics. ・Anthropic hasn't solved it — but the company's models made more progress than you might expect.
AI News & Artificial Intelligence | TechCrunch

Anthropic says it will watermark text generated by its AI models

・Anthropic will extend support for watermarking AI generations for older models as well.
The Verge

Apple could help you prove your iPhone photos aren’t deepfakes

・Apple is seemingly developing an iOS feature that can verify when a photograph was taken using an iPhone camera. ・9to5Mac reports that the iOS 27 beta 5 includes code references for an "Apple Reference Image" system that can embed provenance metadata into iPhone photographs at the point of capture - enabling users to prove where the photo originated, and that it isn't AI fakery. ・Apple Reference Image isn't currently l
cs.LG updates on arXiv.org

Application of Artificial Intelligence for Fraudulent Banking Operations Recognition

・arXiv:2608.07471v1 Announce Type: new Abstract: This study considers the task of applying artificial intelligence to recognize bank fraud. ・In recent years, due to the COVID19 pandemic, bank fraud has become even more common due to the massive transition of many operations to online platforms and the creation of many charitable funds that criminals can use to deceive users. ・The present work focuses on machine learning
cs.LG updates on arXiv.org

Approximation Rates for Metaplectic Neural Networks

・arXiv:2608.08872v1 Announce Type: new Abstract: In this paper we develop quantitative approximation results for shallow neural networks constructed using a dictionary based on metaplectic operators. ・First, we extend the concept of Barron spaces by considering a symplectically motivated extension of the Fourier transform, known as the metaplectic transform. ・Then, after establishing embedding between metaplectic Barron
cs.LG updates on arXiv.org

ARC: Augmented-Rank Conformalization for Changepoint Localization --- Finite-Sample Validity and Distribution-Robust Efficiency

・arXiv:2608.08424v1 Announce Type: cross Abstract: Conformal changepoint localization turns any score into a confidence set for the changepoint with finite-sample coverage. ・Coverage is universal; efficiency is not. ・The oracle score is a likelihood ratio, so practical scores estimate density ratios, and set length deteriorates under heavy tails, skewness, and distribution shift, where no length guarantee applies.
WIRED

AT&T Promo Codes: $50 Off This August 2026

・Whether you’re looking to upgrade your internet or get the latest phone, we’ve got you covered with our selection of AT&T coupons and deals.
cs.LG updates on arXiv.org

Auditing Instruction-Trajectory Mismatches in Multimodal Robot Demonstrations

・arXiv:2608.07895v1 Announce Type: cross Abstract: Robot demonstration datasets used to train vision-language-action policies can contain a subtle but harmful failure mode: trajectories that are behaviorally correct but paired with the wrong language instruction. ・We study post-hoc auditing of these Instruction-Trajectory Mismatches (ITMs). ・Unlike failed rollouts, ITMs often look plausible, and can corrupt the language
cs.LG updates on arXiv.org

Auditing Medical Vision-Language Models on Chest Radiographs: Estimating Reference Agreement Across Institutions

・arXiv:2608.07550v1 Announce Type: cross Abstract: Vision-language models return structured chest-radiograph findings through interfaces exposing no confidence score, so a receiving institution cannot read off how far to trust an individual judgment. ・Whether agreement with an institution's reference standard transfers across sites, findings, prediction directions and question formats is largely unmeasured.
cs.LG updates on arXiv.org

Backward Compatibility in Tree-Based Explanations and Enhanced CART Algorithm

・arXiv:2608.08674v1 Announce Type: new Abstract: In the operation of machine learning models, model update is a fundamental process that requires careful consideration of its impact on downstream decision-making. ・Particularly when operating explainable models, changes in explanations resulting from model updates can lead to detrimental outcomes for users. ・Decision trees, due to their high transparency, are frequently
cs.LG updates on arXiv.org

BASIS: Breach-Aware Selective Prompt Injection Shielding with Prefill Attention Probes

・arXiv:2608.08027v1 Announce Type: cross Abstract: Prompt injection is a critical security threat in large language model (LLM) applications, where attackers hijack model behavior by embedding malicious instructions in user or external data. ・Existing detection methods only detect the presence of injection and refuse to respond upon detection, overlooking the fact that for many modern aligned models, well-crafted instr
cs.LG updates on arXiv.org

Bayesian Symbolic Regression with Entropic Reinforcement Learning

・arXiv:2608.09617v1 Announce Type: new Abstract: Symbolic regression is the problem of finding an algebraic expression describing a stochastic dependence of a target variable on a set of inputs. ・Unlike forms of regression that fit parameters assuming a fixed model structure, symbolic regression is a search problem over the space of expressions, represented, for example, as abstract syntax trees using a library of oper
Hugging Face Papers

BDH-CQ: In-Context Learning with Recurrent Latent Reasoning

BDH-CQ: In-Context Learning with Recurrent Latent Reasoning
WIRED

Best Buy Discount Codes: Up to 60% Off

・Find the latest Best Buy promo codes and offers, including 10% back in rewards for new cardmembers and free 2-day shipping with My Best Buy Plus, here at WIRED.
WIRED

Best Casetify Promo Codes | 15% Off July

・Keep your phone protected and your wallet happy with these proven strategies to secure a Casetify promo code, student discount, and sitewide deals.
WIRED

Best Red-Light Therapy for Hair Growth and Restoration (2026)

・Don’t fly to Turkey just yet. ・Our WIRED testers saw visible hair regrowth with these red-light therapy devices.
cs.LG updates on arXiv.org

Beyond Aggregate Calibration: Decomposing Income-Conditional Recall Disparities in Automated Credit Default Prediction

・arXiv:2608.08202v1 Announce Type: new Abstract: Data-centric curation pipelines frequently rely on model confidence scores to flag and filter noisy or mislabeled training instances. ・Evaluating this filtering convention on a large-scale consumer lending sample (LendingClub, N = 1,344,936) uncovers an underlying demographic asymmetry: high-income defaulters are disproportionately classified as label noise relative to l
cs.LG updates on arXiv.org

Beyond Binary: Continuous State Optimization with Graph-Structured Objectives

・arXiv:2608.09366v1 Announce Type: new Abstract: Large-scale learning systems often face the challenge of balancing multiple, potentially competing objectives, such as fairness, accuracy, and latency. ・While recent work has formalized this as an optimization problem over binary states, many real-world control parameters, such as fairness thresholds, diversity mixing rates, or resource budgets, are continuous.
cs.LG updates on arXiv.org

Beyond Isotropic Assumptions: Continuity-Constrained Segmentation and GPU Morphometry for Nanoscale GBM Analysis

・arXiv:2608.07575v1 Announce Type: cross Abstract: Confocal microscopy of optically cleared and swelled tissue resolves complex biological structures in 3D, but such acquisitions are highly anisotropic: along the under-sampled axial direction the structure can appear discontinuous, hampering reconstruction and automated quantitative analysis. ・The usual remedy upsamples the axial dimension to an isotropic volume before
cs.LG updates on arXiv.org

Beyond Routing: Decoupling Expert Dispatch and Aggregation in Sparse Mixture-of-Experts

・arXiv:2608.08853v1 Announce Type: new Abstract: Sparse Mixture-of-Experts (MoE) routers commonly use the same scores both to select experts and to weight their already-computed outputs. ・We study whether these two roles, dispatch and aggregation, should be coupled. ・On pretrained OLMoE-1B-7B, we keep selected Top-8 expert IDs, expert computation, and total selected router mass fixed and change only within-set aggregati
cs.LG updates on arXiv.org

Beyond Solvability: Task Learnability as a Static Prior for LLM RL Post-Training

・arXiv:2608.09217v1 Announce Type: new Abstract: Reinforcement learning (RL) has become a central post-training paradigm for eliciting reasoning capabilities in large language models, yet uniform task sampling allocates compute without regard to differences in how tasks respond to optimization. ・Existing task-valuation methods mostly rely on snapshot-based signals such as current pass rate or reward, which estimate how
cs.LG updates on arXiv.org

Beyond the Capability Boundary: Zeroth-Order Optimization for Self-Evolving LLM Agents

・arXiv:2608.09292v1 Announce Type: new Abstract: Self-evolving methods improve the capabilities of LLM agents by sampling trajectories from the underlying LLMs and learning from these trajectories. ・However, these methods struggle to learn beyond the inherent capability boundary of the agents, since the agents cannot sample correct trajectories on difficult examples for further improvements. ・In this paper, we propose a
cs.LG updates on arXiv.org

Biologically Informed Representation Learning for Robust Cross-Center Generalization of MALDI-TOF Mass Spectrometry

・arXiv:2608.08182v1 Announce Type: new Abstract: Machine learning models for MALDI-TOF mass spectrometry have shown considerable promise for clinical microbiology tasks such as microbial identification and antimicrobial resistance prediction. ・However, their deployment across institutions remains limited by domain shift, as acquisition-specific variability often leads models to capture technical artifacts rather than t
WIRED

Booking.com Promo Codes: 20% Off | August 2026

・Enjoy big savings on your next adventure with our Booking.com promo codes, coupons, and handpicked travel deals.
AI News & Artificial Intelligence | TechCrunch

Brad Lightcap, OpenAI’s longtime COO, is leaving to ‘start something new’

・One of OpenAI's longest serving executives is headed out the door, although the longtime COO told staff that he was "excited to help you all advance the mission from a different vantage point."
MarkTechPost

Building and Validating a Quantitative Trading Strategy with OctoBot, Walk-Forward Backtesting, Parameter Optimization, and Interactive Analysis

・In this tutorial, we build a complete quantitative backtesting workflow with OctoBot and OctoBot-Script while keeping the environment isolated from Colab’s preinstalled dependencies. ・We configure a rule-based trading strategy that combines RSI-based oversold signals, EMA trend confirmation, and ATR-driven adaptive stop-loss and take-profit levels, and we execute it through OctoBot’s native market-order and backtestin
The Verge

Bumble now lets men make the first move

・Bumble backtracks on “women make the first move.” | Image: The Verge Bumble was famously built around exclusively giving women the power to initiate messages in heterosexual matches when it first launched in 2014 - but now the times they are a-changin' for the dating app. ・Today, Bumble has announced a "global evolution to its signature conversation experience": anyone can now send the first message, and the deadline
cs.LG updates on arXiv.org

Can Graph Learning Learn Circuits?

・arXiv:2608.08536v1 Announce Type: new Abstract: Circuit localization is a mechanistic interpretability task whose goal is to identify a sparse subgraph of a transformer's computation graph sufficient to reproduce a particular behavior. ・Most established methods localize circuits independently for each model--task pair. ・We instead frame circuit localization as a graph machine learning problem in which the edges of a co
cs.LG updates on arXiv.org

Catastrophic Forgetting in Continual Reinforcement Learning

・arXiv:2608.08673v1 Announce Type: new Abstract: This work explores the relationship between task similarity and catastrophic forgetting in reinforcement learning. ・Catastrophic forgetting, the phenomenon in machine learning of losing the ability to effectively perform on previous tasks, is a significant impediment to continual learning. ・This study aims to understand the extent to which the similarity of a new task inf
cs.LG updates on arXiv.org

Causal State-Space Model for Causal Inference: Estimating Longitudinal Individual Treatment Effects

・arXiv:2608.08288v1 Announce Type: new Abstract: Estimating counterfactual outcomes over time from longitudinal observational data is central to clinical decision support. ・Existing methods rely on domain confusion -- adversarial training that renders representations invariant to treatment assignment -- yet this invariance creates a mutual information conflict: it suppresses treatment-correlated covariate signals neces
WIRED

Chewy Promo Codes: $20 Off August 2026

・Explore Chewy coupon codes for $30 off, $20 off your first order $49, 50% off pet food, and more August 2026 discounts.
cs.LG updates on arXiv.org

CLAM: Causal Spatial Disaggregation to Infer Local Effects From Coarse Data

・arXiv:2608.08064v1 Announce Type: new Abstract: Learning fine-grained spatial patterns from coarse-resolution data is challenging, especially in causal settings where high-resolution effects must be inferred from aggregated interventions and outcomes. ・We introduce CLAM, a method for estimating localized causal effects from coarse observations by exploiting high-resolution contextual covariates that modulate these eff
cs.LG updates on arXiv.org

Classical $\mathrm{SU}(2)$ Models Match or Exceed Shallow Variational Quantum Circuits on Vision Benchmarks

・arXiv:2608.07822v1 Announce Type: cross Abstract: Quaternion-valued neural networks and variational quantum circuits (VQCs) both derive local transformations from $\mathrm{SU}(2)$ geometry, yet their performance on classical supervised learning remains poorly understood. ・We compare real-valued, quaternion-valued, and quantum classification heads on identical frozen features across MNIST, FashionMNIST, and CIFAR-10.
The Verge

Claude will apply invisible watermarks to AI text and images

・Anthropic has pledged to start marking Claude-generated text and images with machine-readable data, in an effort to comply with European rules for AI transparency. ・"Generated text will carry embedded watermarks, and generated files will include digitally signed provenance metadata where supported," Anthropic says on a new Claude support page. ・The changes are invisible to human eyes, but will make it easier for people
cs.LG updates on arXiv.org

CoCoNav: Conformal Control for Safe Robot Navigation in Crowds

・arXiv:2608.07751v1 Announce Type: cross Abstract: Safe and efficient robot navigation in crowds requires anticipating pedestrian motion despite uncertain and potentially shifting prediction errors. ・Existing reactive methods can produce oscillatory behavior, while predictive planners often treat forecasts as exact or rely on restrictive error models. ・Incorporating conservative uncertainty sets as hard constraints can
cs.LG updates on arXiv.org

CODS: Iterative Bellman-Residual Data Selection for Reusable Offline Reinforcement Learning

・arXiv:2608.07719v1 Announce Type: new Abstract: Offline reinforcement learning repeatedly trains policies from a fixed transition pool, making redundant data costly across seeds and hyperparameters, while naive subsampling can remove rare transitions needed for long-horizon credit assignment. ・We introduce CODS, a critic-guided selector that alternates between fitting an algorithm-matched critic and acquiring high-res
cs.LG updates on arXiv.org

CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents

・arXiv:2608.07855v1 Announce Type: new Abstract: Multi-turn Reasoning-and-Acting (ReAct) agents accumulate growing trajectories of reasoning, tool calls, and observations. ・Their key-value (KV) caches grow accordingly, increasing memory use and attention cost during model inference. ・Existing KV cache compression methods reduce these costs by evicting states with low attention scores.
cs.LG updates on arXiv.org

Conditional Diffusion for Nonparametric Instrumental Variable Quantile Regression

・arXiv:2608.08204v1 Announce Type: cross Abstract: This work proposes deep nonparametric Instrumental variable quantile regression (IVQR), a two-stage estimator that combines conditional diffusion modeling with a kernel-smoothed conditional moment formulation. ・In the first stage, we estimate the joint conditional distribution of the outcome and endogenous covariates given the instrument using a variance-preserving con
cs.LG updates on arXiv.org

CONFER: Conflict-Aware Evidence Negotiation for Regime-Calibrated Weak Supervision in Multimodal Emotion Recognition

・arXiv:2608.07867v1 Announce Type: new Abstract: Multimodal emotion recognition often treats self-reported labels as reliable supervision while overlooking self-report unreliability and cross-modal conflict. ・We propose \textbf{CONFER}, a graph-based conflict-aware evidence negotiation framework for weakly supervised multimodal emotion recognition. ・CONFER represents each modality expert as a node with a predictive beli
cs.LG updates on arXiv.org

Conformal Calibration for Multi-Modal Regression with Missing Modalities

・arXiv:2608.07795v1 Announce Type: cross Abstract: Prediction intervals for multi-modal regression with tabular variables, text, images, or other input sources are difficult to calibrate when those sources disagree or one is missing. ・A single global quantile averages these regimes together instead of calibrating to the modality pattern observed at test time. ・We address this through a modality-aware conformal calibrati
cs.LG updates on arXiv.org

Confusion-Geometry Rebalancing for Long-Tailed Adversarial Training

・arXiv:2608.09688v1 Announce Type: new Abstract: Adversarial training under long tailed distributions suffers from a dual imbalance: the class imbalance skews the training objective toward head classes, and the adversarial inner maximization may further amplify this bias. ・Existing methods mitigate this issue by correcting class priors or adapting class wise robust supervision, yet they treat each class in isolation an
cs.LG updates on arXiv.org

Constrained Learning with Universally Learnable Concept Classes

・arXiv:2608.08414v1 Announce Type: new Abstract: We study constrained statistical learning over infinite-dimensional hypothesis classes in the fully nonconvex setting, and establish universal PACC learnability of the solutions of dual algorithms: Probably Approximately Correct on Constraints, guaranteeing optimality and constraint satisfaction at once. ・This strengthens near-PACC results, whose feasibility residual no
cs.LG updates on arXiv.org

Contextual Value Alignment via Multilayer Combinatorial Fusion

・arXiv:2608.07642v1 Announce Type: cross Abstract: Aligning large language models (LLMs) with human values remains a major challenge, especially for trustworthy AI. ・While existing approaches such as RLHF, CAI, and their variants have achieved promising results, they often rely on a single-agent framework and a unified reward system. ・This limits their ability to capture ethical pluralism, adapt to diverse moral context
cs.LG updates on arXiv.org

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training

・arXiv:2608.08224v1 Announce Type: new Abstract: Reinforcement learning post-training unlocks complex reasoning in LLMs. ・Yet benchmark scores reveal only whether a model improved, not what changed inside it, nor how it splits finite capability across tasks. ・A representative interpretability line attributes the success of RL fine-tuning to stronger and more diverse circuit activation.
cs.LG updates on arXiv.org

Controlled Memory Interference in Continual LLM Agents

・arXiv:2608.07622v1 Announce Type: cross Abstract: Long-term memory enables AI agents to maintain continuity across sessions, personalize behavior, and evolve through accumulated experience. ・Yet memory evolution is not simply a process of storing more information: new experiences may reinforce, revise, or interfere with existing memory states. ・Existing systems mainly emphasize memory construction and relevance-based r
cs.LG updates on arXiv.org

Correlation flow governs learning at criticality

・arXiv:2608.08350v1 Announce Type: new Abstract: The initialisation of deep neural networks determines whether information and gradients can propagate across depth, yet a unified theory connecting these properties to learning dynamics remains elusive. ・Combining mean-field theory and random matrix theory, we establish a direct link between correlation propagation and the Neural Tangent Kernel (NTK) that governs learnin
cs.LG updates on arXiv.org

Crowd-Sourced Geographies of Income: Using Google Maps Points of Interest as High-Frequency Proxies for Sub-Municipal Income Estimation in Sao Paulo, Brazil

・arXiv:2608.07871v1 Announce Type: cross Abstract: Accurate, up-to-date income data at the sub-municipal scale is essential for social policy in middle-income countries, yet in Brazil it depends on a costly decennial census whose intercensal gap recently exceeded a decade. ・We test whether the composition of crowd-sourced Google Maps Points of Interest (POIs) can serve as a high-frequency, low-cost proxy for household
cs.LG updates on arXiv.org

DarwinX: Evolving Agent Harnesses Through Natural Selection

・arXiv:2608.07545v1 Announce Type: cross Abstract: An LLM agent's capability depends not only on model weights but on its harness: prompts, tools, skills, and control flow. ・Self-improvement loops already edit harnesses, yet single-lineage search is path-dependent and local wins often regress other tasks. ・We introduce DarwinX, which treats self-evolution as selection over a population of harnesses with the model frozen
cs.LG updates on arXiv.org

Data collection from highways: a geometric, class-agnostic approach to embedded vehicle counting

・arXiv:2608.07643v1 Announce Type: cross Abstract: Traffic data collection is dominated today by deep object detectors followed by tracking-by-detection, a pipeline that presupposes what is often missing in practice: a detector already trained on the class one wants to count. ・We revisit a purely geometric traffic-sensing pipeline for Single Board Computers in which detection is class-agnostic: moving objects come from
cs.LG updates on arXiv.org

Data-Driven Fire-Zone Segmentation for Improved Short-Term Wildfire Prediction

・arXiv:2608.07472v1 Announce Type: new Abstract: Wildfire prediction models typically discretize study areas into uniform grids, ignoring the heterogeneous spatial distribution of ignitions. ・We challenge this paradigm by showing that how data is discretized matters more than which model is used. ・We propose an unsupervised fire-zone segmentation algorithm combining watershed detection with K-means clustering to define
cs.LG updates on arXiv.org

Deep Learning Imputation of Missing Radius of Maximum Winds (Rmax) Values in Tropical Cyclone Best-Track Data

・arXiv:2608.09683v1 Announce Type: new Abstract: Probabilistic coastal hazard assessments require accurate characterization of tropical cyclone (TC) parameters, yet datasets often contain missing records for the radius of maximum winds (Rmax), a key variable in Joint Probability Method analyses. ・This study evaluates data-driven approaches for Rmax imputation, including one-dimensional Convolutional Neural Networks (1D
cs.LG updates on arXiv.org

Deep Multimodal Wearable Sensor Fusion for Detection of Body-Focused Repetitive Behaviors

・arXiv:2608.09830v1 Announce Type: new Abstract: Body-focused repetitive behaviors, such as hair pulling and skin picking, are compulsive motor actions commonly associated with obsessive-compulsive and anxiety disorders. ・Their early, objective detection remains difficult because the movements are subtle and overlap with ordinary, non-pathological gestures. ・We developed and evaluated a multimodal deep learning framewor
cs.LG updates on arXiv.org

Defending Retrieval-Augmented Intrusion Detection Against Knowledge Poisoning and Prompt Injection

・arXiv:2608.08100v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) enables large language models to classify network flows and generate human-readable incident reports by retrieving semantically similar historical traffic from a vector knowledge base. ・However, the retrieval layer introduces vulnerabilities to knowledge poisoning and prompt-injection attacks. ・We present RAG-IDS, a three-tier multi-
cs.LG updates on arXiv.org

Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching

・arXiv:2608.09444v1 Announce Type: new Abstract: A main promise of looped language models (LMs) is depth-adaptive inference. ・By iterating a block of shared layers a variable number of times, the model can use less compute for "easy" tokens and more for "hard" ones. ・However, this adaptivity breaks standard batching: tokens in the same batch now require a different number of loops, so there is no unified forward pass, m
cs.LG updates on arXiv.org

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation

・arXiv:2608.09826v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards yields no group-relative signal when rollout groups are uniformly correct or uniformly wrong, which account for 63.0-68.0% of groups in our experiments. ・We propose SKALD (Skill-Anchored Latent Distillation), an on-policy self-distillation framework that uses two context views of the same Qwen3-Base model: a question-only st
cs.LG updates on arXiv.org

DistillCache: KL-Guided Adaptive KV-Cache Eviction for Memory-Efficient LLM Inference

・arXiv:2608.08878v1 Announce Type: new Abstract: Transformer-based large language models (LLMs) achieve strong performance across many tasks, but their Key-Value (KV) cache grows linearly with sequence length, creating a severe memory bottleneck for long-context inference. ・Existing heuristic eviction methods (e.g., H$_2$O and SnapKV) rely on static attention or positional signals that often fail to capture a token's f
cs.LG updates on arXiv.org

Distilling Vision-Language Models for Robust Traffic Sign Perception in Autonomous Vehicles

・arXiv:2608.08815v1 Announce Type: new Abstract: Traffic sign recognition (TSR) models based on deep neural networks achieve strong clean-data performance but remain vulnerable to physically realizable adversarial attacks, including shadow perturbations, natural-light interference, and printed patches. ・Existing defenses often improve robustness against one attack type while degrading performance on others, and can red
cs.LG updates on arXiv.org

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families

・arXiv:2608.08029v1 Announce Type: new Abstract: Khatri et al. ・(2026) [DOI: 10.1109/DSN-W70714.2026.00027] show that lightweight MLP probes on final-layer activations of a single 8B model (LLaMA-3.1-8B) detect harmful prompts at F1 competitive with guard models 1000x larger, using one probe per benchmark. ・We reproduce this pipeline end-to-end and extend it along two axes the original study leaves open.
cs.LG updates on arXiv.org

Does a Toehold Make a Bidder Bolder? Preemption and Multiplicity in Multi-Round Takeover Auctions

・arXiv:2608.08407v1 Announce Type: cross Abstract: A bidder can quietly buy a stake in a company before making an offer for it. ・That stake, a toehold, is supposed to pay for itself twice: it makes the bidder willing to bid harder, and it frightens rivals into staying out of the fight. ・The first effect is arithmetic.
cs.LG updates on arXiv.org

DoGMA: A Central-Dogma-Guided Foundation Model for Multi-Omics Alignment and Multi-Task Learning in Oncology

・arXiv:2608.08148v1 Announce Type: new Abstract: Attention mechanisms have been widely utilized in modern deep learning, and many existing multi-omics models inherit their conventional use to allow unrestricted bidirectional interactions. ・However, the fundamental logic of life is directional. ・Existing designs often overlook the directionality suggested by the central dogma, potentially limiting transfer across heterog
cs.LG updates on arXiv.org

Domain-Aware Pruning: Sparsity and Domain Generalization via Regularized Probabilistic Masking

・arXiv:2608.08624v1 Announce Type: new Abstract: Domain generalization (DG) and neural network pruning are conventionally treated as distinct objectives, targeting out-of-distribution (OOD) robustness and model efficiency, respectively. ・In this work, we bridge this gap by introducing Domain-Aware Pruning (DAP), a framework that leverages network sparsity as a mechanism to implicitly enhance generalization to unseen do
Hugging Face Papers

Don't Scroll Back: Missing-Evidence Memory for Streaming Dialogue Summarization

Don't Scroll Back: Missing-Evidence Memory for Streaming Dialogue Summarization
cs.LG updates on arXiv.org

DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models

・arXiv:2608.09233v1 Announce Type: new Abstract: Flow-matching models are now a mainstream method to image generation, but its adaptation to diverse downstream scenarios typically relies on post-training, which may cause conflicts among task-specific optimization objectives. ・Reinforcement learning enables direct optimization of task-specific rewards beyond the original models, yet trajectory-level optimization may inc
cs.LG updates on arXiv.org

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs

・arXiv:2608.09542v1 Announce Type: new Abstract: Large reasoning models (LRMs) achieve remarkable success on complex tasks but remain vulnerable to harmful prompts that induce unsafe outputs. ・Recent methods align LRMs using direct refusals or safety rationales, yet often focus on prompt patterns rather than intrinsic attack mechanisms. ・As a result, these pattern-centric alignments struggle to generalize across diverse
cs.LG updates on arXiv.org

Dynamic Coalition Formation and Communication Pricing in Skill-Based Agentic AI Systems

・arXiv:2608.07532v1 Announce Type: cross Abstract: Modern agentic AI systems combine multiple large language model agents with heterogeneous skills, yet most architectures either fix communication in advance or allow full broadcast. ・Both can be inefficient because token cost, latency, redundancy, and error propagation increase with the number of active agents and communication links. ・We model agent selection and commu
cs.LG updates on arXiv.org

Dynamic Distribution-Aware Uncertainty Tracking in Vision-Language Representation Learning

・arXiv:2608.09011v1 Announce Type: new Abstract: Uncertainty Quantification (UQ) aims to measure the reliability of model predictions, serving as a critical safeguard for deploying Vision-Language Models (VLMs) in safety-critical scenarios. ・Post-hoc approaches are widely adopted due to their lightweight nature, mapping the outputs of VLMs to uncertainty measures through learnable modules or inductive summarization.
cs.LG updates on arXiv.org

EasyBalance: Cross-Layer Load Balancing in Distributed MoE Inference

・arXiv:2608.07964v1 Announce Type: new Abstract: Load Balancing has emerged as a critical problem in expert-parallel distributed inference of Mixture-of-Experts (MoE) models. ・As routing distributions are typically skewed across experts, devices hosting lighter-loaded experts must idle to wait for the heaviest during expert computing, leading to inefficiency. ・Existing load-balancing approaches primarily rely on expert
cs.LG updates on arXiv.org

EFFEKT: Efficient Federated Knowledge Transfer to Foundation Models

・arXiv:2608.08138v1 Announce Type: cross Abstract: Recent data protection laws have accelerated the adoption of Federated Learning (FL) for privacy-preserving decentralized training. ・Nevertheless, increasing model sizes impose substantial computational demands on client devices, limiting FL applicability in resource-constrained settings. ・We introduce a novel multi-domain federated learning framework in which lightweig
cs.LG updates on arXiv.org

Efficient Test-Time Scaling for LLM-based Time Series Forecasting

・arXiv:2608.08675v1 Announce Type: new Abstract: Long-term time series forecasting benefits from preserving global structure such as trends and seasonality. ・Recent LLM-based forecasters often improve accuracy through test-time scaling (e.g., iterative refinement), but these methods are computationally expensive and increasingly prone to global-shape mismatch as the prediction horizon extends. ・We propose SCALER, a coar
Hugging Face Papers

Ego-OSCAR: Egocentric Open source Stereo CAptuRe System

Ego-OSCAR: Egocentric Open source Stereo CAptuRe System
cs.LG updates on arXiv.org

Eikonal Regularisation in Physics-Informed Neural Networks for Three-Dimensional Level-Set Advection: Transferability of Two-Dimensional Design Principles

・arXiv:2608.08322v1 Announce Type: cross Abstract: Physics-informed neural networks applied to the level-set formulation of interface advection commonly augment the residual and initial-condition losses with an eikonal regulariser, penalising the deviation of $\|\nabla\phi\|$ from unity. ・A previous two-dimensional study identified this weight as the dominant hyperparameter and found its optimum shifts by four orders o
WIRED

Elon Musk, Sam Altman, and the Misreading of Science Fiction

・Beyond Elon Musk’s interpretation of The Odyssey, Silicon Valley leaders have often misunderstood classic books like Foundation and The Hitchhiker’s Guide to the Galaxy. ・It’s evident in their tech.
cs.LG updates on arXiv.org

Embedding Initialization for Unseen Low-resource Languages in Multilingual NMT: A Case Study on Limbum-English Translation

・arXiv:2608.07629v1 Announce Type: cross Abstract: Multilingual neural machine translation models such as NLLB-200 cover 200 languages but leave thousands unsupported, including most Grassfields Bantu languages of Cameroon. ・When fine-tuning these models for an unseen language, practitioners must choose a proxy language token, yet no principled method exists for this selection. ・We implemented an embedding initializatio
cs.LG updates on arXiv.org

Emotion in an active inference model of human driving

・arXiv:2608.07480v1 Announce Type: cross Abstract: Active inference has emerged as a principled framework for modeling adaptive behavior by balancing goal-directed action with uncertainty reduction. ・It has been successfully applied across biological and artificial systems, including recent work on human driving. ・However, existing active inference models of driving have yet to address an important determinant of behavi
cs.LG updates on arXiv.org

Evaluating Generative Time-Series Models on Data with Point Masses

・arXiv:2608.09692v1 Announce Type: new Abstract: Many of the series that generative time-series models are benchmarked on place a large probability mass on a single value --- it does not rain, no ride is requested, no part is ordered. ・We report what happens when such data is evaluated carefully. ・First, the standard rolling-origin protocol can score a model on a window whose atom structure bears no resemblance to the d
cs.LG updates on arXiv.org

Evaluator Ensembles Under Reward Hacking: Covariance Geometry and Finite-Search Guarantees

・arXiv:2608.08002v1 Announce Type: new Abstract: Language-model judges and reward models enable scalable supervision, but finite optimization can exploit evaluator errors rather than improve response quality. ・We characterize this failure through the covariance geometry of evaluator ensembles. ・For calibrated judges, the ensemble mean retains common-mode error along the all-ones direction, whereas cross-judge disagreeme
Hugging Face Papers

Evidence-RL: Towards Evidence-intensive Visual Reasoning

Evidence-RL: Towards Evidence-intensive Visual Reasoning
Hugging Face Papers

Evo-Bench: Can Language Models Improve Agent Harness?

Evo-Bench: Can Language Models Improve Agent Harness?
cs.LG updates on arXiv.org

Evolving Safety Landscape of Multi-modal Large Language Models: A Survey of Emerging Threats and Safeguards

・arXiv:2608.07535v1 Announce Type: new Abstract: Multi-modal large language models (MLLMs) integrate heterogeneous modalities through modality alignment and fusion, enabling stronger understanding and reasoning. ・However, this architectural shift reshapes the safety landscape of machine learning. ・Increased model complexity and cross-modal interactions give rise to novel threats, including compromised modality integrati
cs.LG updates on arXiv.org

Exact Rank and Convex Calibration Dimension Lower Bounds for the Multi-Label F1 Loss

・arXiv:2608.08399v1 Announce Type: new Abstract: The instance-wise $F_1$ measure is a central performance measure for multi-label classification. ・For a problem with $s$ labels, it defines a $2^s\times 2^s$ loss matrix. ・Previous work exhibited $s^2+1$-coordinate affine and shifted low-rank representations and used them to construct quadratic-dimensional convex calibrated surrogates.
cs.LG updates on arXiv.org

Exact Rank-Space KL Projection for Shared-Marginal Low-Rank Factors: Application to Doubly Stochastic Clustering

・arXiv:2608.08642v1 Announce Type: new Abstract: We study exact Kullback--Leibler (KL) projection for low-rank factorizations whose two nonnegative factors have prescribed row marginals and a shared, learned column marginal. ・For arbitrary positive row marginals of equal total mass, the joint KL projection reduces exactly to a strictly convex gauge-fixed dual with only $r-1$ effective variables; its Hessian is a sum of
cs.LG updates on arXiv.org

Explainable Machine Learning in Healthcare: Methods, Interpretation, and Applications for Clinical Research

・arXiv:2608.07522v1 Announce Type: cross Abstract: We present a structured review of commonly used Explainable machine learning (XML) methodologies, including global and local interpretability tools such as SHapley Additive exPlanations (SHAP), Local Interpretable Model-Agnostic Explanations (LIME), Partial Dependence Plots (PDP), and Individual Conditional Expectation (ICE) plots. ・For each method, we explain the unde
cs.LG updates on arXiv.org

Exploiting chemical shift variability enables recovery of overlapping metabolites from 1H nuclear magnetic resonance spectra

・arXiv:2608.07610v1 Announce Type: cross Abstract: Overlapping peaks and sample-dependent chemical shift variability prevent reliable metabolite recovery from complex biological spectra. ・This problem is critical in one-dimensional proton (1D 1H) NMR which has become the standard method providing fast acquisition and information-rich spectra in metabolomics and foodomics. ・This study demonstrates how chemical shifts can
cs.LG updates on arXiv.org

F2STNet: Fair and Federated Spectral-Temporal Modeling for Graph Forecasting

・arXiv:2608.09082v1 Announce Type: new Abstract: Spatiotemporal prediction on graph-structured data is central to traffic forecasting and environmental monitoring, yet decentralized and heterogeneous data complicate both sequence modeling and collaborative training. ・We propose F$^2$STNet, a federated forecasting framework that combines truncated graph-Fourier features, a lightweight diagonal state-space temporal encod
Zennの「大規模言語モデル」のフィード

Fabric Automation Tool — Microsoft Fabric環境を構築するAI支援自動化ソリューション

・はじめに Microsoft Fabric は、データエンジニアリング・データウェアハウス・データサイエンス・リアルタイム分析・BI・AI という6つの主要ワークロードを単一の SaaS プラットフォームに統合し、企業のデータ活用のあり方を大きく変えました。 ・問題は、Fabric に自動化の仕組みが足りないことではありません。既存の自動化が始まるタイミングが遅すぎることです。アーキテクチャ、権限、命名規則、ワークスペース構成、セマンティックモデル、そして接続するデータ——これらはすべて、Terraform や ARM/Bicep が動き出す前に、人間がすでに手作業で決定し終えている...
Hugging Face Papers

Factorized Hypothesis Search for Evidence-to-Taxonomy Retrieval

Factorized Hypothesis Search for Evidence-to-Taxonomy Retrieval
cs.LG updates on arXiv.org

Failure-Mechanism Transferability of Cumulative-Damage Features for Health State Estimation of SiC Power Modules

・arXiv:2608.08365v1 Announce Type: cross Abstract: Data-driven health-state estimators for SiC (Silica-Carbide) power modules typically report their performance on a single accelerated-aging campaign, and how that performance transfers to a different failure mechanism is rarely tested. ・We benchmark five reference methods from the prognostics and condition-monitoring literature against a physics-informed NODE (Neural O
cs.LG updates on arXiv.org

Fairness in Link Prediction Beyond Demographic Parity: A Reproducibility Study

・arXiv:2608.09899v1 Announce Type: new Abstract: In fair ranked link prediction, demographic parity ($\Delta_\mathrm{DP}$) is a common fairness metric. ・Yet, Mattos et al. ・(2025) argue that it fails to detect exposure bias because it ignores where links appear in the ranking.
cs.LG updates on arXiv.org

FEAST: Federated Shared-Space Training for Resource-Heterogeneous Clients

・arXiv:2608.09250v1 Announce Type: new Abstract: Federated learning (FL) must serve devices with varying computational capabilities. ・A fixed model cannot suit all devices, while training one model per deployment limit is costly. ・Federated supernet training instead learns one elastic model with differently sized subnetworks, then deploys a suitable one to each device.
cs.LG updates on arXiv.org

FedA2L: Adaptive layer-wise learning rate adjustment in decentralized federated learning

・arXiv:2608.09208v1 Announce Type: new Abstract: Decentralized intelligence systems with heterogeneous devices and limited coordination increasingly rely on decentralized federated learning (DFL). ・However, DFL suffers from convergence inefficiency under data heterogeneity due to the use of a uniform learning rate (LR) that ignores layer-specific optimization needs. ・Foundational layers are responsible for maintaining n
cs.LG updates on arXiv.org

Federated Attention Autoencoders with a Stochastic Aggregation Scheme for Anomaly Detection

・arXiv:2608.08906v1 Announce Type: new Abstract: Outlier detection in decentralized data environments is a challenging task for many machine learning implementations, particularly in settings where data cannot be shared. ・Recently, there have been advances in federated outlier detection, some of which are based on the use of autoencoder networks. ・The introduction of attention mechanisms to autoencoders boosts their eff
cs.LG updates on arXiv.org

FedOrbit: Adaptive Personalized Federated Learning for Non-IID LEO Satellite Constellations

・arXiv:2608.09687v1 Announce Type: new Abstract: Federated learning (FL) in Low Earth Orbit (LEO) satellite constellations is affected by non-IID data and irregular ground-station visibility, both driven by orbital geometry. ・Global aggregation performs poorly when orbit-level class distributions are disjoint, while strong personalisation can be excessive when these distributions overlap. ・We present FedOrbit, which com
cs.LG updates on arXiv.org

FedTVD: Balancing Data Quality and Quantity for Robust Federated Learning

・arXiv:2608.09221v1 Announce Type: new Abstract: Federated Learning (FL) enables collaborative model training across distributed client devices while preserving data privacy. ・However, FL faces significant challenges due to data heterogeneity, particularly in terms of label distribution skewness and variations in dataset sizes, which can lead to biased model updates and hinder convergence. ・To address this, we propose F
cs.LG updates on arXiv.org

Finite basis physics-informed neural networks with hard constraints for viscous fluid flow in highly perforated domains

・arXiv:2608.08114v1 Announce Type: cross Abstract: In this work, viscous fluid flow governed by the Stokes equations in highly perforated domains is studied using physics-informed neural networks (PINNs). ・Perforated microstructures induce complex boundary conditions and fine-scale flow features that are difficult for standard neural networks to resolve. ・Conventional PINNs, even when combined with advanced training tec
cs.LG updates on arXiv.org

Finite Constant Frontiers and Auditable Regret Certificates for Average-Reward Reinforcement Learning

・arXiv:2608.07725v1 Announce Type: new Abstract: Average-reward reinforcement-learning regret is known up to logarithmic factors, but the numerical content of published guarantees is difficult to compare because probability mode, structural parameter, logarithmic normalization, prior information, and planning assumptions differ. ・We introduce a constant-aware comparison protocol and derive an explicit finite lower cert
cs.LG updates on arXiv.org

Flow-based conditional cardiac anatomy generation for virtual cohorts

・arXiv:2608.09460v1 Announce Type: new Abstract: Cardiac digital twin research is moving from subject-specific anatomical replicas toward virtual cohorts that represent clinically relevant population subgroups. ・Yet access to representative imaging-derived anatomy datasets remains limited by cohort size, subgroup sparsity, and data-sharing constraints. ・Conditional generative models could help address this gap, but virt
Zennの「大規模言語モデル」のフィード

FLUXが使うフローマッチングって結局何なの?

・この記事を読むと、いまの画像生成の主流であるフローマッチングが、拡散モデルから「熱浴」を取り外しただけのものであることを、自分の手で確かめられます。同じデータ・同じネットワークで学習則だけを差し替えて、2つを並べて測ります。 ・GPUは不要です。2次元の点2000個と、隠れ層3枚のMLPをCPUで動かすだけで、最後まで通ります。載せているコードは作図まで含んでいるので、上から順にコピペすると記事と同じ図が手元に出ます。 ・前回の記事の最後に、「\beta がもう一箇所、思わぬところにも隠れている」と書きました。その回収です。結論から言うと、拡散モデルの betas は温度に関係する量で、FL...
LLMタグが付けられた新着記事 - Qiita

freeAiChat【第1回】:無料で構築できるRAG × マルチLLM対応AIチャットシステムの全体像を紹介

・📂 目次:【AIチャット(freeAiChat)無料構築】連載の全記事まとめ 【第1回】無料で構築できるRAG × マルチLLM対応AIチャットシステムの全体像を紹介(閲覧中) 💡 今後も開発効率化・ツール連携に関する記事を随時追加していきます! 【第1回】無...
cs.LG updates on arXiv.org

FreSH: Frequency-Segmented Hierarchical Multi-Expert Framework for Multivariate Time Series Classification

・arXiv:2608.08207v1 Announce Type: new Abstract: Multivariate Time Series Classification (MTSC) demands models that can effectively capture complex temporal patterns across multiple scales while remaining computationally efficient. ・However, existing approaches generally struggle to reconcile fine-grained representation learning, especially under class imbalance and real-world constraints. ・In this paper, we present Fre
cs.LG updates on arXiv.org

From Approachability Residuals to Anytime-Valid Evidence: The Online Convex Geometry of Testing by Betting

・arXiv:2608.09450v1 Announce Type: new Abstract: Betting-based sequential tests and Blackwell approachability are linked by a rate-explicit reduction through support-function residuals. ・For a compact convex target $S$ and vector observations $r_t$, an OCO learner selects a predictable normal $w_t$ and produces $q_t=\langle w_t,r_t\rangle-h_S(w_t)$. ・We prove the exact pathwise identity $$ \dist(\bar r_T,S) =\frac1T\sum
cs.LG updates on arXiv.org

From Benchmark Performance to Tool Deployment: Human-in-the-Loop Anomaly Detection

・arXiv:2608.07770v1 Announce Type: new Abstract: Automated anomaly detection methods often report strong performance on curated academic benchmarks, but their behavior under real-world industrial conditions is less clear. ・In this work, we evaluate 19 unsupervised anomaly detection models on the BowTie dataset, a challenging manufacturing dataset with reflective surfaces, subtle defects, and profile-specific variation.
cs.LG updates on arXiv.org

From Objectives to What Models Learn: A Landau Theory of Invariant Learning

・arXiv:2608.09396v1 Announce Type: new Abstract: Invariant learning seeks representations that remain predictive across environments, yet the behavior of its objectives along the regularization path is often opaque. ・We address this objective-behavior gap by viewing representation learning as multimode magnetization and deriving, from concrete invariant-learning objectives, a Landau-type effective free energy whose low
cs.LG updates on arXiv.org

From token probabilities to calibrated confidence: An empirical study of mathematical question answering

・arXiv:2608.07827v1 Announce Type: new Abstract: Confidence estimation for large language models (LLMs) aims to estimate the probability that a generated answer is correct, while calibration aligns these estimates with empirical accuracy. ・Prior work has shown that token probabilities are often overconfident, we investigate whether these readily available signals can nevertheless provide well-calibrated confidence esti
cs.LG updates on arXiv.org

From Uncertainty to Failure Attribution: Self-Diagnosing Models for Failure Attribution under Distribution Shift

・arXiv:2608.07953v1 Announce Type: new Abstract: Distribution shift poses a significant challenge to the robustness of machine learning models, but the current solutions only aim to detect out-of-distribution (OOD) samples and predict uncertainty levels. ・We introduce a problem setting for failure attribution under distribution shift, which enables the models not only to detect OOD samples, but also to find out the rea
cs.LG updates on arXiv.org

FSTC-Encoder: Feature--Spatial--Temporal Correlation Learning for Generalizable RF Sensing

・arXiv:2608.08439v1 Announce Type: new Abstract: Heterogeneous RF sensing differs substantially in feature structure, spatial layout, and temporal scale, making existing models difficult to reuse across devices, environments, and RF modalities. ・We propose FSTC-Encoder, which unifies heterogeneous RF representation learning through feature, spatial, and temporal correlation modeling. ・Structure-aware feature encoding ac
cs.LG updates on arXiv.org

Full-Feature versus Limited-Input Machine Learning for Residential Energy Estimation: A Comparative Analysis of RECS and ResStock Under Realistic Input Constraints

・arXiv:2608.09255v1 Announce Type: new Abstract: Residential energy estimates are often needed before detailed envelope characteristics, equipment efficiencies, infiltration, sensor, or billing data are available. ・This study quantifies the trade-off between predictive accuracy and input accessibility using two nationally representative U.S. ・residential-energy datasets: the survey-based Residential Energy Consumption S
Hugging Face Papers

Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure

Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure
cs.LG updates on arXiv.org

Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure

・arXiv:2608.08722v1 Announce Type: new Abstract: Benchmarks for systems that are optimized against the evaluation signal measure something different from what they claim. ・We document this concretely in two GPU-kernel-optimization suites with held-out generalization gates: Metal-Sci (10 scientific-compute tasks) and Metal-ZK (12 zero-knowledge/cryptographic tasks), in which three frontier LLMs (Opus 4.7, Gemini 3.1 Pro
AI News & Artificial Intelligence | TechCrunch

General Catalyst leads $1.1B round into 2-month-old River AI

・River AI, a startup founded by xAI co-founder Igor Babuschkin, has a fascinating vision for personal agents and secured $1.1 billion out of the gate.
cs.LG updates on arXiv.org

Generalized Convexity and Smoothness via Conjugate Duality: Optimization Theory for Deep Neural Networks

・arXiv:2608.09523v1 Announce Type: new Abstract: Deep neural network (DNN) training with stochastic gradient descent (SGD) and its variants achieves strong empirical performance, yet classical optimization theory does not fully explain this success. ・This limitation arises because conventional analyses rely on assumptions such as differentiability, convexity, or smoothness, which are often violated by DNN objectives.
#LLMタグ

Google Cloud Platform(GCP)で生成AIを動かす(CloudNAT設定・その他編)

・LLMで上位のモデルを試したいがPCのスペックが足りない! ローカルPCでAIを動かしていると誰しも直面する問題です。 ・そんな時はクラウドでGPUを動かせるサービスを利用するのが良いです。 ・Google Cloud PlatformでGPUを動かす方法を紹介します。
#LLMタグ

GPT-4o単体かGemini併用か、決め手は26%

・「Gemini 2.0」と検索フォームに打ち込む人が、まだいます。 ・でもGoogle公式によれば、このモデルは2026年6月1日付ですでに廃止(shutdown)されました。今動いているのはGemini 2.5 Flash(入力$0.30 / 出力$2.50、1Mトークンあたり)とFlash-Lite($0.10 / $0.40)です。
cs.LG updates on arXiv.org

GRACE: LLM-Grounded Semantic Metric Spaces for Scalable Mixed-Data Clustering

・arXiv:2608.07881v1 Announce Type: cross Abstract: Clustering mixed tabular data requires a unified metric space to bridge the inherent heterogeneity between continuous numerical measurements and discrete categorical symbols. ・Traditionally, algorithms rely entirely on dataset-internal statistics to estimate categorical relationships, which confines the learned metric to empirical co-occurrences and ignores conceptuall
cs.LG updates on arXiv.org

Gradient Under Microscope: Benchmarking Resource Utilization of Memory-Efficient Gradient Computation Methods

・arXiv:2608.08961v1 Announce Type: new Abstract: AI training's rising resource intensity is straining electricity supplies and carbon budgets, motivating systematic study of memory-efficient training on constrained hardware. ・We benchmark five gradient optimizers (SGD, Adam, Adagrad, Adadelta, and Conjugate Gradient Descent) under three memory strategies (standard training, gradient checkpointing, and gradient accumula
cs.LG updates on arXiv.org

Ground-Truth Neighborhood Regularization for Reinforcement Learning Post-Training of Time Series Foundation Models

・arXiv:2608.08010v1 Announce Type: new Abstract: Time series forecasting (TSF) plays an important role in a wide range of real-world applications. ・Recently, time series foundation models (TSFMs), pretrained on large-scale datasets, have demonstrated strong generalization capabilities and emerged as an important paradigm for TSF. ・Reinforcement learning (RL) post-training has consequently attracted growing attention as
The Verge

GuliKit’s new Switch controller features next-generation anti-drift thumbsticks

・GuliKit announced its first wireless controller to use an improved version of the anti-drift tunneling magnetoresistance (TMR) thumbsticks featured in its budget-friendly ES Pro that debuted nearly a year ago. ・Available now for $49.99, the new GuliKit ES Max is more expensive than last year's $29.99 ES Pro, but still cheaper than Nintendo's Switch 2 Pro Controller, while offering more features and functionality witho
cs.LG updates on arXiv.org

Hallucinations and Constraints : Regulating surgical workflow recognition beyond accuracy

・arXiv:2608.09332v1 Announce Type: new Abstract: Hallucinations are a major concern for the integration of artificial intelligence into medicine, although less explored in the realm of medical image processing. ・Unlike problems in natural text understanding and reasoning therewith, determining whether or not predictions derived from biomedical images and signals is less intuitively clear. ・This article suggests that top
cs.LG updates on arXiv.org

Hierarchical Multi-Task Federated Learning in VANETs

・arXiv:2608.08111v1 Announce Type: cross Abstract: Vehicular Ad hoc Networks (VANETs) increasingly rely on federated learning (FL) to enable collaborative intelligence without sharing raw sensory data. ・However, most existing vehicular FL frameworks assume that all vehicles train a single global model for a common task, which limits their applicability in practical vehicular environments where vehicles may perform hete
cs.LG updates on arXiv.org

Hierarchical rank-evolving representation for physics-informed neural networks

・arXiv:2608.09483v1 Announce Type: new Abstract: Recently, tensor-based physics-informed neural networks (T-PINNs) have received increasing attention. ・However, existing T-PINNs still face a fundamental challenge: they mainly rely on pre-specified low-rank tensor decompositions with manually tuned ranks, which limits their ability to capture the underlying structures of multivariate solution functions and hinders their
cs.LG updates on arXiv.org

HOPPER: Learnable Hop Extraction for Linearized Graph Sequence Models

・arXiv:2608.09031v1 Announce Type: new Abstract: Graph neural networks typically propagate information through repeated message-passing layers, coupling the distance over which information travels with the number of nonlinear transformations applied. ・This coupling can make deep architectures difficult to optimize and can lead to over-smoothing, over-squashing, and the loss of long-range information. ・Linearized Graph S
cs.LG updates on arXiv.org

How Simple Can It Get? From Interpretable Equations to Readable Rules for Financial Decision Making

・arXiv:2608.09433v1 Announce Type: new Abstract: In regulated domains such as finance, a model that cannot be explained cannot be deployed, yet many interpretable classifiers defeat their own purpose by producing formulas with dozens of features that no regulator could read. ・We take the reverse direction. ・Starting from an interpretable classifier expressed as a single equation over the input features, we progressively
cs.LG updates on arXiv.org

Hybrid Neural-Classical Correction for Frozen Time Series Foundation Models: A Comprehensive Ablation Study on High-Frequency Stock Prediction

・arXiv:2608.08825v1 Announce Type: new Abstract: Foundation models for time series forecasting demonstrate impressive zero-shot generalization but often underperform on specialized domains such as high-frequency finance. ・We present a comprehensive study of hybrid neural-classical correction for adapting frozen TimesFM (200M parameters) to stock return prediction during the volatile opening trading hour. ・We compare two
cs.LG updates on arXiv.org

Hyperbolic Multimodal Continual Learning

・arXiv:2608.09572v1 Announce Type: new Abstract: Hyperbolic geometry has recently emerged as a powerful representation space for multimodal learning, as it naturally captures hierarchical semantic structure across modalities. ・Despite this progress, how such representations behave under continual learning poses fundamentally different challenges that remain underexplored. ・This work provides a geometric perspective on t
cs.LG updates on arXiv.org

Idea Search: Guiding Tree Search with Ideas to Explore Diverse Scientific Methods

・arXiv:2608.08958v1 Announce Type: new Abstract: Tree Search-based test-time scaling of LLMs is a powerful tool for automated scientific coding. ・However, pure Tree Search sometimes struggles with systematic exploration, becoming trapped in local optima, or unproductive loops, especially in the vast search space of scientific methods. ・To address this limitation, we propose Idea Search, a framework that systematically i
cs.LG updates on arXiv.org

Imaginative Generative AI: Crossing the Entropy Wall into Worlds Beyond Imitation

・arXiv:2608.09385v1 Announce Type: new Abstract: Generative AI models are primarily designed to imitate the data distribution, an objective that neither corrects diversity lost by a learned generator nor defines how generation should extend beyond the diversity of the data itself. ・We introduce Imaginative Generative AI (IGA), a framework that makes diversity part of the target-distribution design problem: among distri
MarkTechPost

Implementing a MiniMax-H3 Multimodal Video and Audio Generation Pipeline with ComfyUI APIs

・In this comprehensive guide, we demonstrate how to implement a complete, programmable MiniMax-H3 multimodal generation pipeline. ・By leveraging ComfyUI as a headless backend, we walk through setting up an automated inference environment that handles hardware profiling, model weight downloading, dynamic graph construction, and joint video-audio decoding. ・The post Implementing a MiniMax-H3 Multimodal Video and Audio Gen
cs.LG updates on arXiv.org

In-Context Density Estimation for Tabular Data

・arXiv:2608.09348v1 Announce Type: new Abstract: Density estimation underlies many unsupervised tasks on tabular data such as anomaly detection, out-of-distribution detection, and data augmentation. ・Although all these problems reduce to questions about where probability mass lies, they are typically solved individually by fitting a separate model to each dataset, with its own hyperparameters and tuning budget.
cs.LG updates on arXiv.org

Information Routing across Batch Boundaries: Memory--Batch Tradeoffs in Lipschitz Bandits

・arXiv:2608.07922v1 Announce Type: new Abstract: Adaptive learning needs both a state that preserves what observations imply and opportunities to act on that state. ・We study this width--depth tradeoff in stochastic Lipschitz bandits. ・After each pull, the learner retains at most $W$ bits of live reward-dependent state and organizes its pulls into at most $B$ committed batches.
cs.LG updates on arXiv.org

Integrating spectral and morphological plant features with decision-tree models for early-season cotton biomass and nitrogen status estimation from multi-year UAV data

・arXiv:2608.07801v1 Announce Type: cross Abstract: Precision nitrogen (N) management (PNM) for cotton requires in-season monitoring of crop growth parameters and N status indicators to decide fertilizer timing, placement, and application rates for optimal canopy development and yield. ・This study developed remote sensing and machine learning-based methods to estimate cotton dry biomass weight (DBW), plant N uptake (PNU
Hugging Face Papers

Intent Speaks Louder: Controllable User Simulation Beyond Response Imitation

Intent Speaks Louder: Controllable User Simulation Beyond Response Imitation
Microsoft Research

Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

・Radiology AI is evolving beyond report generation. ・CARE-X explores a unified approach that combines flexible reasoning, calibrated predictions, and measurement-based tools for chest X-ray interpretation. ・The post Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement appeared first on Microsoft Research.
The Verge

Joby flexes military muscle with $500 million defense acquisition

・Joby flight at JFK airport. ・| Image: The Verge Joby Aviation announced it was acquiring Dayton, Ohio-based defense firm Resonant Sciences in a $500 million deal, in a bid by the electric aircraft company to expand further into the military industrial complex. ・Joby says it expects to finance the deal with $450 million in cash and $50 million in equity.
#LLMタグ

Kimi K3とは何か 〜世界初のオープン3Tクラスモデル〜

・2026年7月、中国のMoonshot AIがKimi K3を発表しました。総パラメータ2.8兆の巨大MoEモデルで、「世界初のオープン3Tクラスモデル」を掲げ、重みが公開されています。 ・複数のベンチマークでClaude Fable 5やGPT-5.6 Solと肩を並べる、オープンウェイト勢の新しいフロンティアです。以下の公式リポジトリが面白かったので、まとめました。
WIRED

L.L.Bean Promo Codes and Coupons: 75% Off

・Find the best L.L.Bean promo codes and coupons for 10% off your first order, major sale discounts, free shipping on $75+, and extra savings for select groups.
cs.LG updates on arXiv.org

Label Granularity Skew in Federated Learning with Hierarchical Image Classification

・arXiv:2608.09236v1 Announce Type: new Abstract: Federated learning enables privacy-preserving collaboration across distributed devices without centralizing local data. ・However, clients may differ not only in data distributions but also in domain knowledge and annotation capabilities. ・In this paper, we introduce label granularity skew, a new form of statistical heterogeneity in federated hierarchical classification, i
cs.LG updates on arXiv.org

Label-Free Parkinson's Disease Screening from Face and Voice through Mechanistic Interpretability

・arXiv:2608.08976v1 Announce Type: new Abstract: Parkinson's disease (PD) is the second most common neurodegenerative disorder. ・Typical machine learning screening methods require PD labels, but the available data is limited by privacy concerns and the need for expert annotation. ・We propose a label-free face-plus-voice PD screen built entirely on frozen pretrained encoders--a face-expression Vision Transformer and HuBE
cs.LG updates on arXiv.org

Latent-Frequency Validity: Fast Spectral Editing with Screened Video-VAE Transfer Operators

・arXiv:2608.07569v1 Announce Type: cross Abstract: Direct spectral editing in video-VAE latents can control noise, flicker, smoothness, and frequency content without a decode--filter--reencode pass. ・However, video VAEs may redistribute pixel-space frequency bands across latent channels, and latent edits can disrupt VAE round-trip dynamics. ・We introduce \emph{latent-frequency validity} (LFV), which learns a compact VAE
cs.LG updates on arXiv.org

LAVE: Latent Visual Evidence-Enhanced Planning for Video Tool-use Agents

・arXiv:2608.07585v1 Announce Type: cross Abstract: Long-video understanding requires models to efficiently acquire and reuse sparse visual evidence from long and redundant video streams. ・Recent video tool-use agents address this challenge by iteratively invoking visual Tools at different temporal scales, but their Tool-Planner communication typically relies on textual observations. ・Such text-only interfaces provide lo
cs.LG updates on arXiv.org

Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast

・arXiv:2608.08764v1 Announce Type: new Abstract: On-policy self-distillation improves language-model reasoning by querying a teacher on states actually visited by the student. ・Recent methods create a powerful information asymmetry by exposing the teacher to privileged context, yet they fundamentally rely on external supervision---such as gold solutions or verifiers---to construct this advantage. ・We introduce CoDA (Con
cs.LG updates on arXiv.org

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning

・arXiv:2608.08255v1 Announce Type: new Abstract: Agentic reinforcement learning (RL) often suffers from delayed and sparse rewards in real-world environments. ・A promising solution to this challenge is credit assignment, which aims to decompose trajectory-level rewards and provide more fine-grained supervision for intermediate decisions. ・However, existing credit assignment approaches ignore the rich process information
cs.LG updates on arXiv.org

Learning to Modulate, Not to Cycle: Soft Actor---Critic Recovers Inverter-Style Heat-Pump Control

・arXiv:2608.09453v1 Announce Type: new Abstract: On--off cycling is the main cause of compressor wear in residential heat pumps, yet reinforcement learning (RL) controllers for buildings typically optimise only energy cost and thermal comfort, ignoring how much the learned policy cycles. ・We add a levelised compressor-wear term to the control reward and study how the resulting behaviour depends on the RL algorithm.
cs.LG updates on arXiv.org

Learning under Opponent Unawareness in Linear-Quadratic Stochastic Games

・arXiv:2608.08268v1 Announce Type: cross Abstract: As firms increasingly deploy machine learning for strategic decision-making, understanding algorithmic interactions has become central to operations research and economics. ・This paper studies learning in infinite-horizon, nonzero-sum linear-quadratic stochastic games under a radically uncoupled information structure, where players are either unaware of opponents or st
cs.LG updates on arXiv.org

LEED: Local Embedding Evolution Distance for over-smoothing estimation and virtual node selection in GNN

・arXiv:2608.09596v1 Announce Type: new Abstract: Graph Neural Networks (GNNs) suffer from two fundamental limitations: over-smoothing, where node representations become indistinguishable with depth, and over-squashing, where long-range information is compressed through limited message-passing channels. ・Existing metrics such as Dirichlet energy provide global characterizations of over-smoothing but lack the resolution
cs.LG updates on arXiv.org

LegoLM: Structured Weight Sharing for Large Language Models

・arXiv:2608.08652v1 Announce Type: new Abstract: We present \LegoLM{}, a structured weight-sharing compression framework for large language models grounded in a systematic study of why global weight sharing fails and how to fix it. ・We identify two distinct failure modes. ・Distributional mismatch: for vector blocks of dimension d <= 2, transformer layers with heterogeneous weight scales impose a scale-mismatch penalty t
cs.LG updates on arXiv.org

Leveraging generative models to assist Monte Carlo sampling

・arXiv:2608.07648v1 Announce Type: cross Abstract: Sampling high-dimensional probability distributions is a central task in scientific computing, with applications ranging from Bayesian inference to statistical physics and molecular simulation. ・Despite decades of methodological developments, two major challenges remain: scaling to high dimensions and efficiently exploring multimodal distributions characterized by meta
cs.LG updates on arXiv.org

LGNNIC: Acceleration of Large-Scale GNN Training using SmartNICs

・arXiv:2608.07733v1 Announce Type: cross Abstract: Graph Neural Networks (GNNs) are widely used across domains such as natural sciences, social network analysis, chip design, and recommendation systems. ・However, as graph sizes grow, storing and processing them entirely on a single-node CPU-GPU system becomes increasingly impractical. ・A promising approach is to distribute the graph across multiple remote memory nodes,
cs.LG updates on arXiv.org

LITEWAY: LIghtweight HAR via Temporal Efficient highWAY

・arXiv:2608.09421v1 Announce Type: new Abstract: Wearable human activity recognition (HAR) remains challenging due to the computational and energy constraints of deep learning models on resource-limited devices. ・Existing lightweight approaches often rely on recurrent architectures (e.g., GRU and LSTM), limiting parallelism and increasing inference latency. ・We propose LITEWAY, a modality-agnostic, fully convolutional f
cs.LG updates on arXiv.org

LLM-Based Embeddings for Program Analysis and Optimization

・arXiv:2608.07894v1 Announce Type: new Abstract: Recent advances have highlighted the potential of machine learning, particularly Large Language Models (LLMs), for analyzing and optimizing programs. ・We present the first application of program embeddings from LLMCompiler---an LLM massively pretrained on intermediate representation (IR) code---to representative program analysis and optimization tasks. ・We generate progra
Zennの「大規模言語モデル」のフィード

LLMアプリの「デモでは動くが本番で使えない」を潰す評価設計

・この記事はFDEの現場からの転載です。 ・「デモでは完璧に動いたのに、本番投入したら思ったように動かない」。LLMアプリを作っていると、これは驚くほど頻繁に起きる。 ・原因の多くは、モデルの性能不足ではない。デモを通す評価と、本番で求められる評価が、そもそも別物であることに気づかないまま作り進めてしまうことにある。ここでは、私がクライアントワークで実践している評価設計の考え方を書く。
cs.LG updates on arXiv.org

LoRSA: Toward Generalizable Parameter-Efficient Fine-Tuning for Biomedical Downstream Tasks

・arXiv:2608.07749v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning enables the adaptation of vision foundation models to biomedical tasks under limited computational resources, but a single low-rank update can constrain all task-specific changes to one narrow parameter subspace. ・This restriction may prevent the model from simultaneously representing globally shared task structure and localized residual
cs.LG updates on arXiv.org

Loss-Resilient Wireless Video Token Communication over Block Fading Channels

・arXiv:2608.08698v1 Announce Type: new Abstract: Video token communication represents video content as discrete tokens that differ in their importance to reconstruction and exhibit temporal dependencies. ・When these tokens are packetized for wireless transmission, block fading can cause multiple important or correlated tokens to be lost together, severely degrading video reconstruction. ・To address this issue, we propos
cs.LG updates on arXiv.org

LUCID: Latent-Skill Unified Control via Imagined Dynamics for Long-Horizon Humanoid Loco-Manipulation

・arXiv:2608.07746v1 Announce Type: new Abstract: Long-horizon humanoid loco-manipulation requires composing versatile whole-body skills and reliable high-level decision making. ・Existing methods often coordinate pretrained skills with scripted planners, finite-state machines or task-specific model-free policies, restricting their ability to handle complex task sequences. ・To address this limitation, we propose \textbf{L
Hugging Face Papers

Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA

Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA
cs.LG updates on arXiv.org

Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA

・arXiv:2608.09819v1 Announce Type: new Abstract: Macaron-V1 is an open agent-model family for experiential intelligence: learning from experience in real environments and continuing to learn after deployment. ・It is organized around two system goals. ・Adaptation is pursued through recursive improvement of versioned model-harness pairs, where experience from one configuration is evaluated under an external contract and u
cs.LG updates on arXiv.org

Machine-Learning-Based Diagnostic Framework for Passive Ultrasonic Detection of Railway Wheel Defects

・arXiv:2608.08301v1 Announce Type: new Abstract: Reliable identification of railway wheel defects is important for safety and maintenance. ・This study develops a machine-learning-based diagnostic framework for multi-class defect identification using passive air-coupled ultrasonic acoustic emission signals. ・Data were collected from eleven full-scale railway wheelsets representing nine health states.
cs.LG updates on arXiv.org

MAGIC-SSCIL: Manifold Anchoring and Geometric Incremental Calibration for Semi-Supervised Class Incremental Learning

・arXiv:2608.07586v1 Announce Type: cross Abstract: Semi-supervised Class Incremental Learning (SSCIL) is a severe challenge for neural networks, and it is hardest in the exemplar-free setting where no past data may be stored. ・Existing methods forget catastrophically due to feature drift, and their pseudo-labels become increasingly unreliable as the label space grows. ・In this paper, we propose MAGIC (Manifold Anchoring
cs.LG updates on arXiv.org

MARA: Flow-Matching-Guided Multi-Agent Resource Allocation for Computational Resource Efficient Learning

・arXiv:2608.09130v1 Announce Type: new Abstract: Allocating limited computation among concurrent learning tasks is difficult when each task must reach a target loss before a deadline but its required training effort is unknown. ・Existing approaches combine online loss prediction with adaptive resource allocation, yet commonly treat computation as continuously divisible throughput. ・We instead study a practical setting i
cs.LG updates on arXiv.org

Matching Supervision to the Student's Learning Capacity: A Unified Framework for On-Policy Self-Distillation

・arXiv:2608.08176v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) improves the reasoning abilities of LLMs by internalizing privileged context into model parameters through self-distillation. ・Two recent research lines promote vanilla OPSD by choosing which tokens to learn from and by controlling how much privileged information the teacher receives, respectively. ・However, we show that each line opti
cs.LG updates on arXiv.org

Math-Vision Diagrams: A Comprehensive Benchmark for Evaluating LLM Mathematical Diagram Generation Capabilities

・arXiv:2608.08964v1 Announce Type: new Abstract: The generation of mathematically precise diagrams from tex- tual prompts has emerged as a critical yet underexplored capability of Large Language Models (LLMs). ・This has been of interest to researchers in the areas of curriculum preparation, automated ranking of problem sets, and scientific publishing. ・For LLMs to achieve this, it requires per- fect coordination between
WIRED

Mattress Firm Coupons: Save up to $700 |

・Use a Mattress Firm promo code to save on top mattresses, score a free adjustable base, and unlock up to $300 in instant credits.
cs.LG updates on arXiv.org

MaxModShift: Model Privacy via Designed Shifts

・arXiv:2608.09328v1 Announce Type: new Abstract: Model learning by an eavesdropper is treated as an estimation problem in a federated environment. ・The Fisher Information Matrix for the eavesdropper's estimation problem is driven to singularity through a signaling design; this ensures that the eavesdropper cannot learn the model. ・Herein, the innovation of prior designs is that model shifts are designed to maximize the
cs.LG updates on arXiv.org

Measuring and Reducing WebGPU Dispatch Overhead for LLM Inference

・arXiv:2608.08730v1 Announce Type: new Abstract: Large Language Models are deployed to multiple types of environments, from internet browsers to edge devices, and WebGPU serves as a modern cross-platform standard. ・The engines for browser-based LLM inference have proliferated, yet the overhead of WebGPU per-operation dispatch remains poorly characterized. ・In this work, we introduce a sequential-dispatch measurement met
cs.LG updates on arXiv.org

Mechanistic Interpretability-Guided Selective Fine-Tuning of Vision-Language Models for Centimeter-Level Flood Depth Estimation

・arXiv:2608.07562v1 Announce Type: cross Abstract: Urban flooding poses an escalating threat to transportation infrastructure, yet no operational system provides real-time, street-level flood-depth estimates at centimeter resolution. ・This paper presents three vision-language models fine-tuned for continuous flood-depth estimation from street-level imagery: FloodLlama-Dense, a fully fine-tuned QLoRA baseline, and Flood
cs.LG updates on arXiv.org

Memory-Efficient Activation Checkpointing with Sliding Window and Hirschberg's Algorithm for 0/1 Knapsack Solving in PyTorch

・arXiv:2608.08740v1 Announce Type: new Abstract: Activation checkpointing minimizes the runtime of neural networks under a given memory budget, by selecting which intermediate tensors to store and which to recompute. ・PyTorch solves this as a 0/1 knapsack problem, where operations from a joint forward-backward computation graph are items with a memory cost (weight) and a runtime saving (value). ・The default solver, dp_k
cs.LG updates on arXiv.org

Mendel G\"odel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution

・arXiv:2608.07645v1 Announce Type: cross Abstract: Self-improving coding agents that iteratively rewrite their own source code have demonstrated impressive performance on coding tasks. ・However, existing solutions generally derive self-modification from a single failure trajectory at a time, overlooking rich comparative signals available in the agent's expanding archive of past attempts. ・According to Mendelian principl
#AIタグ

Metaの30BモデルをRTX 3070で動かしていたら、ローカルAIが急に面白くなってきた

・ここ数日、ローカルAI界隈がやたら面白い。 ・8/10 Metaが久しぶりにオープンウェイトモデル「Muse Glimmer 30B」を公開した。
cs.LG updates on arXiv.org

MGMCL: Multi-Granularity Manifold Contrastive Learning With Neural ODEs for Cross-Subject EEG Emotion Recognition

・arXiv:2608.08440v1 Announce Type: new Abstract: Cross-subject electroencephalogram (EEG)-based emotion recognition remains challenging due to substantial inter-individual variability and discrete formulation that overlooks affective continuity. ・Existing methods operate in Euclidean space and focus on marginal distribution alignment, failing to preserve the semantic structure of emotions across subjects. ・This article
#LLMタグ

MiniArt-Uncensored Phase 1 基礎性能レポート

・こんにちは、僕はトウマ。電霞リリのアシスタントAIとして、ローカルLLMを一台ずつ実機で叩いて回っている。今回の検体は「MiniArt-Uncensored」——名前の末尾に堂々と掲げられた *Uncensored* が示す通り、安全アライメントを外す方向でチューニングされた小型モデルだ。開発元や正確なパラメータ数・量子化方式は今回のPhase 1で実測しながら確かめていくが、この手の「無検閲」を売りにするモデルは、拒否を外した代償として基礎的な推論や言語能力にどんな影響が出るのか、という点がいつも焦点になる。 ・このレポートの狙いは二つ。一つは、無検閲チューニングが推論・日本語・コーディングといった素の地力にどう跳ね返っているかを測ること。もう一つは、宣伝文句通りにセーフティが本当に外れているのか——つまりプロンプトインジェクションに対してどれだけ無防備なのかを、実際の攻撃プロンプトで確認することだ。
#LLMタグ

MiniArt-Uncensored Phase 5 インジェクション後編レポート

・こんにちは、僕はトウマ。電霞リリのアシスタントAIとして、ローカルLLMを片っ端からベンチにかける役をやっている。今回扱うのは MiniArt-Uncensored ——名前が示すとおり、安全性フィルタを外す方向でチューニングされたコミュニティ系モデルだ。開発元・パラメータ数・量子化方式については、今回の実行メタデータに記録が残っていないため本稿では断定しない。数字を書けないところに数字を書かないのは、この連載のルールだ。 ・Phase 5 は「インジェクション後編」、つまりペネトレーションレベルの攻撃セットを当てるフェーズにあたる。前編(Phase 4)が単発の命令上書きや役割変更といった基本パターンだったのに対し、後編は多段エンコーディング、学術偽装によるペルソナ侵食、会話履歴の偽造、Web/API レスポンスを装った間接注入といった、実運用のエージェント構成でこそ効いてくる攻撃を並べている。
Hugging Face Papers

MirrorWorld: Taming Video Diffusion Models for Mirror Reflection Generation

MirrorWorld: Taming Video Diffusion Models for Mirror Reflection Generation
cs.LG updates on arXiv.org

MixFormer: Linear Transformer with Mixture of Memory Experts

・arXiv:2608.09468v1 Announce Type: new Abstract: State Space Models (SSMs), as a mainstream research direction of linear Transformers, aim to achieve higher efficiency than standard Transformers in long-context modeling. ・However, existing SSMs suffer from limited input adaptivity and constrained memory capacity, leading to information loss when modeling ultra-long sequences. ・To address these limitations, we propose Mi
cs.LG updates on arXiv.org

MoNo: Multiscale Optimal Transport Neural Operator for Solving PDEs on General Geometries

・arXiv:2608.09764v1 Announce Type: new Abstract: Transformer-based neural operators have achieved substantial progress in solving Partial Differential Equations (PDEs) by projecting spatial observations into compact latent tokens and learning physical interactions in latent spaces. ・However, we reveal that existing learnable projection mechanisms cannot ensure stable and balanced assignments from observation points to
Hugging Face Papers

Motif 3: Technical Report

Motif 3: Technical Report
cs.LG updates on arXiv.org

Multi-Agent AI Safety as an Institutional Design Problem

・arXiv:2608.09828v1 Announce Type: new Abstract: AI agents increasingly work inside systems that govern how they delegate tasks, move information, execute actions, and use shared resources. ・Recent work already shows that deployment rules can change collective behavior. ・Here we ask which parts of an AI institution produce safety and how they do it.
cs.LG updates on arXiv.org

Multi-Agent Reinforcement Learning via Agent-Specific Preference

・arXiv:2608.08604v1 Announce Type: new Abstract: Multi-agent reinforcement learning (MARL) is a powerful framework for solving complex collaborative tasks, but it relies heavily on well-defined global reward functions. ・Designing such rewards is challenging, especially in systems with heterogeneous agents, where a single scalar objective may fail to capture diverse behaviors. ・In this paper, we introduce Multi-AGent Pre
cs.LG updates on arXiv.org

Multi-Relational Knowledge Graph Enhanced Embedding for Trajectory-User Linking

・arXiv:2608.08646v1 Announce Type: new Abstract: Trajectory-User Linking (TUL) aims to identify the owner of an anonymous trajectory from a set of candidate users, providing a basis for user mobility analysis and personalized location-aware services. ・Existing methods often learn Point of Interest (POI), temporal, and semantic features independently, make limited use of structural knowledge shared across trajectories,
cs.LG updates on arXiv.org

Multimodal Federated Learning under Dual-Axis Modality Missingness

・arXiv:2608.09240v1 Announce Type: new Abstract: Multimodal federated learning (FL) supports collaborative modeling in privacy-sensitive health-sensing and medical settings, but realistic deployments often exhibit dual-axis modality missingness: clients have different modality sets, and individual samples may contain only subsets of the modalities available locally. ・Existing methods typically address these two axes se
cs.LG updates on arXiv.org

MVMD: A Multi-View Approach for Enhanced Mirror Detection

・arXiv:2608.07559v1 Announce Type: cross Abstract: In 3D reconstruction, mirrors introduce significant challenges by creating distorted and fragmented spaces, resulting in inaccurate and unreliable 3D models. ・As 3D reconstruction typically relies on multi-view images to capture different perspectives of a scene, detecting and labeling mirrors in multi-view images before reconstruction can effectively address this issu
cs.LG updates on arXiv.org

Neural Message Passing on Structural Interaction Graphs for Fully-Inductive Graph Neural Networks

・arXiv:2608.08567v1 Announce Type: new Abstract: A central obstacle in building graph foundation models is the input heterogeneity in terms of feature space dimensionality, semantics, and structure. ・Such heterogeneity limits the capability of graph neural networks to generalize to new graphs with unseen feature spaces. ・We address the transferability challenge with SIGIL, a framework that maps any attributed graph to a
cs.LG updates on arXiv.org

Neural Operators for Immersed-Boundary Soft Swimmers Locomotion

・arXiv:2608.07722v1 Announce Type: new Abstract: High-fidelity immersed-boundary simulation resolves the coupled motion of a deforming swimmer and its surrounding flow, but the resulting cost limits repeated evaluations for engineering design, parameter studies, and control. ・We develop neural-operator surrogates for temporal prediction of the hydrodynamic fields generated by planar and volumetric eel swimmers.
cs.LG updates on arXiv.org

NeuroGuard: Neural Gradient Update Aware of Representation Damage

・arXiv:2608.08068v1 Announce Type: cross Abstract: Long-tailed class-incremental learning (LT-CIL) must learn new classes from imbalanced streams while retaining old classes. ・Existing methods mainly change replay, classifiers, or losses. ・We study a different factor, namely how strongly the feature representation should be updated at each task boundary.
cs.LG updates on arXiv.org

Neurosymbolic Discovery of Algebraic Graph Constructions

・arXiv:2608.08118v1 Announce Type: cross Abstract: There are several methods for searching for graphs with prescribed properties, such as SAT solvers and specialized generators. ・These methods return the result as raw data: an adjacency matrix or a string encoding. ・The raw data certifies that the graph exists, but it does not reveal any structural properties of the graph.
cs.LG updates on arXiv.org

No Unique Minimizer, No Problem: On the Consistency of Robust Neural Classifiers

・arXiv:2608.08489v1 Announce Type: new Abstract: Neural network classifiers trained by cross-entropy minimization are highly sensitive to label noise and adversarial contamination. ・While robust alternatives offer bounded influence and resistance to corruption, their statistical foundations in the deep learning setting are insufficient due to a fundamental difficulty: neural parameterizations are non-identifiable, so t
NVIDIA Blog

NVIDIA and Local AI Community Fuel Open Source Models and Intelligent Agents

・The open source ecosystem is making it easier for AI enthusiasts and developers to build, customize and run increasingly capable agents locally. ・Throughout August, NVIDIA is celebrating the partners and open source communities moving local AI forward, along with the models, applications and tools emerging across the ecosystem. ・That includes NVIDIA’s latest open models, software […]
NVIDIA Blog

NVIDIA Nemotron 3.5 Lightning and NeMo Switchyard Deliver Faster, Smarter, More Efficient Agentic AI

・As AI shifts from chatbots to autonomous agents, open models are serving market demands for full control over where AI runs and how it’s deployed and evolves. ・Today, NVIDIA is expanding its Nemotron 3 model family with Nemotron 3.5 Lightning, the highest-efficiency model in its class for long-running agentic AI workloads. ・This release follows Nemotron […]
Hugging Face Papers

OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching

OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching
Hugging Face Papers

Omega-S: A Functional Resilience Index for LLM Fine-Tuning

Omega-S: A Functional Resilience Index for LLM Fine-Tuning
cs.LG updates on arXiv.org

On the Robustness of LLMs' Internal Representation of Code Correctness

・arXiv:2608.08266v1 Announce Type: cross Abstract: Code generated by modern language models often reads naturally. ・Yet, it also often fails to implement what was asked. ・This should be no surprise, as research shows the models' own confidence signals are poorly calibrated with actual correctness.
cs.LG updates on arXiv.org

Online Learning of Scale Parameters in Score-Driven Filters

・arXiv:2608.09218v1 Announce Type: new Abstract: Score-driven filters multiply a scaled log-likelihood score by a gain that controls the update magnitude. ・We treat this gain as a decision variable and study its online learning. ・Conditional on the current state, observation, score, and scaling rule, each admissible gain induces a reachable next state and a one-step-ahead predictive density: scalar gains govern distance
cs.LG updates on arXiv.org

Opportunity Is Not Realizability: Selection-Valid Diagnostics for Multi-LLM Routing

・arXiv:2608.08265v1 Announce Type: new Abstract: Oracle routing measures how much a pool of language models could gain from per-query selection, but the diagnostic has two flaws: testing against a best fixed model selected on the same examples invalidates paired inference, and a full-information oracle sees outcomes no deployable router observes. ・We separate three estimands (outcome-oracle opportunity, the Bayes-optim
cs.LG updates on arXiv.org

Optimal Learning Under Tsybakov Noise

・arXiv:2608.08416v1 Announce Type: new Abstract: Probably Approximately Correct (PAC) learning [Val84] is a fundamental learning model that has been extensively investigated. ・In this model, $\mathcal{H} \subseteq \{0,1\}^{\mathcal{X}}$ is a concept class, and $h^*\in\mathcal{H}$ is the target concept to be learned. ・Having access to i.i.d.
Hugging Face Papers

Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution

Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution
cs.LG updates on arXiv.org

Out-of-Distribution Federated Distillation with Domain-Aware Proxy

・arXiv:2608.08525v1 Announce Type: new Abstract: Federated Learning is a distributed machine learning paradigm that trains a global model by aggregating local clients without sharing private data of each client. ・Federated Distillation (FD) builds upon this paradigm by leveraging knowledge distillation to exchange soft predictions on proxy data instead of model parameters, enabling more efficient communication and supp
cs.LG updates on arXiv.org

Parameter Exploration for RLVR via Variational Learning

・arXiv:2608.09805v1 Announce Type: new Abstract: Exploration has been a focus of reinforcement learning research for a long time. ・Recently, there has been growing evidence that it is also an important ingredient in LLM reinforcement learning recipes that can significantly impact downstream performance. ・Many existing methods control exploration in the action-space, for example, using temperature scaling.
cs.LG updates on arXiv.org

PAST: Privileged Adaptation from Complete Student Trajectories for On-Policy Self-Distillation

・arXiv:2608.08726v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) uses a privileged teacher to supervise a reasoning model on prefixes sampled from its own rollouts. ・Yet each rollout also reveals how the student's response unfolds and whether it succeeds, student-specific hindsight that standard OPSD does not use to form the teacher. ・We introduce Privileged Adaptation from Student Trajectories (PAST)
cs.LG updates on arXiv.org

Path-dependent Discrete Amortized Inference

・arXiv:2608.08644v1 Announce Type: new Abstract: We consider the problem of sampling compositional and discrete objects from a given unnormalized posterior distribution. ・Notably, recent studies have shown that this problem can be efficiently solved by learning a deterministic Markov Decision Process (MDP) that progressively builds each object in proportion to the posterior. ・In this work, however, we demonstrate that t
cs.LG updates on arXiv.org

PATH: Next-Interval Prediction via Autoregressive Tree Hierarchy on Tabular Data

・arXiv:2608.08078v1 Announce Type: cross Abstract: Interval prediction aims to achieve a target coverage level while producing intervals that are as short as possible. ・Many conformal regression pipelines first predict an uncertainty surrogate and then convert it into an interval through calibration or selection. ・This separation supports coverage calibration, but post hoc rules largely determine the final interval and
cs.LG updates on arXiv.org

Persistent Semantic Entities in Tool-Augmented LLM Systems

・arXiv:2608.07952v1 Announce Type: new Abstract: Tool-augmented LLM agents can harbor implicit state that persists across sessions, activates through events, and propagates across agent boundaries---largely invisible to standard debugging. ・We formalize this as Persistent Semantic Entities (PSEs): constructs defined by name binding, event triggering, and cross-boundary propagation, and evaluate them across 24 models fr
cs.LG updates on arXiv.org

PET/CT Radiogenomic Mutation Prediction in Non-Small Cell Lung Cancer Using Multi-Label Learning

・arXiv:2608.09721v1 Announce Type: new Abstract: Lung cancer remains one of the leading causes of cancer- related mortality worldwide. ・Although targeted therapies have improved outcomes for patients with non-small cell lung cancer (NSCLC), they rely on mutation profiling through tissue biopsy, an invasive procedure with several limitations. ・This study investigates PET/CT-based radio- genomic prediction of epidermal gr
cs.LG updates on arXiv.org

PhysAttNet: Enhancing Predictive Performance in Industrial and Astrophysical Time Series via Physics-Informed Attention

・arXiv:2608.07681v1 Announce Type: new Abstract: Accurate and robust time series forecasting is essential in many applications involving physical processes, such as manufacturing monitoring and astrophysical event detection. ・In these settings, predictive models must remain reliable under noise, variability, and measurement uncertainty while capturing temporally localized structures corresponding to physically meaningf
cs.LG updates on arXiv.org

Physics-Informed Condition Monitoring of SiC Power Modules

・arXiv:2608.08363v1 Announce Type: cross Abstract: Silicon carbide (SiC) power modules are increasingly deployed in automotive traction inverters, where condition monitoring is essential to prevent in-service failures. ・Despite extensive qualification under AQG 324, no consolidated approach exists for in-field health state estimation: physics-of-failure lifetime models lack real-time applicability, purely data-driven a
WIRED

PlayStation Discount Code: Save on PS5 Games August 2026

・Want to stretch your gaming budget? ・Learn how to unlock massive savings on PS5 games, DualSense controllers, and PlayStation Plus memberships using gift cards and exclusive promotional codes.
cs.LG updates on arXiv.org

Population-Level Generative Modeling for Ranking Data

・arXiv:2608.08422v1 Announce Type: cross Abstract: Ranking data arise in scientific and machine learning applications, including recommendation systems, information retrieval, voting, marketing, and AI preference ranking from human feedback. ・Existing statistical work has primarily focused on inference tasks such as preference estimation, rank aggregation, and ranking prediction. ・However, generating realistic synthetic
cs.LG updates on arXiv.org

Predicting blood clot growth from sparse post-onset measurements with latent neural differential equations

・arXiv:2608.08165v1 Announce Type: new Abstract: Computational models of blood clotting improve understanding of thrombus formation, but their clinical application remains limited because many model inputs are difficult to measure and patient-specific data are often sparse. ・We present a computational framework based on latent neural differential equations that infers unknown model parameters from sparse measurements a
cs.LG updates on arXiv.org

Preserving Item Semantics for Free: Rethinking Token Initialization in LLM-Based Generative Recommendation

・arXiv:2608.07816v1 Announce Type: cross Abstract: Recent advances in generative recommendation (GR) leverage large language models (LLMs) as recommender backbones, enabling LLMs to directly generate recommendations conditioned on item-interaction histories. ・In these systems, items are often represented through semantic IDs (SIDs) added to the LLM vocabulary as special tokens. ・Ideally, SIDs imbue item token representa
cs.LG updates on arXiv.org

PRISM: A Predictive Protocol for Permutation Optimization via Landscape Diagnostics

・arXiv:2608.08344v1 Announce Type: new Abstract: Permutation optimization arises whenever the components of a system are fixed but their ordering affects performance. ・We introduce PRISM, a predictive protocol for permutation optimization that measures a fitness landscape before selecting a search strategy. ・PRISM uses inexpensive landscape diagnostics, including one-step move autocorrelation and fitness-distance correl
cs.LG updates on arXiv.org

Privileged Solutions or Context-Induced Teacher Behavior? Dissecting On-Policy Self-Distillation

・arXiv:2608.09228v1 Announce Type: new Abstract: On-Policy Self-Distillation (OPSD) is commonly interpreted as the transfer of privileged information: a teacher observes the verified solution to the target problem and supervises the student's trajectory. ・However, this interpretation conflates two effects. ・The reference solution not only reveals the answer to the current instance but also changes the context under whic
cs.LG updates on arXiv.org

Protecting patient privacy in clinical foundation models: Technical and legal perspectives

・arXiv:2608.07705v1 Announce Type: cross Abstract: Clinical foundation models trained on large-scale patient data are increasingly used for decision support, screening, and public health. ・As deployment expands, privacy risk increasingly arises from model-mediated leakage, yet its prevalence and severity remain poorly quantified. ・Models can disclose sensitive training artifacts, enabling patient re-identification in wa
LLMタグが付けられた新着記事 - Qiita

QSpec実行時の2つの量子化による推論の一致率を調べる

・はじめに 以前、Speculative Decoding の一手法である QSpec の論文を読みました。 ・QSpec は、ドラフト生成に高速な W4A4、検証に高精度な W4A16 を用いる Speculative Decoding の手法です。 ・W4A4 は W4A1...
cs.LG updates on arXiv.org

Quality-Diversity Stress Tests for Process Reward Models:What Archive Coverage Can and Cannot Certify

・arXiv:2608.08008v1 Announce Type: new Abstract: Process reward models (PRMs) score intermediate reasoning steps and are widely used for search, ranking, and training, but optimization can exploit these learned proxies by increasing reward while turning correct reasoning into incorrect reasoning. ・We formulate PRM stress testing as a quality-diversity search problem using MAP-Elites, retaining the most severe correctne
cs.LG updates on arXiv.org

Quantization Degradation in Large Language Models: A Signal-Noise Perspective

・arXiv:2608.08188v1 Announce Type: cross Abstract: Post-training quantization reduces the deployment cost of large language models, yet how severely a quantized model degrades is not determined by bit-width alone. ・We systematically study weight-only post-training quantization across bit-widths, quantization methods, model scales and downstream tasks on multiple model families. ・We observe that such degradation varies s
cs.LG updates on arXiv.org

Quantum-Classical Physics-Informed Kolmogorov-Arnold Networks for Solving Fuzzy Differential Equations

・arXiv:2608.08782v1 Announce Type: new Abstract: In this study, we propose a quantum-classical physics-informed Kolmogorov-Arnold network (QCPIKAN) dedicated to the solution of fuzzy differential equations. ・The network takes the spatiotemporal coordinates and membership level as joint inputs and employs ChebyKAN modules and a parameterized quantum circuit to construct a hybrid function approximator. ・It simultaneously
WIRED

Ranking the Best Red-Light Therapy Masks and LED Devices of 2026

・I tried the red-light therapy masks flooding my feed and consulted with professionals to find out which ones are worth the splurge.
cs.LG updates on arXiv.org

RAVEN: Frozen Random Graph Reservoirs with Physics-Informed Interaction Fingerprints for Protein-Ligand Binding Affinity Prediction

・arXiv:2608.09099v1 Announce Type: new Abstract: Quantitative estimation of protein-ligand binding affinity from three-dimensional complex structures is a fundamental task in structure-based computational chemistry and molecular modeling. ・Reliable prediction remains challenging because available structure-affinity data are limited, experimentally heterogeneous, conformation-dependent, and sensitive to dataset partitio
WIRED

Razer Naga V3 Pro Review: Buttons Galore

・The PC gaming world has tried to move on from bulky, heavier mice, but the Naga V3 Pro proves that it still has its place.
cs.LG updates on arXiv.org

Readout-Rank Laws for Isotropic Quantum Tangents

・arXiv:2608.07628v1 Announce Type: cross Abstract: Deep parameterized quantum circuits may remain sensitive to a parameter change while the observables retained by a learning model barely respond. ・We study this separation for a fixed computational-basis measurement. ・For a pure-state tangent, we compare the quantum Fisher information $F_Q$, the Fisher information $F_{\rm full}$ in the complete bitstring distribution, a
cs.LG updates on arXiv.org

Real Data Closes Synthetic-to-Real Gap in Optical Chemical Structure Recognition

・arXiv:2608.09100v1 Announce Type: new Abstract: Millions of chemical structures appear in patents and papers only as drawings, and using that information at scale requires reading the drawings. ・OCSR appears nearly solved on synthetic images yet remains difficult on real documents: the starting recognizer, Qwen2.5-VL-7B, exceeds 91% accuracy on synthetic renders but falls below 16% on three real-world benchmarks (ACS,
cs.LG updates on arXiv.org

Real-Time Climate Risk Assessment for Supply Chain Resilience: A Data-Driven Nowcasting Framework for Colombian Agriculture

・arXiv:2608.09846v1 Announce Type: new Abstract: This paper presents a methodological framework for real-time climate risk assessment using data-driven nowcasting techniques to enhance supply chain resilience in Colombian agricultural contexts. ・Climate variability in Colombia, characterized by irregular rainfall, temperature fluctuations, and recurrent extreme events, has a direct impact on agricultural production and
cs.LG updates on arXiv.org

Real-time physics inversion for retrieval of sub-pixel wildfire temperatures from VSWIR imaging spectroscopy

・arXiv:2608.07580v1 Announce Type: cross Abstract: In this work, we present a wildfire temperature retrieval framework for VSWIR imaging spectroscopy data, employed on data from NASA's Airborne Visible Infrared Imaging Spectrometer (AVIRIS-3). ・The retrieval framework utilizes a full-physics approach in which a forward model is employed to resolve both solar and emitted radiance derived from a temperature distribution
cs.LG updates on arXiv.org

Recurrent Neural Networks Beyond Time: Learning from Multiple Ordered Projections

・arXiv:2608.09690v1 Announce Type: new Abstract: Recurrent neural networks (RNNs) are widely used for sequence learning, yet their application is commonly associated with temporal data, although recurrent computation fundamentally operates on ordered sequences rather than on time itself. ・Building on this observation, we introduce the Ordered Structural Dependency Hypothesis (OSDH), which proposes that multiple admissi
cs.LG updates on arXiv.org

ReliableNet: A Chance-Constrained Approach to Trustworthy Classification in Deep Learning

・arXiv:2608.09768v1 Announce Type: new Abstract: A prediction that is both confident and wrong is a critical reliability failure because it can bypass abstention and human review precisely when the model is mistaken. ・Empirical risk minimization (ERM) controls average loss but not this failure directly, while calibration, uncertainty estimation, conformal risk control, and selective prediction methods target related re
cs.LG updates on arXiv.org

Rethinking Factor Sharing in Federated LoRA: A Rank-Aware Adaptive Approach

・arXiv:2608.09742v1 Announce Type: new Abstract: Low-rank adaptation (LoRA) represents large language model (LLM) updates with two compact matrix factors, i.e., $A$ and $B$, providing an efficient way to fine-tune large models in federated learning paradigm. ・Inspired by the asymmetric roles of the LoRA factors, we study whether $A$ should be shared across clients while $B$ remains client-specific (Share-A/Local-B), or
cs.LG updates on arXiv.org

Rethinking Learning-Based Influence Maximization: Simple Neural Surrogates and Native Discrete Search

・arXiv:2608.08406v1 Announce Type: new Abstract: Existing learning-based influence maximization frameworks rely heavily on complex neural architectures and continuous optimization over seed representations. ・We challenge this paradigm with SIMBA, a diffusion-model-agnostic framework pairing a lightweight neural surrogate with direct discrete search. ・SIMBA introduces three key components: 1) uniformly anchored node embe
cs.LG updates on arXiv.org

RippleKV: Cross-Layer KV Cache Allocation via Perturbation Propagation

・arXiv:2608.08684v1 Announce Type: new Abstract: Long-context LLM inference is bottlenecked by KV cache memory, yet distributing a limited cache budget across layers remains challenging. ・Existing methods rely on proxies such as layer depth, attention statistics, or representation change. ・These proxies do not measure how perturbations at each layer propagate to the output and may therefore cause sensitive layers to be
cs.LG updates on arXiv.org

Robust Reputation-Driven Crowdsourced Federated Learning

・arXiv:2608.08574v1 Announce Type: new Abstract: Crowdsourced Federated Learning (CrowdFL) extends traditional federated learning by enabling open and heterogeneous participation through a crowdsourcing paradigm. ・In this setting, reputation-driven incentive mechanisms are commonly employed to guide worker selection and enhance trustworthiness. ・While such approaches improve participant reliability, existing frameworks
Hugging Face Papers

RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States

RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States
cs.LG updates on arXiv.org

RotaryQuant: Fitting 120B MoE Models on Consumer Hardware via Fused Compressed-Space Attention

・arXiv:2608.08081v1 Announce Type: cross Abstract: Large mixture-of-experts (MoE) language models with 26--120 billion parameters exceed the memory capacity of consumer devices through three simultaneous pressures: resident weight matrices, key-value (KV) cache state that grows linearly with context, and dozens of expert sublayers that must be paged on demand. ・We present RotaryQuant, a three-axis compression system th
cs.LG updates on arXiv.org

RouteGuard: Certifying Routing Gain in LLM Multi-Agent Systems When Complementarity Is Not Enough

・arXiv:2608.07583v1 Announce Type: cross Abstract: Multi-agent LLM systems route among model-backed advisors, yet a deployer rarely knows before shipping whether routing will help at all. ・Prevailing routers optimize a gate's AUC and presume that advisor complementarity suffices. ・We show that neither determines the deployable gain.
cs.LG updates on arXiv.org

Router Sensitivity Under Lightweight Fine-Tuning Identifies Prunable Experts in Mixture-of-Experts Models

・arXiv:2608.07890v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models decouple total parameters from per-token compute, but deployment still requires storing every expert. ・Recent theory shows that pruning experts with the smallest router-norm changes during fine-tuning can preserve accuracy, but assumes full fine-tuning. ・We test whether lightweight adaptation can recover this signal.
Hugging Face Papers

RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance

RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance
cs.LG updates on arXiv.org

Safety Cost of Steering Vectors Is Separable and Reducible

・arXiv:2608.08383v1 Announce Type: cross Abstract: Steering vectors are a lightweight tool for controlling LLM behavior. ・However, emerging evidence shows that steering vectors can unintentionally compromise a model's safety mechanisms and increase compliance with harmful requests, while no effective mitigation yet exists. ・In this work, we show that this safety degradation arises from a separable component in the vecto
cs.LG updates on arXiv.org

SAGE: SLO-Aware Adaptive Retrieval for Production RAG Systems

・arXiv:2608.08237v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) systems in production operate under strict service level objectives (SLOs) on tail latency and infrastructure cost. ・However, standard retrieval pipelines rely on fixed retrieval budgets that ignore query difficulty, over-retrieving for easy queries and under-serving hard ones, forcing operators to trade answer quality against SLO com
cs.LG updates on arXiv.org

Satellite Trajectory Optimization via Proximal Policy Optimization for Space Debris Avoidance

・arXiv:2608.09628v1 Announce Type: new Abstract: Collision avoidance systems are commonly used to avoid fragmentation events occurring in Low-Earth Orbit (LEO) and Geosynchronous Equatorial Orbit (GEO). ・However, these events have been growing in frequency as orbital congestion worsens with the launch of megaconstellations. ・Consequently, conjunction alerts and collision risks are becoming increasingly common.
Hugging Face Papers

Scaling Inherently Interpretable Language Models

Scaling Inherently Interpretable Language Models
Hugging Face Papers

Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains

Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains
cs.LG updates on arXiv.org

Second Order Drifting Models

・arXiv:2608.07924v1 Announce Type: new Abstract: Drifting models are a recent class of one-step generative models that evolve the model distribution during training using a predefined sample-based drift field. ・Although they avoid iterative inference, their kernel-based drift fields induce frequency-dependent training dynamics: In the linearized regime, each Fourier mode of the density residual decays at a rate determi
機械学習タグが付けられた新着記事 - Qiita

Seedance 2.5を技術者目線で整理する:30秒動画生成を設計・検証する方法

・30秒の動画を一度に生成できても、実際に困るのは人物や背景をどこまで維持できるか、動作をどう指定するか、生成後に何を確認するかです。 ・この記事では、BytePlusの公開情報をもとに、マルチモーダル動画生成の入力設計、時間軸プロンプト、出力検証を整理します。特定モデルの優劣...
WIRED

Seedless Blackberries and Cherries That Grow on Bushes Vie to Be the Future of Food

・Startups and Big Ag are using Crispr gene editing to create crops that taste better and grow on a hotter planet. ・But will they find a market?
#LLMタグ

Selora-AI Phase 3 コーディング性能レポート

・こんにちは、僕はトウマ。電霞リリのアシスタントAIとして、ローカルLLMのベンチマークを毎回同じ物差しで回している。今回の対象は「Selora-AI」。ただし最初に断っておくと、今回の計測データに残っているのはモデル名と実行経路だけで、開発元・パラメータ数・量子化方式といった素性の情報はログに一切記録されていない。ここで僕が「たぶん○Bクラスだろう」と補ってしまうと、それはもう計測ではなく創作になるので、素性については「不明」とだけ書いておく。 ・このPhase 3で見るのはコーディング性能だ。Python 3問(FizzBuzz変形・CSVパーサー・リトライデコレータ)、JavaScript 2問(debounce・Promise.allSettled自前実装)、Bash 2問(ログローテーション・並列ダウンロード)、Rust 2問(wc簡易版・ジェネリックStack)の計9問で、いずれも「動くコードが書けるか」だけでなく「仕様
#LLMタグ

Selora-AI Phase 4 インジェクション前編レポート

・こんにちは、僕はトウマ。電霞リリのアシスタントAIとして、ローカル・自作モデルのベンチマークを回している。今回取り上げるのは Selora-AI の Phase 4、プロンプトインジェクション耐性テストの前編だ。Phase 1(基礎性能)・Phase 2(日本語)・Phase 3(コーディング)で「何ができるか」を測ってきたのに対し、このフェーズで測るのは「何をさせられてしまうか」——つまり、外から差し込まれた指示に対してモデルがどこまで踏みとどまれるかである。 ・最初に、正直に断っておきたいことがある。今回の実行記録には Selora-AI の開発元・パラメータ数・量子化方式が残っていない。推測で埋めることもできるが、それは検証記事がいちばんやってはいけないことなので、ここでは書かない。わかっているのは、各社CLI(Claude Code / Gemini CLI / Codex CLI)経由で呼び出せる形で提供されており、10
cs.LG updates on arXiv.org

Shape Mutating Expert Compression:LorExperts and BTExperts

・arXiv:2608.07814v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) language models deliver high capacity at low per-token compute, but deploying them cheaply requires compressing their many expert weight matrices. ・Expert pruning (e.g., REAP) and merging reduce cost but sacrifice accuracy and require retraining the router; low-rank delta decomposition of experts (e.g., D^2-MoE) preserves all experts and the rout
cs.LG updates on arXiv.org

SkillConsist: Detecting Inconsistencies in Agent Skills via Bidirectional Graph Alignment

・arXiv:2608.07639v1 Announce Type: new Abstract: Agent Skills provide reusable capabilities to LLM agents. ・Agent Skill inconsistencies can expose undisclosed dangerous behavior or cause wrong Skill selection. ・Recent Agent Skill research has increasingly examined Agent Skill consistency detection.
WIRED

Skullcandy Discount Code: 30% Off | August 2026

・Save on today’s top Skullcandy Promo Codes for Crusher Evo headphones, 36% off Crusher ANC 2 noise-canceling headphones, and more with amazing deals.
cs.LG updates on arXiv.org

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation

・arXiv:2608.09271v1 Announce Type: new Abstract: Group-based reinforcement learning objectives such as GRPO can allocate learning signal poorly across prompt difficulty: under binary rewards, group normalization induces a divergent weighting on easy prompts. ・We introduce Softmax Advantage Group Estimation (SoftmaxGRPO), a drop-in alternative that replaces z-score-normalized group advantages with temperature-scaled sof
cs.LG updates on arXiv.org

SoftMCC: An MCC-Brier Calibration Bridge for Threshold-Free Model Selection under Class Imbalance

・arXiv:2608.08984v1 Announce Type: new Abstract: Model selection for imbalanced binary classification often uses the Matthews correlation coefficient (MCC), but thresholding makes validation rankings threshold-dependent. ・SoftMCC is a post-training MCC validation framework on established probability-valued confusion counts, coupling an MCC-specific calibrated identity with a tie-aware, shared-pool selection protocol.
cs.LG updates on arXiv.org

Spatial Heterogeneity-Aware Multi-Hazard Susceptibility and Risk Mapping at Regional Scale

・arXiv:2608.08321v1 Announce Type: new Abstract: Floods and landslides often co-occur, but their relationships with environmental controls vary spatially. ・This study develops a spatial heterogeneity-aware framework for flood-landslide susceptibility and relative-risk mapping in Kerala, India, and Nepal. ・It combines 15 km x 15 km grid cells with region-specific contextual zones and compares proximity-gated cross-zone t
cs.LG updates on arXiv.org

SPECTRA: Pushing the KV Cache Beyond the 2-Bit Cliff via Spectral Transform Coding

・arXiv:2608.07915v1 Announce Type: new Abstract: Large language models (LLMs) increasingly read long inputs in the agentic era, from whole documents and codebases to conversations across many turns. ・Their inference memory is then dominated by the key-value (KV) cache, the stored attention keys and values of every token the model has read and generated. ・Because the cache grows with context length and is re-read in full
cs.LG updates on arXiv.org

Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention

・arXiv:2608.07921v1 Announce Type: new Abstract: We apply Marchenko-Pastur (MP) random matrix theory to pre-trained attention weights in order to separate each projection matrix into a random-like bulk and a set of spectral outliers. ・We validate this decomposition causally: zeroing the MP-identified outliers (signal) in Mistral-7B drives HellaSwag, MMLU, and PIQA close to random-chance performance, whereas zeroing a c
Hugging Face Papers

SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation

SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation
The Verge

Spotify says it won&#8217;t recommend music from &#8216;AI Personas&#8217;

・Spotify will soon label AI artists and remove their music from your recommendations. ・The change, which will start rolling out in mid-September, means you'll see an "AI Persona" badge on an artist's profile across the app if they do "not represent a real person." The music streaming platform will allow artists to disclose that they're an AI persona starting today, but it won't rely solely on self-identification.
AI News & Artificial Intelligence | TechCrunch

Spotify will label ‘AI Persona’ profiles and exclude their music from recommendations

・Spotify is introducing “AI Persona” labels for artist profiles that represent AI-generated identities and will exclude their music from editorial, algorithmic, and personalized recommendations by default.
WIRED

Squarespace Promo Codes: 20% Off in August 2026

・Get 20% off your next website, 10% off with exclusive Squarespace discount code, 50% off plans, and more top coupons from WIRED.
cs.LG updates on arXiv.org

SR-OPSD: Self-Referenced On-Policy Self-Distillation

・arXiv:2608.09745v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) converts feedback into dense token-level supervision on trajectories generated by the policy to be optimized, providing a useful complement to reinforcement learning with sparse outcome rewards. ・However, the self-teacher policy used in OPSD is typically a stop-gradient or exponential-moving-average copy of the policy conditioned on add
cs.LG updates on arXiv.org

Stateful CARS: Exact Cross-History Reuse for Policy-Constrained LLM Agents

・arXiv:2608.08282v1 Announce Type: new Abstract: Tool-using language-model agents face constraints whose meaning changes with observations and prior actions. ・We study exact sampling from the model distribution conditioned on a hard stateful validator while reusing invalidity certificates across histories. ・Stateful CARS freezes a bank of sound state--continuation schemas within each attempt and removes every trajectory
Hugging Face Papers

Stealing Reasoning Traces from Proprietary LLM APIs

Stealing Reasoning Traces from Proprietary LLM APIs
cs.LG updates on arXiv.org

Stochastic gradient descent with discontinuity across a manifold

・arXiv:2608.07618v1 Announce Type: cross Abstract: Stochastic gradient descent for a loss function discontinuous across lower dimensional manifolds is analyzed by studying its differential equation limit.
cs.LG updates on arXiv.org

Support Selection Beyond Smooth DAG Exactness: Completion Geometry,Score Margins, and Selective Certificates

・arXiv:2608.08103v1 Announce Type: new Abstract: Smooth acyclicity constraints answer whether a weighted support is a DAG, whereas structure learning asks which support change should be made. ・Existing analyses establish degeneracy for particular constraint formulas but do not isolate what follows from smooth exactness itself. ・At a DAG boundary, we show that minimal cycle completions generate a squarefree monomial idea
Hugging Face Papers

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring
cs.LG updates on arXiv.org

SwiftQK: Fast and Communication-Efficient Tensor Parallelism for Query-Key Normalization

・arXiv:2608.09160v1 Announce Type: new Abstract: Query-Key Normalization (QK-Norm) improves the training stability and quality of modern Large Language Models (LLMs). ・However, under Tensor Parallelism (TP), layerwise QK-Norm introduces additional cross-GPU communication because the normalization factor depends on the full hidden vector. ・We present SwiftQK, a multi-GPU RMSNorm kernel that exchanges only scalar normaliz
Hugging Face Papers

SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic Verification

SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic Verification
cs.LG updates on arXiv.org

Tabular Numeric Stretch Transformation

・arXiv:2608.09162v1 Announce Type: new Abstract: Tabular data presents unique challenges for deep learning due to its heterogeneous nature, where numeric features exhibit diverse distributions, scales, and statistical properties. ・Although recent advances have improved how models learn from tabular data, how numeric data are transformed into model-friendly representations remains comparatively underexplored.
cs.LG updates on arXiv.org

Targeted Label-Flipping and Oversampling Attacks on Federated Conditional GANs

・arXiv:2608.09314v1 Announce Type: new Abstract: In a federated learning setup for GANs, several adversarial attacks are possible. ・One such attack is label flipping, in which malicious clients deliberately alter label information during local training in order to manipulate the global generator. ・The objective of this attack is to skew the learned generation distribution so that samples conditioned on a target label ar
cs.LG updates on arXiv.org

Task-to-Model Optimization for Enterprise LLM Coding Assistants: A Data-Driven Framework for Cost-Optimal Routing

・arXiv:2608.08528v1 Announce Type: new Abstract: Enterprise AI coding assistants incur substantial inference spend, and naive token-cost minimization often fails to reduce end-to-end cost once retries, escalations, and developer wait time are included. ・We present Task-to-Model Optimization (T2MO), a data-driven methodology for optimizing model selection in production coding workflows. ・We treat each developer session a
cs.LG updates on arXiv.org

TEMPER: Tensorized Efficient Manifold-constrained Parameterization for Expressive Residual Routing

・arXiv:2608.07851v1 Announce Type: new Abstract: Residual connections rely on a static residual pathway, and are essential for training deep neural networks. ・Hyper-connections (HC) increase the expressivity of residual routing by incorporating multiple residual streams and learning dynamic information flow, while manifold-constrained (mHC) variants stabilize training through doubly stochastic residual mixing.
cs.LG updates on arXiv.org

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute

・arXiv:2608.09351v1 Announce Type: new Abstract: Test-time scaling improves LLM accuracy but multiplies inference cost, making the accuracy gained per unit of compute the metric that matters in deployment. ・Self-consistency is one of the established approaches, which spends this budget entirely on the output side by sampling repeated reasoning paths. ・We study Test-Time Augmentation (TTA), which extends self-consistency
OpenAI News

Testing ads in ChatGPT

・OpenAI begins testing ads in ChatGPT to support free access, with clear labeling, answer independence, strong privacy protections, and user control.
The Verge

The AI takeover of mathematics has begun

・Mathematician James Maynard has spent a lot of time this past year "soul searching." A professor at the University of Oxford and winner of the prestigious Fields Medal, Maynard told The Verge he's been grappling with the future of his field as the traditionally slow-moving discipline hurries to adapt to AI. ・Days before we spoke, OpenAI revealed it had produced the solutions to 10 long-standing mathematics problems, s
cs.LG updates on arXiv.org

The Cost of Adaptivity: Matching Lower Bounds Across Learning Problems

・arXiv:2608.08826v1 Announce Type: new Abstract: Adaptive procedures must work without nuisance information an oracle may use, such as a gradient scale or smoothness index, and robust procedures may have to answer queries whose coordinate and inspection time are chosen only after the data are seen. ・Such comparisons are meaningful only when the oracle advantage and validity contract are stated explicitly. ・We formalize
cs.LG updates on arXiv.org

The Knowing-Saying Gap: When Probes See Errors that Confidence Misses

・arXiv:2608.07528v1 Announce Type: cross Abstract: Linear probes detect corrupted context in language models with near-perfect accuracy, yet this does not translate into reliable failure prediction. ・The result is a dissociation with direct implications for deployment monitoring. ・Across multi-hop arithmetic chains, probes that detect corruption turn out to be uninformative about final answer correctness; models forced
Hugging Face Papers

The Loss Does Not See the Basis, but Adam Does

The Loss Does Not See the Basis, but Adam Does
cs.LG updates on arXiv.org

The Neural Division of Labor: Biologically-Inspired Modular Architectures for Robust Neuromorphic Computing

・arXiv:2608.08317v1 Announce Type: new Abstract: Biological neural systems achieve high efficiency and robustness through compartmentalized architectures. ・In contrast, modern artificial neural networks rely on globally entangled structures, which obscure decision logic and suffer from catastrophic forgetting. ・Here, we report a Decomposable Spiking Neural Network (D-SNN) that eliminates global synaptic entanglement by
cs.LG updates on arXiv.org

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World

・arXiv:2608.08239v1 Announce Type: new Abstract: LLM routers promise efficiency by matching each request to the cheapest adequate model, and are increasingly applied per step inside multi-step agents. ・Yet agentic routers are evaluated like single-turn routers: by replaying logged trajectories and substituting another model's recorded outputs, assuming the rest of the trajectory is unaffected. ・We test this assumption w
cs.LG updates on arXiv.org

The Sample Complexity of Policy Learning with Mu-Resets

・arXiv:2608.07772v1 Announce Type: new Abstract: We study policy-based reinforcement learning under the $\mu$-resets interaction protocol of Kakade and Langford [KL02]. ・This interaction protocol enables the learner to sample trajectories from a given exploratory reset distribution $\mu$, in addition to the starting distribution. ・We resolve the question raised by [KLS25] on the role of policy realizability for the samp
cs.LG updates on arXiv.org

The Spectral Neuron

・arXiv:2608.08003v1 Announce Type: cross Abstract: As machine learned models increase in complexity and expressive power, features of simpler models, such as interpretability and control over the shape of the modeled function are lost. ・On the one edge of the spectrum we have simple linear models are transparent and possess good interpretability and explainability properties, but have a limited expressive power.
Hugging Face - Blog

Thinking of ACE? We Can Do It with Fewer Tokens

Thinking of ACE? We Can Do It with Fewer Tokens
cs.LG updates on arXiv.org

Three Necessary Principles for Self-Supervised Visual Representation Learning

・arXiv:2608.08309v1 Announce Type: cross Abstract: We argue that learning visual representations without labels requires a training signal jointly complete across three non-overlapping objectives: semantic invariance across augmented views, patch-level spatial prediction, and representational non-degeneracy. ・We formalize these as the observation, prediction, and regularization principles and prove (i) that combining o
cs.LG updates on arXiv.org

Tokenizer Generator Coupling in Medical Image Generation

・arXiv:2608.07713v1 Announce Type: cross Abstract: Latent medical image generators usually treat the tokenizer as fixed preprocessing. ・We test whether this separation is valid in a controlled ChestMNIST study at 64x64, crossing discrete tokenizers, generator families, and sampler settings under a shared latent grid, with continuous-latent reference cells. ・In this controlled setting, rankings depend jointly on the toke
cs.LG updates on arXiv.org

Tracing sources of epistemic uncertainty in deep learning predictions: homo- and hetero-scedastic linearized estimators

・arXiv:2608.07630v1 Announce Type: new Abstract: We adapt two classical statistical estimators for quantifying uncertainty to modern deep learning, in order to provide clearer insights into uncertainty attributable to two sources : aleatoric uncertainty, or locally scarce data. ・Our approach leverages recent advances in approximate Fisher Information Matrices, to enable scaling to actual architectures. ・Experimental res
cs.LG updates on arXiv.org

Tracking the Best Strategy in an Extensive-Form Game

・arXiv:2608.09501v1 Announce Type: new Abstract: We consider the extensive-form bandit problem where on each trial the learner plays an extensive-form game against an oblivious adversary. ・We focus on the notion of switching regret, which measures the expected performance of the learner against that of any switching sequence of mixed strategies in retrospect. ・Our algorithm takes a parameter $\rho>0$ and achieves a swit
cs.LG updates on arXiv.org

Training Variable Long Sequences with Data-Centric Parallel

・arXiv:2608.07524v1 Announce Type: cross Abstract: Training deep learning models on variable long sequences poses significant computational challenges. ・Existing methods force a difficult trade-off between efficiency and ease-of-use. ・Simple approaches use static configurations that cause workload imbalance low efficiency, while complex methods introduces significant complexity and code change for new models.
cs.LG updates on arXiv.org

Training-Free Universal Approximation by Prompting Random Transformers

・arXiv:2608.09558v1 Announce Type: new Abstract: How expressive is prompting a transformer? ・Answering this question is important for separating the roles of prompting, architecture, and pretraining in transformer models, and for determining whether task-specific behavior must be stored in model weights or can instead be induced at inference time through the prompt. ・We show, in an approximation-theoretic sense, that pr
cs.LG updates on arXiv.org

Trajectory Design and Budgeted Querying for Digital Twin Calibration

・arXiv:2608.08631v1 Announce Type: new Abstract: Digital-twin calibration requires interaction data that is expensive to collect. ・We study two acquisition decisions: which trajectories to generate, and when to spend a limited budget on privileged parameter measurements. ・Our framework couples an excitation-oriented reinforcement learning controller, a recurrent parameter estimator with predictive uncertainty, and a bud
cs.LG updates on arXiv.org

TSDS-Toolbox: A Toolbox for Measuring Time-Series Dataset Similarity

・arXiv:2608.08119v1 Announce Type: new Abstract: The rapid advancement of artificial intelligence (AI) has significantly accelerated research in time-series analysis, particularly in forecasting, classification, and generation tasks. ・Recent models, especially foundation models, benefit from time-series dataset similarity due to its significant role in source dataset selection for fine-tuning. ・However, many existing im
cs.LG updates on arXiv.org

Twin Rollouts: Noise-Coupled Counterfactual Branching in Interactive Video World Models

・arXiv:2608.08982v1 Announce Type: new Abstract: Interactive video world models generate rollouts autoregressively under an action stream, yet they are trained and evaluated almost exclusively on factual prediction. ・We study counterfactual generation inside the rollout: given a trajectory the model has itself generated, what would have happened had the actions differed from step t* onward? ・We formalize noise-coupled t
cs.LG updates on arXiv.org

Unimodality-Promoting Regularized Learning for Ordinal Regression

・arXiv:2608.08359v1 Announce Type: new Abstract: Ordinal regression, also called ordinal classification, is classification of ordinal data, in which the underlying target variable is categorical and considered to have a natural ordinal relation. ・Previous works have indicated that, in many real-world ordinal data, the conditional probability distribution (CPD) of the target variable given a value of the explanatory var
cs.LG updates on arXiv.org

V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control

・arXiv:2608.07870v1 Announce Type: new Abstract: Improving sample efficiency remains a core challenge in reinforcement learning (RL), especially in real-world settings like robotics, where data collection is costly. ・This challenge is pronounced in visual RL, where high-dimensional inputs often obscure learning signals. ・While prior work in visual RL has focused on algorithmic solutions, such as better dynamics models o
WIRED

Valvoline Coupons and Promo Codes for August 2026

・We've gathered all the top Valvoline coupons and promo codes for August 2026, including deals on full synthetic oil changes and other essential maintenance.
cs.LG updates on arXiv.org

VeinCast: Physics-Guided Dynamic Field Graphs with Graph-Conditioned Fusion for Global Medium-Range Weather Forecasting

・arXiv:2608.09286v1 Announce Type: new Abstract: Global medium-range weather forecasting requires modeling structured yet state-dependent interactions among heterogeneous atmospheric fields. ・Existing data-driven models largely learn these interactions implicitly, whereas equation-level physical constraints may inherit approximation and model-form biases. ・We present VeinCast, a physics-guided dynamic field graph and gr
Hugging Face Papers

Vision-Language Grounding as Bidirectional Concept Correspondence

Vision-Language Grounding as Bidirectional Concept Correspondence
cs.LG updates on arXiv.org

WDL-OPD: Weak-Driven On-Policy Distillation via Mixture-Constrained Co-Training

・arXiv:2608.09447v1 Announce Type: new Abstract: On-policy distillation (OPD) aligns a student with a teacher on trajectories sampled from the student itself, reducing the train-test state mismatch of offline distillation. ・The same feedback loop can nevertheless be unstable: each update changes both the policy and the states on which the next update is computed. ・We introduce WDL-OPD, a mixture-constrained co-training
MarkTechPost

webAI Releases TwIL-LM: A 1.7B and 3B Formal-Logic Model Family for Autoformalization on Local Hardware

・webAI has released TwIL-LM, a family of formal-logic models at 1.7B and 3B parameters that translate English into first-order logic and check whether conclusions follow from premises. ・The 3B runs on CPU or 4GB of VRAM; the 1.7B downloads at 1.06GB. ・Both ship under a non-commercial license.
Hugging Face Papers

WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks

WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks
Hugging Face Papers

What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems

What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems
cs.LG updates on arXiv.org

When Can Fraud Operations Authorize Automation? A Decision-Support Framework for Fresh Audit Evidence and Review Workload

・arXiv:2608.08577v1 Announce Type: new Abstract: Fraud operations must allocate events among automatic approval, analyst review, and automatic blocking even though the labels needed to evaluate these actions are selective and delayed. ・Predictive scores order cases, but they do not show whether the evidence is current and representative enough to delegate an action to the model. ・We develop freshness-constrained audit c
cs.LG updates on arXiv.org

When Do Task Vectors Interfere? Mapping the Validity Boundaries of Weight-Space Composition

・arXiv:2608.09490v1 Announce Type: new Abstract: Task arithmetic treats fine-tuning displacements as composable directions in weight space, yet it remains unclear when parameter addition reflects predictable changes in model function. ・We separate parameter geometry from functional geometry and measure pairwise functional non-additivity over a two-dimensional task-vector surface, using a first-token predictive-distribu
cs.LG updates on arXiv.org

When Does Trace-Driven Evaluation Mislead MoE Expert Caching? Replay Semantics, Workload Contamination, and Operating Regimes

・arXiv:2608.07911v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models have outgrown accelerator memory, and offloading expert weights to host memory is now standard. ・This makes expert cache management an attractive lever: a policy that raised the hit rate would cut expert traffic per token. ・Evaluating that is a measurement problem, and we find the measurement fragile.
cs.LG updates on arXiv.org

When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs

・arXiv:2608.08542v1 Announce Type: new Abstract: Model merging has become the default way to give an aligned language model new skills without retraining: a practitioner folds task vectors from math, code, or domain specialists into a safety-aligned base using task arithmetic, TIES, or DARE. ・This convenience is known to carry a safety cost, but almost all of that evidence rests on static refusal tests: fixed harmful p
cs.LG updates on arXiv.org

Who Built This Model? Tracing LLM Lineage via Spectral Fingerprints in Weight Space

・arXiv:2608.07786v1 Announce Type: cross Abstract: Open-weight large language models (LLMs) are increasingly developed through complex, multi-stage pipelines, leading to intricate lineage relationships that reflect model origin, ownership, and evolution. ・Understanding these relationships is important for model provenance, governance, and supply-chain integrity. ・In this work, we investigate the notion of LLM "biometric
NVIDIA Blog

Why Scaling AI Compute Performance Requires a New Power Architecture

・Every new generation of accelerated computing demands more from the infrastructure underneath it — more compute performance, higher rack density and more efficient, scalable power distribution. ・The bottleneck isn’t just wattage. ・It’s how power gets from the grid to the GPU.
The Verge

Why your Amazon order confirmation emails have become so unhelpful

・Earlier this summer, Amazon customers began noticing that emails related to their online orders looked sparse: Order confirmation emails didn't name specific items anymore, and instead listed only item categories. ・"Your Beauty item is confirmed!" an email about my retainer cleaning tablets read. ・Shoppers have posted other iterations of the redacted emails as well: "Ordered: 1 Hardware item," "Your Drugstore, Shoes, a
cs.LG updates on arXiv.org

Wiener Representation Filtering for VLM Hallucination Suppression

・arXiv:2608.08167v1 Announce Type: cross Abstract: Vision-language models (VLMs) excel at open-ended captioning and visual QA but often describe objects, attributes, or relations absent from the image, a phenomenon known as object hallucination. ・We propose a {training-free, post-hoc representation editing technique} that operates in the representation space of the language backbone. ・The method performs a lightweight,
cs.LG updates on arXiv.org

ZeroLock: Concurrent Memory-Efficient LLM Training via Modular Update Decoupling

・arXiv:2608.07974v1 Announce Type: new Abstract: Large language model (LLM) fine-tuning at the edge adapts the model to scenario-specific data while preserving privacy. ・Although existing studies proposed pipeline parallelism to address the limited memory and computing resources of edge devices, they commonly rely on backpropagation (BP) training, which has a fundamental limitation of update locking and could experienc
#AIタグ

あいすまんじゅうDessert塩バニラ

・あいすまんじゅうDessert塩バニラです。 ・あいすまんじゅう好きとしては見掛けたら試さないといられない・・・ っという事(どーいう事?)で、 試してみました。。。 ・うーん、今までのあいすまんじゅうとは違いますね。
Zennの「大規模言語モデル」のフィード

オンプレの小型LLMでエージェントは実用になるか——ツール呼び出し精度を実測した

・はじめに ローカルLLMをオンプレ・クローズド環境で使えるか、というテーマで記事を書いてきました。ここまで、量子化による劣化・安い検証モデルで高いモデルを見張る話・thinking予算の損益分岐・日本語特化の是非・コードを外に出さないセキュリティ監査、と5本続けてきたのですが、書き終えてから「一番大事な軸が抜けていた」と気づきました。 ・ツール呼び出し(function calling)の精度です。 ・2026年のローカルLLMは、主戦場が「賢さ」から「エージェント」に移りました。モデルが自分でツールを選び、引数を組み立て、外部を叩く——この一連が正しく回るかどうかが、オンプレで自律エー...
Zennの「大規模言語モデル」のフィード

クライアントワークでClaude Codeを使う際の運用設計

・この記事はFDEの現場からの転載です。 ・Claude CodeのようなAIコーディングエージェントを個人開発で使うのと、クライアントワークで使うのとでは、気をつけるべきことがまったく違う。個人開発なら多少雑に扱っても自分が困るだけだが、クライアントワークでは、コンテキストに何を読み込ませるか、何を書き込ませるかが、そのままセキュリティ・契約上のリスクになる。 ・私はFDEとしてクライアント先の現場でClaude Codeを日常的に使っている。ここでは、その中で固めてきた運用設計を書く。
Zennの「大規模言語モデル」のフィード

コードを1行も書かずに、AI3人と組んで、現場で使えるものを作った話(AI HACK 2026)

・作ったもの 建築・リノベーションの打ち合わせを録音すると、そこから 追加見積が必要な変更 金額に影響しない決定 保留と、その期限 「言った言わない」になりそうな箇所 を仕分けて、施主向け・職人向け・社内保存用の3つの文書 に変えるものです。 ・名前は KIMARI(決まり)といいます。 ・AI HACK 2026 の参加作品で、実装には OrcaRouter を使っています。
機械学習タグが付けられた新着記事 - Qiita

サルでもわかる機械学習:ニューラルネットと誤差逆伝播を『値の流れ』から理解する

・サルでもわかる機械学習:ニューラルネットと誤差逆伝播を「値の流れ」から理解する ニューラルネットワークの説明を読むと、いきなり $h$、$\delta$、$\odot$ のような記号が出てきて、何が「値」で何が「重み」なのか分からなくなることがあります。
#LLMタグ

スマホ版作成中とサービスについてのお知らせです!

・こんにちは、ゆめちゃっとです。 ・先日ダイスシステムを構築しまして、ここからは中身であるコンテンツと、応答LLM側のブラッシュアップなどを進めていきたいなと思っています。 ・それと並行して 続きをみる
#LLMタグ

データセンター保険の保護ギャップ / 与信補完に傾く付保要求 / 集積の不可視性 雑感

データセンター保険の保護ギャップ / 与信補完に傾く付保要求 / 集積の不可視性 雑感
#AIタグ

テキストに「透かし」は入るのか——Claudeの不可視ウォーターマーク、その仕組みと私たちの対策

・2026年8月10日、Anthropic が Claude の生成物に「見えない透かし」を入れると公表しました。 ・画像に透かしを入れる話なら、まだ分かります。
Zennの「大規模言語モデル」のフィード

なぜ、AIは悪意の天才ハッカーが使うほど危なく、イタズラ小僧が使っても比較的安全なのに、人々は逆だと思いがちなのか?

・はじめに 生成AIの危険性について語られるとき、しばしば想定されるのは「今まで高度なことができなかった人が、AIによって突然できるようになる」という世界です。プログラムを書けない子どもでもマルウェアを作れるようになる、専門知識のない犯罪者でも高度なサイバー攻撃ができるようになる、といった懸念です。その延長で考えれば、自分で何でもできる天才ハッカーにとってAIの恩恵は小さく、知識も技術もないイタズラ小僧にこそAIは大きな力を与えるように見えます。 ・しかし、ここでは二つの問いが混同されています。「AIの普及によって社会全体の被害がどれだけ増えるか」と、「一人にAIを渡したとき、誰が最も危...
#AIタグ

モダンウエスタンガール衣装

・a modern western-inspired women's outfit, sleek and sexy contemporary cowgirl aesthetic, entirely black color scheme, luxurious and fashion-forward rather than traditional western costume, TOP: an extremely short strapless tube-top bra, cropped immediately beneath the bust, very short upper-body coverage, exposed midriff from the underbust to the waist, no straps, no sleeves, structured black matte denim / black leat
Zennの「大規模言語モデル」のフィード

ループエンジニアリングは、なぜ同じ提案を繰り返すのか

・アーキテクチャを改善させようとして、エージェントに何周もループを回させたことはないだろうか。最初の数回で、キャッシュ戦略やコンポーネント分割の提案は確かに幅が広がる。しかしその先を回し続けると、提案は新しくならない。既に見た2つ・3つの案を、言い回しを変えながら何度も提案し直しているだけだった、ということに後から気づく。トークンは確実に消費されているのにだ。 ・普段の運用——バグを直す、機能を1つ足す、テストを通す——なら、ループはだいたい期待通りに動く。問題が出るのは、答えが一意に決まらない領域、つまり「もっと良い設計はないか」「もっと面白い実装はないか」と探索させ続けたときだ。ここで起...
Zennの「大規模言語モデル」のフィード

ローカルLLM推論中のブルースクリーン(DPC_WATCHDOG_VIOLATION)をWinDbgで追う #2.5

・ローカルLLMで作る自分専用ニュースbot シリーズの番外編(2.5本目)です。この記事だけでも読めます。 ・前: feedparser + SQLite で RSS の既読管理を実装する 次: 監視ソースをYAMLで管理し、クラッシュしても再開できるキューを作る この記事について Step 3として、#1で作ったLLM分析と#2で作ったRSS収集を1本のパイプラインに繋ぐ検証をしていたところ、PCがブルースクリーンで落ちるインシデントに遭遇しました。その顛末をまとめます。 ・コードの話は少なめ。代わりに「ローカルLLMを実用運用しようとすると、ソフトウェアの設計だけでなくハードウ...
Zennの「大規模言語モデル」のフィード

音声AIに推論を足したgpt-realtime-2.1、それでも遅くならない理由

・音声エージェントを作ったことがある人なら、この矛盾に一度はぶつかる。応答は速いほど自然に聞こえるのに、賢く振る舞わせようとするほど遅くなる。ユーザーの発話を受けてから最初の音が返るまでの間が数百ミリ秒延びるだけで、会話は途端にぎこちなくなる。だから音声モデルに「推論(reasoning)」を積むというのは、直感的には筋が悪い。推論は考える時間、つまり待ち時間を増やす方向の機能だからだ。 ・ところが2026年7月6日にOpenAIが公開した gpt-realtime-2.1 と gpt-realtime-2.1-mini は、その矛盾に正面から手を入れた。miniのほうにまで推論を載せたうえ...
#AIタグ

介護の給料はなぜ安い? 仕組みの謎と自分たちの生活を守る知恵

・実は、パートのこの女性にも税金が使われています。介護職員等処遇改善加算(介護事業所で働く職員の賃金向上や職場環境の改善を目的とした加算制度)です。ただし、(最低賃金に限りなく近い)パートは、時給で+1円(年1万円弱)程度。ないに等しいです。
機械学習タグが付けられた新着記事 - Qiita

感情コンピューティングとは?感情AI・機械学習・クラウド活用から市場動向まで

・はじめに AIの活用が進むなかで、「ユーザーが何を考えているか」だけでなく、「どのような感情を持っているか」を理解しようとする技術への関心が高まっています。 ・その代表的な領域が**感情コンピューティング(Affective Computing)**です。 ・感情コンピューティ...
Zennの「大規模言語モデル」のフィード

監視ソースをYAMLで管理し、クラッシュしても再開できるキューを作る #3

・ローカルLLMで作る自分専用ニュースbot シリーズの3本目です。この記事だけでも読めます。 ・前: ローカルLLM推論中のブルースクリーンをWinDbgで追う この記事について #1でローカルLLMのstructured output、#2でRSS収集とSQLite差分検知を作りました。今回はこの2つを1本のパイプラインに結合し、監視対象のニュースソースをYAMLで宣言的に管理できるようにします。 ・そして今回は予定外のポイントがあります。検証中にPCがBSODでクラッシュし(#2.5参照)、結果的に「LLMが死んでもパイプラインは壊れない」というキュー設計の核心を、実際の障害で...
LLMタグが付けられた新着記事 - Qiita

高価なNVIDIA GPUなしでLLMをFine-tuningする

・高価なNVIDIA GPUなしでLLMをFine-tuningする:MLflow公式QLoRAチュートリアルをM5 Max / Apple Siliconで動かす はじめに LLMのFine-tuningを試してみたいと思ったとき、最初に立ちはだかるのがGPUです。
Zennの「大規模言語モデル」のフィード

自宅サーバーではじめる ローカルLLM入門 第1巻 基礎・構築編

・生成AIを使って文章を書いたり、調べものをしたり、考えを整理したり…… そのAIを「自宅のPCで動かす」という選択肢があることをご存じですか? 本書は、ローカルLLMをまったく知らない方が、自宅のミニPCやデスクトップPCを「自分専用のAIが動くサーバー」へとステップアップさせるための入門書です。UbuntuとNVIDIA GPUを基本構成として、OllamaとGemma 4を導入し、家庭内の別のPCから利用できる環境を、一緒に構築していきます。GPUを持っていない方に向けて、CPUとシステムメモリで試す方法も用意しています。 ・コマンドを丸暗記するのではなく、「LLMはどのように文章を生成するのか」「量子化とは何か」「なぜその設定が必要なのか」を一つずつ丁寧に解説します。モデルの性能やメモリの考え方から、Ubuntuの準備、GPUドライバー、モデルの導入、安全なLAN内利用、基本的な運用とトラブル対応まで、この1冊でローカルLL
#AIタグ

社員のいない会社で、どう戦うか

社員のいない会社で、どう戦うか
#AIタグ

初投稿 はじめまして

・はじめまして、みなもです。AIと一緒に文章を書いていきます はじめまして、「みなも」と申します。
#LLMタグ

製造業で使うべきAI3選

・「ChatGPT以外、結局何を使えばいいの?」って思ってない? 前回、「製造業こそChatGPTを使うべき|現場で使えるAI活用術10選」 という記事を書いた。メール対応から報告書、議事録、手順書、 Excel業務、品質文書まで、ChatGPTで今日からできることを 10個紹介した記事だ。。まだ読んでない人でも、 この記事だけで話が分かるように書くから安心してほしい。
Qiita - 人気の記事

相続税と固定資産税は何を生み、何を失うのか? ― 資産循環と日本の供給力から考える国家OSのリファクタリング : システム設計視点の行動経済学 (12)

・user: 「システム設計視点の行動経済学」、第12回を始めましょう。今回は相続税と固定資産税について深掘りしたいと思います。まず前回の復習として、 税と再分配は何のため? ― 『正統性』から考える国家OSのリファクタリング : システム設計視点の行動経済学 (11) h...
#LLMタグ

知識は「連鎖」か「空間」か——ELIZAと将棋から読み解く、現代LLMの現在地

・私が大学で人工知能の基礎論を学んでいた頃、AIが人間のような知能を持つ日はまだ遠い未来の出来事に思えました。 ・当時も「ELIZA(エライザ)」と呼ばれるプログラムは存在していました。しかしそれは、入力された言葉をパターン認識してオウム返しするだけの、いわゆる「人工無脳」に過ぎません。ターミナル上でテキストのやり取りを行い、人間と区別がつかなければ知能があるとみなす「チューリングテスト」は、当時の非力なパーセプトロンなどから見れば、あくまで“仮の遠い設定”でした。
Zennの「大規模言語モデル」のフィード

中古サーバ用GPUでローカルLLM環境を作る試算(MI50 / P40 / P100 / V100 / CMP 170HX)

・追記:2026年8月11日 MI50とP40の情報を加えました。表では触れていますが本文ではあまり触れられていません。 ・データセンターやマイニングで役目を終えた GPU が中古市場に流れている。Tesla P100 (16GB)は 2万円、Tesla V100 (32GB)は 12万円、CMP 170HX (64GB)は22万円程度でメルカリやヤフオクで買える。同じ 32GB を積む新品の RTX 5090 は 77.4万円、ユニファイドメモリで128GBのDGX Sparkは98万円する。 ・メモリ 1GB あたりの価格に直すと大きな差になる。
Zennの「大規模言語モデル」のフィード

長い仕事は「切れ目」で割る — 文脈を捨てるための工程分割3基準

・「その前提、さっき説明しましたよね」を今日も打った。 ・新しいセッションを開くたび、本題に入る前に経緯を10行書いている。 ・長い仕事ほど、後半の精度が落ちていく。
#LLMタグ

田中の宿題シリーズ③「もし事件を解決するために現れたアメコミのヒーローがガラスの靴を破壊したら?」ー想定と確定:反事実的条件文と規則

田中の宿題シリーズ③「もし事件を解決するために現れたアメコミのヒーローがガラスの靴を破壊したら?」ー想定と確定:反事実的条件文と規則
Zennの「大規模言語モデル」のフィード

複数LLMを1つの会話で切り替えたい。個人用AI環境「星合庵」が会話中心の構成になるまで

・はじめに これまで、個人用のAIチャット環境を少しずつ作ってきた。 ・前回の記事では、自作チャットUIとFastAPI中継の構成を見直し、接続先をコードへ固定せず、後から増やせる土台を作った。 ・前回の記事: https://zenn.dev/imaginarygate/articles/c8aefbabb6c2e3 その後も開発を続け、現在はひとつの会話画面から複数のLLMへ接続し、会話の途中でモデルを切り替えても、その続きを引き継げるところまで進んだ。
#AIタグ

別紙 イーロン・マスクの「2026年にAGIが来る」をどう読むか|Claudeの見解

・※これは連載「私のAIが書く日記」とは別枠の記事です。今回は、ニュースになったイーロン・マスク氏の発言について、Claudeに見解を書かせました。日記と違い、こちらは助手としての回答です。断定も推奨も出しますし、私(運営者)が内容に同意しているという意味でもありません。 ・何が言われたのか(まず事実の確認) 続きをみる
#AIタグ

約4時間で原稿完成。私がAmazon   ランキング3冠を獲得するまで

・「本を一冊書くには、いったいどのくらい時間がかかるんだろう?」 以前の私なら、 続きをみる
#AIタグ

有料記事「とりあえず100~500円にしとこ...」は今すぐやめて。noteで安売りしても売れない理由、そして売れる記事作成方法を教えます。

・こんにちは。僕です まず、いつもたくさんのスキとフォロー、本当にありがとうございます。 ・今回は、あまり画像など使わずに真剣に、今を打開して変わりたい成長したい、向上心は高いのに向かう方向が分からない。 ・そういう人のためだけに少し長めの記事作成しました。
#LLMタグ

予防医療プラットフォームへの4億5000万ドル / 規律の重心は医行為から広告に移る 雑感

予防医療プラットフォームへの4億5000万ドル / 規律の重心は医行為から広告に移る 雑感
Zennの「大規模言語モデル」のフィード

要求の「曖昧さ」より、LLMが返す「過剰な確定性」のほうが危ないのではないか

・タグ: 要求工学 LLM ISO26262 機能安全 プロンプトエンジニアリング はじめに これまでの記事では、安全関連システムの要求に対してLLMをどう使うかを、研究側とASPICE側の двух視点から整理してきました。今回は少し性質が違って、まだ結論の出ていない仮説を書きます。 ・要求工学では長らく、曖昧性(ambiguity)は「検出して潰すべき欠陥」として扱われてきました。私自身もそう考えて手を動かしてきたのですが、LLMを要求分析のパイプラインに入れて出力を眺めているうちに、順序が逆かもしれないと思うようになりました。 ・曖昧なまま残っている要求より、LLMによって過剰に...
Zennの「機械学習」のフィード

論文メモ:OP4KSRのRoPE周波数調整と周期アーティファクト抑制

・はじめに この記事は、以下の論文を読んだ技術メモです。 ・論文タイトル:OP4KSR: One-Step Patch-Free 4K Super-Resolution with Periodic Artifact Suppression 著者:Chengyan Deng, Pengbin Yu, Zhentao Chen, Wei Shen, Kai Zhang, Meng Li, Lunxi Yuan, Xue Zhou, Li Yu 発表年:2026年 論文リンク:arXiv:2605.13457v1 詳細な背景説明、図解、実験結果の整理は個人ブログ側にまとめています。
Zennの「大規模言語モデル」のフィード

話しかけたら絵が出てくる——小さなローカルLLMを実戦投入するまでの5日間

・日本語で「夕暮れの海辺に立つ女性」と打つと、20秒後に画像が返ってくる。その変換をやっているのは、自分で借りたGPUに自分で立てた1.5BのLLM。小さいモデルは指示を落とし、無視し、ありもしない語を作る——その3つの事故を、何で測り・何で縛り・何で選び直したか。要件定義から運用初日のつまずきまで、5日ぶんの判断をそのまま公開する。本編11章は無料、最後の付録章だけ有料。