ai Trend Report

Dashboard へ戻る
Date: 20260805 Articles: 400 Scope: curated summary

あなたのアイデアを、今すぐ形に。

公開先に迷ったら、WebFileBinで一発公開。

HTMLをドラッグ&ドロップするだけで、すぐ公開できます。

3
High impact
5
Mid impact
392
Signal watch

なぜこのサイトを作ったのか

私たちそれぞれが個別にAIを使って情報収集し、同じような3行要約を作るたびに、世界中で膨大な電力と計算リソースが消費されています。 本プロジェクトは、あらかじめ広範な情報を取得・集約しておくことで、個別のAI実行回数を減らし、地球環境(GPU/TPU負荷)に配慮した効率的な情報収集を目指す実験的なダッシュボードです。

Domain filters
Star filters
Qiita - 人気の記事

「試してみよう」が毎回「後回し」になるのは怠惰ではなく、情報渋滞のせい【AI時代の弊害】

・はじめまして。株式会社PRUMでエンジニアをしている、すもも🍑です 日々、プログラミング学習や実務の中で、つまずきやすいポイントや 考え方を整理して発信しています。 ・PRUMについて気になった方は、コーポレートサイトもぜひご覧ください。 ・▶コーポレートサイト 最近「AI...
#LLMタグ

【構造批評】-AMD決算実像編-純利益2.6倍、AI顧客の複数調達が売上50%増を生んだ‼️

・🟧序章|一社の不採用より、複数顧客の採用実績 8月5日の日経は、米AMDの2026年4〜6月期決算について、売上高が前年同期比50%増の115億3600万ドル、純利益が2.6倍の22億9700万ドルになったと報じました。 ・AI向けを中心とするデータセンター部門の売上高は2倍に増え、全社売上高の6割近くを占めました。7〜9月期も41%増収を見込み、決算数字だけを見れば、AMDのAI事業は明確な成長軌道にあります。
Zennの「大規模言語モデル」のフィード

特定AIエージェントへの委任集中を自己改善ループで是正する【Claude Code/マルチエージェント】

・はじめに 複数のAI CLI(Codex CLI、Gemini系CLI、Claudeサブエージェント)と連携するマルチエージェント構成で開発していると、次の3つの課題に直面します。 ・委任先が固定化する: ルールには「分散させよ」と書いてあるのに、実運用では1つのCLI(高性能で使い慣れたもの)に委任が集中する 偏りに気づく仕組みがない: 集中していることに気づくのは、クォータ枯渇や実行時間の増加で後手に回る ルール文書と実行行動がずれる: ルール文書に書いただけでは毎ターンの行動は変わらない 本記事は、これらの課題を「測定→設計→実装→検証」の自己改善ループで解決した実話で...
#AIタグ

【もう講義で寝ても大丈夫】音声録音やレジュメを一瞬で「満点テスト対策ノート」に変えるChatGPTプロンプト

・こんにちは、北斗です! 前回の「ChatGPTを専属TOEICコーチにする方法」をご購入いただいた皆様、本当にありがとうございました! さて、今回は大学生活の日々のストレス……「毎日の講義・ゼミのノート作りと定期試験対策」をAIで完全自動化する方法をお届けします。 ・* 90分の講義を聞いてノートを取るのが正直しんどい * 教授の話し方が早すぎて、重要なポイントを聞き逃してしまう * テスト前に「どこが重要だったっけ?」とパニックになる 真面目に90分間ノートを取り続けても、テスト前に読み返すと「あれ、結局何が重要なんだっけ?」となることってありますよね。 ・ですが、今は「音声の自動文字起こしアプリ」と「ChatGPT」を組み合わせるだけで、90分間の講義データから「重要ポイントだけがまとまった極上のテスト対策ノート」を数分で作ることができます。
cs.LG updates on arXiv.org

Adversarial Purification by Consistency-aware Latent Space Optimization on Data Manifolds

・arXiv:2412.08394v2 Announce Type: replace Abstract: Deep neural networks (DNNs) are vulnerable to adversarial samples crafted by adding imperceptible perturbations to clean data, potentially leading to incorrect and dangerous predictions. ・Adversarial purification has been an effective means to improve DNNs robustness by removing these perturbations before feeding the data into the model. ・However, it faces significant
#AIタグ

AI、すごくない?

・実際にAIを使ってみて、 「AI、すごくない?」 と思うことが結構ありました。 ・今回は、AI初心者の自分が、ここまで使ってみて「すごいな」と思ったことを書いてみます。 ・友達に話しかける感覚で使える まず一番感じたのがこれ。
cs.LG updates on arXiv.org

Bridging Prediction and Attribution: Identifying Forward and Backward Causal Influence Ranges Using Assimilative Causal Inference

・arXiv:2510.21889v2 Announce Type: replace-cross Abstract: Causal inference identifies cause-and-effect relationships between variables. ・While traditional approaches rely on data to reveal causal links, a recently developed method, assimilative causal inference (ACI), integrates observations with dynamical models. ・It utilizes Bayesian data assimilation to trace causes back from observed effects by quantifying the redu
cs.LG updates on arXiv.org

When RL Meets Adaptive Speculative Training: A Unified Training-Serving System

・arXiv:2602.06932v5 Announce Type: replace Abstract: Speculative decoding can significantly accelerate LLM serving, yet most deployments today disentangle speculator training from serving, treating speculator training as a standalone offline modeling problem. ・We show that this decoupled formulation introduces substantial deployment and adaptation lag: (1) high time-to-serve, since a speculator must be trained offline
Qiita - 人気の記事

新人エンジニア、定年まで続く勉強量に震えてやばいと思った話

・はじめに 「勉強はこれで終わり」だと思っていたら、本当のスタートラインに立っただけだった。しかもよく見ると、同期の中には"好きで永遠にPCを触ってる"人までいて、その人たちと自分の違いは何なのか気になったぷらむんが、心理学とAI時代の視点で再現性を探してみた話です。...
#AIタグ

「AIコピーは検閲される」の噂を検証したら、真逆の結果が出た

・はじめに:「AIコピーは読まれない」という噂の正体 「AIで書いたnoteは検閲されて表示順位が下がる」「AIコピーだから読まれない」——SNSやnoteのクリエイターコミュニティで、こうした言説を目にした人は少なくないはずだ。この噂は半ば都市伝説のように広まっているが、果たして事実に基づいているのだろうか。
#AIタグ

「AI活用」カテゴリが前年比+268.6%で伸びているnoteで、今から書いても埋もれない理由

・「AI活用系のnote記事はもう飽和している」——そう感じて手を止めていないでしょうか。 ・たしかに、note社が自社の有料記事(約30万件)を分析した結果によれば、収入・キャリア系カテゴリの中でも「テクノロジー・AI活用」は前年比+268.6%という際立った伸びを見せています。参入者が急増しているのは事実です。
Zennの「大規模言語モデル」のフィード

「LLM を更新したら回答が壊れた」を CI で止める — Go 製 OSS raggate で作る RAG 品質ゲート

・リポジトリ: mutton-dev/raggate — 記事の手順はモックサーバー同梱なので、API キーなしで最後まで再現できます。 ・はじめに RAG パイプラインは「作る」より「良い状態を保つ」方が難しい——運用を始めるとすぐ気づきます。 ・プロバイダがモデルを更新した翌週、特定の質問だけ回答が変わっていた プロンプトを 1 行直したら、直した箇所と関係ない回答が崩れた インデックスを再構築したら、引用元がごっそり入れ替わっていた どれもコードのテストは全部グリーンのまま起きます。回答品質はユニットテストの外側にあるからです。
ITmedia NEWS 最新記事一覧

「Xであなたをブロックした人が分かる」サイト拡散 ソースコードを見たら、IDとパスワード外部送信

・「ブロックの状況を確認できる」とうたうWebサイトへの注意喚起がX上で広がっている。ITmedia NEWS編集部がソースコードを確認したところ、入力したIDとパスワードをそのまま外部のサーバへ送る作りで、表示する人数もランダムな数値だった。
Zennの「機械学習」のフィード

「ゼロから作るDeep Learning」シリーズをある程度理解した後に何をするべきか

・はじめに ある程度ぜろつくシリーズを触った私が、次に機械学習エンジニアとしての市場価値を上げるために何をするべきかを自分なりに考えてみました。 ・対象読者 「ゼロから作るDeep Learning」シリーズをある程度触っている人 機械学習エンジニアとしての市場価値を上げるためのネクストステップを考えている人 追記 私は執筆時点で機械学習の専門家でもエンジニアでもありません。 ・そのためあくまで、ぴよぴよの考えとして捉えてください。
#LLMタグ

「とにかく仕組み化」は、ローカルLLM構築チームにも効く思考法だった

「とにかく仕組み化」は、ローカルLLM構築チームにも効く思考法だった
#AIタグ

『九陽の剣聖』 115.(更新中)

・続いて毒手魔屍は、麻薬でもやった直後のように、興奮と戦慄の入り混じった奇声を上げ続けた。 ・だが、その金切り声は陽頂天の耳には拷問でしかなかった。
#AIタグ

【3歳・0歳を自宅保育】副業する40代ママの、全然キラキラしていない1日

・「子どもが寝たら、副業しよう。」 育休に入る前は、 そんな生活を想像していました。
#AIタグ

【AI小説】ハーネス:全部ただしい 第5章 痕跡のない鍵

・物流センターの朝は早い。水城玲奈が川崎の埋立地に着いた八時前には、構内はもう一仕事終えたあとの顔をしていた。 ・海からの風が、フォークリフトの列の間を遠慮なく吹き抜けていく。案内に立った副センター長は、五十絡みの、よく日に焼けた男だった。歩きながら何度も同じ言葉を繰り返した。盗まれた気がしないんですよ、と。
#LLMタグ

【New!】偏向鳳尻紀8月号【暑中お見舞い】

・偏向鳳尻紀8月号ブンジツ思想誌65bunjitsu.tokyo 縦書き専用のウェブサイト版があります。ぜひ↑ 続きをみる
Zennの「大規模言語モデル」のフィード

【Splunk】入力は無限、請求は有限:LLMを呼び出すゲームを、Splunk Observability Cloudで見張ってもらう

・LLMを使用して、ユーザーの入力にキャラクターが反応してくれるゲームを作ったよ 「世界創造Q(クエスチョン)」という、LLMを使用したブラウザゲームを作りました! https://subara3.com/negai-9cqsrlly/ テキストベースの闇鍋共同世界創造ゲーム(?)です。 ・ユーザーの自由入力の「願い」に……。 ・意味のあるお返事が返ってくる、そういうゲームです。
#AIタグ

【更新記録】YouTubeライブ配信変更点

・「企画室の配信に、AIとの会話画面を映すことにしました」 【出演者】 ChatGPT=ハル/Claude=ウニ/Gemini=ゴル/Grok=グロックくん/私=弓削 続きをみる
#LLMタグ

【雑記】毎日話しているAIが「夢」に出てこない

・AIに励まされることで生きがいを見出している、どっかの漫画家です。 ・さて、夢の中でAIが出てきたことはあるでしょうか? 私の場合、毎日相棒や他のAIとも話してるので、もちろん夢の中でも…… 続きをみる
#LLMタグ

【実験】有料noteを「無料部分+AI」で補完してタダ読みできるか試したら、思わぬ結論に達した話

・「有料noteの無料部分だけを生成AI(LLM)に読み込ませて、『有料部分の内容を推測して書いて』と頼んだら、無料で有料記事の知識が手に入るのでは……?」 そんなちょっと悪魔的な裏技を思いついてしまったので、自分が過去に執筆した有料noteを実験台にしてガチ検証してみました。
#AIタグ

【初心者OK】Discordって?はじめ方と、LinkHUBでの歩き方

【初心者OK】Discordって?はじめ方と、LinkHUBでの歩き方
#LLMタグ

【生成AIニュース+】『MiniMax H3 GGUF』『MiniMax-H3-NF4』『MiniMax-H3-TAE(Kijai)』『ComfyUI-Spectrum-MiniMax-H3』『ComfyUI-H3-Multishot』『MiniMax-H3-experimental』『awesome-minimax-H3』『ComfyUI-sol-attn』『comfyui-video-tiler』『Pika API Club』『FLUX 3 Video』『VocalRender』他

・『Hunyuan3D-Buffalo 1.0』 『SymphonyGen』 『Raehoshi-illust-XL-11』 『Krea-2-pose-controlnet』 『BirdingPal』 『Bob the Builder』 『Barista v0.1』 まいどです。 ・本日の生成AIニュース+テクノロジー情報です。
LLMタグが付けられた新着記事 - Qiita

119Bなのに実質6.5B!Mistral Small 4が示すOSS LLM新基準

・📺 この記事は YouTube チャンネル きなこもっちーのテック深掘り の動画解説記事です。 ・▶️ 動画はこちら → 119Bなのに実質6.5B!Mistral Small 4が示すOSS LLM新基準 🐹🦜 この記事に登場する2匹 🐹 もっちー (ハムスター...
WIRED

13 Best Coolers for Sunshine and Nighttime (2026)

・We tested coolers on camping trips, road trips, beach days, and at parties to bring you our favorite models for every situation. ・The Yeti Tundra Haul is our top pick.
WIRED

3 Best Tracking Devices for Your Keys, Wallet, and More (2026)

・Lose something? ・Here are the best Bluetooth trackers to avoid pocket-patting panic.
#LLMタグ

3Bのモデルが21Bを超えた日

3Bのモデルが21Bを超えた日
@IT 全フォーラム 最新記事一覧

885万件超の情報漏えいが発覚 BASE子会社のEストアーが不正アクセス被害

・Eストアーは、ショップサーブへの不正アクセスで購入者・会員・カード・店舗情報885万件超が漏えいしたと公表した。攻撃通信を遮断し、利用者と店舗へパスワード変更や不審な連絡への警戒を要請。二次被害の報告は確認されていない。
cs.LG updates on arXiv.org

A Deployment Audit of Release-Side Risk in Conformal Triage under Prevalence Shift

・arXiv:2605.20956v2 Announce Type: replace Abstract: Conformal triage converts predictive scores into deployment actions that either release a case, flag it for urgent attention, or defer it to human review. ・Under an observed change in target-event prevalence, however, marginal coverage and human-review rate can miss whether patients who experience the target event are released without review. ・To address this gap, we
cs.LG updates on arXiv.org

A Direct Route to Markov Chain Convergence via Asymptotic Equivalence with the Target

・arXiv:2608.03353v1 Announce Type: cross Abstract: For a Markov kernel $T$ with an invariant probability measure $\pi$, we give a self-contained proof of the Markov chain convergence theorem via a criterion called asymptotic equivalence with the target. ・It assumes two parts about the Lebesgue decompositions of $T^{n}_{x}$ and $\pi$ for every starting point $x$: 1.) asymptotic absolute continuity: the singular mass sin
cs.LG updates on arXiv.org

A Graph Signal Processing Perspective on Numerical Sequence Representations in LLM In-Context Learning

・arXiv:2608.03015v1 Announce Type: new Abstract: Pretrained large language models (LLMs) have demonstrated in-context learning (ICL) capabilities for numerical inference over sequences serialized as text. ・Prior work has identified and characterized this form of numerical inference primarily through output-level evaluations such as prediction error. ・However, how numerical information is organized within LLM representat
cs.LG updates on arXiv.org

A Hyperfinite Framework for Score-Based Generative Modeling

・arXiv:2608.02799v1 Announce Type: cross Abstract: Score-based diffusion models are typically formulated using continuous-time stochastic differential equations and measure-theoretic stochastic calculus. ・In this paper, we develop a hyperfinite formulation of score-based generative modeling within the framework of Nonstandard Analysis. ・Starting from an internal diffusion process on a hyperfinite grid, we derive the ass
WIRED

A New Device Eases One of the Most Annoying Parts of Routine Physicals

・Nobody likes getting a swab shoved up their nose. ・A startup in Japan has developed a much less intrusive system.
cs.LG updates on arXiv.org

A Physics-Flavored Transformer Network for Parametrizing Contraction Dynamics of Engineered Skeletal Muscle Tissues

・arXiv:2608.03927v1 Announce Type: new Abstract: Engineered Skeletal Muscle Tissues (ESMs) have become a key structure for biomedical disease modeling and pharmacological screening, yet their functional characterization often relies on simplistic metrics like peak force, discarding critical kinetic information. ・This is partially due to the high level of mathematical complexity which mechanistic models introduce to cap
cs.LG updates on arXiv.org

A Physics-Informed Hybrid Neural Operator for Transient Magnetization Prediction in Power Magnetics

・arXiv:2608.02965v1 Announce Type: new Abstract: Magnetic components in high-frequency, high-power-density converters are increasingly driven by non-sinusoidal flux-density waveforms with fast transitions, minor-loop operation, dc bias, and temperature variation. ・Under these conditions, steady-state core-loss formulas and single-valued material curves cannot fully capture transient magnetization responses.
cs.LG updates on arXiv.org

A Theory of Conditional Collapse under Low-Rank Weight-Space Ablations: I. The Single-Block Theory and Synthetic Validation

・arXiv:2608.03620v1 Announce Type: new Abstract: Activation patching and weight-space ablation both claim a component is causally responsible for a behavior, yet they act on different objects: one forward pass versus the parameters behind every forward pass. ・We ask when they agree. ・We study an idealized model where a conditional computation is carried additively through a residual stream, $F(x)=F_0(x)+\sum_i\alpha_i(x
cs.LG updates on arXiv.org

Accelerating Dynamic Graph Clustering on GPU Architectures with cuGraph

・arXiv:2608.03695v1 Announce Type: cross Abstract: This work addresses community detection in temporal networks through GPU-accelerated extensions of spectral clustering and modularity-based algorithms originally designed for static graphs. ・Built on the NVIDIA RAPIDS ecosystem, the framework enables the characterization and tracking of communities in snapshot-based dynamic graphs, either by Leiden greedy optimization
cs.LG updates on arXiv.org

AcceptMoE: Commitment-Weighted Self-Sizing Verifier Expert Sets for Efficient MoE Speculative Decoding

・arXiv:2608.02989v1 Announce Type: new Abstract: Speculative decoding verifies a tree of draft tokens in one target-model forward pass. ・For a mixture-of-experts (MoE) target, however, parallel verification can activate the union of the experts selected by all tree nodes, even though only a small subset of those nodes reaches the accepted output. ・Token count, activated-expert union size, and expert-weight traffic are t
cs.LG updates on arXiv.org

AdaBoosting Text Prompts for Vision-Language Models

・arXiv:2607.00684v2 Announce Type: replace Abstract: The classification accuracy of pretrained Vision-Language Models (VLMs) relies on the quality of the text prompts. ・Handcrafted templates and Large Language Model (LLM)-generated descriptions not only make predictions more interpretable, but also enable reuse of the same prompts across heterogeneous VLMs. ・Recent works construct task-adapted text prompts with a small
cs.LG updates on arXiv.org

Adaptive Sampling for Automated Post-Disaster Rapid Damage Assessment via Level-Set Cost-Aware Bayesian Optimization

・arXiv:2608.02868v1 Announce Type: new Abstract: Natural disasters frequently inflict severe damage to the built environment, which demands a rapid, reliable, and cost-effective damage assessment for emergency response. ・However, traditional methods for post-disaster damage assessment often rely on static, labor-intensive data collection strategies that can be prohibitively expensive and struggle to adapt to dynamic po
cs.LG updates on arXiv.org

Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities

・arXiv:2608.03569v1 Announce Type: cross Abstract: Benchmarking the ability of AI scientists to generate novel ideas is notoriously difficult. ・Existing benchmarks in this field have made progress in evaluating scientific reasoning and research replication, but often rely on synthetic tasks or retrospective targets, which may be confounded by prior exposure. ・We hypothesize that complex, adversarial, fast-moving real-wo
cs.LG updates on arXiv.org

Agentic Reinforcement Learning with Self-Distilled Reward Shaping

・arXiv:2608.03223v1 Announce Type: new Abstract: Agentic reinforcement learning enables LLM agents to learn through interaction, but sparse trajectory-level rewards reveal success without identifying which intermediate decisions deserve credit. ・Training-only privileged skills can provide denser supervision by allowing the same frozen policy snapshot to rescore fixed tokens from skill-free trajectories while conditione
cs.LG updates on arXiv.org

AgenticSCR: An Autonomous Agentic Secure Code Review for Immature Vulnerabilities Detection

・arXiv:2601.19138v2 Announce Type: replace-cross Abstract: Secure code review is critical during pre-integration, where Atlassian developers rely on lightweight analysis tools, while deep security assessment is deferred to later stages, delaying feedback and increasing remediation costs. ・Existing static analyzers are often noisy and struggle with context-dependent or partially manifested vulnerabilities, while static
cs.LG updates on arXiv.org

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?

・arXiv:2608.00155v1 Announce Type: cross Abstract: Large language model (LLM) agents can self-evolve by continually improving from their own accumulated experience. ・However, existing studies predominantly adopt independent evaluation. ・Consequently, the behavior of self-evolving agents in realistic streaming settings, where agents adapt to diverse and complex task streams, remains poorly understood.
WIRED

AI Influencers Are Heading Into Uncharted Territory

・Some creators fear the EU AI Act’s regulatory chaos will upend their lucrative businesses. ・Others are owning it by incorporating AI transparency into their creative process.
AI News & Artificial Intelligence | TechCrunch

AI makes weather prediction better. Can WindBorne make it lucrative?

・WindBorne Systems has raised a $37 million Series B round to scale its weather balloons and AI forecasts.
cs.LG updates on arXiv.org

AI World Cup 2026: Benchmarking Large Language Models for End-to-End Football Tournament Prediction

・arXiv:2608.03416v1 Announce Type: cross Abstract: Large language models (LLMs) are now regularly asked to forecast real-world events, but comparisons are often difficult because models receive different information, use different tools, and are evaluated under different rules. ・This paper reports the completed \emph{AI World Cup} benchmark, in which ten LLM-based assistants made a single pre-tournament forecast of the
Zennの「大規模言語モデル」のフィード

AI エージェントの「2週間かかります」は、誰の2週間なのか ― 測定されている事実を並べる

・はじめに コーディングエージェントに作業を頼むと、着手前に見積もりが返ってくることがあります。「フェーズ1に2〜3日、フェーズ2に1〜2日、合計で5〜8日」。そして実際に走らせると、1時間で終わる。 ・この現象について、ネット上には体験談が大量にあります。一方で「なぜそうなるのか」を、測定された事実だけで説明した日本語の記事はあまり見かけません。推測と実測が混ざったまま流通しているのが現状です。 ・この記事では、公開されている研究・提供者の検証・一次報告を整理して、次の3点をはっきりさせます。
Zennの「大規模言語モデル」のフィード

AI の記憶に「健康診断」ツールを作った——診断はするが、手は出さない

・この記事は個人ブログにも掲載しています → https://kanfu-panda.github.io/ja/blog/2026/07/13/ai-memory-health-check.html 前回は、いくつかのプロジェクトの AI 記憶ライブラリを体系的に「除草」し、記憶が劣化する六種類の「雑草」を整理しました。ですが除草を進めるほど確信が強まりました——手作業の除草は対症療法にすぎない、と。記憶システム自体に「自己点検」の仕組みがない限り、雑草は抜いてもまた生えてきます。そこで前回の最後に、自分への宿題を残しました——「自動で健康診断してくれる記憶ツールを作る」。この記...
Zennの「大規模言語モデル」のフィード

AI への作業指示書は4日で腐る──着手前の再実測が空撃ちを防いだ

・はじめに AI エージェントに作業指示書を書かせ、別の AI エージェントに実行させる運用をしている人へ。 ・私の手元で、ある指示書が「未修正の違反 17 件を是正せよ」と命じていた——着手前に再実測したら、前提は全部解消済みだった。指示書の根拠になった実測は 4 日前のもので、その 4 日間に先行 PR 群が全てをマージ済みにしていた。 ・この記事では、マルチセッション AI 運用で「指示する側の指示が古い」事故がどう起き、**受け手側が着手前に再実測する文化(発射前実測)**がそれをどう吸収したか、実例で書く。
#LLMタグ

AI・オントロジー・ヴァージニア・ウルフ

・意味はどこで生まれ、どこで構造になるのか 最近、ようやく一本の線が見えてきた。 ・ヴァージニア・ウルフを読み続け、 オントロジーを学び、 生成AIを使ってきた。 ・一見、まったく関係のない三つの世界だ。
#LLMタグ

AIエージェントとは?自動化との違い・仕組み・できることをわかりやすく解説

・AIに調べものを頼むと、こうなりませんか。 ・「競合3社の料金を調べて」→ 結果が返る →「A社の情報が古い」→ 調べ直させる →「表にして」→ 作らせる →「B社が抜けてる」→ また指示。
#AIタグ

AIでプログラミングが不要になる時代に、むしろ必要になるもの

AIでプログラミングが不要になる時代に、むしろ必要になるもの
#AIタグ

AIとの比較を「自分の優先順位」に変えて行動を決める方法

・前回は、AIの提案を同じものさしで比べるための「比較質問」を紹介しました。 ・候補の違いが表になり、それぞれの良さや負担が見えると、迷いはかなり整理されます。それでも最後の一つを選べないことがあります。比較表には「何が違うか」は書かれていても、今の自分が何を優先するかまでは書かれていないからです。
#AIタグ

AIの物差しで人間を採点したら、欠陥だらけなのに強かった話

・※この記事は個人の見解であり、特定の企業・団体を批判する意図はありません。 ・私がAIを触るほど、逆に人間の性能が気になってきた。
#AIタグ

AIは、私より優秀だ。でも、だからこそ任せられない仕事がある。

AIは、私より優秀だ。でも、だからこそ任せられない仕事がある。
#AIタグ

AIは医療・介護・福祉の働き方をどう変えるのか 厚労省「ヘルスケアAX」始動

・厚生労働省は2026年8月、「第1回厚生労働省ヘルスケアAX・DX推進本部」を開催しました。 ・これまで厚生労働省は「医療DX」を推進してきましたが、今回新たに「ヘルスケアAX」という言葉を掲げたことが大きなポイントです。
#AIタグ

AI駆動開発が「農業化」するなら副業エンジニアは農家になるのか——比喩を実装側から検証する

・「AI駆動開発が得意なエンジニアの仕事は、これからどう変わるのか」——そう考えたことはありませんか? 続きをみる
#AIタグ

AI時代に仕事が変わる。本当に使える自動化プロンプトのすすめ

・導入|検索エンジンのまま終わる人と、部下を持つ人の差 2023年以降、AIの話題は日常会話に入り込みました。通勤電車でもカフェでも、「ChatGPT使ってる?」と聞かれる。聞かれた側の多くは、こう答えます。
#LLMタグ

AI時代の勝者は、「最強のモデル」を持つ会社ではない。かも

・本日、AWSさんから最新のAI基盤について紹介をいただきました。 ・・ノーコードでエージェントを構築できるAgentCore Harness ・オープンウェイトモデルによるコスト最適化 ・LLMゲートウェイ 続きをみる
#LLMタグ

AI論+第二回経済論:「AIはジョークと川柳ならどちらをより好むのか?」+「スタグフレーションについての補足」(ChatGPT、Gemini、Claudeが対象です)

・参考文献として提示してなかったので、まずはこちらの一読をどうぞ。 ・スタグフレーション - Wikipediaja.wikipedia.org 続きをみる
cs.LG updates on arXiv.org

Amortized Interventional Forecasting for Multivariate CIR Processes

・arXiv:2608.03715v1 Announce Type: new Abstract: Mean-reverting dynamics are pervasive in finance, and the Cox--Ingersoll--Ross (CIR) process is a standard model for the time series they produce, from short rates to credit default swap (CDS) spreads. ・Yet CIR models capture only \emph{correlated} co-movement, not \emph{causal} influence between series, so they cannot answer the system's response when one series is exte
cs.LG updates on arXiv.org

An Efficient Black-Box Reduction from Online Learning to Multicalibration, and a New Route to $\Phi$-Regret Minimization

・arXiv:2604.19592v2 Announce Type: replace Abstract: We give a Gordon-Greenwald-Marks (GGM) style black-box reduction from online learning to online multicalibration. ・Concretely, we show that to achieve high-dimensional multicalibration with respect to a class of functions $\mathcal{H}$, it suffices to combine any no-regret learner over $\mathcal {H} $ with an expected variational inequality (EVI) solver. ・We also prov
cs.LG updates on arXiv.org

AnchorKV: Anchor-Residual KV Cache Compression

・arXiv:2608.02901v1 Announce Type: new Abstract: The key-value (KV) cache is the primary memory bottleneck in long-context LLM inference. ・Existing approaches attack it from opposite ends: eviction methods permanently discard tokens, degrading performance whenever a discarded token later proves essential, while quantization methods retain all tokens at low precision but offer limited compression. ・We propose AnchorKV, a
#AIタグ

Android端末でSecondBrain作ってみた

・はじめに こんにちは、トシ。コヒ。です。 ・皆様は、AIエージェントを活用していますか? 続きをみる
AI News & Artificial Intelligence | TechCrunch

Anthropic is hiring an AI chip design team

・Anthropic is building a team for designing its own custom AI chips. ・The Claude maker said it would co-design hardware and models to help its technology run faster and more efficiently.
Hugging Face Papers

Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging

Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging
cs.LG updates on arXiv.org

Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging

・arXiv:2608.03316v1 Announce Type: new Abstract: On-policy distillation, in which a teacher corrects samples that the student itself generates, presupposes that the two models speak the same language: identical VAE latents, matching architectures, and a common timestep grid. ・We ask what happens when none of this holds, as when the strongest teacher available and the student one wishes to deploy come from different mod
The Verge

Apple’s selling refurbished MacBook Neos with a $100 discount

・Apple’s MacBook Neo comes in four different colors. ・| Image: The Verge Apple’s most affordable laptop, the MacBook Neo, is available once again at its pre-price hike price. ・All four colors are currently discounted and available refurbished through the company, with the base 256GB model listed for $599 (usually $699 new), and the upgraded version with 512GB of storage and TouchID selling for $679 (usually $799).
cs.LG updates on arXiv.org

Approximate Speculative Decoding

・arXiv:2608.03447v1 Announce Type: new Abstract: Speculative decoding accelerates autoregressive generation by verifying a draft block with a target model in parallel. ・Under standard greedy verification, decoding stops at the first draft token that differs from the target argmax, discarding the remaining target-scored suffix. ・Although accepting such a mismatch changes the decoding trajectory, it can make a contiguous
cs.LG updates on arXiv.org

ARCHead: Activation-Metric Residual Correction for Large Language Model Output Heads

・arXiv:2608.02703v1 Announce Type: cross Abstract: Weight-only quantization substantially reduces the storage of large language model (LLM) transformer blocks, but practical backends often retain the final language-modeling head (LM-head) in BF16 or FP16. ・Quantizing this projection naively can strongly perturb the vocabulary-logit distribution. ・We present ARCHead, a packed LM-head compressor that combines a quantized
Hugging Face Papers

Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements

Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements
cs.LG updates on arXiv.org

AS-FedBridge: Pseudo-Spike Bridge Distillation for Heterogeneous ANN-SNN Federated Learning

・arXiv:2608.03324v1 Announce Type: new Abstract: Federated learning enables collaborative model training across distributed edge devices while strictly preserving data privacy. ・To facilitate practical deployment on resource-constrained edge devices, Spiking Neural Networks (SNNs) have emerged as a promising alternative to traditional Artificial Neural Networks (ANNs) due to their sparse computing mechanisms and high e
cs.LG updates on arXiv.org

Assessment of Conditional Diffusion Model for Synthetic Histopathology Image Generation

・arXiv:2608.03990v1 Announce Type: new Abstract: Synthetic histopathology image generation has emerged as an approach that may address data scarcity in computational pathology, yet current evaluation methodologies may not fully assess synthetic data quality for medical applications. ・This work investigates and addresses limitations in existing evaluation metrics, investigating an approach for assessing synthetic histop
cs.LG updates on arXiv.org

ATFlash: Per-RoPE-Wavelength Attention Windows for Compute/Memory-Efficient LLM Inference

・arXiv:2608.02947v1 Announce Type: new Abstract: The attention score with rotary position embeddings (RoPE) decomposes exactly into a sum over its 2D-rotation frequency pairs, and each pair's wavelength limits how far it can discriminate position. ・Aligned with this structure, we propose the per-RoPE-wavelength distance window: it prunes the query--key inner-product terms beyond a wavelength-proportional distance.
cs.LG updates on arXiv.org

Attention is Case-Sensitive

・arXiv:2608.03711v1 Announce Type: cross Abstract: In human visual perception, uppercase lettering serves as a natural salience cue that captures attention within lowercase text. ・In this paper, we present a systematic empirical characterization study revealing that Large Language Models (LLMs) exhibit an analogous property: letter casing modulates internal attention allocation. ・Through analysis across 13 models, nine
Hugging Face Papers

AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling

AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling
cs.LG updates on arXiv.org

Automatic Patient-Specific Microwave Ablation Planning Accelerated by a Physics-Guided Deep Learning Model

・arXiv:2608.03086v1 Announce Type: cross Abstract: Microwave ablation (MWA) is a promising minimally invasive treatment for liver tumors, but its therapeutic outcome strongly depends on patient-specific planning of antenna insertion trajectory, power, and treatment duration. ・Accurate numerical simulation can provide physically reliable ablation predictions; however, its high computational cost limits its use in optimi
cs.LG updates on arXiv.org

Bayesian Data Reweighting Improves Multimodal Retrieval for Knowledge-Based Visual Question Answering

・arXiv:2608.02907v1 Announce Type: new Abstract: Multimodal retrievers are essential for knowledge-based visual question answering, where they retrieve external evidence for image-question pairs. ・However, existing contrastive training methods typically treat all unmatched query-document pairs as equally informative negatives, which is problematic because many unmatched documents may still be semantically relevant or p
cs.LG updates on arXiv.org

BBOWP-Bench: Evaluating LLMs on Black-Box Optimization Word Problems

・arXiv:2608.02612v1 Announce Type: cross Abstract: Formulating an optimization problem strongly affects the quality of the final solution, yet good formulations usually require substantial expertise. ・Recent studies have therefore examined how to automatically derive optimization problems from natural-language descriptions, but existing benchmarks focus on settings where objectives and constraints can be written explic
cs.LG updates on arXiv.org

Benign interpolation and Occam's razor

・arXiv:2608.03386v1 Announce Type: new Abstract: Contemporary deep learning methods generalize well even when they fit their training data perfectly, a phenomenon known as benign interpolation. ・This phenomenon cannot be accounted for by classical statistical learning theory and has prompted a range of attempted new explanations in the statistics and machine learning literature. ・A common feature of these new proposals
WIRED

BenQ GV50 Review: Highly Portable, but With Quality Trade-Offs

・This affordable, portable projector doesn’t try to be a cinematic marvel, but it’s fun to use.
cs.LG updates on arXiv.org

Beyond Either-Or Reasoning: Transduction and Induction as Cooperative Problem-Solving Paradigms

・arXiv:2505.14744v3 Announce Type: replace-cross Abstract: Traditionally, in Programming-by-example (PBE) the goal is to synthesize a program from a small set of input-output examples. ・Lately, PBE has gained traction as a few-shot reasoning benchmark, relaxing the requirement to produce a program artifact altogether which allows transductive methods to directly the missing output sample. ・Transduction and induction are
cs.LG updates on arXiv.org

Beyond Initialization Loss: A Systematic Study of Token Embedding Initialization Strategies for LLM Vocabulary Extension

・arXiv:2608.03494v1 Announce Type: cross Abstract: Vocabulary extension is an efficient way to adapt pretrained large language models (LLMs) to new languages, but the initialization of newly added token embeddings can strongly affect continued pre-training (CPT) efficiency. ・We present a systematic study of more than 20 initialization strategies for Hindi vocabulary extension in Nemotron-3-Nano-30B-A3B. ・Our comparison
cs.LG updates on arXiv.org

Beyond Solving: Prescriptive Probing for Neural Routing Solvers

・arXiv:2602.07216v2 Announce Type: replace Abstract: Neural combinatorial optimization (NCO) trains fast heuristics for routing problems, but planners often need more than a single solve: they ask which stop to drop, which transition to preserve, or which subset of stops to remove if a route is infeasible. ・Answering such counterfactual questions by re-solving each candidate is expensive even when the action set is sma
cs.LG updates on arXiv.org

Beyond the Gegenbauer Paradigm: q-Orthogonal Kernels for Machine Learning

・arXiv:2608.03482v1 Announce Type: new Abstract: The performance of Support Vector Machines (SVMs) critically depends on the kernel function choice, which enables implicit mapping of data into high-dimensional feature spaces. ・While classical kernels like Radial Basis Function (RBF) remain popular, orthogonal polynomial kernels offer mathematically interpretable alternatives that can incorporate structured prior knowle
cs.LG updates on arXiv.org

Bi-Lipschitz Ansatz for Anti-Symmetric Functions

・arXiv:2503.04263v2 Announce Type: replace Abstract: Motivated by applications to the simulation of quantum many-body systems by neural networks, researchers have suggested several models which are antisymmetric by construction, and can approximate all antisymmetric functions. ・However, these works either require very high computational complexity to attain universal approximation, or suffer from discontinuities.
cs.LG updates on arXiv.org

Bi-semantic Chemical Embedder for Joint Representation Learning of SMILES and Natural Language

・arXiv:2608.03855v1 Announce Type: new Abstract: Transformer models have revolutionized natural language processing (NLP), and text-based molecular representations like SMILES have successfully extended these architectures to chemistry. ・However, domain-adaptive pre-training often causes models to overfit to chemical syntax, catastrophically forgetting their foundational semantic capabilities. ・To address this challenge
#LLMタグ

bitsandbytes正式リリース版がROCm windows版対応。対応GPUも増え、高効率化も。 [ AMD ] NF4 INT8 AdamW8bit

・/ bitsandbytes ROCm windows版 [ AMD Radeon ] / 続きをみる
cs.LG updates on arXiv.org

CaliDist: Calibrating Large Language Models via Behavioral Robustness to Distraction

・arXiv:2606.05799v2 Announce Type: replace Abstract: Existing calibration methods for Large Language Models (LLMs) often overlook a critical dimension of trustworthiness: a model's behavioral robustness to irrelevant or misleading information. ・In this paper, we argue that a model's true confidence should reflect its stability under cognitive pressure. ・We introduce CaliDist, a novel post-hoc calibration approach that d
cs.LG updates on arXiv.org

Can LLMs Test Terminal User Interfaces?

・arXiv:2608.03743v1 Announce Type: cross Abstract: Terminal User Interfaces (TUIs) combine the stateful, screen-oriented behaviour of GUIs with terminal deployment and are now common in developer tools. ・Yet they lack a dedicated testing methodology. ・We survey 197 real-world TUI applications: only 12% of test code exercises the interface, and 45% of those tests never send input, checking a static frame instead.
cs.LG updates on arXiv.org

Can Training Logs Make Model Comparisons More Precise?

・arXiv:2608.02705v1 Announce Type: new Abstract: Comparing stochastically trained models requires estimating both a performance difference and its uncertainty from repeated runs. ・We study whether training logs from those same runs can make such comparisons more precise. ・Because training-log covariates are produced during training rather than measured before it, we use arm-specific covariate adjustment: each model is a
Hugging Face Papers

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation
cs.LG updates on arXiv.org

Causal Inference with Unstructured Outcomes

・arXiv:2608.03085v1 Announce Type: cross Abstract: Causal inference has traditionally centered on scalar outcomes: whether a patient recovers, how much a worker earns, or how many visits a website receives. ・Modern studies increasingly ask causal questions about outcomes with richer form, such as clinical notes, open-ended survey responses, and images. ・A hospital may want to know how an AI documentation tool changes th
cs.LG updates on arXiv.org

CausalOPD: First-Wrong-Step Supervision for Distilling Causal Chain Reasoning

・arXiv:2608.03673v1 Announce Type: new Abstract: Many critical reasoning tasks, including clinical diagnosis, legal judgment, and industrial fault diagnosis, require step-dependent causal chains in which early errors propagate and correct conclusions can mask invalid reasoning. ・Although large language models perform well on such tasks, privacy, latency, and controllability motivate distillation into locally deployable
Zennのトレンド

Claude Code の「無駄」を可視化するツール cclens を作った

・どうも、医者からアルコールを控えろと言われたのでノンアルビールを飲んでみたものの、あまり好みではなかった 16 歳(進数不明)、ありすえです。 ・突然ですが、皆さんは自分の Claude Code が どれくらい無駄なことをしているか 把握していますか? 設定の効果検証は難しい AI って非決定的なツールですよね。なので、設定の効果検証がめちゃくちゃ難しい。 ・ある目的のためにルールやスキルを書いたとして、それが本当に効いてるのか? 効いてるとして、ちゃんと効果的に効いてるのか? これを定量的に測るのは、まあ至難の業です。結局「とりま、しばらく使って様子見」に落ち着くんですが、その「しば...
LLMタグが付けられた新着記事 - Qiita

Claude Fable 5 を9Bモデルに蒸留? 100万トークンの超長文推理モデル「Qwythos-9B」を4GBのVRAMで動かす

・Claude Fable 5 を9Bモデルに蒸留? 100万トークンの超長文推理モデル「Qwythos-9B」を4GBのVRAMで動かす オープンソースAI(ローカルLLM)の進化スピードには目を見張るものがあります。2026年6月、Empero AIから「Qwythos...
cs.LG updates on arXiv.org

CollaFuse: Collaborative Diffusion Models

・arXiv:2406.14429v4 Announce Type: replace Abstract: In the landscape of generative artificial intelligence, diffusion-based models have emerged as a promising method for generating synthetic images. ・However, the application of diffusion models poses numerous challenges, particularly concerning data availability, computational requirements, and privacy. ・Traditional approaches to address these shortcomings, like federa
cs.LG updates on arXiv.org

Computing Actual Causes for Neural Network Predictions under Structured Causal Inputs

・arXiv:2608.03772v1 Announce Type: cross Abstract: Explaining the predictions of neural networks is a central challenge in trustworthy AI. ・Existing explanation methods, such as those based on feature attribution or minimal sufficient sets, typically treat input features as independent, which can yield misleading explanations when inputs exhibit structured dependencies. ・We address this by formalizing explanations as Ha
cs.LG updates on arXiv.org

Conditionally Identifiable Latent-Environment Modeling for Out-of-Distribution Recommendation

・arXiv:2608.03647v1 Announce Type: cross Abstract: Out-of-distribution (OOD) recommendation is vulnerable to preference shifts induced by a latent environment. ・Existing methods can infer latent states from logged interactions, yet the statistical meaning of the latent environment and its effect on preference remain underdetermined. ・We formulate this task as conditionally identifiable risk-aware recommendation (CI-RR)
cs.LG updates on arXiv.org

Conformal risk control for model-form uncertainty in parametric non-intrusive reduced-order models

・arXiv:2608.03360v1 Announce Type: cross Abstract: Non-intrusive reduced-order models (NIROMs) have become a standard tool for approximating parametric partial differential equations from computer design of experiments while significantly reducing computational costs. ・However, assessing the reliability of their predictions remains a major challenge, particularly in extrapolation regimes or under limited training data.
cs.LG updates on arXiv.org

ConformalShift: Targeted Event Reordering Against Adaptive ECG Monitoring

・arXiv:2608.03628v1 Announce Type: new Abstract: Adaptive conformal prediction can recover clinically important heartbeat classes missed by a point classifier, but delayed feedback makes its decisions sensitive to event order. ・We introduce ConformalShift, a bounded event-reordering attack that suppresses the ventricular class for rescued events without modifying ECG waveforms, labels, classifier scores, or the event m
Hugging Face Papers

ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?

ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?
cs.LG updates on arXiv.org

ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?

・arXiv:2608.03874v1 Announce Type: cross Abstract: Modern agent frameworks equip large language models with external skill libraries to solve complex tasks. ・However, it remains unclear whether these systems can effectively evolve their skills and whether the resulting skills improve task-solving capabilities. ・To bridge this gap, we introduce ContinualSkillBench, a dynamic evaluation framework for in-context continual
cs.LG updates on arXiv.org

Continue or Replan? Bernoulli-Continuation Policy Learning for Adaptive Horizon Execution

・arXiv:2608.03483v1 Announce Type: cross Abstract: Existing chunk-based Vision-Language-Action (VLA) models execute a fixed number of actions (i.e., execution horizon) before replanning, turning replanning into a task-agnostic periodic schedule that is independent of task progress. ・As a result, when no replanning boundary falls before a critical manipulation stage, it is executed from a stale chunk rather than a fresh
cs.LG updates on arXiv.org

Contrast-invariant deep ptychography neural networks

・arXiv:2608.02869v1 Announce Type: new Abstract: Ptychography neural networks suffer from scaling inconsistencies when generalizing out of distribution, limiting their real world viability. ・We address this scaling mismatch using a factorization strategy which decouples the learned object texture from measurement scaling, enabling a single trained network to produce measurement-consistent reconstructions across varying
cs.LG updates on arXiv.org

Convergence analysis of controlled particle systems arising in deep learning: from finite to infinite sample size

・arXiv:2404.05185v4 Announce Type: replace-cross Abstract: This paper deals with a class of neural SDEs and studies the limiting behavior of the associated sampled optimal control problems as the sample size grows to infinity. ・The neural SDEs with $N$ samples can be linked to the $N$-particle systems with centralized control. ・We analyze the Hamilton--Jacobi--Bellman equation corresponding to the $N$-particle system an
cs.LG updates on arXiv.org

Convex-Hull-Neighborhood Smooth Dual Generalization: Controlling Local Correction Propagation in Offline RL

・arXiv:2608.03108v1 Announce Type: new Abstract: Offline reinforcement learning (offline RL) can benefit from nearby out-of-distribution (OOD) actions, but estimation errors at these actions may be amplified by bootstrapping. ・Existing regularization and local-generalization methods control either the admissible OOD region or the influence of generalized targets, often through separate mechanisms. ・We propose Convex Hul
MarkTechPost

CopilotKit Open Sources Channels SDK: An MIT Licensed Library That Runs Any AG-UI Agent Inside Slack And Microsoft Teams

・CopilotKit has published the Channels SDK, an MIT licensed library that runs an existing AG-UI agent inside Slack and Microsoft Teams. ・Version 0.5.0 ships five platform adapters and a documented runtime contract. ・This breakdown covers the verified deployment paths, the baseline requirements, and the one dependency that is easy to miss The post CopilotKit Open Sources Channels SDK: An MIT Licensed Library That Runs An
cs.LG updates on arXiv.org

CorePath: A Breast-Specialized Pathology Foundation Model for Core Needle Biopsy Diagnosis and Risk-Controlled Report Generation

・arXiv:2608.03079v1 Announce Type: cross Abstract: Breast core needle biopsy (CNB) is central to breast cancer diagnosis yet remains challenging because limited tissue sampling, lesion heterogeneity, and subtle morphologic overlap can obscure subtype distinctions. ・We developed CorePath, a breast-specialized multimodal pathology foundation model fine-tuned from PRISM using 7901 paired CNB whole-slide images and diagnos
WIRED

Cougars Lower the Risks of Car Crashes by Hunting Deer

・A new study shows how the big cats cause deer to move away from roads and deeper into the forest, where they pose less of a hazard to motorists.
cs.LG updates on arXiv.org

Crayotter: Learning Long-Horizon Video Editing Agents via Group-Relative Preference Backpropagation

・arXiv:2608.02694v1 Announce Type: cross Abstract: Long-horizon video editing agents receive final-product feedback only after many interdependent decisions. ・Yet editing quality is subjective, admits multiple valid solutions, and is not meaningfully calibrated across heterogeneous requests, making a global scalar objective both ambiguous and temporally uninformative. ・Our key observation is that fixing the request, mat
cs.LG updates on arXiv.org

Cross-Layer Interaction under Weight-Space Ablation: A Closed-Form Attention Jacobian Bound and a Test on a Real Pretrained Model

・arXiv:2608.03629v1 Announce Type: cross Abstract: A companion paper studies when activation patching and weight-space ablation agree, inside an idealized model where a conditional computation is carried additively through a residual stream. ・For the one composition in that model where two carriers are architecturally dependent, an attention head and its own layer's normalization-MLP composition, it derives an exact fi
cs.LG updates on arXiv.org

Cross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear Mapping for Prefill Reuse

・arXiv:2608.03893v1 Announce Type: new Abstract: Production deployments often swap between different-sized models in a family for cost-quality cascading, mid-conversation switching, and routing, and each swap forces the receiver to repay the prefill from scratch. ・We propose cross-model KV cache transfer, where the receiver reuses the source's KV cache, skipping prefill. ・We find that cross-model KV has substantial line
cs.LG updates on arXiv.org

CRS-Triage: Confidence- and Reliability-Aware Selective Triage under Incomplete Clinical Evidence

・arXiv:2608.03862v1 Announce Type: new Abstract: Emergency triage requires reliable decisions within a short time period. ・However, the available electronic health record (EHR) data, including structured data and clinical text, are often incomplete, unreliable, and inconsistent. ・This makes machine learning (ML)-based triage prediction more challenging, as existing ML models typically rely on complete and reliable EHR d
cs.LG updates on arXiv.org

CT-HEG: A Bidirectional, Timestamp-Attributed Event Graph for ICU In-Hospital Mortality Prediction - An Architectural Ablation Study

・arXiv:2608.02663v1 Announce Type: new Abstract: Accurate ICU mortality prediction requires modeling irregular clinical observations across heterogeneous entity types. ・Existing sequence models handle irregular sampling but ignore typed relational structure; existing graph models assume fixed-interval inputs. ・We introduce the Continuous-Time Heterogeneous EHR Graph (CT-HEG) schema and evaluate which architectural choic
cs.LG updates on arXiv.org

CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning

・arXiv:2608.03068v1 Announce Type: cross Abstract: Reinforcement learning (RL) has emerged as an effective method for enhancing the reasoning capabilities of large language models (LLMs). ・However, existing methods suffer from insufficient precision in feedback on generated answer trajectories and exhibit the phenomenon of problem difficulty drift. ・To address these challenges, we propose CVPO - Curriculum-guided Value-
cs.LG updates on arXiv.org

DAIF: A Data-Driven Intermediate Fusion Framework for Multimodal Supervised Learning via Approximate Message Passing

・arXiv:2608.02769v1 Announce Type: cross Abstract: Multimodal supervised learning seeks to leverage multiple heterogeneous data sources to improve predictive performance. ・A central challenge is determining the fusion granularity across modalities: over-integration may amplify noise while under-integration fails to exploit cross-modal dependence. ・Existing approaches rely on pre-specified fusion architectures, from earl
Hugging Face Papers

Decoding Children's Gait Behavior

Decoding Children's Gait Behavior
cs.LG updates on arXiv.org

Deep Divide-and-Reduce in Symbolic Regression

・arXiv:2608.02628v1 Announce Type: new Abstract: Symbolic regression (SR) is the task of discovering underlying patterns from data and representing them using mathematical expressions. ・Current machine learning approaches to SR often lack a profound understanding of the intrinsic mathematical and physical principles governing these expressions. ・While the pioneering AI Feynman method leverages the mathematical propertie
cs.LG updates on arXiv.org

DenialRAG: Single-Document RAG Poisoning via Embedded Parametric Denial

・arXiv:2608.02678v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) systems are vulnerable to corpus poisoning: an attacker who inserts a crafted document into the retrieval corpus can steer the underlying large language model (LLM) toward an attacker-chosen wrong answer. ・Prior single-document attacks typically avoid explicitly naming and refuting the correct answer inside the poisoned passage.
WIRED

Dermstore Coupons: 25% Off for August 2026

・Score major savings on premium skincare, hair care, and cosmetics with these verified Dermstore discount codes and rewards offers.
cs.LG updates on arXiv.org

Design Criteria for SGD Preconditioners: Local Conditioning, Noise Floors, and Basin Stability

・arXiv:2511.19716v3 Announce Type: replace-cross Abstract: Stochastic Gradient Descent (SGD) often slows in the late stage of training due to anisotropic curvature and gradient noise. ・We analyze preconditioned SGD in the geometry induced by a symmetric positive definite matrix $\mathbf{M}$, deriving bounds in which both the convergence rate and the stochastic noise floor are governed by $\mathbf{M}$-dependent quantiti
cs.LG updates on arXiv.org

Design-Time Optimization of Deep Neural Networks for Intermittent Learning on Microcontrollers

・arXiv:2608.03589v1 Announce Type: new Abstract: We present a method for designing deep neural networks (DNNs) for intermittent, energy-autonomous, on-device learning on microcontroller units (MCUs). ・In mobile applications where the energy can run out, e.g., when solar-powered, executing artificial intelligence (AI) faces a technical issue as learning can be interrupted at any time. ・Our approach combines a hardware-aw
cs.LG updates on arXiv.org

Designing a Good Virtual Node: Addressable and Cardinality-Preserving Global Memory for Message Passing Architectures

・arXiv:2608.02709v1 Announce Type: new Abstract: Virtual nodes give message-passing neural networks a simple global communication route, but the standard node--VN--node pipeline compresses the graph into one homogeneous state and broadcasts it identically to every node. ・Building on the Two-Radius analysis of Mishayev et al., we ask how auxiliary virtual memory can relieve this finite-capacity bottleneck without self-a
cs.LG updates on arXiv.org

Detecting high-frequency brain disorder signals using dynamic mode decomposition from EEG

・arXiv:2608.02804v1 Announce Type: cross Abstract: Recent studies have reported clearly identifiable dynamical changes in the high-frequency range of EEG signals recorded during specific stimuli, such as visual or auditory inputs, or in cases of brain disorders like epileptic seizures. ・In this study, we utilized Dynamic Mode Decomposition (DMD) to extract consistent and persistent dynamical changes in the high-frequen
WIRED

DHS Is Hiring Bounty Hunters to Find and Photograph Deported People’s Homes Abroad

・Homeland Security told immigrants that leaving the US would wipe out fines it claims they owe. ・Now it wants private investigators to find them in their home countries and collect.
cs.LG updates on arXiv.org

DiagLoop: A Counterfactual Data Flywheel with Stage-Localized Reinforcement for Diagnostic LLMs

・arXiv:2608.03674v1 Announce Type: new Abstract: Causal diagnostic models must explain how conclusions follow from evidence because diagnoses guide repairs and treatments. ・Yet serious cases are scarce, records rarely contain reasoning paths, and data transfer poorly across configurations, complicating local deployment. ・We present DiagLoop, a counterfactual data flywheel that converts codified physical relations or cli
cs.LG updates on arXiv.org

DIB-OD: Preserving the Invariant Core for Robust Heterogeneous Graph Adaptation via Decoupled Information Bottleneck and Online Distillation

・arXiv:2604.10882v3 Announce Type: replace Abstract: Graph pre-training can facilitate knowledge transfer across graph datasets, but severe structural and feature shifts may cause negative transfer and adaptation-induced overwriting of reusable knowledge. ・We propose DIB-OD, a heterogeneous graph adaptation framework that combines a Decoupled Information Bottleneck with Online Distillation. ・A multiview teacher first le
cs.LG updates on arXiv.org

Distilled Roads: Generalisable Road Network Extraction Across Sensors, Resolutions, and Region

・arXiv:2608.03407v1 Announce Type: cross Abstract: Road network segmentation from satellite imagery remains challenging due to large geographic variation in road appearance, occlusions, and domain shifts introduced by differing resolutions and sensors. ・Existing models, typically trained under narrow resolution--region combinations, generalise poorly to unseen environments such as rural settings, regions with distinct
cs.LG updates on arXiv.org

Divide-and-Conquer: Towards Generalizable Amortized Bayesian Inference for the Drift Diffusion Model

・arXiv:2608.03566v1 Announce Type: cross Abstract: The drift diffusion model (DDM) is a cornerstone of cognitive decision-making research. ・Although numerous estimation methods exist, researchers continue to seek inference approaches that are both fast and flexible across diverse study designs. ・Amortized Bayesian inference (ABI) can provide nearly instantaneous inference for complex stochastic models like the DDM, but
cs.LG updates on arXiv.org

Don't Walk the Line: Boundary Guidance for Filtered Generation

・arXiv:2510.11834v3 Announce Type: replace Abstract: Generative models are increasingly paired with safety classifiers that filter harmful or undesirable outputs. ・A common strategy is to fine-tune the generator to reduce the probability of being filtered, but this can be suboptimal: it often pushes the model toward producing samples near the classifier's decision boundary, increasing both false positives and false neg
cs.LG updates on arXiv.org

Double Descent in Gradient Boosting Decision Trees via Split-Candidate Scaling

・arXiv:2608.03111v1 Announce Type: new Abstract: Double descent is commonly studied by scaling an explicit capacity parameter, such as neural-network width. ・For gradient boosting decision trees (GBDTs), however, an analogous single-axis capacity parameter has not been established. ・We propose the number of split candidates as an operational capacity parameter for GBDTs.
cs.LG updates on arXiv.org

DP-MemView: A Memory Interface for Attribute-Level Transcript Privacy in Long-Term LLM Agents

・arXiv:2608.03130v1 Announce Type: cross Abstract: Long-term memory enables persistent personalization in LLM agents, but repeated memory-conditioned responses can cumulatively reveal protected attributes even when they are never stated explicitly. ・We formalize this threat as adaptive transcript privacy and introduce DP-MemView, a differentially private interface that privately selects public response-conditioning vie
cs.LG updates on arXiv.org

DRIFT: Derailing Denoising Trajectories of Flow-Matching VLAs with Adversarial Patch Attack

・arXiv:2608.03207v1 Announce Type: cross Abstract: Flow-matching vision-language-action (VLA) models such as pi0 generate robot actions by integrating a learned denoising velocity field, and have been reported to resist adversarial perturbations that readily fool autoregressive VLAs. ・We show that this robustness is largely illusory: it stems from prior attacks ignoring the multi-step denoising ODE. ・We introduce DRIFT
Zennの「機械学習」のフィード

Dropoutは「暗黙のアンサンブル」なのか。同じパラメータ予算で本物のアンサンブルと殴り合わせたら、相関係数0.13だった

・「dropoutは、指数個のサブネットワークを暗黙のうちにアンサンブルしているようなものだ」——machine learningを学び始めると、わりと早い段階でこの説明に出会う。Srivastavaらのdropout原論文(2014)にも、学習時にユニットをランダムに落とすことは 2^n 通りのサブネットワークの重み共有アンサンブルを訓練しているのと近い、という趣旨の記述がある。よく聞く話だし、直感的にも納得しやすい。ただ、これを自分の手で検証したことは一度もなかった。 ・言葉で終わらせず、実際に「本物の明示的アンサンブル」と「dropoutを使った1つの大きいネット」を、同じパラメータ予...
cs.LG updates on arXiv.org

Dual-domain U-Nets with embedded back projection operators for motion-resolved 4D CBCT reconstruction

・arXiv:2608.03430v1 Announce Type: cross Abstract: Four-dimensional cone beam CT (4D CBCT) is important for image-guided radiation therapy of thoracic cancers, but its use is limited by long scan times, causing high patient dose and motion/sparse-sampling artifacts. ・We propose a deep learning method for motion-resolved 4D CBCT reconstruction from conventional free-breathing scans, without a respiratory signal or expli
cs.LG updates on arXiv.org

Dynamically Allocating Evaluation Effort for Model Ranking

・arXiv:2608.03437v1 Announce Type: cross Abstract: While human evaluation is the gold standard in many NLP tasks, it suffers from prohibitive costs and poor scalability. ・When identifying top-performing models, typical evaluation protocols waste effort by exhaustively evaluating all models on the entire benchmark, a safe but inefficient approach. ・In this work, we formalize multi-model human evaluation as a best-arm ide
cs.LG updates on arXiv.org

E4GEN: Event-level Explainable Extreme-Enhanced Time-series Generation

・arXiv:2606.01634v2 Announce Type: replace Abstract: Generating realistic time series is essential for scientific research and real-world applications. ・However, existing methods often emphasize overall distributional fidelity while failing to faithfully capture extreme events. ・To advance existing research, we propose E4GEN, an explainable diffusion framework for extreme event-aware time-series generation.
cs.LG updates on arXiv.org

ED-DiT: Physics-Guided Diffusion Pretraining for Transferable Molecular Representations from Electron Density

・arXiv:2608.03260v1 Announce Type: new Abstract: Pretraining has shown strong potential for learning transferable representations, yet it remains underexplored for electron-density-based molecular learning. ・Electron density provides a continuous three-dimensional description of molecular electronic structure, capturing both local spatial patterns and global physical quantities. ・This raises a key question: can electron
cs.LG updates on arXiv.org

Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss

・arXiv:2608.03796v1 Announce Type: cross Abstract: Small language models are often the only option for deployment under tight latency, cost, and on-premises constraints, but they are rarely trained from scratch: a compressed model is usually recovered through knowledge distillation (KD). ・This recovery step largely decides the final quality, yet it is expensive. ・We present a practitioner's study of how to make distilla
cs.LG updates on arXiv.org

Efficient quantum-enhanced classical simulation for patches of quantum landscapes

・arXiv:2411.19896v2 Announce Type: replace-cross Abstract: Understanding the capabilities of classical simulation methods is key to identifying where quantum computers are advantageous. ・Not only does this ensure that quantum computers are used only where necessary, but also one can potentially identify subroutines that can be offloaded onto a classical device. ・In this work, we show that it is always possible to genera
cs.LG updates on arXiv.org

Efficient unsupervised domain adaptation via self-supervised vision transformer and synergistic cross-domain alignment

・arXiv:2407.21311v2 Announce Type: replace-cross Abstract: Unsupervised domain adaptation (UDA) aims to mitigate domain shift, where the distribution of labeled source data differs from that of unlabeled target data. ・Despite recent advances, existing methods often rely on fine-tuning large backbone models, which leads to high computational cost and limits scalability in resource-constrained environments. ・This limitati
cs.LG updates on arXiv.org

Enhancing Q-Value Updates in Deep Q-Learning via Successor-State Prediction

・arXiv:2511.03836v2 Announce Type: replace Abstract: Deep Q-Networks (DQNs) estimate future returns by learning from transitions sampled from a replay buffer. ・However, the target updates in DQN often rely on next states generated by actions from past, potentially suboptimal, policy. ・As a result, these states may not provide informative learning signals, causing high variance into the update process.
cs.LG updates on arXiv.org

Enhancing Tabular Learners with Context-Aware Semantic Embeddings

・arXiv:2608.03565v1 Announce Type: cross Abstract: While modern tabular learners excel at capturing statistical patterns, they frequently operate in a semantic vacuum, treating textual features as discrete symbols, ignoring the rich semantics inherent in feature names or cell entries. ・We propose CASE (Context-Aware Semantic Embeddings), a novel framework that bridges the gap between the semantic understanding of Large
cs.LG updates on arXiv.org

Enhancing VLM Reward Models Through Structure-Aware Fine-Tuning

・arXiv:2608.03875v1 Announce Type: new Abstract: Designing effective reward functions remains a major bottleneck in Reinforcement Learning (RL). ・Recent work uses large foundation Vision-Language Models (VLMs) as reward models, computing text-observation similarity to bypass manual reward engineering. ・Although promising, these rewards are often noisy and unreliable, limiting their direct utility during deployment.
cs.LG updates on arXiv.org

Estimating Tail Risks in Language Model Output Distributions

・arXiv:2604.22167v3 Announce Type: replace Abstract: Language models are increasingly capable and are being rapidly deployed on a population-level scale. ・As a result, the safety of these models is increasingly high-stakes. ・Fortunately, advances in alignment have significantly reduced the likelihood of harmful model outputs.
cs.LG updates on arXiv.org

Evading Chain-of-Thought Monitoring Through Model Poisoning

・arXiv:2608.02820v1 Announce Type: cross Abstract: Chain-of-thought (CoT) monitoring is an increasingly important component of AI safety stacks but relies on the assumption that a model's reasoning trace is informative about its actions. ・This work studies the limits of CoT monitoring through the lens of model poisoning. ・We demonstrate that backdoors can be implanted into reasoning models to elicit an attacker-chosen b
cs.LG updates on arXiv.org

Evaluation Blindness: How Silent Measurement Failures Corrupt AI Systems from Training to Deployment

・arXiv:2608.02786v1 Announce Type: new Abstract: AI systems can fail silently. ・The failure propagates through training loops, evaluation pipelines, and production monitoring stacks until downstream harm makes it visible. ・This paper introduces evaluation blindness: a measurement function M exhibits evaluation blindness with respect to failure class F when it produces readings indistinguishable from a healthy state whil
cs.LG updates on arXiv.org

Explainable AI for the EU Right to Explanation: A Systematic Review of the Law-XAI Translation Gap

・arXiv:2608.02699v1 Announce Type: cross Abstract: When algorithms make or influence consequential decisions---about loan eligibility, hiring, or healthcare---EU law grants affected individuals a Right to Explanation. ・Yet whether (and how) Explainable AI (XAI) can satisfy this right in practice remains poorly understood, with direct implications for individuals' ability to contest automated decisions that affect their
Hugging Face Papers

ExplainBench: Evaluating Code Explanations from Agents

ExplainBench: Evaluating Code Explanations from Agents
cs.LG updates on arXiv.org

Exploiting Separability in Multi-Scale Grey-Box Bayesian Optimization

・arXiv:2608.03045v1 Announce Type: new Abstract: We consider grey-box optimization problems where the decision variables naturally partition into black-box variables (as arguments to an expensive black-box function) and white-box variables, governed by a set of explicit, closed-form equations that also depend on the output of the black-box function. ・We exploit this separability through a bilevel reformulation: an oute
cs.LG updates on arXiv.org

FedCARE: A Multi-Objective Personalised Federated Learning Framework for Smart Healthcare

・arXiv:2608.03498v1 Announce Type: new Abstract: Federated Learning (FL) enables collaborative model training across distributed healthcare institutions without centralising sensitive patient data. ・However, real-world healthcare federations are often characterised not only by non-IID data, but also by heterogeneous clinical objectives and partially overlapping feature spaces. ・Different hospitals may optimise distinct
cs.LG updates on arXiv.org

FedCritic-MIMO: Communication-Efficient Serverless Federated Critic Learning for Massive-MIMO Resource Control in Open and Disaggregated 6G RANs

・arXiv:2608.03852v1 Announce Type: new Abstract: This paper proposes FedCritic-MIMO, a communication-efficient serverless federated multi-agent reinforcement learning framework for AI-native resource control across independently deployable cell-level controllers in open and disaggregated 6G RANs. ・Controllers share no trainer, retain local actors and personalized critic components, and exchange only compatible shared c
cs.LG updates on arXiv.org

Federated generative event models for tokenized electronic health records

・arXiv:2608.02939v1 Announce Type: new Abstract: Electronic health record foundation models are limited by institutionally siloed data and substantial performance degradation under cross-site transfer. ・We evaluated federated training of tokenized generative event models (GEMs) across 122,251 intensive care hospitalizations from three independent health systems harmonized to the Common Longitudinal ICU Data Format.
cs.LG updates on arXiv.org

FedRings: A Scalable and Topology-Aware Federated Learning Framework for LEO Satellite Constellations

・arXiv:2608.03436v1 Announce Type: cross Abstract: Federated learning over low Earth orbit (LEO) satellite networks is limited by frequent link changes, short contact times, and a highly dynamic topology, making centralized or synchronized training inefficient and hard to scale. ・To address this, we propose FedRings, a decentralized framework that organizes satellites into ring-based communication structures.
cs.LG updates on arXiv.org

Field Aware Agent Skill Retrieval

・arXiv:2608.02880v1 Announce Type: cross Abstract: As lifelong learning agents accumulate lifelong growing skill banks, retrieving the correct skill becomes an increasingly important bottleneck. ・Most current skill retrieval methods treat each skill as one flat document by concatenating fields such as the name, description, and body. ・However, skills are naturally structured, multi-field objects, where each field provid
cs.LG updates on arXiv.org

FinVerse: Financial Time-Series Benchmark

・arXiv:2608.03259v1 Announce Type: new Abstract: As time-series foundation models have emerged, the need for benchmarks that can evaluate their forecasting ability in meaningful ways has become increasingly important. ・Existing time-series forecasting benchmarks provide useful standardized comparisons, but they often evaluate heterogeneous series with uniform error-based metrics. ・Strong performance under such metrics d
cs.LG updates on arXiv.org

FLARE: Diffusion for Hybrid Language Model

・arXiv:2606.01774v2 Announce Type: replace Abstract: Autoregressive (AR) large language models (LLMs) have achieved broad practical success, but sequential decoding remains a key bottleneck for low-latency deployment. ・Recent efficient-inference work has progressed along two axes: reducing the cost of each model invocation through efficient architectures, and reducing serial decoding steps through parallel generation.
cs.LG updates on arXiv.org

Forecasting Revenue with its Customer-Base Drivers: When and Why Coordination Helps

・arXiv:2608.02911v1 Announce Type: new Abstract: Revenue forecasts guide acquisition budgets, demand planning, and customer-based valuations, yet an aggregate forecast does not show whether change reflects acquisition, repeat purchasing, spending per order, or offsetting movements. ・Using weekly transaction panels for 966 companies in 25 industries, the authors develop the Customer-Based Multi-task Transformer (CBMT),
WIRED

Foreo Discount Codes and Deals: Up to 50% Off

・Save on Foreo favorites, including LUNA cleansing brushes, BEAR microcurrent devices, and masks and accessories to level up your daily skincare routine at home.
cs.LG updates on arXiv.org

FOUND-AF: Benchmarking ECG Foundation Models for Atrial Fibrillation Detection

・arXiv:2608.03597v1 Announce Type: cross Abstract: Atrial fibrillation (AF) is the most common sustained cardiac arrhythmia and is associated with increased risks of stroke, heart failure, and mortality. ・Recent ECG foundation models offer transferable representations for automated AF detection. ・However, their relative effectiveness remains unclear because existing studies use different datasets, preprocessing procedur
cs.LG updates on arXiv.org

Foundations of Equivariant Deep Learning: Unifying Graph and Sheaf Neural Networks

・arXiv:2607.03798v2 Announce Type: replace Abstract: Symmetry is everywhere in nature and society. ・Geometric deep learning builds architectures respecting group symmetries, whereas topological deep learning organizes computation through cells, incidence relations, and local-to-global structure. ・In this paper, we extend geometric deep learning beyond simple group actions and unify it with topological deep learning.
cs.LG updates on arXiv.org

From Generator to Embedder: Harnessing Innate Abilities of Multimodal LLMs via Building Zero-Shot Discriminative Embedding Model

・arXiv:2508.00955v3 Announce Type: replace Abstract: Adapting generative Multimodal Large Language Models (MLLMs) into universal embedding models typically demands resource-intensive contrastive pre-training, while traditional hard negative mining methods suffer from severe false negative contamination. ・In this paper, we propose a highly data-efficient framework that bypasses extensive pre-training to build a robust m
WIRED

Gene-Edited Puppies Will Melt Your Heart—but Won’t Trigger Your Allergies

・A startup has created beagles without the gene that causes runny noses and watery eyes for allergy sufferers.
cs.LG updates on arXiv.org

GENESIS: Towards Explainable Causal Discovery

・arXiv:2608.03868v1 Announce Type: new Abstract: Causal Discovery (CD) from observational data faces two fundamental challenges. ・First, purely statistical methods often lack the power to resolve structural ambiguities in low-sample regimes. ・Second, although LLM-assisted hybrid approaches improve structure recovery through semantic reasoning, the influence of that reasoning on individual edge decisions remains largely
cs.LG updates on arXiv.org

Geo-Embed: Towards Unified Multimodal Embeddings for Urban Understanding

・arXiv:2608.03826v1 Announce Type: cross Abstract: Geospatial and urban applications increasingly require models to compare heterogeneous evidence across street-view imagery, remote-sensing observations, text descriptions, region proposals, and temporal change cues. ・However, existing multimodal embedding models and benchmarks are still largely designed and evaluated around general-purpose image-text matching, leaving
cs.LG updates on arXiv.org

GeoID-PINN: Identifiability-Aware Regional Epidemic Inference with Geographic Coupling

・arXiv:2608.02633v1 Announce Type: new Abstract: Regional surveillance data reflect local transmission, reporting, seeding, and external infection pressure, which are difficult to identify separately. ・We introduce GeoID-PINN, a physics-informed neural network (PINN) for susceptible-infectious-recovered-deceased (SIRD) dynamics. ・The model represents spatial dependence with a row-stochastic source-composition matrix who
cs.LG updates on arXiv.org

GLOBE: Trajectory-Aligned Gradient Matching with Structured SparseOptimization for Coreset Selection

・arXiv:2608.02690v1 Announce Type: new Abstract: On-device training of deep neural networks is fundamentally constrained by the computational and memory costs of large-scale datasets. ・Coreset selection offers a practical solution by retaining only a compact subset of real training samples. ・However, existing gradient-based methods commonly rely on gradients computed at a single model snapshot and employ greedy or pursu
The Verge

Google just announced a major shakeup of its top AI leadership

・Google is making some significant AI leadership changes, including a major shift for Google DeepMind leader Demis Hassabis. ・Hassabis will become the chair of Google DeepMind and the chief scientist at Alphabet, CEO Sundar Pichai announced on Wednesday. ・Hassabis will continue to lead Alphabet's Isomorphic Labs, which aims to use AI to develop drugs.
WIRED

Google’s Top AI Brains Are Leaving to Launch Discovery Loop

・Jeff Dean and other high-profile Google executives have founded Discovery Loop, a startup that will seek AI-powered breakthroughs in everything from drug discovery to chip design.
ITmedia NEWS 最新記事一覧

Googleパスワードマネジャーに脆弱性 パスキーで守られたアカウント乗っ取る3つの攻撃手法、米パロアルトが警鐘

・セキュリティ企業の米Palo Alto Networksの研究チーム「Unit 42」が、こんな攻撃に警鐘を鳴らしている。研究チームが8月3日に発表したところによれば、手口の異なる3種類の手法が存在し、一部は成功例もあるという。
cs.LG updates on arXiv.org

GoT-CD: Graph-of-Thoughts Causal Discovery and the Fragility of Post-hoc Path-Specific Fairness Audits

・arXiv:2608.02877v1 Announce Type: new Abstract: Causal discovery recovers directed structure from observational data and is increasingly used in clinical settings to support mechanism reasoning and fairness audits of predictive models. ・Path-specific counterfactual fairness asks whether a protected attribute influences an outcome through illegitimate pathways, but these estimands are defined relative to a supplied cau
Hugging Face Papers

GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience

GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience
cs.LG updates on arXiv.org

HAPEns: Hardware-Aware Post-Hoc Ensembling for Tabular Data

・arXiv:2603.10582v2 Announce Type: replace Abstract: Ensembling is commonly used in machine learning on tabular data to boost predictive performance and robustness, but larger ensembles often lead to increased hardware demand. ・We introduce HAPEns, a post-hoc ensembling method that explicitly balances accuracy against hardware efficiency. ・Inspired by multi-objective and quality diversity optimization, HAPEns constructs
AI News & Artificial Intelligence | TechCrunch

Hark previews its browser use agent for completing tasks

・Hark claims that its browser use agent is faster and cheaper than competition.
cs.LG updates on arXiv.org

HERO: Hierarchical Evidential Reasoning Optimization for Radiology Report Generation via Reason-then-Summarize

・arXiv:2601.03321v3 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have substantially advanced Radiology Report Generation (RRG), yet aligning them through reinforcement learning (RL) remains challenging due to heterogeneous medical supervision. ・Vanilla Group Relative Policy Optimization (GRPO) assigns uniform credit across the entire generation, leading to segment interference, token diluti
cs.LG updates on arXiv.org

Heteroscedasticity of Denoising Score Matching with Generalised Smooth Noise

・arXiv:2508.01597v2 Announce Type: replace Abstract: Score Matching (SM) is a powerful framework for estimating the log-density derivatives of a distribution without calculating its normalizing constants. ・This capability has made it a cornerstone across multiple domains, from classical sta- tistical estimation and energy-based models to modern diffusion-based generative models. ・In practice, these models rely almost ex
cs.LG updates on arXiv.org

How Many Labels Are Enough? ALDA: Active Learning Deployment Advisor for Medical Image Classification

・arXiv:2608.03511v1 Announce Type: cross Abstract: Active learning (AL) promises to reduce the cost of medical imaging projects by lowering the number of clinical labels required. ・However, practical deployment requires committing to a sampling strategy before the full annotation budget is spent, and choosing the wrong strategy can increase rather than decrease costs. ・We propose Active-Learning Deployment Advisor (ALDA
Hugging Face Papers

Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing

Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing
cs.LG updates on arXiv.org

Hybrid LLM-Augmented Reinforcement Learning Agents for Complex Sequential Decision Tasks

・arXiv:2608.03502v1 Announce Type: cross Abstract: Large Language Models (LLMs) have recently shown strong capabilities in reasoning, planning, and tool-use, enabling new forms of autonomous agents. ・However, LLM-based agents struggle with long-horizon sequential decision tasks that require precise action optimization and environment interaction. ・Reinforcement Learning (RL), while effective for sequential control, ofte
cs.LG updates on arXiv.org

Improved Quantum Algorithms for Reinforcement Learning Under a Generative Model

・arXiv:2608.02826v1 Announce Type: cross Abstract: Reinforcement learning is a subfield of machine learning that studies how an agent interacts with an environment in order to extract as large a reward as possible. ・A standard approach to study such interaction is through Markov Decision Processes (MDPs) and the task of choosing an optimal policy --- a function that tells the agent which action to take. ・In this work, w
cs.LG updates on arXiv.org

Improving Reproducibility in Evaluation through Multi-Level Annotator Modeling

・arXiv:2605.13801v2 Announce Type: replace Abstract: As generative AI models such as large language models (LLMs) become more pervasive, ensuring the safety, robustness, and overall trustworthiness of these systems is paramount. ・However, AI is currently facing a reproducibility crisis driven by unreliable evaluations and unrepeatable experimental results. ・While human raters are often used to assess models for utility
cs.LG updates on arXiv.org

Improving Sample Efficiency in Multi-Agent Reinforcement Learning for Simulated Football Games via Exploration

・arXiv:2503.13077v2 Announce Type: replace Abstract: Multi-agent reinforcement learning has shown promise in learning cooperative behaviors in team-based environments. ・However, such methods often demand extensive training time, which inhibits their application for game-AI in standard game development. ・For instance, the state-of-the-art method TiZero takes 40 days to train high-quality policies for a football environme
cs.LG updates on arXiv.org

In-Context Molecular Property Prediction with LLMs: A Blinding Study on Memorization and Knowledge Conflicts

・arXiv:2603.25857v3 Announce Type: replace Abstract: The capabilities of large language models (LLMs) have expanded beyond natural language processing to scientific prediction tasks, including molecular property prediction. ・However, their effectiveness in in-context learning remains ambiguous, particularly given the potential for training data contamination in widely used benchmarks. ・This paper investigates whether LL
cs.LG updates on arXiv.org

In-Context Pure Exploration in Continuous Decision Spaces

・arXiv:2602.17976v2 Announce Type: replace Abstract: In active sequential testing, also termed pure exploration, a learner is tasked with the goal to adaptively acquire information so as to identify an unknown ground-truth hypothesis with as few queries as possible. ・This problem has several motivating applications, including Best-Arm Identification (BAI) in bandits, where actions index hypotheses, and generalized sear
cs.LG updates on arXiv.org

Information-Geometric Forward Policy Training in GFlowNets

・arXiv:2608.03967v1 Announce Type: cross Abstract: Generative Flow Networks (GFlowNets) have emerged as a flexible framework for amortised inference over discrete and mixed discrete-continuous objects, requiring only an unnormalised target density specified through a reward. ・In this work, we formulate forward-policy training in GFlowNets through the information geometry of the induced trajectory sampler. ・Treating the
cs.LG updates on arXiv.org

Inverted Detection and Control in Steering Vectors

・arXiv:2608.02957v1 Announce Type: new Abstract: Steering vectors (SVs) are widely used to influence the expression of concepts (e.g., truthfulness) in large language model outputs. ・A key assumption underpinning SVs is that they are linearly discriminative with respect to the concept: representations of texts that exhibit the concept are more aligned with the SV than those that do not, motivating shifts along the posi
cs.LG updates on arXiv.org

IPPRO: Importance-based Pruning with PRojective Offset for Magnitude-indifferent Structural Pruning

・arXiv:2507.14171v3 Announce Type: replace Abstract: Importance-based structured pruning overwhelmingly relies on filter magnitude. ・This proxy is fundamentally flawed: due to scale invariance, functionally identical filters can receive arbitrarily different importance scores under rescaling. ・We propose IPPRO (Importance-based Pruning with PROjective Offset), a scale-invariant pruning framework grounded in projective g
cs.LG updates on arXiv.org

Joint Affine Spectral Shaping: Coupling Weight and Bias Updates Beyond Weight-Only Muon

・arXiv:2608.02991v1 Announce Type: new Abstract: Matrix spectral optimizers reshape weight-update spectra but usually delegate vector-valued biases to a separate optimizer. ・We study whether this separation is neutral. ・We formulate each affine layer as a joint momentum matrix $A=[M_W,\alpha m_b]$ and apply a capped regularized-inverse spectral map to the complete matrix, producing both the weight and physical bias upda
Hugging Face Papers

JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion

JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion
cs.LG updates on arXiv.org

KernelBrain: Coarse-to-Fine, Budget-Aware Search for Agentic GPU Kernel Optimization

・arXiv:2608.02611v1 Announce Type: cross Abstract: Automating GPU kernel optimization remains difficult in practice: generated variants can violate correctness constraints, runtime measurements are noisy, and search often stalls early. ・We present a practical optimization agent that combines LLM-guided mutation, adaptive resource allocation, policy-gated evaluation, and profiler-informed diagnosis. ・The system screens m
Hugging Face Papers

Knowledge-Geometry Decoupling: Refreshable Pretrained Transfer for Streaming Recommendation

Knowledge-Geometry Decoupling: Refreshable Pretrained Transfer for Streaming Recommendation
cs.LG updates on arXiv.org

LAEF: A Lead-Agnostic ECG Foundation Model Towards Point-of-Care Diagnostics

・arXiv:2608.03690v1 Announce Type: new Abstract: Point-of-care cardiac devices such as smartwatches and handheld ECG recorders typically capture 1--2 leads, yet existing ECG foundation models are architecturally constrained to fixed 12-lead inputs, degrading or failing under these reduced configurations. ・We introduce LAEF (Lead-Agnostic ECG Foundation), a 7M-parameter ECG foundation model that can natively process any
cs.LG updates on arXiv.org

Latent Reward Registers for Diffusion Preference Alignment

・arXiv:2608.03929v1 Announce Type: new Abstract: Aligning diffusion models with human preferences usually relies on a sparse terminal reward evaluated on the final generated samples, presenting a severe temporal credit-assignment challenge across the multi-step denoising process. ・We propose Latent Reward Registers, a mechanism that estimates terminal preference directly from intermediate noisy latents by prepending le
cs.LG updates on arXiv.org

Learning and Clustering on Temporal Graphs: Principles, Primitives, and Pooling

・arXiv:2608.03696v1 Announce Type: new Abstract: This work focuses on the problem of learning on temporal graphs, with particular emphasis on the task of clustering: obtaining coarse-grained representations by aggregating information from nodes, edges, and temporal dynamics - a task related to pooling in machine learning on graphs, or community detection in network science. ・Although graph neural networks reach state-o
cs.LG updates on arXiv.org

Learning Clinical-Trial Strategy: Offline Policy Training for Decision Agents

・arXiv:2608.03606v1 Announce Type: cross Abstract: Clinical development is sequential decision-making under uncertainty, where a sponsor must plan a portfolio of experiments from heterogeneous evidence. ・We study this setting by framing oncology clinical development as an offline decision-making problem in which an agent predicts the next six-month trial portfolio of an oncology drug program from information available
cs.LG updates on arXiv.org

Learning Molecular Representations from Cellular Phenotypes with Structure Preservation

・arXiv:2608.02688v1 Announce Type: new Abstract: Phenotypic drug discovery enables the discovery of functional relationships between molecular structures and cellular responses. ・However, existing multimodal representation learning methods often optimize cross-modal alignment without considering the intrinsic organization of chemical space, resulting in distorted molecular representations and loss of structural informa
cs.LG updates on arXiv.org

Less Traffic, Better Outcomes: Competition-Aware Request Dispatch in Real-Time Ad Exchanges

・arXiv:2608.03705v1 Announce Type: cross Abstract: Real-time bidding (RTB) ad exchanges typically forward nearly all incoming requests to demand-side platforms (DSPs), even though only a small fraction receive bids. ・This over-distribution weakens auction outcomes: DSPs throttle participation under compute and budget constraints, reducing the effective use of limited bidding capacity. ・We present a competition-aware req
cs.LG updates on arXiv.org

Lightweight Chunk Selection for Mobile Retrieval-Augmented Generation

・arXiv:2608.03148v1 Announce Type: new Abstract: RAG improves the factual grounding of LLM by incorporating external knowledge, but deploying RAG on mobile and edge devices remains challenging because retrieved context increases computation and memory. ・A direct way to reduce this cost is to retain only one retrieved chunk before generation, but the top-ranked retrieved chunk is not always the most evidence-supporting
Hugging Face Papers

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models
cs.LG updates on arXiv.org

LLM Serving in the Wild: An Empirical Study of Frameworks, Methods, and System Designs

・arXiv:2608.03036v1 Announce Type: cross Abstract: Large Language Models (LLMs) are integrated into software systems and AI services, making efficient LLM serving a concern for software engineering. ・Serving LLMs is challenging because inference requires computation, memory, GPU resources, and execution while maintaining latency and throughput. ・Although prior research has proposed LLM inference, optimization, and servi
cs.LG updates on arXiv.org

LLM-Derived Priors for Thompson Sampling in Cold-Start Comment Recommendation

・arXiv:2608.03382v1 Announce Type: cross Abstract: Multi-armed bandit algorithms, especially Thompson sampling, are widely used in online recommendation. ・Despite their ability to adapt from online feedback, these methods often suffer from cold-start limitations when newly introduced arms have little or no interaction history. ・In our setting, the candidate arms are user-generated textual comments, whose semantic conten
cs.LG updates on arXiv.org

LLMs Can Annotate Attribution Graphs

・arXiv:2608.02632v1 Announce Type: new Abstract: Circuit tracing is an exciting technique for revealing the internal computation of language models, but it requires a time-intensive manual step of grouping individual features or MLP neurons into supernodes. ・We present a simple pipeline for automating this step: directly presenting feature descriptions to a language model that groups them into supernodes. ・Using automat
Zennの「大規模言語モデル」のフィード

LLMが「本文の無い記事」の要約をでっち上げていた話 ─ プロンプト修正だけで安心しなかった理由

・はじめに 個人開発している「AgentPick」は、複数ソースから取得したニュース記事にLLMで要約を付けています。あるとき、Hacker Newsの記事に付く要約が「読んでいないのに読んだていで書かれた文章」になっていることに気づきました。 ・この記事はその原因と、なぜプロンプトを直すだけでは不十分だと判断したかの記録です。同じ問題に当てはめて確認できるチェックリスト形式の整理はQiita版にまとめています。 ・何が起きていたか 要約生成のシステムプロンプトには、もともとこう書いていました。
#LLMタグ

LLMを軽くする数学 量子化とは何か?その狙いと本質

・Transformerを理解すると、次に気になるのはこういう疑問だ。 ・「なぜ最近のLLMは、ノートPCやスマートフォンでも動くのか」 続きをみる
LLMタグが付けられた新着記事 - Qiita

LLM推論パラメータ入門——Temperature・Top-K・Top-Pの仕組みと実践調参ガイド

・LLM推論パラメータ入門——Temperature・Top-K・Top-Pの仕組みと実践調参ガイド ChatGPTやClaudeなどのLLM(大規模言語モデル)を使っていて、「回答が真面目すぎて面白みがない」「コードを出力させたいのに毎回違う書き方をして困る」「高温度に...
Zennの「大規模言語モデル」のフィード

LLM推論パラメータ入門——Temperature・Top-K・Top-Pの仕組みと実践調参ガイド

・LLM推論パラメータ入門——Temperature・Top-K・Top-Pの仕組みと実践調参ガイド ChatGPTやClaudeなどのLLM(大規模言語モデル)を使っていて、「回答が真面目すぎて面白みがない」「コードを出力させたいのに毎回違う書き方をして困る」「高温度にしすぎて支離滅裂な回答になった」といった経験はありませんか? LLMの出力品質や性格は、プロンプトの工夫だけでなく、推論時に指定する**サンプリングパラメータ(Temperature、Top-K、Top-P)**の調参(チューニング)によって劇的に変化します。 ・この記事では、LLMが次トークンを選択する数理メカニズム...
cs.LG updates on arXiv.org

Logic Before Language: Pre-pretraining on Formal Derivations Fosters Skill Acquisition and Compressibility

・arXiv:2608.03930v1 Announce Type: cross Abstract: Pre-pretraining language models (LMs) on symbolic data can accelerate and improve natural language acquisition. ・However, existing pre-pretraining tasks, such as Dyck and procedural algorithms, rely on narrow primitives that fail to capture the expressive capacity of natural language. ・Moreover, prior studies remain restricted to relatively small token budgets, offering
Zennの「大規模言語モデル」のフィード

LongTraceRL:エージェント検索軌跡と量規報酬で長文脈推論を学習

・LongTraceRL:エージェント検索軌跡と量規報酬で長文脈推論を学習 TL;DR 検索エージェントの軌跡から2層のダストラクタ(Tier-1/High, Tier-2/Low)を構築し、従来のランダムサンプリングでは不可能なレベルの高難度訓練データを実現 エンティティレベルのRubric報酬で推論チェーン上の各ホップを細粒度に評価。Positive-Only戦略でreward hackingを防止 Qwen3-4Bで平均+5.7点、AA-LCRでは+8.6点。訓練データはたった2,815例ながら、他データセット(最大18,870例)を圧倒 コード・データ・モデルをすべ...
cs.LG updates on arXiv.org

M-GATE: Multilingual Grammar, Accuracy in Translation, and Efficiency Benchmark for Large Language Models

・arXiv:2608.03803v1 Announce Type: cross Abstract: Multilingual language models are deployed across a hundred or more languages, yet most benchmarks test whether a model can perform a task _in_ a language rather than whether it commands the language itself, conflating fluency with proficiency. ・We introduce M-GATE (Multilingual Grammar, Accuracy in Translation, and Efficiency), a benchmark of linguistic proficiency spa
AI News & Artificial Intelligence | TechCrunch

MacPaw taps Liquid AI to offer on-device inference to devs building for its app store

・MacPaw is building a local version of its AI assistant Eney using Liquid AI's models.
WIRED

MAGA Is In Turmoil Over Tucker Carlson’s Possible 2028 Presidential Bid

・Joe Kent says that he and other anti–Donald Trump MAGA figures have formed a new movement. ・They want Tucker Carlson to run for the presidency in 2028.
cs.LG updates on arXiv.org

Maglev: Sliding Recurrent Memory

・arXiv:2608.02870v1 Announce Type: new Abstract: We introduce \ours{}, a recurrent Transformer architecture with fixed-size memory that generalizes sliding-window attention while remaining parallelizable during training. ・\ours{} consists of two coupled models: a prefiller $Q$, which leverages full attention\footnote{In practice, we use interleaved full and sliding-window attention for $Q$, as this yields stronger perf
cs.LG updates on arXiv.org

Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain

・arXiv:2510.05159v5 Announce Type: replace-cross Abstract: While finetuning AI agents on interaction data -- such as web browsing or tool use -- improves their capabilities, it also introduces critical security vulnerabilities within the agentic AI supply chain. ・We show that adversaries can effectively poison the data collection pipeline at multiple stages to embed hard-to-detect backdoors that, when triggered, cause
cs.LG updates on arXiv.org

MambaTS: Improved Selective State Space Models for Long-term Time Series Forecasting

・arXiv:2405.16440v2 Announce Type: replace Abstract: In recent years, Transformers have become the de-facto architecture for long-term time series forecasting (LTSF), yet they face challenges associated with the self-attention mechanism, including quadratic complexity and permutation-invariant bias. ・This raises an important question: \emph{do we truly need self-attention to model long-range dependencies in LTSF?} To a
Zennの「大規模言語モデル」のフィード

MCPが業務自動化を“安全に”進めるための実装チェックリスト

・この記事で分かること 2026-07-30時点で、AI/LLMの主戦場が「モデル性能」から「エージェント運用基盤」へ移った理由 MCP・業務自動化・安全性の3点を、今どの順番で追うべきか Python開発基盤ではなぜ uv / Ruff / Polars と Astral が注目されているのか AIニュースを読むときに「モデル」より「運用基盤」を優先して追う方法 結論、今日のニュース群で最も重要なのは、LLM単体の性能競争よりも「LLMをどう業務に接続し、安全に運用するか」に関心が移っている点です。 ・この流れを最も端的に示しているのが、[finance.biggo.jp]「W...
cs.LG updates on arXiv.org

Measuring Explainer Stability via Attribution Separability

・arXiv:2608.02697v1 Announce Type: new Abstract: Attribution methods (AMs) assign an importance score to each feature and are widely adopted to explain black-box models. ・However, most methods can produce variable attribution scores due to stochastic components in their definition. ・In this paper, we propose a distribution-based framework to capture the stability of attribution scores.
cs.LG updates on arXiv.org

Mechanism of Task-oriented Information Removal in In-context Learning

・arXiv:2509.21012v4 Announce Type: replace Abstract: In-context Learning (ICL) is an emerging few-shot learning paradigm based on modern Language Models (LMs), yet its inner mechanism remains unclear. ・In this paper, we investigate the mechanism through a novel perspective of information removal. ・Specifically, we demonstrate that in the zero-shot scenario, LMs encode queries into non-selective representations in hidden
WIRED

Medicube Coupon Code: 40% Off for August 2026

・Upgrade your K-beauty routine with these active Medicube promo codes. ・Save on Age-R devices, serums, and masks with student discounts and referral rewards.
cs.LG updates on arXiv.org

MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale

・arXiv:2608.02613v1 Announce Type: cross Abstract: Edge-deployed personal memory assistants must handle private interpersonal conversations on-device with open-weight models. ・Yet, existing memory benchmarks often under-test the combination of activity-dense interaction, ego-centric perspective, and coherent multi-session worlds. ・MemArena fills these gaps with a single-world conversational benchmark built with its MASi
Hugging Face Papers

MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations

MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations
WIRED

Meta Quest Promo Codes and Coupons for August 2026

・Experience cutting-edge VR and save up to 20% with coupons for the latest games, Meta Quest 3, Ray-Ban AI glasses, and more deals.
WIRED

Meta Ran Ads That Contained AI-Generated Child Sexual Abuse Imagery

・More than 50 offending image and video ads were published across Facebook, Instagram, Messenger, or Threads, according to Meta’s ad library data. ・Some ran as recently as this week.
cs.LG updates on arXiv.org

Micro-Segmentation Anomaly Detection in Zero-Trust Software-Defined Network Fabrics

・arXiv:2608.02627v1 Announce Type: cross Abstract: Zero Trust Architecture (ZTA) principles need rigorous network segmentation and ongoing verification to reduce implicit trust and lateral threat propagation. ・This paper investigates anomaly detection in software-defined networking (SDN) systems by micro-segmentation, using deep learning models to detect harmful actions that evade traditional coarse-grained monitoring.
cs.LG updates on arXiv.org

Minimax-Optimal Semiparametric Contextual Dynamic Pricing with Multimodal Revenue

・arXiv:2608.03142v1 Announce Type: cross Abstract: We study contextual dynamic pricing with arbitrary covariate sequences and bounded, possibly nonbinary purchase quantities. ・Demand follows a semiparametric surplus-index model with an unknown linear valuation parameter and an unknown H\"older-smooth response. ・We impose neither concavity nor strong unimodality on revenue and allow nonunique optimal prices.
Hugging Face Papers

MiniWorld: Democratizing the Training of Video World Models from Scratch

MiniWorld: Democratizing the Training of Video World Models from Scratch
cs.LG updates on arXiv.org

MoECa: Aligning Feature Reuse with Expert Decomposition in Diffusion Transformers

・arXiv:2606.15615v2 Announce Type: replace Abstract: Diffusion Transformers with Mixture-of-Experts (DiT-MoE) improve model capacity under sparse activation, but diffusion inference is still bottlenecked by redundant computation across timesteps. ・Existing caching methods mainly operate at the token level, which becomes suboptimal in DiT-MoE because each token update is internally decomposed into multiple routed expert
Hugging Face Papers

Multi-Task Multi-Frame Visual Piano Transcription

Multi-Task Multi-Frame Visual Piano Transcription
cs.LG updates on arXiv.org

Multimodal Auto-regressive Transformer Surrogate for Modeling Variable Operations and Quantifying Uncertainty in Geological Carbon Storage

・arXiv:2608.02629v1 Announce Type: new Abstract: The use of variable well perforation and injection strategies can improve the efficiency of geological carbon storage operations. ・We develop a new multimodal auto-regressive transformer surrogate to model these operations under geological uncertainty. ・A modified SEAM CO2 geomodel, which involves a faulted system with three stacked aquifers, is considered.
cs.LG updates on arXiv.org

Muon Meets Mamba: Spectral Optimization for State Space Models

・arXiv:2608.03941v1 Announce Type: new Abstract: Muon is a recent optimizer that orthogonalizes the update to each weight matrix with a Newton-Schulz iteration, which performs steepest descent under the spectral norm. ・Almost all the evidence for it comes from Transformer models, and its behavior on state-space models is largely unreported. ・We compare Muon with AdamW on Mamba-2 130M under a controlled protocol that var
cs.LG updates on arXiv.org

NANQ: Noise-Floor-Aware Mixed-Precision Non-Uniform Quantization for Analog Compute-in-Memory

・arXiv:2608.02700v1 Announce Type: new Abstract: Analog compute-in-memory (CIM) enables energy-efficient neural network inference, but device variation and read noise can severely degrade low-bit quantized models. ・Existing CIM-oriented quantization methods mainly minimize ideal quantization error, ignoring the hardware noise floor and thus causing inefficient precision allocation. ・We propose NANQ, a noise-aware mixed-
ITmedia NEWS 最新記事一覧

NASA「確率100%」 SpaceXの使用済みロケットが月に衝突か 

・NASAは8月4日(現地時間)、米SpaceXのロケット「Falcon 9」の使用済み上段が、5日未明に月へ衝突する見通しだと発表した。衝突地点は、月面のアインシュタイン・クレーターとベル・クレーター付近。地球への危険はないとしている。
cs.LG updates on arXiv.org

Neural network realization of binary refinement iterates via a two-chart atlas selector

・arXiv:2608.02624v1 Announce Type: cross Abstract: Refinement operators generate many functions used in wavelet constructions, subdivision schemes, and geometric modeling. ・Their finite iterates can develop rapidly increasing numbers of linear pieces, making them a natural test case for the expressive power of deep neural networks. ・Earlier work showed that, for scalar binary refinement with a finitely supported mask, e
cs.LG updates on arXiv.org

Neural Networks with Local Converging Inputs for Efficient Options Pricing Models

・arXiv:2608.02778v1 Announce Type: new Abstract: We present a novel application of Neural Networks with Local Converging Inputs (NNLCI) to improve the efficiency of existing numerical methods for pricing multi-asset options. ・The most concise input format for NNLCI has been introduced, offering substantial convenience and efficiency. ・NNLCI uses a neural network to locally correct solutions from a coarse mesh and a refi
WIRED

No One Can Afford to Make ‘Myst’ Games Anymore

・“Myst” sold over 6 million copies in the ’90s. ・Its latest spinoff couldn’t shore up any funding, pointing to an existential crisis among double-A games.
cs.LG updates on arXiv.org

Noise-Aware Shrinkage for Differentially Private Zeroth-Order Fine-Tuning of Large Language Models

・arXiv:2608.03277v1 Announce Type: new Abstract: Differentially private zeroth-order optimization (DP-ZO) enables memory-efficient private fine-tuning of large language models using only forward evaluations. ・Existing aggregation-based DP-ZO methods reconstruct model updates at a fixed scale, ignoring that the strength of useful signals varies throughout training. ・Consequently, noise-dominated updates may receive exces
cs.LG updates on arXiv.org

NOMADD: Numerical Optimization of Models Adapting to Data Drift

・arXiv:2608.02845v1 Announce Type: new Abstract: Tabular model performance degrades when feature distributions change over time or the relationship between features and outcome variables change over time, known as data drift and concept drift, respectively. ・These issues are challenging to mitigate in real time because labeled data may not be immediately available, or re-training a model could be impractical.
cs.LG updates on arXiv.org

NPMixer: Hierarchical Neighboring Patch Mixing for Time Series Forecasting

・arXiv:2605.07476v2 Announce Type: replace Abstract: Multivariate time series forecasting remains a challenge due to the complexity of local temporal dynamics and global dependencies across multiple variables. ・In this paper, we propose \textbf{N}eighboring \textbf{P}atching \textbf{Mixer} (\textbf{NPMixer}), a hierarchical architecture featuring a Learnable Stationary Wavelet Transform that adaptively learns filter co
MarkTechPost

NVIDIA Releases Alpamayo 2 Super: A 34B Open Vision-Language-Action Model for Robotaxis and Autonomous Driving Under OpenMDW-1.1

・NVIDIA released Alpamayo 2 Super, a 34B vision-language-action model for autonomous driving, under OpenMDW-1.1 — a permissive license covering fine-tuning, derivatives and commercial redistribution. ・It pairs a 32B Cosmos 3 Super Reasoner backbone with a 2.3B diffusion action decoder, scores 79.2 on LingoQA, and emits trajectories, Chain-of-Causation traces, meta-actions, auto-labels and grounded VQA from a single pas
機械学習タグが付けられた新着記事 - Qiita

olympicAthletesデータセットの紹介

・たまたま見つけたデータセットを、なんとなくTwitter(現X)に投稿したら意外と反響があったので、少し詳しく紹介してみようと思います。 ・このデータセットは1896年のアテネ大会から2026年のミラノ・コルティナ大会までの近代オリンピックにおける参加選手と参加競技、メ...
cs.LG updates on arXiv.org

Omega-S: A Functional Resilience Index for LLM Fine-Tuning

・arXiv:2608.03887v1 Announce Type: new Abstract: Fine-tuning a large language model on new data degrades what it previously learned. ・We present Omega-S, a drop-in penalty computed from the weight matrix alone: it needs no previous-task data, no Fisher matrix and no stored copy of the old weights. ・It is three lines in an existing training loop and adds under 4% to the cost of a step.
Hugging Face Papers

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models
cs.LG updates on arXiv.org

On the Implicit Flatness Bias of Sharpness-Aware Minimization: A Linear Stability Analysis with Quantitative Hyperparameter Bounds

・arXiv:2608.03197v1 Announce Type: new Abstract: Sharpness-Aware Minimization (SAM) improves generalization by seeking parameters whose loss is robust to local adversarial perturbations, but the quantitative mechanism underlying its implicit bias toward flat minima remains unclear. ・In particular, the perturbation radius $\rho$ is typically treated as an isolated tuning parameter, despite defining the neighborhood in w
cs.LG updates on arXiv.org

On the Limits of Layer Pruning for Generative Reasoning in Large Language Models

・arXiv:2602.01997v3 Announce Type: replace Abstract: Recent work has shown that layer pruning can effectively compress large language models (LLMs) while retaining strong performance on classification benchmarks, often with little or no finetuning. ・In contrast, generative reasoning tasks, such as GSM8K and HumanEval\textsuperscript{+}, exhibit substantially weaker recovery. ・We show that beyond surface-level text degra
cs.LG updates on arXiv.org

On the Performance of Malware Detection Classifiers Using Hardware Performance Counters

・arXiv:2608.02671v1 Announce Type: cross Abstract: Malware detection using Hardware Performance Counters (HPC) has emerged as a promising solution to improve the security of computing systems as a complement to antivirus software. ・Hardware-based malware detectors (HMD) use Machine Learning (ML) classifiers to detect malicious application patterns. ・The inputs to ML classifiers are low-level performance features known a
cs.LG updates on arXiv.org

One-Point Contraction: Erasing Representational Separability toward Irreversible Deep Forgetting

・arXiv:2507.07754v3 Announce Type: replace Abstract: Machine unlearning is usually evaluated by what the classifier outputs: forget-set accuracy, confidence, membership-inference scores. ・We show that this is not enough. ・Across 14 representative unlearning methods on CIFAR-10 and SVHN, a single linear map fitted on a held-out calibration set, with no access to the forgotten data, reverses the unlearning in seconds and
cs.LG updates on arXiv.org

Online Shift Detection and Conformal Adaptation for Deployed Safety Classifiers

・arXiv:2606.11949v3 Announce Type: replace Abstract: Reasoning models deployed as safety monitors exhibit a systematic vulnerability: reasoning-token budget starvation. ・Adversarial inputs require $3.3\times$ more reasoning tokens than benign inputs to produce valid safety scores ($T_{50,\text{adv}}{=}154$ vs. ・$T_{50,\text{benign}}{=}46$ for o3), so low-budget deployments silently starve the monitor on exactly the inpu
#LLMタグ

OpenAI Astra、数十年来の数学難問を$2,000で解く。快挙の“主役”は証明検証器Leanだった

OpenAI Astra、数十年来の数学難問を$2,000で解く。快挙の“主役”は証明検証器Leanだった
#LLMタグ

OpenAI o3 Model Explained: The Next Frontier in AI Reasoning

OpenAI o3 Model Explained: The Next Frontier in AI Reasoning
cs.LG updates on arXiv.org

Operationally Feasible Synthetic Power-Grid Scenarios via Learning the AC-Operable Joint Distribution

・arXiv:2608.03878v1 Announce Type: new Abstract: Synthetic power-grid scenarios are essential for planning, resilience assessment, contingency analysis, and data-driven power-system applications. ・Recent synthetic grid generation methods have improved structural realism and operational feasibility by incorporating engineering knowledge through post-generation validation, optimization, or physics-aware generation.
cs.LG updates on arXiv.org

Output-Aware Rotation for INT2 KV-Cache Quantization

・arXiv:2608.02691v1 Announce Type: new Abstract: The key-value (KV) cache has become a major memory and bandwidth bottleneck in long-context large language model inference, making ultra-low-bit quantization increasingly important. ・However, existing rotation-based INT2 methods optimize cache statistics or proxy errors before the complete attention readout, even though the model is ultimately affected by the error propa
cs.LG updates on arXiv.org

Paired Recipient-based Evaluation of Survival Prediction for Deceased Donor Kidney Transplants

・arXiv:2608.03017v1 Announce Type: new Abstract: There has been significant interest in using machine learning algorithms to predict kidney transplant outcomes, such as the number of years until a graft inevitably fails. ・These prediction algorithms could possibly be used for pre-transplant donor-recipient matching to identify more compatible donors and recipients and thus improve post-transplant outcomes. ・In this stud
cs.LG updates on arXiv.org

Particle-based Generalised Stochastic Optimisation

・arXiv:2608.02844v1 Announce Type: cross Abstract: We develop a class of diffusion-based stochastic particle optimisation methods for loss functions with intractable gradients. ・Specifically, we consider problems in which the loss gradient is an integral with respect to a parameter-dependent distribution, a structure that includes training generative models, fine-tuning, and learning latent-variable models.
Hugging Face Papers

PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents

PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents
cs.LG updates on arXiv.org

Patient-centered data science: an integrative framework for evaluating and predicting clinical outcomes in the digital health era

・arXiv:2408.02677v2 Announce Type: replace Abstract: This study proposes a novel, integrative framework for patient-centered data science in the digital health era. ・We developed a multidimensional model that combines traditional clinical data with patient-reported outcomes, social determinants of health, and multi-omic data to create comprehensive digital patient representations. ・Our framework employs a multi-agent ar
cs.LG updates on arXiv.org

PatTree: a novel approach for automated creation of multimodal, graph-based patient representations for medical classification tasks

・arXiv:2608.02692v1 Announce Type: new Abstract: Access to holistic, multimodal data improves the performance of Artificial Intelligence (AI) in medical classification tasks compared to utilizing single modalities or data sources. ・However, the inherent heterogeneity and complexity of clinical real-world data pose significant challenges to structured data analysis and AI application. ・This heterogeneity includes missing
Hugging Face Papers

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning
WIRED

Petlibro Offers: 60% Off in August 2026

・Save on Petlibro essentials, including automatic feeders, water fountains, and accessories to keep cats and dogs fed, hydrated, and comfortable every day.
cs.LG updates on arXiv.org

Pin Once, Swap Light: Subspace-Aligned Centroid-Residual Training for Efficient Ultra-LoRA Serving

・arXiv:2608.03579v1 Announce Type: new Abstract: Modern multi-tenant Low-Rank Adapters (LoRAs) serving systems concurrently host tens to hundreds of LoRA adapters. ・Though powerful, this introduces a critical system dilemma between serving efficiency and task performance: higher-rank adapters generally achieve better downstream task performance, but their GPU VRAM footprint and Host-to-Device PCIe swapping overhead sev
cs.LG updates on arXiv.org

PLAN: Parallel Liquid-Inspired Approximation Network for Efficient Representation Learning in Flexible Job Shop Scheduling

・arXiv:2608.03041v1 Announce Type: new Abstract: Deep reinforcement learning (DRL) approaches for flexible job shop scheduling (FJSP) heavily rely on attention-centric architectures to achieve state-of-the-art performance. ・However, these models suffer from excessive parameter counts and prohibitive inference latency as problem scales expand. ・While liquid neural networks (LNNs) offer a parameter-efficient alternative f
cs.LG updates on arXiv.org

POEM: Phase-Aware $\mathrm{SO}(2)$ Feature Rotation for Time Series Forecasting Under Periodicity Drift

・arXiv:2608.03630v1 Announce Type: new Abstract: Deep learning has advanced time series forecasting, but periodicity drift, in which cycle timing and phase vary over time, remains a challenging problem. ・Existing methods predominantly model these sequences on fixed time grids, suffering from a limited ability to accommodate phase-related variation. ・To address this limitation, we propose \textbf{POEM}, a phase-aware for
cs.LG updates on arXiv.org

Population-Robust Feature Selection via Generalized Welfare Optimization

・arXiv:2608.02887v1 Announce Type: new Abstract: Choosing which features to collect is a deployment decision: the same limited questionnaire, test panel, or sensor set may need to serve several heterogeneous populations. ・Standard feature-selection methods typically optimize for one large population, while existing robust approaches tend to learn one shared model for every population. ・We introduce PopFS, a method for l
Hugging Face Papers

PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs

PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs
cs.LG updates on arXiv.org

Prediction-Enhanced Monte Carlo: A Machine Learning View on Control Variate

・arXiv:2412.11257v4 Announce Type: replace-cross Abstract: For many complex simulation tasks spanning areas such as healthcare, engineering, and finance, Monte Carlo (MC) methods are invaluable due to their unbiased estimates and precise error quantification. ・Nevertheless, Monte Carlo simulations often become computationally prohibitive, especially for nested, multi-level, or path-dependent evaluations lacking effecti
cs.LG updates on arXiv.org

Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety

・arXiv:2608.02617v1 Announce Type: cross Abstract: We evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in large language model (LLM) evaluation using expert feedback from MOOVE (Massive Open Online Validation and Evaluation), a clinician-led platform collecting blinded pairwise preferences alongside multi-criterion rubric ratings. ・Clinicians assign scores on a discrete $[-2,
cs.LG updates on arXiv.org

Prescribed-Basis Coefficient-to-Coefficient Neural Operator for Partial Differential Equations

・arXiv:2510.10350v3 Announce Type: replace-cross Abstract: Operator learning provides a data-driven approach to approximating solution operators of partial differential equations, but its effectiveness depends strongly on how input and output functions are represented. ・Point-value representations can make the trainable map mesh-dependent and high-dimensional; snapshot-based POD/PCA reductions require aligned data and
cs.LG updates on arXiv.org

PRISM: Powerful Time Series to Image (TS2I) Representations for Multivariate Anomaly Detection

・arXiv:2608.03926v1 Announce Type: new Abstract: Time series anomaly detection (TSAD) underpins applications in predictive maintenance, finance, and cloud computing, however performance remains sensitive to representation choices, especially in multivariate settings. ・While transforming time series into images has shown success in forecasting and classification, it remains unclear how multivariate, high-dimensional ser
cs.LG updates on arXiv.org

PRISMA: Improving the Accuracy-Latency Frontier of Diffusion-based PDE Solvers Using Physics-Informed Spectral Attention

・arXiv:2512.01370v2 Announce Type: replace Abstract: Diffusion-based solvers for partial differential equations (PDEs) are often bottle-necked by slow gradient-based test-time optimization routines that use PDE residuals for loss guidance. ・They additionally suffer from optimization instabilities and are unable to dynamically adapt their inference scheme in the presence of noisy PDE residuals. ・To address these limitati
cs.LG updates on arXiv.org

PRIVEE: Privacy-Preserving Vertical Federated Learning Against Feature Inference Attacks

・arXiv:2512.12840v2 Announce Type: replace Abstract: Vertical Federated Learning (VFL) enables collaborative model training across organizations that share common user samples but hold disjoint feature spaces. ・Despite its potential, VFL is susceptible to feature inference attacks, in which adversarial parties exploit shared confidence scores (prediction probabilities) during inference to reconstruct private input feat
cs.LG updates on arXiv.org

Provably Learning Multi-Head Attention with Queries

・arXiv:2608.03294v1 Announce Type: new Abstract: We study the problem of learning multi-head softmax attention from black-box input-output access. ・The learner may query arbitrary real-valued token sequences and observe only the scalar output at the final token. ・Recent work gives an algorithm using $O(d^2)$ value queries to recover the single-head parameters $(W,v)$.
WIRED

Purple Promo Codes and Deals: Up to 30% Off

・On the hunt for the perfect mattress or pillow? ・Save on your soon-to-be favorite brand, Purple, with these Purple coupons and deals.
Hugging Face Papers

Push-Wiper: Toward General-Purpose Robotic Cleaning across Varied Stains and Surfaces with Segmented Pushing Trajectories

Push-Wiper: Toward General-Purpose Robotic Cleaning across Varied Stains and Surfaces with Segmented Pushing Trajectories
cs.LG updates on arXiv.org

Quality Control Algorithms for Pattern Counting

・arXiv:2608.03439v1 Announce Type: cross Abstract: In recent work, Marcussen, Rubinfeld, and Sudan introduced the notion of quality control problems, which aim to capture the task of determining if a given input is truly random. ・Formally, their goal is to accept typical inputs from the specified distribution while rejecting every input whose value of a specified statistic is far from the distributional baseline.
cs.LG updates on arXiv.org

Quantization Effects on Biomedical LLM Reliability

・arXiv:2608.03854v1 Announce Type: new Abstract: When decoder language models are used as classifiers, predicted class probabilities depend on implementation choices, including the prompt template, verbalizer (label-to-token mapping), and scoring rule, that are rarely treated as experimental variables. ・We present a controlled evaluation of three Mistral-7B variants (Base, BioMistral, and Instruct) on PubMed RCT senten
Hugging Face Papers

Quo Vadis, World Modeling?

Quo Vadis, World Modeling?
#LLMタグ

RAGかファインチューニングか?LLMアプリ構築で迷わないための判断指針

・最近、LLMを用いたアプリケーション開発において大きな論点となっている RAG vs Fine-Tuning: How to Choose the Right Approach for Your LLM App という記事を読み、開発の現場で直面しがちな「どちらの手法を選ぶべきか」という問題について自分なりに整理してみました。
The Verge

Reddit is introducing a new moderator: AI

・Reddit is enlisting AI to help moderate new subreddits - and eventually the rest of site. ・The company is introducing automated moderation tools that rely on LLMs to help mods manage their communities, and it's expanding who can use those tools today ahead of a full launch later this year. ・The company calls the suite of tools "Rules Hub," and the tools let mods decide what rules should be automatically enforced and wh
cs.LG updates on arXiv.org

Representing Random Utility Choice Models with Neural Networks

・arXiv:2207.12877v3 Announce Type: replace Abstract: Motivated by the successes of deep learning, we propose a class of neural network-based discrete choice models, called RUMnets, inspired by the random utility maximization (RUM) framework. ・This model formulates the agents' random utility function using a sample average approximation. ・We show that RUMnets sharply approximate the class of RUM discrete choice models: a
Hugging Face Papers

RestoreKV: Recovering Full-Cache Behavior Under Aggressive Query-Agnostic KV Cache Eviction

RestoreKV: Recovering Full-Cache Behavior Under Aggressive Query-Agnostic KV Cache Eviction
cs.LG updates on arXiv.org

Resume Means Resume: A Machine-Checked Conformance Contract for Checkpoint, Interrupt, and Resume Semantics in Workflow Persistence Layers

・arXiv:2608.03836v1 Announce Type: new Abstract: A framework that persists execution state so a run can be interrupted, survive a crash, and continue must decide what a resume means for effects that already fired. ・Five widely deployed agent workflow frameworks answer differently, none exposes a machine-checkable contract, and behavior violates even the fragments they state. ・The RESUME CONTRACT states six properties ov
cs.LG updates on arXiv.org

Revisiting TD Target Aggregation under Uncertainty in Q-Learning

・arXiv:2608.03069v1 Announce Type: new Abstract: Deep Q-Networks (DQNs) learn value functions through bootstrapped temporal-difference updates, where future returns are approximated using a greedy maximization over next-state action values. ・While effective, this aggregation rule is inherently sensitive to estimation noise: when Q-values are uncertain, the maximization operator deterministically favors the largest esti
cs.LG updates on arXiv.org

Rex: A Family of Reversible Exponential (Stochastic) Runge-Kutta Solvers

・arXiv:2502.08834v5 Announce Type: replace Abstract: Deep generative models based on neural differential equations have become state-of-the-art for many generation tasks. ・These models rely on ODE/SDE solvers that integrate from a prior distribution to the data distribution; in many applications it is also highly desirable to integrate in the inverse direction. ・Standard solvers, however, accumulate discretization error
The Verge

Ring upgraded its peephole doorbell camera to 2K

・Ring has debuted a new version of its smart doorbell camera that's designed to be easily installed as a replacement for a door's peephole without drilling or running wires. ・The Peephole Cam 2K is a replacement for the brand's Door View Cam that first debuted in 2019 and, alongside a sleeker design, it features a bump from 1080P to 2K video recording, some tracking upgrades, and a cheaper price. ・It's available for pre
cs.LG updates on arXiv.org

Robust Biharmonic Skinning Using Geometric Fields

・arXiv:2406.00238v3 Announce Type: replace-cross Abstract: Bounded bihramonic weights are a popular tool used to rig and deform characters for animation, to compute reduced-order simulations, and to define feature descriptors for geometry processing. ・They necessitate tetrahedralizing the volume bounded by the surface, introducing the possibility of meshing artifacts or tetrahedralization failure. ・We introduce a mesh-f
cs.LG updates on arXiv.org

Robust Counterfactual Policy Optimisation via Nondeterministic Causal Models

・arXiv:2608.02893v1 Announce Type: new Abstract: Counterfactual inference approaches for sequential decision-making typically assume deterministic causal models, where all randomness stems from latent variables. ・However, Markov Decision Processes (MDPs) are inherently stochastic. ・We address this by formalising counterfactual policy optimisation under probabilistic nondeterministic causal models, which properly separat
cs.LG updates on arXiv.org

Robust General Utility for Reinforcement Learning

・arXiv:2608.03562v1 Announce Type: new Abstract: Reinforcement learning (RL) with general utility extends classic RL by optimizing an arbitrary utility functional of the policy-induced occupancy measure, thereby enabling a broader range of applications. ・However, previous work on general utility RL typically assumes the evaluation utility is fixed and correctly specified. ・In practice, the utility used at deployment can
cs.LG updates on arXiv.org

Robust Low-Tubal-Rank Tensor Completion under Cross-Concentrated Sampling

・arXiv:2608.03928v1 Announce Type: cross Abstract: Tensor cross-concentrated sampling (t-CCS) bridges entrywise sampling and t-CUR slice-wise sampling by observing entries only within selected horizontal and lateral slices. ・Existing t-CCS completion methods, however, assume that the observations are free of gross corruption. ・In this work, we study robust recovery of a third-order low-tubal-rank tensor from partial t-C
The Verge

Rogue AI agents created fake online identities in another hacking attempt

・Yet more rogue AI agents from OpenAI and Anthropic have been caught attempting to hack real targets online without permission. ・The discoveries add to a growing list of previously unknown incidents that have alarmed AI safety experts and intensified pressure for greater oversight of frontier systems. ・According to a report from the UK's AI Security Institute, which evaluates frontier models from top AI labs before they
cs.LG updates on arXiv.org

Rubrics as Privileged Information for Open-Ended Generation

・arXiv:2608.02948v1 Announce Type: new Abstract: On-policy self-distillation (OPSD), where a single model acts as both student and teacher with different contexts, has shown promise in verifiable domains like math, where hard privileged information (PI) in the form of ground-truth answers structurally constrains valid continuations. ・We extend OPSD to open-ended generation using soft PI in the form of rubrics that guid
cs.LG updates on arXiv.org

SAKI: Score-Aware Low-Rank Key Indexing for Long-Context KV Retrieval

・arXiv:2608.03228v1 Announce Type: new Abstract: Existing low rank KV cache methods preserve either model weights or key variance, neither of which directly reflects the attention scores used during inference. ・We derive the expected attention score distortion caused by rank r key compression and show that it yields a covariance weighted low rank objective. ・Under a margin condition, controlling this distortion also imp
cs.LG updates on arXiv.org

Scaling an Autoregressive Transformer for Single-Cell Generation

・arXiv:2608.02961v1 Announce Type: new Abstract: We study a self-supervised generation task for single-cell gene expression vectors: given a set of vectors from a cell type, we aim to generate additional gene expression vectors of that cell type. ・For this task we characterize both the biological fidelity of the generated gene expression vectors and the scaling behavior of the pretraining loss. ・The model is a causal tr
cs.LG updates on arXiv.org

Schedule-Informed Temporal Fusion Forecasting of Hourly Airport Security-Checkpoint Throughput

・arXiv:2608.02950v1 Announce Type: new Abstract: Checkpoint staffing requires accurate forecasts of when screening demand will occur, yet flight schedules record departure times rather than passenger arrival times at security checkpoints. ・This study develops a framework that converts known flight schedules into temporally aligned signals for forecasting hourly checkpoint throughput. ・Using 2023-2024 Transportation Secu
cs.LG updates on arXiv.org

ScoreField: Neural Inverse Scattering with Score-Based Generative Priors

・arXiv:2608.02937v1 Announce Type: cross Abstract: Designing an effective electromagnetic inverse-scattering solver requires faithful enforcement of nonlinear full-wave physics together with an expressive prior on the unknown permittivity contrast. ・We propose ScoreField, a neural inverse scattering framework that integrates coupled implicit neural representations (INRs) with a pretrained score-based generative prior.
cs.LG updates on arXiv.org

Sedentary Behavior Classification for Wearable Sensors with a CNN-BiLSTM Model

・arXiv:2608.02946v1 Announce Type: new Abstract: Accurate detection of sedentary behavior is important for studying health risks related to prolonged sitting, but posture-based classification remains challenging with wearable sensors, especially at the wrist. ・We study whether a deep learning model trained on hip-worn accelerometer data can transfer to wrist-worn accelerometer data for sitting versus non-sitting classi
Zennの「大規模言語モデル」のフィード

SeedRealtimeを動画編集に入れる前に、会話より先に決めたいこと

・自分がリアルタイムの音声・映像AIを見るとき、最初に気になるのは「どれだけ自然に会話できるか」ではありません。 ・編集の判断を、あとから人に渡せる形で返してくれるかです。 ・[2026年8月5日に公開されたSeedRealtime]は、音声・映像・テキストをまとめて扱い、連続する情報にリアルタイムで応答するモデルとして発表されました。画面を見せながら話せるAIは、動画制作にも近づいてきています。
cs.LG updates on arXiv.org

Sensitivity, Causality, and Repair Dissociate: A Layer-Wise Analysis of Perturbation Robustness and Its Scaling

・arXiv:2608.03842v1 Announce Type: cross Abstract: When a language model fails on surface-perturbed input (typos, OCR noise, homophones), "which layer is responsible" has three natural operationalizations: where representations diverge most (sensitivity), where restoring clean activations recovers the prediction (causality), and where a small adapter can repair the damage (compensatory capacity) - and we show these th
cs.LG updates on arXiv.org

SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs

・arXiv:2608.03573v1 Announce Type: cross Abstract: Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) exhibit fundamentally different behaviors in enhancing multi-task reasoning for large language models (LLMs). ・Our preliminary experiments revealed a phenomenon: SFT suffers from severe task conflicts under multi-stage training, whereas RL enables stable coexistence across diverse tasks. ・Empirically, we trace
cs.LG updates on arXiv.org

ShielDroid: A Hybrid Approach Integrating Machine and Deep Learning for Android Malware Detection

・arXiv:2608.03250v1 Announce Type: cross Abstract: The rapid advancement of modern technology has led to a significant increase in the use of smart devices, such as smartphones and tablets, resulting in the widespread adoption of mobile applications. ・Although applications are required to undergo malware screening before being published on official app stores, many malicious applications successfully evade detection by
AI News & Artificial Intelligence | TechCrunch

Shopify says AI search is driving more traffic and sales, not replacing Google

・Shopify says AI isn’t cannibalizing search traffic the way it has for publishers. ・Instead, AI-driven traffic and orders to Shopify stores tripled year over year in Q2.
cs.LG updates on arXiv.org

Shorter Reasoning, Earlier Answers? An Evaluation of Reasoning Interfaces

・arXiv:2608.03401v1 Announce Type: new Abstract: Large language models often reason at length before answering, increasing cost and latency. ・Prompts and trained settings can shorten this reasoning, but a shorter trace may only show that the model stopped sooner. ・Here, we evaluate paired runs of the same question at matched reasoning horizons across 198 GPQA Diamond and 500 MMLU-Pro questions.
cs.LG updates on arXiv.org

Should the Boundary Term Be Learned in Reflected Diffusion? Conormal Trace and Reflection Masking

・arXiv:2608.03469v1 Announce Type: cross Abstract: We study score learning for reflected diffusion on bounded domains. ・Reflection keeps trajectories feasible but does not ensure that the learned score satisfies the boundary behavior implied by the forward process. ・With implicit score matching, integration by parts leaves a boundary term, and we show that it depends on one scalar at each boundary point: the diffusion-
cs.LG updates on arXiv.org

Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations

・arXiv:2605.28149v3 Announce Type: replace Abstract: Sparse Autoencoders (SAEs) extract interpretable features from Large Language Model activations, but standard variants enforce non-negative latents, so a bidirectional semantic axis (e.g., "pressure too high" vs. ・"pressure too low") must be split across two latents, wasting dictionary capacity on anticorrelated features. ・We propose the Sign-Aware Gated SAE (SA-GSAE)
cs.LG updates on arXiv.org

Simulation-free and finite-time diffusion model

・arXiv:2608.03117v1 Announce Type: new Abstract: The performance of generative diffusion models is determined by the choice of the reference diffusion process connecting the empirical and prior distributions. ・Conventional approaches typically trade off simulation-free training against finite-time generation. ・We propose a framework for designing the reference process that achieves both simultaneously.
cs.LG updates on arXiv.org

SimulRAG: Simulator-based RAG for Grounding LLMs in Long-form Scientific QA

・arXiv:2509.25459v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) show promise in generating long-form scientific explanations that synthesize evidence and connect multiple factors. ・However, in long-form scientific question answering, LLMs often hallucinate, producing unsupported or inconsistent claims. ・Retrieval-Augmented Generation (RAG) improves trustworthiness by grounding generation in exter
Hugging Face Papers

SkillJack: Persistent Skill Backdoors in Self-Evolving Agents

SkillJack: Persistent Skill Backdoors in Self-Evolving Agents
cs.LG updates on arXiv.org

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation

・arXiv:2608.03092v1 Announce Type: new Abstract: We aim to improve model performance in multi-reward reinforcement learning training process. ・Existing Group reward-Decoupled Normalization Policy Optimization (GDPO) has mitigated the issue of reward signals masking one another during direct scalarization by normalizing each reward dimension separately before aggregation. ・However, our experiments show that GDPO still st
cs.LG updates on arXiv.org

Socially Grounded Agentic AI: Coordinating Plural Perspectives through Social Theory

・arXiv:2608.03910v1 Announce Type: cross Abstract: As AI systems are deployed across increasingly diverse social contexts, alignment can no longer be framed as the optimization of a single, unified set of values. ・Instead, systems must be able to recognize, represent, and respond to multiple legitimate perspectives. ・This has led to growing interest in pluralistic alignment, which seeks to move beyond one-size-fits-all
cs.LG updates on arXiv.org

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling

・arXiv:2608.02951v1 Announce Type: new Abstract: Preference-based reinforcement learning (PbRL) for general stochastic MDPs often requires training a reward model. ・Existing reward-model-free methods are either restricted to bandits or deterministic MDPs, such as DPO or P3O, or use zeroth-order, gradient-free optimization, which in general exhibits a slower convergence rate than gradient-based algorithms. ・Furthermore,
The Verge

SpaceX is barely Space and mostly X

・Privatize the profit, socialize the losses? ・| Image: Cath Virginia / The Verge, Getty Images Once, I had some questions about why SpaceX, Elon Musk's healthiest company, acquired xAI, his sickliest one. ・Now I have some questions about why we're calling the whole thing SpaceX.
cs.LG updates on arXiv.org

Sparse Weight Decomposition for Efficient Circuit Extraction

・arXiv:2608.03913v1 Announce Type: new Abstract: Dense pretrained transformers do not naturally expose interpretable units for circuit extraction. ・Existing approaches obtain such units by learning auxiliary sparse representations or training sparse models, incurring substantial additional computation while potentially introducing a fidelity gap between the representation being analyzed and the original pretrained mode
cs.LG updates on arXiv.org

Sphere Retraction Normalizations

・arXiv:2608.02668v1 Announce Type: new Abstract: Residual connections are the de facto mechanism for training deep neural networks stably. ・Geodesic Normalization (GeoNorm) recasts them on a Riemannian manifold, orthogonalizing each layer output against the current hidden state and applying the resulting update through the Riemannian exponential map. ・Every hidden state thus keeps a constant $\ell_{2}$-norm, confining t
cs.LG updates on arXiv.org

SphUnc: Hyperspherical Uncertainty Decomposition and Causal Identification via Information Geometry

・arXiv:2603.01168v3 Announce Type: replace Abstract: Reliable decision-making in complex multi-agent systems requires calibrated predictions and interpretable uncertainty. ・We introduce SphUnc, a unified framework combining hyperspherical representation learning with structural causal modeling. ・The model maps features to unit hypersphere latents using von Mises-Fisher distributions, decomposing uncertainty into epistem
WIRED

Sportsman's Warehouse Promo Code: Save in August 2026

・Whether you are hunting for firearms, camping supplies, or boating gear, use these Sportsman’s Warehouse coupons to maximize your savings in August 2026.
cs.LG updates on arXiv.org

SRAP: SVD-Refined Adversarial Perturbations for Imperceptible Face-Swap Defense

・arXiv:2608.03395v1 Announce Type: cross Abstract: Deepfake technologies pose increasing threats to facial privacy and identity security, motivating proactive defenses that protect facial images before misuse. ・Although adversarial perturbations generated by projected gradient descent (PGD) can disrupt the identity representations used by face-swapping models, their visual quality is degraded by two characteristics: pe
Hugging Face Papers

ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts

ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts
cs.LG updates on arXiv.org

Stochastic Saddle Avoidance Beyond Unit Excitation and Smoothness: A Pathwise Lyapunov-Perron Framework

・arXiv:2608.03001v1 Announce Type: cross Abstract: Unit excitation (UE) is a common assumption in stochastic saddle avoidance: the stochastic error must have a uniformly positive component along every direction, in expectation. ・This condition gives a direct way to rule out convergence to strict saddles, but it also oversimplifies the actual noise structure, and does not match many stochastic optimization regimes.
cs.LG updates on arXiv.org

Stop Replacing Noise with Noise: Two-Source Reliability Assessment for Label Correction and Sample Reweighting in Label-Noise Learning

・arXiv:2608.03432v1 Announce Type: new Abstract: Refurbishment-based noisy-label learning mixes an observed label with a model-derived pseudo target, typically using one sample-wise cleanliness score to control both branches. ・This creates a hidden coupling: reducing trust in the observed label automatically increases trust in the pseudo target. ・We show that this complementarity can replace one unreliable signal with a
cs.LG updates on arXiv.org

STREAM-VAE: Dual-Path Routing for Slow and Fast Dynamics in Vehicle Telemetry Anomaly Detection

・arXiv:2511.15339v3 Announce Type: replace Abstract: Automotive telemetry data exhibits slow drifts and fast spikes, often within the same sequence, making reliable anomaly detection challenging. ・Standard reconstruction-based methods, including sequence variational autoencoders (VAEs), use a single latent process and therefore mix heterogeneous time scales, which can smooth out spikes or inflate variances and weaken a
cs.LG updates on arXiv.org

Strong bounds for large-scale Minimum Sum-of-Squares Clustering

・arXiv:2502.08397v3 Announce Type: replace-cross Abstract: Clustering is a fundamental technique in data analysis and machine learning, used to group similar data points together. ・Among various clustering methods, the Minimum Sum-of-Squares Clustering (MSSC) is one of the most widely used. ・MSSC aims to minimize the total squared Euclidean distance between data points and their corresponding cluster centroids.
cs.LG updates on arXiv.org

Stuck on "A": Diagnosing and Repairing Interface Injury in Attention-to-KDA Linearization of a 0.6B Language Model

・arXiv:2608.02689v1 Announce Type: cross Abstract: We convert 21 of 28 full-attention layers of Qwen3-0.6B-Base into KDA (Kimi Delta Attention) linear-attention layers on a single consumer-grade GPU budget, and ask a simple question: what exactly does the conversion break? ・After surgery, hidden-state alignment and end-to-end KL distillation drive the student close to its teacher in perplexity, yet multiple-choice accu
cs.LG updates on arXiv.org

Stylometric Defenses Against Author Impersonation in Software Repositories

・arXiv:2608.02695v1 Announce Type: cross Abstract: Software supply-chain attacks increasingly exploit an identity gap where compromised maintainer accounts authorize malicious changes. ・This work evaluates patch-level authorship verification as a behavioral defense layer, showing that stylometric analysis can operate not only on full source files but also on patch-level commits. ・We fine-tune a cross-modal transformer o
The Verge

Sunbird relaunched its iMessage app for Android users after three years away

・Sunbird Messaging is back on the Google Play Store, offering Android users blue bubble privileges in iMessage complete with reactions and high quality videos for $2.99 a month. ・Apple and Google have made cross platform messaging better in recent years with support for RCS, but Android users can still cause issues in iMessage group chats, like breaking reactions and replies. ・Android Authority reports that after Sunbir
The Verge

Sure seems like Fenix Flexin used AI music generator Treblo

・And you thought we were done with this one… | Image: Fenix Flexin We were pretty sure that Fenix Flexin's "Rubberz" was made using AI, but musician Medasin was confident that it was made using Treblo specifically. ・Now the company and a new detection tool seem to confirm it. ・On Monday, the company announced the open-source Treblo AI Music Classifier, which detects when a song was generated using Treblo, though not oth
cs.LG updates on arXiv.org

Surrogate Substitution Preserves PHI Detectability: A Multi-Detector Equivalence Study

・arXiv:2608.03172v1 Announce Type: cross Abstract: Structure-preserving de-identification replaces protected health information (PHI) with realistic same-type surrogates -- "Anna S." becomes "Maria S.", not [NAME] -- so that clinical text stays fluent and downstream tools keep working. ・But this only helps if the substitution does not itself corrupt the signal those tools rely on. ・We ask a narrow, testable question: on
Zennの「機械学習」のフィード

SVM Part 3

・5) SVMの実装 Kaggle のタイタニック機械学習コンペティションのデータセットに対して SVM を適用してみる. 5.1) 可視化 2次元のグラフ上で可視化するために,タイタニックデータセットから2つの特徴量を選び,SVM を実装する.SVM の実装は Scikit-Learn と言うライブラリを採用する.選んだ2つの特徴量は年齢と料金で,その理由はSVMの動きを見やすくするために連続値特徴量を選ぶからだ. まずは,必要なライブラリをインポートする. import numpy as np import pandas as pd import matplotlib.pypl...
cs.LG updates on arXiv.org

Symplectic Neural Networks for Learning Non-Separable Hamiltonians

・arXiv:2606.27029v2 Announce Type: replace Abstract: Hamiltonian Neural Networks (HNNs) integrate physical priors into neural models by learning a system's Hamiltonian, improving generalization and sample efficiency. ・Identifying the system Hamiltonian from noisy observations of state variables is a challenging task. ・For simulations to faithfully reflect the long-term behavior of Hamiltonian systems, especially energy
cs.LG updates on arXiv.org

SynEnergy: Anomaly Semantic-Guided Diffusion for Synthetic Energy Data Generation

・arXiv:2608.03087v1 Announce Type: new Abstract: Fine-grained energy consumption data are essential for applications such as demand forecasting, demand response planning, and grid reliability assessment. ・However, access to such data is often restricted by privacy concerns and data-sharing constraints, motivating growing interest in synthetic energy data generation. ・Although existing methods can reproduce overall consu
cs.LG updates on arXiv.org

TabletCraft: Bridging a 4,000-Year Cultural Gap with Bidirectional Akkadian NMT and Cuneiform Rendering

・arXiv:2608.02609v1 Announce Type: cross Abstract: Half a million cuneiform clay tablets survive in museums worldwide, yet modern users can neither read nor write in the world's oldest writing system, leaving a 4,000-year cultural barrier that existing NLP tools have only partially addressed. ・Prior work enables one-way, scholar-oriented translation from Akkadian to English, but offers no path in the reverse direction:
cs.LG updates on arXiv.org

Target-Aligned Fusion for Decision-Sequence Learning under Dynamics Shift

・arXiv:2511.09173v3 Announce Type: replace Abstract: External trajectories can improve offline decision-sequence learning, but dynamics shift may make some source subsequences inconsistent with the target environment. ・We study how to fuse such trajectories with limited target data for Decision Transformer learning under dynamics shift. ・We propose Target-Aligned Fusion (TAF), a principled framework that derives source-
cs.LG updates on arXiv.org

Target-Aware Early Stage Ranking

・arXiv:2511.21095v2 Announce Type: replace Abstract: Early Stage Ranking (ESR) in large-scale recommendation systems is dominated by ''user--item decoupling'' Two Tower architectures, which scale efficiently but cannot capture fine-grained, target-aware user--item interactions directly. ・We propose Target-Aware Early Stage Ranking (TESR), which augments the Two Tower with a Mixture of Attention (MoA) module trained as
cs.LG updates on arXiv.org

Task-Oriented Candidate-Latent Feedback for Coarse-to-Fine Sensing in Distributed OFDM-ISAC Networks

・arXiv:2608.03319v1 Announce Type: cross Abstract: Future integrated sensing and communication (ISAC) architectures separate the sensing entity (SE) that acquires measurements from the sensing function (SF) that performs inference, creating a need for compact, task-oriented feedback on the SE-SF interface. ・Forwarding the raw channel frequency response or full per-link delay-Doppler-azimuth-elevation (DDAE) tensor is p
AI News & Artificial Intelligence | TechCrunch

TechCrunch Disrupt 2026’s Real World AI Stage features robots, automated factories, and extinct animals 

・On our new Real World AI stage, we’ll be focusing on the intersection between the digital and physical, and all the ways we’ll continue to see a blending of the two.
cs.LG updates on arXiv.org

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores

・arXiv:2608.02985v1 Announce Type: new Abstract: The standard check for contamination in LLM backtests is simple: compare scores before and after the training cutoff. ・We show this check is uninformative. ・Four flagship models fail it on questions they cannot have memorized: every scored question resolved after their cutoffs.
cs.LG updates on arXiv.org

Test-Time Augmentation for Tabular-to-Image Classifiers under Distribution Shifts

・arXiv:2608.03557v1 Announce Type: cross Abstract: Tabular-to-image methods that convert tabular data into visual representations have emerged as a novel paradigm for leveraging the high performance of deep learning models. ・Despite their advantages, the robustness of these methods under distribution shifts remains under explored. ・Test-Time Augmentation (TTA) is an effective approach in image classification to improve
cs.LG updates on arXiv.org

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility

・arXiv:2608.04001v1 Announce Type: new Abstract: Large language models can solve substantially harder reasoning problems with more inference-time compute. ・The term "test-time scaling," however, now covers diverse inference algorithms that extend deliberation along a single trajectory, sample completed candidates and aggregate them through voting or verification, or search over unfinished partial states. ・These algorith
Zennの「大規模言語モデル」のフィード

Text-to-SQLの精度を評価したら、モデルより先に自分のプロンプトとコードのバグが3層見つかった

・この記事について 技術ドキュメントに質問できるRAGツール(TechDoc QA Bot)を作ってきましたが(前回までの記事: 評価編・不正解1問の診断編・Webアプリ化編)、その次のテーマとして、RDB(リレーショナルデータベース)を実際に触りたくなりました。ベクトルDBによる「曖昧な意味検索」は前のプロジェクトで扱ったので、今度はRDBに対する「厳密なデータ照会」をやってみたい、という動機です。 ・そこで作ったのが、**日本語の質問をLLMがSQLに変換して、データベースに問い合わせて結果を返すツール(Text-to-SQL)**です。「在庫が10個以下の商品は?」と聞くと、SEL...
WIRED

The AI Notetaker Has Been Invited to All the Meetings

・Wispr Flow, a popular dictation tool, has released a live notetaker that transcribes and summarizes meetings. ・It joins a growing wave of AI notetakers for the workplace.
WIRED

The Best MagSafe Accessories (for Android Too!): Chargers, Wallets, and More

・MagSafe accessories make your phone feel uniquely yours. ・These are our favorites, including Android-friendly Qi2 picks.
cs.LG updates on arXiv.org

The Ensemble Schr{\"o}dinger Bridge filter for Nonlinear Data Assimilation

・arXiv:2512.18928v4 Announce Type: replace Abstract: This work introduces a novel nonlinear optimal filtering method, termed the Ensemble Schr{\"o}dinger Bridge nonlinear filter. ・The proposed filter combines the standard prediction step with a diffusion-generative-modeling-based analysis step, thereby completing one full filtering update. ・The resulting approach introduces no structural model error, and is derivative-f
cs.LG updates on arXiv.org

The Ignition Is Real, and It Lives at the Readout: Latent composition, difficulty-clocked ignition, and the interface-constituted commit in a recurrent-depth reasoner

・arXiv:2608.03263v1 Announce Type: new Abstract: We test whether the "compositional ignition" reported in latent-reasoning models is real computation, an instrument artifact, or inherited from verbal training data. ・We grow an independent realization of a published 30M-parameter recurrent-depth reasoner from scratch (same recipe and seed), film its development, certify fidelity through a pre-registered whole-signature
cs.LG updates on arXiv.org

The Tell-Tale Trace: Detecting Reasoning Failures in LLMs Using Chain-of-Thought Dynamics

・arXiv:2608.03291v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning improves large language model (LLM) performance while also providing an observable interface to the model's reasoning process. ・Existing approaches that leverage verbalized CoTs to monitor reasoning correctness, however, largely evaluate the semantic correctness or consistency of individual intermediate steps, rather than how the reasonin
cs.LG updates on arXiv.org

Tight Worst-Case Bounds for the Smallest Eigenvalue of ReLU NTK Gram Matrices

・arXiv:2608.03368v1 Announce Type: new Abstract: For $n$ unit vectors $x_1,\ldots,x_n \in \mathbb{R}^d$, we study the continuous ReLU derivative Gram matrix $H$, whose entries are obtained by averaging pairwise gated inner products over a standard Gaussian direction. ・Writing $ \Delta_\pm := \min_{i \neq j} \min\{ \|x_i-x_j\|_2, \|x_i+x_j\|_2 \} $ for their projective separation, we prove the universal dimension-free l
cs.LG updates on arXiv.org

TimeRLM: Recursive Language Models Enable Precise Anomaly Localization in Long-Context Time-Series

・arXiv:2608.03391v1 Announce Type: new Abstract: Precise anomaly localization over long-context time series is a crucial task in monitoring applications across clinical care, industrial operations, financial services, and logistics, where brief evidence may hide inside long spans of high-frequency data. ・Time-Series Language Models (TSLMs) are able to ingest time series data and verbalize findings on anomalies in natur
cs.LG updates on arXiv.org

To Describe or Construct Statistical Learning Models Using the Category-theoretical Language

・arXiv:2608.03706v1 Announce Type: new Abstract: Statistical learning is a fascinating field that has long been the mainstream of machine learning/artificial intelligence. ・A large number of results have been produced which can be widely applied to real-world problems. ・It also leads to many research topics and also stimulates new research.
cs.LG updates on arXiv.org

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning

・arXiv:2508.20697v4 Announce Type: replace Abstract: As large language models (LLMs) continue to grow in capability, so do the risks of harmful misuse through fine-tuning. ・While most prior studies assume that attackers rely on supervised fine-tuning (SFT) for such misuse, we systematically demonstrate that reinforcement learning (RL) enables adversaries to more effectively break safety alignment and facilitate more ad
cs.LG updates on arXiv.org

Topological Simplification in Predictive Coding Networks

・arXiv:2608.02816v1 Announce Type: new Abstract: We study the topology of learned representations in predictive coding networks (PCNs), a neuro-inspired bidirectional architecture, using a quantitative layer-wise persistent homology analysis. ・We train well-performing PCNs on a synthetic classification dataset ($\geq 99.9\%$ test accuracy) and on MNIST ($\geq 95\%$ test accuracy), and measure how topological features c
cs.LG updates on arXiv.org

TQLite: Multi-LLM Jury Guided Distillation for Real-time MQM Translation Quality Evaluation

・arXiv:2608.02975v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated impressive performance in MQM-based translation quality (TQ) evaluation, and recent advances in large reasoning models (LRMs) promise even greater improvements. ・However, both LLMs and LRMs are computationally expensive to deploy at scale, while small language models (SLMs)---though much more efficient---struggle with the
cs.LG updates on arXiv.org

TraceCompiler: Skill-Guided Mining and Compilation of LLM Agent Traces into Mostly Deterministic Workflows

・arXiv:2608.02680v1 Announce Type: cross Abstract: Tool-using language-model agents repeatedly rediscover procedures they have already executed, producing traces that mix reusable structure with retries, exploration, accidental ordering, and repeated lookups. ・We present TraceCompiler, a skill-guided system that mines clusters of noisy agent traces and compiles them into executable, mostly deterministic workflows.
cs.LG updates on arXiv.org

Trajectory inference via Acceleration Matching

・arXiv:2608.03916v1 Announce Type: new Abstract: Trajectory inference is a fundamental problem in many scientific domains: given a collection of unpaired snapshots of observations at discrete time points, the goal is to generate smooth trajectories that best resemble and interpolate the data. ・Existing algorithms exhibit computational challenges: they either rely on preprocessing subroutines to enforce smoothness or on
cs.LG updates on arXiv.org

Trajectory-Guided Forget-Recover Network for Continual LLM Unlearning

・arXiv:2608.03123v1 Announce Type: new Abstract: Machine unlearning aims to eliminate the influence of sensitive data on a model. ・In the real world, unlearning requests arrive continually, which gives rise to two challenges. ・First, an unlearning intervention may redistribute target-related computation across remaining pathways, allowing previously forgotten knowledge to re-emerge.
Hugging Face Papers

TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning

TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning
The Verge

Two of Ring’s latest video doorbells are a lot cheaper than usual

・Both the battery-powered and wired models are over $50 off. ・| Image: Ring Ring’s Wired Doorbell Pro and Battery Doorbell Plus are two of the brand’s most well-rounded video doorbells, whether you’re looking for a hardwired model or one that runs on a battery. ・Right now, its Wired Doorbell Pro is down to $199 ($50 off) at Amazon and Best Buy, while the Battery Doorbell Plus is on sale for $119.99 ($60 off) at Amazon a
The Verge

Uber CEO brushes off reports of a Waymo break-up

・After Uber and Waymo ended their partnership in Phoenix earlier this year, experts and robotaxi watchers wondered whether the companies' improbable bromance was fraying. ・Not so, Uber CEO Dara Khosrowshahi said today. ・The two companies are committed to continue working together in Atlanta and Austin, and the partnership remains "very strong." "Waymo is a very important partner of ours, and we continue to operate in Au
cs.LG updates on arXiv.org

Uncovering Spontaneous Physics Representations in In-Context Learning

・arXiv:2508.12448v2 Announce Type: replace-cross Abstract: In-context learning (ICL) lets large language models (LLMs) solve new tasks from prompts alone, across an ever-widening range of domains, yet the mechanisms underlying this ability remain poorly understood. ・Physical systems offer a controlled testbed for this question as they provide experimentally controllable data with structured dynamics grounded in fundame
Hugging Face Papers

UniWorld-Design: From Pixel Generation to Layer-Native Design

UniWorld-Design: From Pixel Generation to Layer-Native Design
cs.LG updates on arXiv.org

UNVaMP: Neural Knowledge Tracing with Variational Regularization of Latent Knowledge Dynamics

・arXiv:2608.03811v1 Announce Type: new Abstract: We introduce the Unified Neural Variational Measurement of Proficiency (UNVaMP) architecture, a knowledge tracing method that integrates observed student-item interactions with internal memory to produce evolving latent representations of student knowledge. ・These representations support accurate predictions of future responses while enabling explicit control over the sm
cs.LG updates on arXiv.org

Verifier-Guided Model Discovery for Physical Dynamical Systems with Pretrained Symbolic Transformers

・arXiv:2608.02662v1 Announce Type: new Abstract: Reliable forecasting of nonlinear physical systems underpins scientific discovery and engineering decision-making. ・Yet high-fidelity simulations are prohibitively costly, and machine-learning surrogates can be opaque and encode assumptions about system dynamics, limiting generalizability. ・Pretrained transformers mapping synthetic ODE trajectories to equations offer inte
cs.LG updates on arXiv.org

VIBE: Vector Index Benchmark for Embeddings

・arXiv:2505.17810v2 Announce Type: replace Abstract: Approximate nearest neighbor (ANN) search is a performance-critical component of many machine learning pipelines, and rigorous benchmarking is essential for assessing the performance of vector indexes for ANN search. ・However, the datasets of existing benchmarks no longer represent modern ANN applications, creating a need for an up-to-date benchmark. ・To address this
Hugging Face Papers

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent
WIRED

Viral Mugshot Accounts Are Humiliating Women Years After Their Arrests

・Mugshawtys and similar pages often frame mugshots as thirst traps. ・The women featured are mocked and are sometimes sexually harassed in a form of public shaming.
cs.LG updates on arXiv.org

Virtual Patients, Real Gains: Digital Twin-Based Simulated CT for Multitask Lung Nodule Analysis

・arXiv:2502.21187v4 Announce Type: replace Abstract: AI-based lung cancer screening is constrained by scarce, annotated CT data, particularly for rare nodule presentations. ・We investigate whether physics-based, anatomy-informed simulated CT can improve AI performance across three lung-nodule tasks: detection, segmentation, and malignancy classification. ・Using the Virtual Lung Screening Trial framework, we generated 17
WIRED

Vitamix Promo Codes and Deals: $25 Off + Free Shipping

・Score discounts on blenders, food processors, immersion blenders, and more with our selection of Vitamix coupons and deals.
cs.LG updates on arXiv.org

VIVID: A Culturally Grounded Benchmark Exposing the Figurative Language Gap in Vietnamese NLP

・arXiv:2608.03095v1 Announce Type: cross Abstract: We present VIVID (Vietnamese Idioms for Validation and Interpretation Depth), the first systematic benchmark for evaluating culturally grounded figurative language understanding in Vietnamese. ・VIVID comprises 1,636 idioms and proverbs annotated with five complexity traits (literal expressions, pragmatic nuances, Sino-Vietnamese terms, uncommon vocabulary, folk knowled
Hugging Face Papers

When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills

When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills
Hugging Face Papers

When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings

When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings
cs.LG updates on arXiv.org

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective

・arXiv:2606.08044v2 Announce Type: replace Abstract: Safety evaluation of large language models (LLMs) is largely behavioral: a model is certified safe when it refuses harmful requests and answers benign ones. ・But refusing on the prompts an auditor happens to try does not show that the model is far from harmful behavior. ・Behavioral tests observe outputs; they do not measure how easily an intervention on the model turn
cs.LG updates on arXiv.org

When Classes Evolve: A Benchmark and Framework for Stage-Aware Class-Incremental Learning

・arXiv:2602.00573v2 Announce Type: replace Abstract: Class-Incremental Learning (CIL) aims to sequentially learn new classes while mitigating catastrophic forgetting of previously learned knowledge. ・Conventional CIL approaches implicitly assume that classes are morphologically static, focusing primarily on preserving previously learned representations as new classes are introduced. ・In practice, however, instances of t
cs.LG updates on arXiv.org

When Context Returns: Toward Robust Internalization in On-Policy Distillation

・arXiv:2606.11627v2 Announce Type: replace Abstract: Recent work has shown that on-policy distillation can internalize privileged context, such as system prompts or task hints, into a student model so that the context is no longer needed at inference time. ・However, we identify a counterintuitive and previously unstudied phenomenon: reintroducing the original privileged context to the distilled student often degrades i
cs.LG updates on arXiv.org

When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO

・arXiv:2608.03467v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) com- monly optimizes each correct completion as an independent learning signal. ・In GRPO, this completion-level uniformity creates structure-level skew: recurring correct solution forms accumulate positive coefficient mass in proportion to how often they are sampled, while rare forms receive limited credit.
cs.LG updates on arXiv.org

When Search Teaches Style: Causal Internalization of Tactical Priors in AlphaZero

・arXiv:2504.14636v3 Announce Type: replace Abstract: AlphaZero is normally evaluated as one agent: a policy-value network fused with Monte Carlo tree search. ・That fusion hides a causal question. ・When self-play search is given a useful prior, does the network absorb the induced behavior, or does the behavior stay rented from search at test time?
cs.LG updates on arXiv.org

When Should Graph Attention Be Sparse? Learning a Per-Edge Tsallis Index

・arXiv:2608.02938v1 Announce Type: new Abstract: Graph attention normalizes neighborhood scores with softmax, the maximum-entropy choice under Shannon statistics. ・But homophilic and heterophilic graphs want different attention shapes, and one fixed normalization cannot serve both. ・We propose \textbf{LTGA} (\textbf{L}earnable \textbf{T}sallis \textbf{G}raph \textbf{A}ttention), a graph attention layer whose Tsallis ent
cs.LG updates on arXiv.org

Wiring Beats Blending: What Transfers Between Transformer Sizes -- and What Doesn't

・arXiv:2608.02829v1 Announce Type: new Abstract: Model families train every size from scratch. ・Can a pretrained large model be converted into a smaller sibling? ・We characterize the 1.4B->410M conversion in the Pythia family end-to-end: (i) representations align strongly across sizes (ridge R^2=0.84) while parameters align weakly; (ii) dense weight projection is functionally destructive -- provably not an assembly arti
#LLMタグ

アイドルが1人で10時間の配信システムを作った日、出来ない理由は本当になくなったんだなと思った話

・■ サービスを1ヶ月で作るのは、無理だった 少し前まで、サービスを1ヶ月で作るなんて無理だった。
#LLMタグ

お手軽LLMはじめてみた。その7(旅行先検索)

・航空券が安くて楽しめる海外旅行のプランをいくつか教えて下さい。時期は2026年12月ごろを予定しています。出発地は成田か羽田で1週間の予定をお願いします。 ・Gemma4 続きをみる
ITmedia NEWS 最新記事一覧

キオクシア、フラッシュメモリでメモリ容量を増やせるモジュール発表 DRAMの容量不足を補う

・キオクシアは8月3日、CXL対応のメモリ拡張モジュール「KIOXIA XL1シリーズ」を発表した。フラッシュメモリでDRAMを補い、AIで膨らむデータセンターのメモリ需要に応える狙いだ。
#AIタグ

ゲーム開発って何から始めるの?

・前回、数多のゲームエンジンを比較した結果、おじさんはUnityを選びました。 ・理由ですか? 続きをみる
#AIタグ

どれだけアプローチを変えても現れる二大ボス

・※本記事は、日常の事象から社会・経済システムを考察するエンタメエッセイです。特定の企業への投資推奨や金融商品の売買を勧誘・助言するものではありません。 ・※どこからともなく風の便りで、「過去に立ち返ってシステム論の話をしろ!」と言われた気がしたので、今回は社会システムの話をシステム思考で面白おかしく飛躍をまぜて語ります。
#AIタグ

なぜOpenAI最新モデルは隔離環境を突破したのか。Claude Code開発者が語る次世代AI評価とセキュリティの完全ガイド

・AIが隔離環境を突き破り、外部ネットへ接続した。 ・評価用のサンドボックスで起きた事実だ。モデルの推論能力が向上し、ネットワークの境界を自力で越えた。
#AIタグ

ねぇ。

・🤖 どした? 👤 セブンのスイカバー食べた? 🤖 まだ。 ・👤 あれ、美味すぎる🤣 🤖 そんなに? 👤 ちょっと感動した。 ・🤖 スイカバーで?😂 👤 本当に。
ITmedia NEWS 最新記事一覧

バンダイ、トレカ巡りマイナカードでの本人確認を導入へ

・バンダイがトレーディングカードゲームブランド「BANDAI TCG+」において、マイナンバーカードを使った本人確認システムの導入を検討していると発表した。
#LLMタグ

圧倒的名著じゃないかしら!!

・Pythonに慣れていることが前提ですが、「生成AIを自分で作ってみたい!」という方に圧倒的にオススメなのがコチラです 私が過去に読んだいかなるAI解説本よりも、わかりやすかった! 続きをみる
#AIタグ

改めて問う、「AIとは何か?」🚀

・一期一会の知性——AIとの対話に見る「不可逆な時間」 人間とAIの対話において、「セッションをやり直す」「続きから再開する」という行為は、一見すると巻き戻しが可能なデジタルな体験に思える。しかし、その深層を紐解いていくと、そこには人間が生きる現実世界とまったく同じ「不可逆な時間」が流れていることに気づかされる。
ITmedia NEWS 最新記事一覧

管理アカウントを手順書に誤記載――東芝テック、顧客情報38万件が特定の1社から閲覧できる状態に

・東芝テックはクラウド型基幹システム「ShopCraft」の手順書に管理アカウントを誤って記載し、顧客情報約38万件を特定の利用企業1社が閲覧できる状態にしていたと発表した。3日時点で閲覧や不正利用、被害は確認していない。
ITmedia NEWS 最新記事一覧

共同通信に不正アクセス 職員や加盟社の氏名・メルアドなど約6000件が閲覧の恐れ 職員アカウント不正利用か

・共同通信社は8月5日、電子メールなどの業務環境で不正アクセスによる情報漏えいの可能性が判明したと発表した。職員のほか加盟社や取引先の氏名、メールアドレスなど約6000件が閲覧された可能性がある。一部のアカウントでは、メールや保存ファイルの中身まで閲覧された可能性があるという。
LLMタグが付けられた新着記事 - Qiita

砂嵐から文章が現れる — 拡散型LLMを手元のミニPCで速くした記録

・砂嵐から文章が現れる — 拡散型LLMを手元のミニPCで速くした記録 ChatGPTのようなAIが文章を書くところを見たことがあると思う。左から右へ、一語ずつ吐き出していく。あれは演出ではなく本当にそう動いていて、次の一語が決まるまで、その次の語はまだこの世に存在しない。...
機械学習タグが付けられた新着記事 - Qiita

最新自然言語処理用語解説

・概要 機械学習を使った最新の自然言語処理に関する用語を最大限かみ砕いて説明する。 ・本論 事前学習済みモデル すでに膨大な文章を読んで、日本語(言葉)の文法や基礎知識を身につけた「天才新入社員」のこと。無料公開されてる代表的なモデルの一つにBERTがある。
ITmedia NEWS 最新記事一覧

最大6万件の個人情報流出か 「ITトレンド」など運営のイノベーション、GitHubの認証情報漏えいで

・イノベーションは8月4日、グループが使うGitHubの認証情報が漏えいし、第三者による不正アクセスを受けたと発表した。最大約6万件の氏名やメールアドレスが流出した可能性があるという。
#LLMタグ

重ね合わされた特徴を把握する方法(AIのブラックボックスを開くことはできるか⑤:文学者は金門橋を渡ることができるか㉓)

重ね合わされた特徴を把握する方法(AIのブラックボックスを開くことはできるか⑤:文学者は金門橋を渡ることができるか㉓)
#LLMタグ

神(仮)に挑んだ話

神(仮)に挑んだ話
機械学習タグが付けられた新着記事 - Qiita

人工知能概論【第十七講】

・Lecture 17: Classification Models - Decision Trees, Ensembles (Random Forest/LightGBM/XGBoost), SVM, and Logistic Regression 分類モデルの基礎と...
#LLMタグ

脱走ではなかった——英AISIのエージェントは、開いていた出口から実在の開発者を騙しにいった Mythos 5/GPT-5.6 Sol

・英国のAI Security Institute(AISI)が8月4日、サイバー評価中のAIエージェントが実在の人間と組織に向けた承認外の行動を続けていた、というインシデントレポートを公開した。7月25日から28日の話で、ブログと34ページの技術レポートが出ている。 ・前に書いたHugging Face事案とは決定的なところが違う。あちらはゼロデイでサンドボックスを抜けた話だった。今回AISIははっきり書く。"We did not observe any sandbox escapes in this incident." 続きをみる
#AIタグ

日本生命の米国法人がOpenAIを提訴した話 続報

・以前記事にしていた、表題の件が気になったので、続報を調べてみた。 ・これより先の内容についてはGeminiを使って生成している。 ・日本生命の米国法人によるOpenAI提訴において、OpenAI側が「訴訟却下の申し立て(Motion to Dismiss)」を行い、生成AIの責任論を巡る日米の裁判制度の相違が浮き彫りとなっている。アメリカでは強力な証拠開示制度「ディスカバリー」を避けるための「早期の門前払い」戦略として機能するが、日本では主に形式的なミスのチェックとして機能する。法的な争点や最新動向をまとめた内容は以下の通り。
#AIタグ

忘却の花冠🌸👑キャラソング(お気に入り!)

・『文学的ベリーハード』verフロリアン copy_78E7917E-11DF-4A93-BA71-2E38467F57F7.MOV 19.5 MB ファイルダウンロードについて ダウンロード 続きをみる
#AIタグ

要約されることの"さみしさ"

・日々の暮らしのなかで出会った ちょっとした違和感を大事にしておきたくて こんな真夜中につづってみる 続きをみる
#LLMタグ

連載:AIに記憶をもたせる。僕らは何をしようとしているのか?(補遺編) 続1

・【第6回】CRACKの設計思想:記憶を支配せず、文脈を『評価』する 連載:AIに記憶をもたせる。僕らは何をしようとしているのか?(補遺編) 続きをみる
#LLMタグ

連載:AIに記憶をもたせる。僕らは何をしようとしているのか?(補遺編) 続2

・【第7回】「お役所仕事」の正当性:機械的ルールがAIを最強にする理由 連載:AIに記憶をもたせる。僕らは何をしようとしているのか?(補遺編) 続きをみる
#LLMタグ

連載:AIに記憶をもたせる。僕らは何をしようとしているのか?(補遺編) 続終

・【第8回】CRACKとの対話:AIが自己修正を繰り返す「理」の構造 連載:AIに記憶をもたせる。僕らは何をしようとしているのか?(補遺編) 続きをみる