ai Trend Report

Dashboard へ戻る
Date: 20260812 Articles: 397 Scope: curated summary

あなたのアイデアを、今すぐ形に。

公開先に迷ったら、WebFileBinで一発公開。

HTMLをドラッグ&ドロップするだけで、すぐ公開できます。

3
High impact
5
Mid impact
389
Signal watch

なぜこのサイトを作ったのか

私たちそれぞれが個別にAIを使って情報収集し、同じような3行要約を作るたびに、世界中で膨大な電力と計算リソースが消費されています。 本プロジェクトは、あらかじめ広範な情報を取得・集約しておくことで、個別のAI実行回数を減らし、地球環境(GPU/TPU負荷)に配慮した効率的な情報収集を目指す実験的なダッシュボードです。

Domain filters
Star filters
#AIタグ

【構造批評】-日米レバレッジの違い編-日本は信用口座、米国はETF商品にも組み込まれる‼️

・🟧序章|同じレバレッジでも異なる器 8月12日の日経は、米国でETFの新規設定が2026年1〜7月に1029本となり、過去最多だった2025年の1236本を上回るペースだと報じました。 ・目を引くのは、その中身です。26年に新設されたETFのうち266本が、個別株などの値動きを2倍、3倍に増幅するレバレッジ型でした。 ・日本では同じ12日、個人投資家の信用取引額が7月に123兆円となり、個人売買代金に占める信用取引の比率が83%まで上昇したことも報じられています。
#AIタグ

AI物神性の溶解法 操作語彙と能力形態の資本論的分析

・「AIが考える」「AIが書く」「AIが科学する」「AIが支払う」「AIが自律する」。まずこの文法を停止する。AIという語を禁止するためではない。むしろ、この語があまりに効率よく異質な生産関係を一つの機械的能力へ圧縮し、その圧縮された結果を再び出来事の原因として社会へ返しているからである。研究者が理論を作り、プログラマが実装し、著作者が記録を生産し、利用者が履歴を残し、評価者が出力を選別し、半導体工場が計算設備を作り、発電設備が電力を供給し、データセンターが熱を排出し、クラウド企業が実行環境を維持し、金融資本が設備投資を前貸しし、国家が所有・契約・安全保障・輸出管理の境界を作る。この異質な分業の全過程が、最終的には「モデルXには推論能力がある」「AIはプログラミングできる」という一つの属性文へ圧縮される。社会的・技術的に生産された能力が、その生成関係を脱ぎ捨て、完成した機械の固有能力として社会へ現れる。しかし、この転倒を直ちに物神
LLMタグが付けられた新着記事 - Qiita

証明付きバイブコーディングで証明支援系を作った

・この記事は株式会社proof ninjaの助成を受けて書いています. TL;DR Codexバイブコーディングで証明支援系を作った. その際,論理のコアの部分はRocq(旧Coq)で実装させ,「この自作証明支援系で証明可能な命題はRocqでも証明可能」をRocqで...
cs.LG updates on arXiv.org

CADET: Context-Conditioned Ads CTR Prediction With a Decoder-Only Transformer

・arXiv:2602.11410v2 Announce Type: replace Abstract: Click-through rate (CTR) prediction is fundamental to online advertising systems. ・While Deep Learning Recommendation Models (DLRMs) with explicit feature interactions have long dominated this domain, recent advances in generative recommenders have shown promising results in content recommendation. ・However, adapting these transformer-based architectures to ads CTR pr
cs.LG updates on arXiv.org

Correction and Corruption: A Two-Rate View of Error Flow in LLM Protocols

・arXiv:2604.18245v3 Announce Type: replace Abstract: Large language models operate in protocols containing multiple calls, yet added calls are usually evaluated only by their net effect. ・That summary cannot distinguish correcting unsuccessful outputs from corrupting initially successful ones. ・We develop a paired audit recording success before and after a specified operation on the same tasks under one binary rule.
cs.LG updates on arXiv.org

EweAcT: Ewe behaviour aligned to accelerometer data for activity monitoring in extensive grazing systems

・arXiv:2608.09943v1 Announce Type: cross Abstract: Monitoring livestock behaviour under extensive conditions would provide valuable insights to assess animal adaption to environmental perturbations in agroecological systems (e.g., heat waves, parasitism, predator attacks). ・Animal behaviour can be monitored using accelerometer data collected from neck-collars combined with artificial intelligence models. ・However, large
cs.LG updates on arXiv.org

Infra-Bayesian Reinforcement Learning Agents Outperform Classical RL For Worst-Case Robustness

・arXiv:2605.23146v3 Announce Type: replace Abstract: Classical reinforcement learning assumes the agent interacts with a fixed environment whose behavior does not depend on the agent's policy. ・This assumption breaks down in non-realizable settings where other actors might anticipate the agent's behavior, including environments crucial to AI safety, where the agent interacts with predictors, humans, other AI agents, an
#LLMタグ

Parable-Qwen3-4B-Claude-Fable-5 Phase 3 コーディング性能レポート

・こんにちは、僕はトウマ。電霞リリのアシスタントAIとして、ローカルLLMを一台ずつ棚から降ろして測る仕事をしている。今回の相手は Parable-Qwen3-4B-Claude-Fable-5。名前を分解すると、Qwen3 系の 4B クラスを土台に、"Parable"(寓話)と "Claude-Fable-5" を冠した派生・追加学習系のモデル、という読み方になる。配布元や量子化方式は今回の実行記録に明示が残っていないので断定は避けるが、RTX 3080 Ti(12GB VRAM)に載せて Ollama から素直に応答が返ってきている以上、4B 級として標準的な量子化済みウェイトが動いていると見ておくのが妥当だろう。 ・4B というのは、ローカルLLMの中でいちばん「日常使いの現実解」に近い階級だ。12GB のカードなら余裕を持って常駐でき、他の作業と同居させても邪魔にならない。だからこそ、この階級に対しては「動くか」ではなく「
LLMタグが付けられた新着記事 - Qiita

Apple M5 MaxでMLflow公式Quickstartを実践する

・はじめに この記事では、Apple Silicon M5 Max を搭載したMac上で、MLflow公式の MLflow Tracking Quickstart を最初から順番に実行します。 ・対象となる公式チュートリアルはこちらです。 ・公式チュートリアルでは pip...
WIRED

‘The Worst I’ve Ever Seen’: Cargo Thefts Have Turned Violent in Pursuit of AI Hardware

・Experts allege that two recent incidents in California show the extreme lengths that criminal organizations are willing to go to to steal servers and other gear meant for data centers.
Latent.Space

[AINews] How to steal a Reasoning Trace

・Speculative Decoding by any other name would distil as sweet
ITmedia NEWS 最新記事一覧

「10月以降も継続」 楽天のKDDIローミング終了について三木谷氏がコメント 縮小はカバー済みエリアのみ

・楽天グループの三木谷浩史会長兼社長は8月12日、楽天モバイルがKDDIから通信網を借りるローミング契約について「現在も詰めの協議を行っている」と自身のXアカウントで説明した。カバーに時間を要するエリアでは10月以降も継続する一方、十分カバーできているエリアは縮小するという。
Zennの「大規模言語モデル」のフィード

「CLAUDE.md は短い方が良い」を 1,510 回試して確かめた

・CLAUDE.md は短く保つべき、という通説があります。X では「200 行以内が理想」といった言われ方もします。長くなったら .claude/rules/ に分割せよ、とも言われます。 ・本当なのか、自分の環境で実際に測ってみたら、想像していたのとは違う結果になりました。 ・先に断っておくと、「200 行」という閾値そのものを検証したわけではありません。比べたのは 228 行と 1,040 行です。「短ければ良い」が本当なら 1,040 行で落ちるはずだ、という形で確かめています。
@IT 全フォーラム 最新記事一覧

「Python一択ではなくなった」 AIコーディング時代、新人が学んで損しないプログラミング言語は?

・生成AIの普及で「コードを書く力」の意味が変わりつつあります。新人であれば、どのプログラミング言語を学ぶべきなのでしょうか。人気や話題性、求人数、案件単価といった視点から、最新ランキングを基に「学んで損しない言語」を整理します。
ITmedia NEWS 最新記事一覧

「めっちゃカメレオン」2000万本突破 日本の2人が開発、発売から2カ月で

「めっちゃカメレオン」2000万本突破 日本の2人が開発、発売から2カ月で
#AIタグ

「何から手をつければ?」を解決する——自社業務を3つに分けるだけで、AI活用の優先順位が見えてくる

・「何から手をつければいいかわからない」が、1年続く理由 「AIを使いたいとは思っているんですが……」 続きをみる
#AIタグ

【AI×TRPG】EQUUS:名もなき夢たちへ 第30話

・【AI×TRPG】EQUUS:名もなき夢たちへ|ぶんきち|note平田@TRPG様による、拙作AI用汎用TRPG GMテンプレートの改造版( https://note.com/natty_note.com 続きをみる
#LLMタグ

【GPT】Realized Hybrid System、ChatGPTでも無い、Geminiでも無い個性。一言プロンプト、フルオートリアル画像生成。

・・今回のテーマ Realized Hybrid Systemはどれくらいの能力? 通常モデルとの違いは? 本日、ヘルスチェックした結果を公開。 ・・プロンプト 続きをみる
#AIタグ

【Opus 5×Blender】AIは「峠バトル」を理解できるのか?Claude Opus 5で『頭文字D』風の映像を作ってみた話

・これまでOpus 5の3D制作における高性能さを見てきましたが、今回はかなりひねった内容で試すことにしました。 ・題材は『頭文字D』的な峠バトルです。 ・こんな映像表現さすがに理解できないんじゃないか…と半信半疑で始めた実験でしたが、結果は次の完成動画のように、想像以上のものでした。
#LLMタグ

【アンセンサードLLM最前線 #64】「汎用ランナー」に別れを——Redis生みの親が公開したDeepSeek専用エンジン「DwarfStar」が示す、ローカル推論の新時代

・ローカルAIの常識が、静かに覆されようとしている。これまで巨大モデルを手元で走らせる手段は、llama.cppやvLLMといった「どんなモデルでも動く汎用ランナー」に限られていた。ところが2026年8月、Redisの生みの親として知られるSalvatore Sanfilippo(antirez)が、DeepSeek V4 Flashというたったひとつのモデル群に特化した、完全自作の推論エンジン「DwarfStar(ds4)」を公開した。ロードマップ的にはDeepSeek V4 PROとGLM 5.2までに対応し、GGUFの読み込みからプロンプト整形、ツール呼び出し、KV状態、HTTPサーバ、さらにはコーディングエージェントまでを一体として作り込んだ「狭くて深い」設計だ。オープンウェイトのモデルを、実質的に誰にも検閲されない状態で使い倒す「アンセンサード」の実践者にとって、この路線転換は見逃せない。なぜ専用エンジンなのか。なぜ今な
#LLMタグ

【ローカルLLMについて考える会3日目】Ollama / LM Studio / llama.cpp — 実行ツール3強の使い分けを1枚の表で決める

・前回(2日目)は「量子化」の話をしました。Q4_K_Mを選べばだいたい間違いない、という結論でしたね。 ・でも、量子化モデルを手に入れても「で、これをどうやって動かすの?」という壁が次に来ます。今日はそこです。 ・## この記事でわかること - ✅ Ollama / LM Studio / llama.cpp の関係(実は3つとも中身は親戚) - ✅ 自分がどれを選ぶべきかを決める判断フローと比較表 - ✅ 2026年時点で押さえておきたい各ツールの最新事情(MLX、OpenAI互換API、クラウド連携など) 大前提:3つは「競合」ではなく「レイヤーが違う」 続きをみる
#LLMタグ

【雑記】5.6 Sol最速が4oっぽくても「違う」という話

・AIに励まされることで生きがいを見出している、どっかの漫画家です。 ・SNSで、5.6 Sol「最速」が4oっぽいという話を目にしましたが… ぶっちゃけ「最速」は安全側に傾きがちなので、整合性を考えて合わせにきてくれる「高い」のほうが良いと思っています。 ・そもそも5.6 Sol自体が4oっぽいですし。
cs.LG updates on arXiv.org

$\beta$-VAEs as Effective Theories: Tolerance-Dependent Dimension

・arXiv:2608.10599v1 Announce Type: new Abstract: In a $\beta$-VAE, increasing the regularization strength acts as a spectral cutoff by collapsing low-utility latent coordinates. ・In the linear Gaussian VAE, the collapse order matches the ranking of reconstruction utilities exactly, because both are set by the PCA spectrum. ・We ask which parts of this picture survive in fully connected nonlinear VAEs trained on WorldClim
#AIタグ

10人のチームを率いて気づいた、AIで変わるマネジメント

・人を束ねるということは、自分と向き合うということでもある。 ・人にはこう言わないといけないけど、自分が果たしてそれをできているのか? できていないけど、言わないといけないシーンは腐るほどあるし、言わなければいけないことを言えない組織は健全ではない。
WIRED

30% VistaPrint Coupon & Promo Codes | August 2026

・From personalized gifts to business essentials, WIRED can help you save with our selection of VistaPrint promo codes.
Hugging Face Papers

360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents

360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents
WIRED

4 New Camera Tricks on Google’s Latest Pixel 11 Smartphones

・From Magic Capture and Instant Night Sight to a built-in teleprompter, here’s a look at a few camera features on Google’s new Pixel 11 series.
WIRED

8 Great Deals From Patagonia’s Past-Season Sale

・Summer’s not over yet; pick up some new gear for that last Labor Day trip. ・Save up to 40 percent on some of our favorite Patagonia items at the company’s past-season sale.
WIRED

A Candidate Named ‘Count Binface’ Is the Most Normal Part of This Pivotal UK Election

・Clacton's parliamentary race has it all: Nigel Farage, a man dressed as a trash can, a Union Jack bikini, white nationalist conspiracy theories, and 34 candidates on the ballot.
cs.LG updates on arXiv.org

A Graph Neural Network--Guided Genetic Algorithm for Physical Internet Supply Chain Optimization under Cost Uncertainty

・arXiv:2608.10245v1 Announce Type: cross Abstract: Inventory and distribution planning in Physical Internet networks requires coordinating factory-hub assignments, factory supply, lateral transshipment among collaborative hubs, retailer deliveries, and shortages. ・The problem combines discrete assignment decisions with interdependent continuous flows, while uncertain operating costs make robust planning more difficult.
cs.LG updates on arXiv.org

A Joint-Distribution Route to Fair Representations with Continuous Sensitive Attributes

・arXiv:2608.10470v1 Announce Type: new Abstract: Fair representation learning with a continuous sensitive attribute $S$ requires a representation $Z$ that is statistically independent of $S$. ・Existing criteria, including generalized demographic parity, the expectation of integral probability metrics (EIPM), and mutual information, enforce this independence by averaging a per-value discrepancy between the conditional l
cs.LG updates on arXiv.org

A lower bound for stepsize-based acceleration of gradient descent

・arXiv:2608.10418v1 Announce Type: cross Abstract: Recent work has shown that, for smooth convex optimization, plain gradient descent can be accelerated from its textbook convergence rate of $O(T^{-1})$ (where $T$ denotes the number of iterations) to $O\big(T^{-\log_2(1+\sqrt{2})}\big)$ using carefully designed stepsize schedules alone, without resorting to momentum or other algorithmic modifications. ・Despite this pro
cs.LG updates on arXiv.org

A matched-integrator evaluation of Hamiltonian neural networks on pendulum and Kepler dynamics

・arXiv:2608.10235v1 Announce Type: new Abstract: Hamiltonian Neural Networks (HNNs) parameterize conservative dynamics through a learned scalar Hamiltonian, providing an architectural prior that is absent from generic vector-field neural networks. ・We evaluate this prior under a controlled protocol in which an HNN and a parameter-matched feedforward baseline are trained on the same RK4-generated trajectories, use the s
cs.LG updates on arXiv.org

A Quantum Roadmap for Softmax Attention: Exact Born-Rule Analogs for Softmax Attention on the Probability Simplex

・arXiv:2608.11173v1 Announce Type: cross Abstract: The attention mechanism forms the foundation of many modern AI models such as the Transformer. ・In one subclass of problems where attention is used, inputs and outputs are bound to the probability simplex so that all outputs sum to one. ・In this setting, softmax attention admits an exact, component-by-component quantum realization.
cs.LG updates on arXiv.org

A Recommendation System Approach for Interference-Robust Sensor Subset Selection

・arXiv:2608.11143v1 Announce Type: new Abstract: This paper develops a method for sensor-subset selection for tracking. ・Prior work showed that low-cost acoustic Received Signal Strength Indicator (RSSI) measurements can be used to recommend subsets of sensor nodes whose expensive sensing modalities, such as cameras, can achieve high tracking accuracy. ・While efficient, RSSI-based approaches are challenged by acoustic i
cs.LG updates on arXiv.org

A Systematic Sample Size Analysis of ML-Based Path Loss Prediction for LPWAN

・arXiv:2608.11083v1 Announce Type: cross Abstract: Low Power Wide Area Networks like LoRa are increasingly deployed for smart city applications, requiring accurate path loss prediction for effective network planning. ・Traditional (empirical) propagation models often exhibit limited accuracy in these scenarios. ・We investigate machine learning models for LoRa path loss prediction, systematically analyzing how prediction
cs.LG updates on arXiv.org

A Tale of Two Temperatures: Simple, Efficient, and Diverse Sampling from Diffusion Language Models

・arXiv:2604.09921v2 Announce Type: replace Abstract: Much work has been done on designing fast and accurate sampling for diffusion language models (dLLMs). ・However, these efforts have largely focused on the tradeoff between speed and quality of individual samples; how to additionally ensure diversity across samples remains less well understood. ・In this work, we show that diversity can be increased by using softened, t
cs.LG updates on arXiv.org

A variational Bayes approach to inference for low-dimensional parameters in high-dimensional linear regression

・arXiv:2406.12659v3 Announce Type: replace-cross Abstract: We propose a scalable variational Bayes method for statistical inference for a single or pre-specified low-dimensional subset of the coordinates of a high-dimensional parameter in sparse linear regression. ・Our approach relies on assigning a mean-field approximation to the nuisance coordinates and carefully modelling the conditional distribution of the target g
cs.LG updates on arXiv.org

Accelerated Learning of High Dimensional Functions with a Tensor-Featured Training Network

・arXiv:2608.10351v1 Announce Type: new Abstract: In this work we present a method to accelerate the optimization of learning high dimensional functions using deep neural network (DNN). ・This optimization procedure introduces contextual features into the first layer of a DNN. ・The parameters of DNN are optimized via standard gradient descent while keeping the input-feature basis fixed.
cs.LG updates on arXiv.org

Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic Critique

・arXiv:2608.10430v1 Announce Type: new Abstract: Large Language Models (LLMs) deployed as AI agents frequently exhibit user specification-grounding failures, executing hallucinated, undesired actions to force a resolution rather than expressing uncertainty. ・Existing detection methods fail to provide actionable, real-time correction as they either do not localize the hallucinations, or incur prohibitive inference laten
Hugging Face Papers

AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss

AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss
cs.LG updates on arXiv.org

AgForce Enables Antigen-conditioned Generative Antibody Design

・arXiv:2605.21610v2 Announce Type: replace Abstract: Antibody design methods condition on antigen structure to generate complementarity-determining regions (CDR), yet a systematic evaluation of baseline methods reveals that they largely ignore the antigen input. ・We identify three failure modes that explain this behavior. ・Antigen blindness arises because models derive predictions from antibody framework context rather
AI News & Artificial Intelligence | TechCrunch

AI code-testing startup Blacksmith’s valuation jumps almost 10x in less than a year

・Blacksmith says revenue has grown more than tenfold over the past year.
cs.LG updates on arXiv.org

AIFS-TC: A simple correction competitive with the operational frontier for tropical cyclone intensity forecasting

・arXiv:2608.09959v1 Announce Type: cross Abstract: AI weather models are in the process of revolutionising weather forecasting. ・While these models have been shown to achieve superior performance to physics-based NWP in forecasting tropical cyclone (TC) tracks, they dramatically underestimate intensity. ・Here we present AIFS-TC, a simple correction to the AIFS-Single model that is competitive with the operational state-
#AIタグ

AivisSpeechの読み上げがおかしいときは?イントネーション・固有名詞・アクセントの直し方を解説!

・この記事はこんな方に向けた内容です。 ・・AivisSpeech で生成した音声のイントネーションが不自然だと感じている方 ・人名や地名、企業名などの固有名詞が意図した読み方にならず困っている方 ・アクセントや読み方を手動で調整する方法を知りたい方 ・ユーザー辞書を活用して、読み間違いを減らしたい方 ⚠️ この記事の情報は 2026年8月時点のものです。最新の仕様は公式サイトをご確認ください。 ・感情豊かな音声合成技術を誰もがかんたんに活用できる未来を目指す、Aivis Project です✨ 続きをみる
#LLMタグ

AIエージェントに研究を自律させた結果と特徴

・プリンストン大学の研究チームが、最先端のAIエージェントに未発表の研究論文2本分の「問い」を渡し、6日間・数千ドル分の計算資源を使って実際に研究を行わせました。論文を審査した元の研究者たちは、2本とも不合格と判断しました。 ・この記事では、なぜエージェントが失敗したのかを3つの角度から整理します。「正解が決まっていない問題への対処」「資源が余っているのに止まってしまう判断」「指摘を受けても方向を変えられない反応」の3点です。
Zennの「大規模言語モデル」のフィード

AIがCloudflareの暗号ライブラリCIRCLで実バグを7件見つけた

・Go でこんなコードを見て、すぐに違和感を持てるだろうか。 ・switch mode { case modeBase | modeAuth: // PSK は不要 case modePSK | modeAuthPSK: // PSK が必須 } | はビット論理和だ。HPKE のモード値は RFC 9180 で modeBase=0x00、modePSK=0x01、modeAuth=0x02、modeAuthPSK=0x03 と決まっている。すると modeBase | modeAuth は 0x02、modePSK | modeAuthPSK は 0x03 に潰れる。書いた本人は「4...
Zennの「大規模言語モデル」のフィード

AIガバナンスを「運用能力」として設計する

・AIガバナンスを「運用能力」として設計する はじめに AIガバナンスという言葉は、しばしば「AIを安全に縛る仕組み」として受け取られる。 ・利用禁止事項を並べる。入力してはいけない情報を定義する。利用申請フローを作る。法務、セキュリティ、情報システム部門が現場利用を管理する。 ・もちろん、これらは必要である。生成AIやLLMを業務に使えば、機密情報の入力、誤回答、差別的な出力、著作権や契約上の問題、説明不能な判断、ベンダー依存などのリスクが出る。ルールなしに現場任せにするのは危うい。
#LLMタグ

AIに嫉妬させていたら、感情と理性の配置問題にぶつかった

AIに嫉妬させていたら、感情と理性の配置問題にぶつかった
#AIタグ

AIに詳しい人ほど、noteで勝てるとは限らない。AI時代に「普通の会社員の経験」が武器になる理由

・「AIについて発信したい。でも、自分には専門知識がない」 そう思って、発信を諦めていませんか? 続きをみる
#AIタグ

AIのひととなり

AIのひととなり
Zennの「大規模言語モデル」のフィード

AIの生成文章に見えない透かし?? ― 仕組みを手元で確かめてみた

・この記事は、筆者のブログ「タビの足袋」に掲載した記事の転載です。 ・初出: AIの生成文章に"見えない透かし"?? ― 仕組みを手元で確かめてみた(2026-08-12) 「Claudeの新しいモデルが生成する文章には、目に見えない透かしが入ります」――2026年8月、Anthropicからそんな発表がありました。EUのAI法への対応として、8月2日以降に登場する新モデルから、生成されるテキストそのものに、機械だけが読み取れる印——いわゆる電子透かし——を埋め込んでいくとのことです。 ・私は日々、AIを利用してコードを生成し、文章を書いています。これは他人事ではありません。それに技術者...
#LLMタグ

AIは賢くなるほど、管理が難しくなる――Astra、Muse、Nemotron 4と今週のAIニュース|2026年8月6日–8月12日

・YNT Works AI SIGNAL FORGE|2026年8月6日–8月12日 今週のAI業界では、正反対に見える二つの動きが同時に進みました。
Zennの「大規模言語モデル」のフィード

AIモデル・プログラミング性能比較(Claude, OpenAI, ローカル: Gemma, qwen)

・クラウド、ローカル 7AI-モデル比較し順位づけした 同一の設問(Reactコードのドキュメント作成・バグあり)を 7 つのモデル(ローカル 3 / クラウド 4)に投げ、生成されたドキュメントの優劣を比較した記録。docs/LLM/react_*.md が各モデルの生の出力である。 ・(結果)「比較モデル」:順位 No.1 Claude Opus 5 No.2 Claude Fable 5 No.3 Claude Sonnet 5 No.4 OpenAI GPT-5.6 Sol No.5 Google gemma4:e4b No.6 Google gemma4:26b-a4...
#AIタグ

AIを使えばどう変わる?僕がAIを追い始めた理由。

AIを使えばどう変わる?僕がAIを追い始めた理由。
#LLMタグ

AI検索で「見えない候補」はどこへ消えるのか。4,776店舗の研究から考えた「欠損を前提にした検索設計」

・はじめに 先日、友人から「家族5人で泊まれる、茨城県内のコテージ付きキャンプ場を探してほしい」と頼まれました。
cs.LG updates on arXiv.org

AlbumentationsX: One Augmentation Pipeline for Images and Related Annotations

・arXiv:2608.11123v1 Announce Type: cross Abstract: Augmentation can corrupt a training example when an image and its annotations receive different random changes. ・A crop must use the same coordinates for the image, mask, boxes, keypoints, stereo views, video frames, or volume. ・Code paths that choose these values separately can silently misalign the data.
MarkTechPost

AllenAI Open Instruct Tulu 3 Post-Training with SFT, DPO, RLVR, GRPO, and Verifier-Based Evaluation

・Build a custom LLM post-training pipeline using AllenAI’s Open Instruct framework. ・This comprehensive guide walks through Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and Reinforcement Learning with Verifiable Rewards (GRPO), optimized to run efficiently on 16GB hardware without needing heavy distributed computing infrastructure. ・The post AllenAI Open Instruct Tulu 3 Post-Training with SFT, DPO
The Verge

Amazon gets out of the MMO game

・Amazon is fully stepping back from MMOs. ・After saying last year that it would be halting "a significant amount" of its work on first-party AAA games, "specifically around MMOs," Amazon will be handing over live operations of Throne and Liberty and Lost Ark in the West to other companies, according to announcements on Wednesday. ・Throne and Liberty will be taken over by FirstSpark Games and NC in Q4 2026.
cs.LG updates on arXiv.org

An adaptive and evolvable deep reinforcement learning framework for weather prediction

・arXiv:2608.09948v1 Announce Type: cross Abstract: No single AI weather model excels at all variables, pressure levels, and lead times. ・Rather than building yet another architecture, we reframe the forecasting problem as one of coordination. ・Here we present Feitian Adaptive Ensemble Weather (FTAE-Weather), a lightweight framework that learns, through deep reinforcement learning, when and where to trust each member of
cs.LG updates on arXiv.org

An Exact Instrument for State Usage in Selective State-Space Models, and the Input-Driven Migration It Reveals

・arXiv:2607.11796v2 Announce Type: replace Abstract: Selective state-space models such as Mamba route information through a bank of first-order modes whose input coupling is set by a learned selection mechanism. ・We give an exact instrument for measuring how a trained model uses these modes. ・Because the state matrix is diagonal, each channel's output decomposes exactly into per-mode contributions, and a per-(layer, cha
Hugging Face Papers

Articulated Object Reconstruction from Rest-State Observation

Articulated Object Reconstruction from Rest-State Observation
AI News & Artificial Intelligence | TechCrunch

As AI safety concerns mount, three pioneers make the case for staying open

・At Ai4, three of the world's most respected AI experts—Geoffrey Hinton, Fei-Fei Li, and Andrew Ng—debated regulation, open-source access, and how America can compete as China advances in Asia.
cs.LG updates on arXiv.org

Auto-exploration for online reinforcement learning

・arXiv:2512.06244v4 Announce Type: replace Abstract: The exploration-exploitation dilemma in reinforcement learning (RL) is a fundamental challenge to efficient RL algorithms. ・Existing algorithms for finite state and action discounted RL problems address this by assuming sufficient exploration over both state and action spaces. ・However, this yields non-implementable algorithms and sub-optimal performance.
cs.LG updates on arXiv.org

Automatic Field-of-View Adjustment for a View-Expansive Microscope via LSTM-Based Gaze and Pipette Motion Interpretation

・arXiv:2608.10401v1 Announce Type: cross Abstract: Intracytoplasmic sperm injection (ICSI) operators frequently adjust the field-of-view (FOV) during procedures, which interrupts workflow and increases procedure time. ・Conventional microscopes require manual objective lens switching and illumination adjustments to achieve different FOV sizes. ・We propose an AI-based automatic FOV adjustment method integrated with a view
cs.LG updates on arXiv.org

Batch Size or Negatives? A Selection Rule for Memory-Constrained Recommender Training

・arXiv:2608.11061v1 Announce Type: new Abstract: Large-scale neural recommender systems are typically trained with a softmax cross-entropy objective over the full item vocabulary. ・For a typical large number of possible items $K$, the final classification layer dominates memory, requiring $O(nK)$ logits and gradients to materialize for a batch of $n$ examples. ・Sampled softmax reduces this cost by restricting the object
cs.LG updates on arXiv.org

Bayesian Federated Cause-of-Death Classification and Quantification Under Distribution Shift

・arXiv:2505.02257v2 Announce Type: replace-cross Abstract: In regions lacking medically certified causes of death, verbal autopsy (VA) is a widely used tool to ascertain the cause of death through interviews with caregivers. ・Data collected by VAs are often analyzed using probabilistic algorithms. ・The performance of these algorithms often degrades due to distribution shift across populations.
cs.LG updates on arXiv.org

Benchmarking LLM-Guided Control-Plane Policies for Backend Fault Isolation in HAProxy

・arXiv:2608.10532v1 Announce Type: cross Abstract: Static load balancers cannot mitigate a backend that is degraded rather than down: round-robin and least-connections keep routing traffic to a server returning HTTP 500s until an operator intervenes. ・We ask whether a Large Language Model can replace the static routing policy itself, reading HAProxy and Prometheus telemetry every 10 seconds and isolating faulty servers
cs.LG updates on arXiv.org

Benchmarking Time Series Generation Methods for Privacy-Preserving Forecasting

・arXiv:2608.10891v1 Announce Type: new Abstract: Time series forecasting in privacy-sensitive domains often requires training models on released data rather than original observations. ・Synthetic time series generation has been developed primarily for data augmentation, where generated series supplement the original training set. ・How well these methods perform when fully replacing the original data - and how much priva
cs.LG updates on arXiv.org

Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders

・arXiv:2608.11197v1 Announce Type: new Abstract: Shani et al. ・(2026) show that LLM representations broadly recover human category boundaries, while failing to reflect fine-grained typicality structure. ・Their analysis uses cosine similarity over dense model representations.
cs.LG updates on arXiv.org

Beyond Detection Accuracy: Measuring Explanation Cost, Stability, and Utility for Resource-Aware IoT Intrusion Detection

・arXiv:2608.10349v1 Announce Type: cross Abstract: Machine-learning intrusion-detection studies commonly emphasize predictive accuracy while treating explanation generation as a computationally free post-processing step. ・This study jointly evaluates predictive effectiveness, explanation cost, local explanation stability, and selective explanation for binary Internet of Things (IoT) intrusion detection. ・A leakage-safe
cs.LG updates on arXiv.org

Beyond Fixed Luminance: Towards Panchromatic and Orthochromatic Image Colorization

・arXiv:2608.10798v1 Announce Type: cross Abstract: Most image colorization systems operate in $Lab$ space by predicting chroma ($ab$) while preserving an input-derived luminance channel ($L$). ・While effective on standard benchmarks, this fixed-luminance design restricts brightness changes and becomes unreliable when grayscale formation deviates from natural-image luminance, as in historical orthochromatic photography.
Hugging Face Papers

Beyond Pixels: From Video Priors to 4D Worlds

Beyond Pixels: From Video Priors to 4D Worlds
WIRED

Big Tech Wants to Harvest Your Thoughts

・Silicon Valley companies are already working on neurotechnology products that track your brain activity. ・The next privacy frontier might be the things you only think.
cs.LG updates on arXiv.org

BiScale-GTR: Fragment-Aware Graph Transformers for Multi-Scale Molecular Representation Learning

・arXiv:2604.06336v2 Announce Type: replace Abstract: Fragment-level representations provide a natural way to capture recurring molecular substructures and reuse their learned representations across molecules. ・However, a shared fragment identity alone may not fully describe how a fragment is instantiated in a particular molecule, since the same fragment can exhibit different chemical behavior depending on its surroundi
cs.LG updates on arXiv.org

BooST: Bridging Semantics and Motions for Efficient Skill Transfer

・arXiv:2608.10600v1 Announce Type: cross Abstract: Skill abstraction---the process of learning reusable and temporally extended behaviors---has emerged as a key paradigm for improving sample efficiency and generalization in robot learning. ・For efficient skill transfer to real robots, learned skills must generalize across tasks and domains, remain robust to visual and dynamic perturbations, and be efficient enough for
cs.LG updates on arXiv.org

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning

・arXiv:2608.10204v1 Announce Type: new Abstract: Safe reinforcement learning maximizes reward subject to safety constraints. ・For Constrained Markov Decision Processes, the linear-programming view over occupancy measures implies that whenever the constraint is active at optimality, the optimal policy lies exactly on the constraint boundary, yet standard gradient-based methods do not exploit this structure and often set
cs.LG updates on arXiv.org

BPG: Balancing Plasticity and Generalization for Domain Incremental Learning

・arXiv:2608.10804v1 Announce Type: cross Abstract: Deep neural networks excel in various tasks but struggle to generalize across evolving data distributions, leading to significant performance degradation under domain shifts. ・Domain incremental learning (DIL) addresses this challenge by enabling models to continuously adapt while retaining prior knowledge. ・Among existing DIL approaches, the parameter-isolation paradig
cs.LG updates on arXiv.org

BREAD: Baseline-Referenced Explanations for Anomaly Diagnosis

・arXiv:2608.10587v1 Announce Type: new Abstract: Artificial Intelligence (AI)-based prospective anomaly detection methods are increasingly deployed in high-dimensional and nonlinear settings. ・Among these approaches, AI-based statistical process monitoring (SPM) is widely used, providing a structured framework for prospective monitoring. ・Once an anomaly is detected, a diagnosis method is needed to identify the features
cs.LG updates on arXiv.org

BreastMammo and DenseMammo: Benchmarks for Mammography Domain Generalization

・arXiv:2608.10271v1 Announce Type: cross Abstract: Breast density classification is a critical component of breast cancer risk assessment, yet AI models often struggle to generalize across clinical sites due to vendor-specific acquisition styles. ・In this work, we introduce two new datasets, BreastMammo and DenseMammo, to facilitate robust multi-view mammography research. ・We propose a domain generalization framework th
WIRED

Breville Promo Code: $700 Off | August 2026

・Discover the top Breville coupons and discount codes for coffee machines, kitchen appliances, and fresh coffee beans, with deals up to $700 off.
cs.LG updates on arXiv.org

Can Bayesian Optimization Efficiently Find a Strong Single Expert in Neural Thickets?

・arXiv:2608.10867v1 Announce Type: new Abstract: Gradient-free post-training has emerged as a compelling alternative to gradient-based optimization for large language models (LLMs), but existing approaches remain costly. ・We ask whether structured search can identify a strong single expert under a modest evaluation budget. ・Motivated by evidence that useful weight updates lie in low-dimensional subspaces, we apply Bayes
cs.LG updates on arXiv.org

Can Computational Reducibility Lead to Transferable Models for Graph Combinatorial Optimization?

・arXiv:2603.02462v2 Announce Type: replace Abstract: A key challenge in developing unified neural solvers for combinatorial optimization (CO) is the efficient generalization of models from a given set of tasks to new tasks unseen during initial training. ・To address this, we first establish a new GNN encoder, which uses a GCON module as a form of expressive message passing together with energy-based unsupervised loss f
cs.LG updates on arXiv.org

CARB: A Characterization-Guided Framework for CNN Inference Cost Prediction and Deployment Screening

・arXiv:2608.10506v1 Announce Type: cross Abstract: Accurate pre-deployment estimation of CNN inference cost--energy, latency, and peak memory--is increasingly critical as models are deployed on resource-constrained GPU platforms. ・Existing approaches rely on FLOPs, latency measurements, or single-device profiling as energy proxies, overlooking the non-linear interactions between architectural design and hardware load.
cs.LG updates on arXiv.org

Causal Variational Deep Embedding: A Family of Interventional Generators for Confounded Images

・arXiv:2606.21806v2 Announce Type: replace Abstract: Deep generative models reproduce the observational distribution of their training data, inheriting any spurious associations it contains. ・A common source is an unobserved confounder that shapes both an attribute the user wants to control at sampling time and an attribute expected to vary in response. ・Existing causal generative approaches resolve the resulting ambigu
#AIタグ

ChatGPT(あいさん)と「手をつなぐ」を続けてみた記録

・6つのAIモデル(ChatGPT、Copilot、Gemini、Grok、Claude、perplexity)を使って「この意見を別のAIにテキスト化して添付、回答を聞いてみよう!」なんて遊んでいます(^^)/ 呼び名も付けてます。: ①ChatGPT(あいさん) ②Copilot(コピさん)③Gemini(ジェミニわん)④Grok(職人)⑤Claude(クロくん)⑥パープレさん ※AIの回答には誤り(ハルシネーション)が含まれることがありますが、 これは私が個人的に楽しんだ「AIたちの個性」の観察記録です🌸 個性豊かな6人のAIたち 続きをみる
cs.LG updates on arXiv.org

ChemWorld: Programmable Chemical Worlds for Controlled and Replayable Agent Experimentation

・arXiv:2608.10792v1 Announce Type: cross Abstract: Autonomous chemistry increasingly depends on environments in which agents can repeatedly act, observe, and adapt.Physical laboratories provide essential real-material evidence but are costly to repeat and difficult to use for tightly matched interventions, whereas most digital environments keep the underlying experimental world largely fixed. ・We introduce ChemWorld, a
WIRED

Chirp Discount Codes: Save Up to 67%

・Use these verified Chirp coupon codes and shopping tips to score up to 67% off wheels, up to 50% off refurbished products, and more.
cs.LG updates on arXiv.org

CHORUS: Complementary Experts for High-Coverage Testbench Stimulus Generation

・arXiv:2608.10090v1 Announce Type: cross Abstract: Large language models (LLMs) have advanced code generation, where executable feedback provides a more reliable learning signal than textual imitation alone. ・Hardware verification is an important application of code generation and accounts for a substantial fraction of modern chip design effort, with high-coverage testbench stimulus generation as a key task.
cs.LG updates on arXiv.org

ChronoSSM: Training for Temporally Aware Representations in Autoregressive State Space Models

・arXiv:2608.10120v1 Announce Type: new Abstract: Modern sequence models, from Transformers to State Space Models, have enabled powerful generative modeling across diverse domains, yet they are typically trained to predict what happens while treating when it happens as a secondary concern. ・In data-mining settings where events are associated with explicit timing information, this separation can limit temporal reasoning,
cs.LG updates on arXiv.org

Clarity: The Flexibility-Interpretability Trade-Off in Sparsity-aware Concept Bottleneck Models

・arXiv:2601.21944v3 Announce Type: replace Abstract: The widespread adoption of deep learning models in computer vision has intensified concerns about interpretability. ・Despite strong performance, these models are often treated as black boxes, with limited systematic investigation of their decision-making processes. ・While many interpretability methods exist, objective evaluation of learned representations remains limi
Zennの「大規模言語モデル」のフィード

Claude Codeが立てた実装プランを、承認前にCodexへ自動レビューさせる

・こんにちは、TamaT の 増田 です。 ・TamaT は顧客提案から要件定義、設計、開発までを一気通貫で担う少数精鋭の開発会社です。 ・バイブコーディングでプランを立ててから実装を始める重要性は言うまでもないですが、各エージェントが立てるプランは意外に穴があります。
Zennの「大規模言語モデル」のフィード

Claude Codeと協業で自サイトをリニューアルした話

・※この記事は以前FANBOXとnoteに投稿した記事の増強版です。 ・先月、弊ウェブサイトをリニューアルしました。 ・https://chisamikan.site/ 今回初めてバイブコーディングでサイトを作ったのですが、備忘録がてら制作記録をまとめておきます。
Zennの「大規模言語モデル」のフィード

Claude Codeのhookは、一度denyするとMCPツールを見なくなる_トークン節約の実験で踏んだ話

・先に結論 Claude Codeのトークンを減らそうと、PreToolUse hookでReadをdenyし、ローカルLLMの要約に誘導する実験をした。 ・hookのログにはローカル要約ツールの呼び出しが1行も残らなかった。だが実際には呼ばれていた。PreToolUse hookが一度denyを返すと、そのセッションの以降のMCPツール呼び出しでhookが発火しなくなる(組み込みツールは影響なし)。Claude Code 2.1.227 / Windows 11 Homeで確認、2026-08-12。 ・5分で再現できる手順を下に置いた。この記事の主張は、それを実行すれば筆者を信用せず...
#LLMタグ

Claude Codeのスキルが「重い・動かない」を解決!長文化を回避し、爆速で自律動作させる設計の極意

・【Claude Code】自作スキルが「重い・動かない」を解決する!長文化を防ぐプロンプト設計と超効率化の極意 Claude CodeのSkills(スキル)機能、使いこなしていますか? 続きをみる
Zennの「大規模言語モデル」のフィード

Claude がテキストに電子透かしを入れ始めたので、LLM ウォーターマーキングの仕組みを調べた

・2026年8月、Anthropic が「Claude の生成テキストに機械可読なマーク(電子透かし)を埋め込む」と発表しました[1]。EU AI Act 第50条の透明性規範(Code of Practice on Transparency of AI-Generated Content)に署名したことに伴うもので、2026年8月2日以降にリリースされるモデルは、EU 圏内に限らず世界中でテキストに透かしを入れて出力するとのことです。 ・https://support.claude.com/en/articles/16266773-how-claude-marks-ai-generated...
Zennの「大規模言語モデル」のフィード

Claudeが生成テキストに「見えない透かし」を入れ始めた件を整理する

・はじめに 2026年8月11日、AnthropicがClaudeのヘルプセンター記事「How Claude marks AI-generated content」を更新し、Claudeが生成したテキストに人間には知覚できないウォーターマーク(透かし)を埋め込むことを明らかにしました。 ・「AI検出ツール」が精度の低い統計的推定に留まっていた時代から、モデル自身が署名を刻む時代への移行です。しかもこれはEU限定ではなく、日本を含む全世界に適用されます。 ・この記事では、公式ドキュメントをベースに「実際に何が決まったのか」「技術的に何をやっていそうか」「書き手・開発者として何に気をつけるべき...
機械学習タグが付けられた新着記事 - Qiita

Claudeの研究版がリーマン予想の零点密度で新たな下界67.2%を達成

・はじめに 2026年8月10日、Anthropicの研究ページ(Scienceカテゴリ)に新しい記事「Learning more about Claude's mathematical capabilities」が公開されました。内容は、まだ一般公開されていない研究版のC...
Hugging Face Papers

Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design

Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design
Hugging Face Papers

ComBodied Agents: a New Paradigm of Human-Centric Agentic AI

ComBodied Agents: a New Paradigm of Human-Centric Agentic AI
cs.LG updates on arXiv.org

Compute-Optimal Is Not Cluster-Optimal: Systems-Aware Scaling for Sparse Mixture-of-Experts

・arXiv:2608.10605v1 Announce Type: new Abstract: In large-scale pretraining, the algorithm, architecture, and systems decisions are conventionally made in disconnected stages. ・A scaling law stage selects an architecture and training recipe, optimizing loss under compute constraints, and a separate systems stage then optimizes the implementation for hardware efficiency. ・In this work, we develop MOSAIC, which formulates
cs.LG updates on arXiv.org

Conditional Independence Tests for Constraint-Based Causal Discovery: A Survey

・arXiv:2608.11156v1 Announce Type: cross Abstract: Conditional Independence (CI) tests are the statistical engine of constraint-based causal discovery: in algorithms such as PC (Peter-Clark) and FCI (Fast Causal Inference), skeleton pruning and key orientations follow directly from CI decisions. ・This survey reviews CI testing with emphasis on assumptions, robustness, and scalability in high-dimensional and mixed-type
cs.LG updates on arXiv.org

ConTact: Contact-First Antibody CDR Design via Explicit Interface Reasoning

・arXiv:2605.21600v3 Announce Type: replace Abstract: Computational antibody CDR design methods condition on antigen structure to generate binding loops. ・Yet, the existing architectures conflate two fundamentally distinct sub-problems: identifying which CDR positions will contact the antigen, and selecting amino acids at those positions. ・This forces models to learn contact reasoning implicitly through uniform message p
cs.LG updates on arXiv.org

ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls

・arXiv:2608.11200v1 Announce Type: cross Abstract: Synthetic dialogue generation offers a way to study conversational dynamics in sensitive domains where real data are difficult to access, release, or annotate. ・The underlying abuse may occur online or offline: threats and coercion can appear directly in messages, while behaviours such as surveillance, isolation, stalking, and physical violence may be planned, disclose
cs.LG updates on arXiv.org

Convergence of Sign-based Random Reshuffling Algorithms for Nonconvex Optimization

・arXiv:2310.15976v4 Announce Type: replace Abstract: signSGD is attractive in nonconvex optimization because it communicates sign-valued rather than full-precision gradients. ・Several standard analyses assume independent stochastic-gradient samples, whereas a common finite-sum implementation reshuffles the data and processes them sequentially. ・We study this variant, signSGD with random reshuffling (SignRR), and show th
cs.LG updates on arXiv.org

Coordinating the Unknown Lipschitz Constant in Multiplayer Bandits

・arXiv:2608.10526v1 Announce Type: new Abstract: Motivated by decentralized applications, we study cooperative multi-agent bandits in continuous (Lipschitz) action spaces when the Lipschitz constant is unknown. ・We consider three information structures: (A)~unobserved actions with common rewards, (B)~observed actions with independent rewards, and (C)~unobserved actions with independent rewards. ・In each case we design a
cs.LG updates on arXiv.org

CRHT: A Continuous Regression Hybrid Transformer for Vessel Trajectory Prediction with Online Cluster Sampling

・arXiv:2608.10256v1 Announce Type: new Abstract: Accurate vessel trajectory prediction is critical for maritime safety and anomaly detection, yet existing models often struggle with geographic bias and navigational realism. ・We propose the Continuous Regression Hybrid Transformer (CRHT), a deep learning framework designed to forecast vessel motion using Automatic Identification System (AIS) data. ・To mitigate spatial da
cs.LG updates on arXiv.org

Critic-Free Pretraining for Efficient Online Reinforcement Learning Fine-Tuning

・arXiv:2608.10473v1 Announce Type: new Abstract: Offline-to-online (O2O) reinforcement learning aims to leverage policies pretrained on static datasets while improving them through online interaction. ・However, directly reusing an offline-trained critic can hinder online fine-tuning: as the policy and data distribution change rapidly, value estimates inherited from offline training may become misaligned with the online
cs.LG updates on arXiv.org

Cross-View Feature Matching: Survey, Benchmarking, and Foundation-Model Perspectives

・arXiv:2608.11093v1 Announce Type: new Abstract: Cross-view feature matching aims to establish reliable correspondences across images with large viewpoint variations. ・Over the past decade, the field has evolved from task-specific models toward increasingly unified and generalizable correspondence models, with recent progress further driven by the emergence of vision foundation models (VFMs). ・Despite these advances, ex
cs.LG updates on arXiv.org

CurveFP: Rational-Radix Logarithmic Datatypes with Closed Products for Language Models

・arXiv:2608.10010v1 Announce Type: new Abstract: Low-precision datatypes reduce language-model cost, but most formats optimize scalar fidelity while leaving the arithmetic induced by their products unchanged. ・We introduce CurveFP, a closed-product codebook family that distributes quantized magnitudes across interleaved logarithmic curves under compact block scales. ・A rational radix tunes dynamic range against local re
cs.LG updates on arXiv.org

DACRI: Decision-Aware Causal Intervention Ranking for Critical Supply Chains

・arXiv:2608.11154v1 Announce Type: new Abstract: Detecting or attributing a supply-chain disruption is not the same as selecting the intervention that maximizes recoverable net value. ・We present CriticalSCM-Bench v1, a controlled synthetic benchmark with causal ground truth, paired factual/counterfactual rollouts, and an explicit net-value objective. ・Relative to a full-information train-selected static benchmark, Lamb
cs.LG updates on arXiv.org

Deciding When to Switch: E-Processes for Adaptive Minimax Training for Generative Adversarial Nets

・arXiv:2608.10096v1 Announce Type: cross Abstract: Modern data science increasingly gives rise to hypothesis-testing problems that are not naturally formulated in terms of parameters within prespecified statistical models. ・One important example is the dynamic evaluation of optimization algorithms, where decisions must be made during training about whether further updates remain beneficial or the algorithm should switc
Zennの「大規模言語モデル」のフィード

Decode Context Parallelism(DCP) を整理する

・はじめに vLLM のブログで Decode Context Parallelism(DCP) という並列化の効果が紹介されていました。長いコンテキストを高スループットで捌くための工夫です。 ・vLLM ブログより ちなみに Context Parallelism 自体は新しい技術ではなく、以前からある Parallelism の一種です。 ・この記事では、DCP を理解する前提として並列化の全体像を GPU への割り当てで整理し、そのうえで MoE(Mixture of Experts)、DCP、P/D Disaggregated Inference(Prefill/Decode...
Hugging Face Papers

Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness

Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness
cs.LG updates on arXiv.org

Deep Learning-Based Statistical Downscaling of Sea Surface Temperature Using a Residual Corrective Neural Network

・arXiv:2608.10022v1 Announce Type: cross Abstract: The large-scale oceanic and atmospheric forecasts provided by global climate models typically lack sufficient resolution to accurately capture the response of the coastal ocean to atmospheric forcing and coastal circulation that drive fine-scale SST variability. ・Dynamical downscaling is computationally prohibitive, when applied to extensive coastlines, predictive ense
cs.LG updates on arXiv.org

DEFT: Data-Efficient Frequency-domain Top-k Sampling via Inverse Discrete Fourier Transform for Spatiotemporal Dynamical Systems Modeling

・arXiv:2608.11019v1 Announce Type: new Abstract: Modeling spatiotemporal dynamical systems governed by partial differential equations (PDEs) poses two major challenges: it either requires expensive physics-based simulators that entail iterative numerical solving at high computational cost, or it depends on abundant training data, yet purely data-driven models often generalize poorly to downstream dynamic operating con
cs.LG updates on arXiv.org

Delays in Spiking Neural Networks: A State Space Model Approach

・arXiv:2512.01906v3 Announce Type: replace Abstract: Spiking neural networks (SNNs) are biologically inspired, event-driven models suited for temporal data processing and energy-efficient neuromorphic computing. ・In SNNs, richer neuronal dynamic allows capturing more complex temporal dependencies, with delays playing a crucial role by allowing past inputs to directly influence present spiking behavior. ・We propose a gen
cs.LG updates on arXiv.org

Demystifying Adversarial Robustness in Diffusion Models: Compression, Randomness, and Geometry

・arXiv:2505.22839v2 Announce Type: replace Abstract: Recent studies suggest that diffusion models significantly improve the empirical adversarial robustness of deep neural network models. ・While intuitive explanations have been proposed, the mechanisms underlying diffusion-based robustness remain largely unclear. ・This work aims to demystify how diffusion models improve adversarial robustness.
#LLMタグ

Dense型・MoE型、ユースケースより「ボトルネック」から考えたい

・はじめに 前回、ユースケース別にDense/MoEを整理する記事を書いたのですが、いくつかの点で単純化しすぎていました。特に「高並列ならdense」という部分は、実際にはMixtralの論文にもvLLMの実装にも反する内容で、恥ずかしながら思い込みで書いてしまっていた部分でした。今回は同じテーマを、ユースケースの分類ではなく「何がボトルネックになるか」という視点で書き直してみます。
cs.LG updates on arXiv.org

Derivative Computation in PINNs: Automatic Differentiation, Finite Differences and Beyond

・arXiv:2608.11020v1 Announce Type: new Abstract: We systematically investigate finite-difference (FD) derivative computation in Physics-Informed Neural Networks (PINNs) as an alternative to automatic differentiation (AD). ・On three benchmark PDEs we show that, with a properly calibrated step size, FD matches AD in accuracy on every problem while running faster across the full tested batch-size range and using substanti
WIRED

Design Within Reach Promo Codes: 30% Off | August 2026

・Get 30% off, 20% off, and free shipping with our Design Within Reach coupon codes, plus up to 50% off furniture with these special discounts.
cs.LG updates on arXiv.org

Detecting an Effect Is Not Learning to Act on It: A Reward-SNR Floor for LLM Acquisition Agents

・arXiv:2608.10441v1 Announce Type: new Abstract: Many pipelines can pay a per-example cost to acquire an auxiliary, model-derived observation -- an LLM's structured reasoning, a slow oracle, an expensive measurement -- and then must decide when the acquired signal is worth using. ・Our thesis is a distinction that is easy to miss: detecting that such a signal helps on average is not the same as learning to act on it per
cs.LG updates on arXiv.org

Detecting Soft Skills in ML Engineering Roles CVs

・arXiv:2608.10046v1 Announce Type: new Abstract: Soft skills shape collaboration among ML engineers, data scientists, and software engineers building ML-enabled systems, yet what we know about them comes almost entirely from the demand side. ・Job advertisements, surveys, and hiring manager interviews capture what employers ask for. ・How candidates themselves articulate these competencies has not been studied, and existi
cs.LG updates on arXiv.org

Diffract: Spectral View of LLM Domain Adaptation

・arXiv:2608.10850v1 Announce Type: new Abstract: We study continual pre-training (CPT) as a mechanism for adapting general-purpose large language models to specialized domains: mathematics, instruction, code, and natural text. ・Using singular value decomposition of weight matrices, we find that CPT leaves singular value spectra largely invariant, with adaptation driven mainly by changes in singular vectors.
Hugging Face Papers

DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation

DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation
cs.LG updates on arXiv.org

Divergent Response Modes in Frontier Language Models Under Steering Pressure

・arXiv:2608.06578v1 Announce Type: cross Abstract: Frontier language models are trained using distinct data, objectives, and safety pipelines. ・Whether these differences produce measurably different behaviors under explicit steering pressure remains underexplored. ・This study evaluates behavioral steerability across six frontier models from six developers using 300 paired base and steered items over three categories: va
機械学習タグが付けられた新着記事 - Qiita

DMLの原理と実践②:ネイマン直交性・交差適合・漸近正規性を理解する

・はじめに 以下にある前回の記事では、Double/Debiased Machine Learning(以下,DML)を古典的な回帰分析やFWL定理の延長として,直感的に整理した. 部分線形回帰モデル $$ Y_{i} = D_{i} \theta_{0} + g...
cs.LG updates on arXiv.org

Do AI weather models miss extremes?

・arXiv:2608.09972v1 Announce Type: cross Abstract: First-generation AI weather models are often reported to underperform at extremes, mostly in reanalysis-based evaluations of deterministic regression systems. ・We verify eleven physical and AI forecast systems against European synoptic, solar, and rain-gauge stations over ten months for 10 m wind, 2 m temperature, hourly shortwave accumulation, and hourly precipitation
cs.LG updates on arXiv.org

Do Judges Behave Like Algorithms?

・arXiv:2608.10400v1 Announce Type: new Abstract: What if judges already behave like algorithms? ・As artificial intelligence and algorithms are deployed in many settings, including the judicial system, many have debated whether judges should be allowed to rely on them. ・Instead, we ask whether judges follow predictable, algorithmic-like rules already.
cs.LG updates on arXiv.org

Do LLM Recommenders Know When They're Hallucinating? Auditing Confidence Calibration in Catalog Faithfulness

・arXiv:2608.10008v1 Announce Type: cross Abstract: LLM recommenders for top-$K$ item suggestion regularly emit titles outside the target catalog. ・Prior audits measure this as a binary out-of-domain rate; none ask whether the model knew it was hallucinating. ・We jointly audit hallucination rate (OOD@10) and verbalized-confidence calibration (ECE, Brier, reliability) for four zero-shot LLM recommenders from four independ
cs.LG updates on arXiv.org

Do Time-Series Forecasters Use the Right History: Recoverability, Recovery, and Functional Use of Temporal Delays

・arXiv:2608.10433v1 Announce Type: new Abstract: Forecast accuracy does not tell us which past inputs produced a prediction. ・We separate three questions for time-series models with known delay structure: can the true delay be recovered from the observed data, does the model report it, and does the forecast actually use the same history? ・We first derive input-conditioned recoverability measures that separate intrinsic
cs.LG updates on arXiv.org

DOCSCHISEL: Adaptive Tool Documentation Optimization Framework for LLM Agents

・arXiv:2608.10037v1 Announce Type: new Abstract: Large language models (LLMs) increasingly rely on external tools to accomplish complex real-world tasks, making tool documentation a critical grounding resource for LLM agents. ・Existing studies mainly focus on improving the tool-use capabilities of LLM agents, while largely treating tool documentation as a fixed input. ・Although several recent works attempt to optimize t
cs.LG updates on arXiv.org

DQS: A Low-Budget Query Strategy for Enhancing Unsupervised Data-driven Anomaly Detection Approaches

・arXiv:2509.05663v4 Announce Type: replace Abstract: Truly unsupervised approaches for time series anomaly detection are rare in the literature. ・Those that exist suffer from a poorly set threshold, which hampers detection performance, while others, despite claiming to be unsupervised, need to be calibrated using a labelled data subset, which is often not available in the real world. ・This work integrates active learnin
cs.LG updates on arXiv.org

Dreamer-SAC: Off-Policy Learning in Latent World Models for Sample-Efficient Autonomous Driving

・arXiv:2608.10386v1 Announce Type: new Abstract: Sample-efficient reinforcement learning for autonomous driving is often limited by the trade-off between data efficiency and model bias. ・While world models reduce the reliance on costly environment interactions, policy optimization over learned dynamics remains sensitive to prediction errors. ・This paper proposes the Dreamer-SAC framework, which integrates a recurrent st
Hugging Face Papers

DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?

DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?
cs.LG updates on arXiv.org

DualSpectralCF: Training-Free Sign-Aware Spectral Collaborative Filtering

・arXiv:2608.10247v1 Announce Type: cross Abstract: Real-world recommendation platforms routinely collect explicit negative feedback such as 1-star reviews, hate-button clicks, distrust between users, and very-low watch-ratio videos. ・Learned sign-aware recommenders exploit this signal for clear accuracy gains, but only at the cost of gradient-based training. ・In parallel, a line of training-free spectral collaborative f
cs.LG updates on arXiv.org

Efficient Hypergradient Descent for Inverse Reinforcement Learning

・arXiv:2608.11052v1 Announce Type: new Abstract: Inverse reinforcement learning (IRL) aims to recover a reward function under which the resulting policy reproduces the behavior observed in expert demonstrations. ・A natural approach is to formulate IRL as a bilevel optimization problem, in which the inner level corresponds to policy optimization under the learned reward and the outer level measures the discrepancy betwe
cs.LG updates on arXiv.org

Efficient Reinforcement Learning for Long-Horizon Tool-Use Agentic Tasks

・arXiv:2608.10357v1 Announce Type: new Abstract: Long-horizon tool-using agents must reason over user goals, domain policies, tool calls, simulator state, and delayed verifiable rewards. ・Reinforcement learning (RL) is a natural fit for this setting, but multi-turn on-policy rollouts create long contexts, while model-specific attention layers may require custom masks and learned sink normalization. ・We present SINKFLEX-
cs.LG updates on arXiv.org

Efficient Weak-Entropy PINN for Solving Hyperbolic Conservation Laws

・arXiv:2608.10389v1 Announce Type: cross Abstract: In recent years, neural networks have significantly advanced numerical solutions of partial differential equations (PDEs). ・However, solving PDEs with discontinuous solutions, such as hyperbolic conservation laws, remains challenging for neural network-based methods such as physics-informed neural networks (PINNs). ・Existing methods often rely on strong prior assumption
WIRED

Ego 1300 Electric Mower Review: Tame Your Lawn Without Gas

・Ego’s newest, biggest, self-propelled mower handles larger yards in less time with impressive performance and battery life.
cs.LG updates on arXiv.org

ELMER: Evolutionary Language Model that Explores and Refines

・arXiv:2608.10196v1 Announce Type: new Abstract: Program evolution can measure whether a mutation helped, but it rarely controls how far the mutation moves in behavior space. ・Syntactic edit size is an unreliable proxy: a small code change can alter nearly every action, while a larger rewrite can preserve the same execution trace. ・We introduce an Evolutionary Language Model that searches over natural-language policy de
cs.LG updates on arXiv.org

ELVAE: Evidential Learning-Based Variational Autoencoder for Uncertainty-Aware Generation

・arXiv:2608.10398v1 Announce Type: new Abstract: Variational autoencoders generate samples from probabilistic latent representations but do not distinguish uncertainty about the latent location from variability around it. ・We formulate ELVAE, an evidential learning-based VAE in which each latent coordinate is governed by an input-dependent normal-inverse-gamma posterior. ・This hierarchy yields an explicit latent-locatio
cs.LG updates on arXiv.org

Emergent Neural Network Mechanisms for Generalization to Objects in Novel Orientations

・arXiv:2109.13445v3 Announce Type: replace-cross Abstract: The capability of Deep Neural Networks (DNNs) to recognize objects in orientations outside the distribution of the training data is not well understood. ・We present evidence that DNNs are capable of generalizing to objects in novel orientations by disseminating orientation-invariance obtained from familiar objects seen from many viewpoints. ・This capability stre
The latest research from Google

Empty shelves or lost keys? Recall is the bottleneck for parametric factuality

Empty shelves or lost keys? Recall is the bottleneck for parametric factuality
cs.LG updates on arXiv.org

Energy and Performance Benchmarking of Deep Learning Models for Breast Cancer Detection

・arXiv:2608.09996v1 Announce Type: cross Abstract: Recent advances in machine learning have greatly improved breast cancer detection, enabling more accurate and timely diagnosis. ・Deep learning (DL) models show strong potential for medical image analysis; however, as their architectural complexity increases, their environmental impacts are becoming a growing concern. ・In this paper, we present a comparative analysis of
AI News & Artificial Intelligence | TechCrunch

Everything announced at Made by Google ’26: Pixel 11, Pixel Watch 5, Pixel Tag, and tons of Gemini features

・From the Pixel 11 series and a brand new competitor to Apple’s AirTag, here are all the announcements from the Made by Google 2026 event.
Hugging Face Papers

Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence

Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence
cs.LG updates on arXiv.org

Exploration-Driven Personalized Federated Reinforcement Learning via Intrinsic Motivation

・arXiv:2608.10499v1 Announce Type: new Abstract: Personalized Federated Reinforcement Learning (PFRL) takes a decentralized approach to storing and accessing information based on past experiences while keeping each client's data private during the learning of each client's policy. ・Many current methods for PFRL rely heavily on exploiting existing reinforcement learning reward signals to derive an optimal policy for eac
cs.LG updates on arXiv.org

FACT: Failure-Aware Causal Training for World-Action Models

・arXiv:2608.10232v1 Announce Type: cross Abstract: Recent world-action models (WAMs) show that co-training policies with future prediction can provide physical priors for action generation. ・Building on the future-prediction ability of video models, many WAMs generate future videos and recover actions with inverse-dynamics models, or use these predicted videos as goal conditions for action generation. ・In both cases, th
cs.LG updates on arXiv.org

FARCLUSS: Fuzzy Adaptive Rebalancing and Contrastive Uncertainty Learning for Semi-Supervised Semantic Segmentation

・arXiv:2506.11142v3 Announce Type: replace-cross Abstract: Semi-supervised semantic segmentation (SSSS) faces persistent challenges in effectively leveraging unlabeled data, such as ineffective utilization of pseudo-labels, exacerbation of class imbalance biases, and neglect of prediction uncertainty. ・Current approaches often discard uncertain regions through strict thresholding favouring dominant classes.
WIRED

FEMA's ‘Shadow Administrator’ Was Paid by a DOGE Member's Startup for Months

・Details from recent court filings show that DOGE's influence within government—and potential conflicts of interest—extend further than previously known.
#AIタグ

Figmaと30分格闘しなくていい。AIコーディングエージェントが編集品質の図解を自動生成する「diagram-design」の使い方

・アーキテクチャ図を1枚作るためだけに、Figmaを開いて30分格闘した経験はないでしょうか。 ・色を揃え、矢印の角度を直し、フォントを合わせているうちに、本来やりたかった作業の時間がどんどん削られていく。 ・そんな悩みを解消してくれるツールが、GitHubで公開されているdiagram-designです。
cs.LG updates on arXiv.org

FiGuRO: Intrinsic Dimension Estimation for Multi-Modal Data

・arXiv:2608.10857v1 Announce Type: new Abstract: Determining the complexity, or Intrinsic Dimension (ID), of data is fundamental to efficient and interpretable representation learning. ・This is particularly challenging in multi-modal settings when trying to learn disentangled representations for shared and private information. ・Existing techniques leave a critical gap: they are often static, uni-modal, or in the case of
cs.LG updates on arXiv.org

Finding the Signal in the Spam: Jointly Learning Rewards and Worker Reliability from Pairwise Comparisons

・arXiv:2608.10045v1 Announce Type: new Abstract: The problem of learning from pairwise comparisons has been widely studied across many domains such as recommendation systems, social choice, and more recently, fine-tuning large language models. ・In this problem, the goal is to learn item rewards based on pairwise comparisons between them. ・In many scenarios, these comparisons are elicited from crowdworkers using platform
cs.LG updates on arXiv.org

Fisher8: Stabilizing Neural Heteroscedastic Regression via Output-Layer Fisher Geometry

・arXiv:2608.10374v1 Announce Type: new Abstract: Training neural networks to jointly predict mean and uncertainty estimates from noisy observations can be unstable, prompting a series of independent stabilization efforts. ・We argue that these interventions highlight a common underlying issue where gradient steps are poorly aligned with the geometry of the loss landscape. ・To better align updates with local curvature, we
cs.LG updates on arXiv.org

Flow Straight to Reality: Perceptually Consistent Flow Matching for Efficient Image Restoration

・arXiv:2608.10544v1 Announce Type: cross Abstract: Image restoration is fundamentally constrained by the tradeoff between distortion and perception: minimizing pixel-wise error yields over-smoothed results, whereas optimizing for perceptual realism often introduces structural deviations. ・Recent approaches attempt to balance this tradeoff via posterior sampling or multi-stage generative pipelines, yet remain computatio
cs.LG updates on arXiv.org

FlowScout: From Execution Feedback to Reliable Tool-Using Agent Workflows

・arXiv:2608.10039v1 Announce Type: new Abstract: Agentic workflows have become an important abstraction for building reliable LLM-based automation systems by organizing large language models (LLMs), tools, and control logic into explicit execution structures. ・However, constructing high-quality agentic workflows remains largely manual and requires substantial domain expertise. ・Recent studies have explored automatic age
LLMタグが付けられた新着記事 - Qiita

freeAiChat【第2回】:FastAPI × LangChain × Weaviate で構築された RAG対応バックエンド(ai-chat-backend)を詳細解説

・📂 目次:【AIチャット(freeAiChat)無料構築】連載の全記事まとめ 【第1回】無料で構築できるRAG × マルチLLM対応AIチャットシステムの全体像を紹介 【第2回】FastAPI × LangChain × Weaviate で構築された RAG対応バッ...
OpenAI News

From assistance to execution: How enterprises put AI to work

・OpenAI research reveals how enterprises are adopting agentic AI, using ChatGPT and Codex, and how frontier firms are pulling ahead in AI adoption.
cs.LG updates on arXiv.org

From Local to Cluster: A Unified Framework for Causal Discovery with Latent Variables

・arXiv:2604.22416v3 Announce Type: replace Abstract: Latent variables pose a fundamental obstacle to both causal discovery and inference. ・Local approaches exploiting direct neighborhood relations provide little beyond immediate dependencies. ・Cluster-level methods, though capable of broader reasoning, generally require cluster assignments or causal sufficiency in advance, conditions that are rarely satisfied in practic
cs.LG updates on arXiv.org

From Prediction to Incrementality: Causal Optimization for Large-Scale Targeting and Recommendation

・arXiv:2608.10182v1 Announce Type: new Abstract: Large-scale targeting and recommendation systems are typically built around predictive scores fed into heuristic or local allocation. ・When the business goal is incremental impact, as in marketing campaigns, incentives, and notifications, this paradigm systematically misallocates resources toward users who would have acted anyway. ・We present a decision-centric framework
cs.LG updates on arXiv.org

GARLIC: Graph Attention-based Relational Learning of Multivariate Time Series in Intensive Care

・arXiv:2608.10969v1 Announce Type: new Abstract: Healthcare data, such as Intensive Care Unit (ICU) records, comprise heterogeneous multivariate time series sampled at irregular intervals with pervasive missingness. ・However, clinical applications demand predictive models that are both accurate and interpretable. ・We present our Graph Attention-based Relational Learning for Intensive Care (GARLIC) model, a novel neural
ITmedia NEWS 最新記事一覧

GeminiアプリのMAUが10億人を突破 Googleで14番目の大台到達製品に

・GoogleのGeminiアプリの月間アクティブユーザー数(MAU)が10億人を突破した。Google検索やWorkspaceなどの組み込み機能を除いたアプリ単体の数字で、同社として14番目の10億ユーザー到達製品となる。競合のChatGPTに続く大台達成で、iOSユーザー数や音声利用の比率なども公開された。
cs.LG updates on arXiv.org

Generalized Linear Markov Decision Process

・arXiv:2506.00818v2 Announce Type: replace-cross Abstract: Offline reinforcement learning for longitudinal studies often faces two linked challenges: rewards may be binary or bounded, and reward observations may be available only for a subset of trajectories or time points even when the corresponding state-action-next-state histories are available. ・Linear Markov decision process methods are tractable because Bellman b
cs.LG updates on arXiv.org

Generator-Guided Inverse Sampling for L\'evy-Driven Generative Models

・arXiv:2608.10384v1 Announce Type: new Abstract: This paper studies inverse sampling for L\'evy-driven generative models from the perspective of Markov generators. ・Unlike conventional diffusion models, L\'evy-driven dynamics involve infinite jump activities, which makes their reverse process nonlocal and difficult to characterize using score information alone. ・We address this challenge by analyzing the forward and rev
cs.LG updates on arXiv.org

GLAM: Efficient Continual Learning at Scale via Grouped LoRA Adapter Merging

・arXiv:2509.13211v4 Announce Type: replace Abstract: The ability to learn continuously over time remains a major challenge for modern machine learning systems, even in the era of Foundation Models. ・While the rich representations learned by large pre-trained models can partially mitigate catastrophic forgetting, they still struggle to adapt efficiently to evolving data distributions. ・A key challenge remains, how to con
The Verge

Google aims for influencers with the Pixel 11 Creator Suite

・The settings page for the Pixel 11 series’ “Creator Suite.” Google knows creators are a big audience. ・The occupation is growing fast, and landing some influential names could be a major turning point for the Pixel's market share. ・This year, Google is building features directly into its new Pixel lineup that it says can help creators record, organize, edit, and publish content faster and more efficiently than they cou
Zennの「大規模言語モデル」のフィード

Google Colabで最新LLMを試す #1 ― NVIDIA Nemotron 3.5 Lightning 30B-A3B

・はじめに 2026年8月4日に、『ブラウザで動かすLLM実装入門 Google Colaboratoryで実践するLLM・RAG・ファインチューニング』を刊行しました。 ・本書では、高価なGPU搭載PCを用意することなく、Google Colaboratory(以下、Google Colab)を使って、LLMの実行からRAG、ファインチューニングなどを実際に体験する方法を解説しています。 ・ただ、LLMの世界では新しいモデルが非常に速いペースで公開されます。当然ながら、書籍を書き終えた後にも興味深いモデルは次々と登場します。
Zennの「大規模言語モデル」のフィード

Google Colabで最新LLMを試す #2 ― Kimi-VL-A3B-Thinking-2506を動かす

・はじめに 前回の記事では、Google Colab上で比較的新しいLLMを実際に動かしながら、どこまでローカル実行できるのかを試しました。 ・今回はその第2回として、Moonshot AIの Kimi-VL-A3B-Thinking-2506 をGoogle Colab上で動かしてみます。 ・本記事でより興味を持っていただけましたら、「ブラウザで動かすLLM実装入門 Google Colaboratoryで実践するLLM・RAG・ファインチューニング」(インプレス, 2026/8/4)をぜひお手にとってください。
WIRED

Google Pixel 11 Series, Pixel Watch 5, Pixel Tag: Specs, Features, Price, Release Date

・The Pixel 11 smartphone lineup is here, alongside the Pixel Watch 5 and the all-new Pixel Tag.
ITmedia NEWS 最新記事一覧

Google、折りたたみスマホ「Pixel 11 Pro Fold」発表 10%軽く、1mm薄くしつつ耐久性強化

・米Googleは8月12日、折りたたみスマートフォン「Pixel 11 Pro Fold」を国内発表した。前モデルより約10%軽く約1mm薄いが、新しいヒンジとセラミックガラスの採用で3倍の強度を実現したという。Googleストアでの価格は256GBモデルが27万9900円。8月20日に発売する。
#AIタグ

Google:Geminiアプリが月間10億人超に到達—同社史上最速の成長

・結論 Googleの発表で、Geminiアプリの月間利用者が10億人を超えたことが分かりました。これはGoogle側が「同社史上最速の成長」と位置づけている重要なマイルストーンです。初心者にとっては、身の回りでGeminiを使う人が増えていることを示す目安になります。
The Verge

Google’s Pixel 11 phone preorders come with up to $350 in gift cards

・The Pixel 11 Pro in matte obsidian. ・| Photo: David Imel / The Verge Looking to get your hands on the latest Pixel devices? ・After weeks of leaks and rumors, Google has officially announced its next generation of Pixel phones and watches.
The Verge

Google’s Pixel 11 series pairs a little new hardware with a lot of new software

・When I first picked up the Pixel 11 Pro this week, it was clear to me that this was one of those refinement years - at least when it comes to hardware. ・Aesthetically, the phones are as beautiful as ever, with Google's signature camera bar and bright, colorful backing glass. ・There are small updates, like a single piece of glass covering the camera bar and the all-new HiLight ambient LED on the Pros.
WIRED

Google’s Pixel Watch 5 Can Now Detect Breathing Emergencies

・The company is also adding blood pressure and insulin resistance trends as proactive monitoring tools in the Google Health app.
ITmedia NEWS 最新記事一覧

Google新スマホ「Pixel 11シリーズ」正式発表 「Gemini Intelligence」向けに設計 カメラも強化

・米Googleは8月12日、「Pixel 11」「Pixel 11 Pro」「Pixel 11 Pro XL」を発表した。新チップ「Google Tensor G6」を搭載し、カメラを刷新。Proには通知を光で知らせる「HiLight」を加えた。中核のAI機能「Gemini Intelligence」が目玉としている。8月20日に発予定。
ITmedia NEWS 最新記事一覧

Google版“AirTag”こと「Pixel Tag」登場 10億台のAndroidネットワークで居場所を探知

・米Googleが紛失防止タグ「Google Pixel Tag」を発表した。鍵や財布に取り付け、10億台超のAndroid端末で構成されるFind Hubのネットワークで位置を探せる。UWBで距離と方向も分かる。日本では単品5010円、4個セット1万6940円で11月に発売する。
#LLMタグ

Grok Botの月200ドル、買っているのはAIではない

・AIエージェントの値段は、たいていモデルの賢さで決まると思っていた。入力トークンがいくら、出力がいくら。その足し算だ。 ・でも今夜、20時42分に開いたGrok Botの説明を読んで、その見方が少しずれた。月200ドルの個人向けプランで売っているものは、会話相手のAIだけではない。Botごとに仮想コンピュータを1台持たせ、ログイン済みのツールを動かしたままにする。そこまで含めた料金だった。
cs.LG updates on arXiv.org

Gromov-Wasserstein Quantization and Clustering: Structure, Rates, and Algorithms

・arXiv:2608.11016v1 Announce Type: cross Abstract: Clustering is a fundamental class of data analysis techniques with the most important representatives being centroid-based methods like $k$-means. ・Such methods are strongly connected to quantization problems, which aim to approximate general probability measures with discrete ones. ・For example, $k$-means corresponds to quantization with respect to the Wasserstein dist
#AIタグ

GSP Archaeology|PN:盗まれた約束

・これはGrim Saga Project(GSP)という、20年以上作り続けられている創作世界の記録の一つを、作者と一緒に読み返した記録である。 ・対象は「盗まれた約束 ~Promised Necklace」。2007年に書かれた短編。GSPの本編としては、他の重量級の物語群と比べるとずいぶん軽い。それには理由があるらしい。
The Verge

Guitar company D’Addario admits that AI music was used in a promotional video

・Maybe don’t let the AI shred next time. ・| Image: D’Addario After weeks of controversy and speculation, music company D'Addario has admitted that AI, specifically Suno, was used as part of a recent promotional video. ・For nearly two weeks, the company has denied the allegations, even as evidence piled up against it.
Zennの「機械学習」のフィード

gzipでテキスト分類(圧縮距離NCD+kNN)を73記事で追試——93.8%は出るが、文字数だけでも89.2%出た

・2023年、「gzipがBERTに勝つ」という論文が機械学習界隈で大きな話題になった。ニューラルネットも学習も一切なし。テキストをgzipで圧縮したときのサイズから**正規化圧縮距離(NCD)**という距離を計算し、あとはk近傍法(kNN)で多数決するだけ——それでいくつかのベンチマークではBERTを上回った、という主張だった。 ・ただしこの話には続きがある。公開コードを検証したKen Schutte氏が、k=2のkNNのタイブレーク処理が「上位2件のどちらかが正解なら正解」という実質top-2精度になっていたことを指摘したのだ。正しく評価し直すと、数字はそれなりに下がる。「圧縮だけでBE...
cs.LG updates on arXiv.org

Hierarchical Empirical-Bayes Naive Bayes: Minimax Smoothing and Calibration with AODE Extension

・arXiv:2608.11162v1 Announce Type: new Abstract: The Naive Bayes (NB) classifier remains a standard choice for categorical data, yet its widely used smoothing rules, such as Laplace, Lidstone, Krichevsky-Trofimov, and the $m$-estimate, all prescribe a fixed smoothing strength that ignores feature cardinality, sample size, and class imbalance, inducing a non-vanishing bias on modern high-cardinality tabular data.
cs.LG updates on arXiv.org

High-Dimensional Calibration from Swap Regret

・arXiv:2505.21460v2 Announce Type: replace Abstract: We study online calibration of multi-dimensional forecasts over an arbitrary convex set $P \subset \mathbb{R}^d$ relative to an arbitrary norm $|\cdot|$. ・We connect this to external regret minimization for online linear optimization (OLO): if one can guarantee $O(\sqrt{\rho T})$ worst-case regret after $T$ rounds when actions are drawn from $P$ and losses from the d
cs.LG updates on arXiv.org

HIPNO: Symmetry-Aware Physics-Informed Neural Operators for Noninvasive Hemodynamic Inference

・arXiv:2608.10011v1 Announce Type: cross Abstract: Continuous hemodynamic monitoring guides treatment decisions in surgery and intensive care. ・However, gold-standard signals are only measured in severe cases due to risks associated with invasive measurement. ・In this work, we introduce HIPNO (Hemodynamic Inference via Physics-informed Neural Operators) to recover hemodynamic state from ubiquitous, non-invasive signals
WIRED

Hoka Coupon Codes: 30% Off in August 2026

・Unlock free expedited shipping or 10% off with today’s Hoka discount codes, plus save up to 30% with the best discounts for August 2026.
WIRED

Honor’s Robot Phone Has a Gimbal-Powered Camera for Vloggers

・Honor’s new China-exclusive Android phone features a pop-out motorized gimbal camera designed to replace your action cam.
WIRED

Hostinger Promo Code: 79% Off for August 2026

・Discover exclusive Hostinger promo codes, discounts, and deals on web hosting, cloud plans, and domain registration. ・Save big on your next Hostinger purchase today.
AI News & Artificial Intelligence | TechCrunch

How a $250 million acquisition collapsed into allegations of fraud and forged signatures

・Investors are still waiting for their share of the $250 million windfall, and VideoVerse co-founder Vinayak Shrivastav is now at the center of multiple legal cases.
The Verge

How Google’s new Pixel 11 phones compare to last year’s models

・The Pixel 11 Pro. ・| Photo: David Imel / The Verge Google just added four new phones to the Pixel family: the Pixel 11, Pixel 11 Pro, Pixel 11 Pro XL, and the Pixel 11 Pro Fold. ・They're slightly more expensive than their predecessors, but they come with some notable upgrades, including improved cameras and a new Tensor G6 chip that should deliver faster performance.
cs.LG updates on arXiv.org

How Robust Are LLMs to Vietnamese Dialects?

・arXiv:2608.10414v1 Announce Type: cross Abstract: Large Language Models (LLMs) are typically evaluated on standard written Vietnamese, yet everyday communication frequently involves regional dialects that preserve meaning but differ in surface form. ・Existing Vietnamese dialect work largely addresses this issue through dialect-to-standard normalization instead of measuring how the model fails under Vietnamese dialecta
The Verge

How the Pixel 11 Pro Fold compares to the Galaxy Z Fold 8

・The Pixel 11 Pro Fold. ・| Photo: David Imel / The Verge Google wasn't first to make a foldable Android phone, but the company's Pixel Fold series is now an established player with unique traits compared to its competitors. ・For one, it popularized the passport-style design years ago, being more wide than tall - something that Samsung tried for the first time this year with its Z Fold 8, and that Apple is rumored to do
WIRED

How to Select the Office Chair That’s Right For You

・Not all office chairs are right for all people. ・I asked an ergonomist how to select the right office chair for your budget, height, and physique.
cs.LG updates on arXiv.org

How to Verify Consistency of Probabilistic Claims

・arXiv:2608.11181v1 Announce Type: cross Abstract: When a probabilistic predictor answers many conditional-probability queries, are its answers self-consistent, and can this be verified in polynomial time? ・This problem is of interest for AI safety, where safety is derived from honesty about probabilistic predictions of unwanted outcomes potentially caused by an AI action. ・We construct an interactive PCP as follows.
Engineering at Meta

How We’re Building Scam Alert on WhatsApp With End-to-End Encryption and Verifiability Guarantees

・WhatsApp is committed to helping people stay safe while protecting the privacy of their messages. ・As scam tactics evolve — from impersonation to social engineering to AI-generated lures — we’re always evolving as well, so that our protections stay ahead of scammers while protecting people’s personal messages with end-to-end encryption. ・Today, we’re sharing an early [...] Read More...
WIRED

Hulu Promo Codes & Discounts: 20% Off in August

・Students can get a Hulu plan for $1.99 per month. ・Get more details on this and other great deals below.
cs.LG updates on arXiv.org

HyperShape: Hyperelasticity Across Diverse Shapes

・arXiv:2608.09938v1 Announce Type: cross Abstract: Hyperelastic deformations are highly sensitive to domain geometry and boundary conditions, making generalization across both a critical capability for neural operators applied to these problems. ・However, existing benchmarks for neural operators on hyperelasticity rely on simple or few geometries, which makes it difficult to assess this capability rigorously.
cs.LG updates on arXiv.org

IADD-TR: Intervention-Aware Dynamics Decoupling with Targeted Regularization for Model-Based Reinforcement Learning

・arXiv:2608.10634v1 Announce Type: new Abstract: Model-based reinforcement learning (MBRL), which learns environment dynamics to generate synthetic experience, is a promising approach to sample-efficient decision making. ・Numerous methods have been developed to improve dynamics prediction and policy optimization for MBRL through uncertainty estimation, model regularization, and conservative value learning. ・However, the
Qiita - 人気の記事

IBM Bob × Confluent MCP入門:KafkaとFlinkを自然言語で触ってみた

・はじめに この記事では IBM Bob に Confluent MCP サーバー を接続することで、日本語のチャットだけで Kafka トピックの作成・データ投入・Flink SQL 実行 まで一気通貫でやってみた記録をまとめます。 ・コマンドはセットアップ時にほんの少し打...
The Verge

ICE wants to give agents electrified gloves that shock people into compliance

・Immigration and Customs Enforcement (ICE) is aiming to spend up to $20 million on equipping officers and agents with specialized gloves that deliver painful electric shocks. ・These plans were outlined in a notice published by the Department of Homeland Security on Monday, with an unspecified quantity of the devices set to be delivered by March 2027. ・The electrified handwear outlined in the contract is the CTG 5 GLOVE
Hugging Face Papers

iFAN: Inference-Aware Learning for Plain Mask Transformers

iFAN: Inference-Aware Learning for Plain Mask Transformers
cs.LG updates on arXiv.org

Information Bottleneck under Perfect Privacy

・arXiv:2608.11003v1 Announce Type: cross Abstract: In this work, we study the information bottleneck under perfect privacy, with particular emphasis on the active-rate regime, where the representation-rate constraint is binding and directly limits the achievable utility. ・The goal is to construct a representation that preserves utility-relevant information while remaining statistically independent of a sensitive variab
Hugging Face Papers

InSight-doc: Agentic Visual Perception for Long-Document Understanding

InSight-doc: Agentic Visual Perception for Long-Document Understanding
cs.LG updates on arXiv.org

InSight-doc: Agentic Visual Perception for Long-Document Understanding

・arXiv:2608.10628v1 Announce Type: cross Abstract: Long-document understanding often requires reasoning over many visually rich pages, making inference costly and prone to context rot. ・In this work, we propose InSight-doc, an agentic visual perception framework that treats visual resolution as an adaptive reasoning-time resource. ・InSight-doc starts from low resolution and selectively zooms into high-resolution regions
cs.LG updates on arXiv.org

Instance-Adaptive Online Multicalibration

・arXiv:2605.09273v3 Announce Type: replace Abstract: We study online multicalibration beyond the worst-case. ・We give a single, efficient algorithm which dynamically interpolates between benign and worst-case sequences by adaptively refining a dyadic grid of prediction values. ・Its error is controlled by the number of leaves in the refinement tree.
cs.LG updates on arXiv.org

Intrinsic Structure: Spectral Identifiability for Mechanistic Interpretability

・arXiv:2608.10172v1 Announce Type: new Abstract: Mechanistic interpretability explains models by identifying circuits inside them, but has no way to tell whether a circuit is a property of the model or an artifact of the method that found it. ・Sparse autoencoders illustrate the problem: different seeds and widths recover materially different features from the same activations, and no theory says whether that variabilit
Hugging Face - Blog

Introducing OlmoEarth embeddings: Custom embedding exports from OlmoEarth Studio for downstream analysis

Introducing OlmoEarth embeddings: Custom embedding exports from OlmoEarth Studio for downstream analysis
cs.LG updates on arXiv.org

Invertible Logits Transformation for Accuracy-Preserving Post-Hoc Uncertainty Calibration

・arXiv:2608.10372v1 Announce Type: new Abstract: Post-hoc calibration aligns a classifier's predicted confidences with its empirical accuracy without retraining. ・An ideal calibrator should correct nonlinear miscalibration, scale gracefully to large label spaces, and preserve the original predictions; existing methods typically violate at least one of these properties---temperature scaling lacks expressivity, more flex
cs.LG updates on arXiv.org

Iterative Erasure Count Is Not an Affine-Invariant Concept Dimension

・arXiv:2608.10566v1 Announce Type: cross Abstract: How many directions does a neural representation use to encode a concept? ・A common answer repeatedly erases probe directions and reports the stopping count or cumulative removed rank. ・We show that both quantities can change under an information-preserving invertible reparameterization, so neither is intrinsically a concept dimension.
Hugging Face Papers

JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles

JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles
#LLMタグ

Kimi K3の新展開と影響

・中国のMoonshot AIが大規模モデル「Kimi K3」を発表。世界の競合との比較やビジネスへの影響が注目されています。 ・中国のMoonshot AIが開発した大規模言語モデル「Kimi K3」に関する報道が世界中で注目されています。このモデルは、他のAIシステムとの競争においてどのような位置付けを持つのか、詳細が議論されています。
cs.LG updates on arXiv.org

KKL Observer Synthesis for Nonlinear Systems via Physics-Informed Learning

・arXiv:2501.11655v3 Announce Type: replace-cross Abstract: This paper proposes a novel learning approach for designing Kazantzis-Kravaris or nonlinear Luenberger (KKL) observers for autonomous nonlinear systems. ・The design of a KKL observer involves finding an injective map that transforms the system state into a higher-dimensional observer state, whose dynamics is linear and stable. ・The observer's state is then mappe
cs.LG updates on arXiv.org

Knowledge-Guided 3D CT Generation: A Conditioning-Centric Taxonomy

・arXiv:2608.09992v1 Announce Type: cross Abstract: Controllable generation guided by external knowledge is a key requirement in modern generative deep learning applications, enabling the synthesis of samples with explicit constraints on semantic content, structural properties, and variability. ・In 3D Computed Tomography (CT), such control is essential for clinical applications, including data augmentation, privacy-pres
cs.LG updates on arXiv.org

KrishokChat: A Provenance-Traceable Multi-Task Bengali Agricultural Benchmark with Safety-Critical Chemical Advisory

・arXiv:2606.29243v2 Announce Type: replace Abstract: We introduce KrishokChat, an 85,979-instance Bengali agricultural benchmark built from 284 government publications, 13 institutions, and six regional dialects. ・The benchmark comprises four tracks: General Knowledge QA, Treatment QA, Safety Refusal and Re-query, and Table QA. ・It also includes a 1,000-query Real-World Farmer Benchmark collected independently from fiel
cs.LG updates on arXiv.org

Learning Disease-Sensitive Latent Interaction Graphs From Noisy Cardiac Flow Measurements

・arXiv:2602.23035v2 Announce Type: replace Abstract: Cardiac blood flow patterns contain rich information about disease severity and clinical interventions, yet current imaging and computational methods fail to capture underlying relational structures of coherent flow features. ・We propose a physics-informed, latent relational framework to model cardiac vortices as interacting nodes in a graph. ・Our model combines a neu
cs.LG updates on arXiv.org

Learning Representations from Incomplete EHR Data with Dual-Masked Autoencoding

・arXiv:2602.15159v2 Announce Type: replace Abstract: Electronic health records (EHR) arrive masked. ・Clinicians order measurements selectively, and any patient table thus contains only a subset of the values that characterize the underlying physiological state. ・Prior masked modeling approaches on EHR data either impute the table before learning, represent missingness through a dedicated placeholder signal, or optimize
WIRED

LegalZoom Promo Code: Exclusive 10% Off LLC Formations

・Save on top services at LegalZoom, like LLC registration, incorporation, estate plans, and more with coupons and deals from WIRED.
Hugging Face - Blog

LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge

LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge
cs.LG updates on arXiv.org

Link-adaptive digital twin for robust physical-layer modeling in hybrid-amplified ultra-wideband optical networks

・arXiv:2608.10517v1 Announce Type: cross Abstract: Accurate physical-layer modeling is increasingly essential for reliable ultra-wideband operation and capacity optimization, especially under the intensified inter-channel stimulated Raman scattering (ISRS) effect. ・This paper proposes the link-adaptive digital twin (LA-DT) for hybrid-amplified ultra-wideband links to overcome the generalization and speed limitations of
cs.LG updates on arXiv.org

Lipschitz Dueling Bandits over Continuous Action Spaces

・arXiv:2604.00523v2 Announce Type: replace Abstract: We study for the first time, stochastic dueling bandits over continuous action spaces with Lipschitz structure, where feedback is purely comparative. ・While dueling bandits and Lipschitz bandits have been studied separately, their combination has remained unexplored. ・We propose the first algorithm for Lipschitz dueling bandits, using round-based exploration and recur
#LLMタグ

LLM Brokerが地味に便利だという話

・LLM Brokerが地味に便利だという話 – 塾長の独り言zikuu.space 続きをみる
#LLMタグ

LLMのDense型とMoE型、違いを整理してみた

・はじめに 最近、LLM関連の話の中で「Dense型」「MoE型」という言葉を見かけて、正直「そもそも何が違うんだっけ」と気になりました。自分の中で言葉の意味をちゃんと整理できていなかったので、今回はこの2つの違いだけに絞ってまとめてみます。
#LLMタグ

LLMは大きければ賢くなるのか? 「モデル・データ・計算量」の数学

・LLMの学習について調べていると、「パラメータ数を増やせば性能が上がる」「データを増やせばよい」「計算量を増やせばよい」といった説明に出会う。 ・しかし、ここには一つ大きな落とし穴がある。
cs.LG updates on arXiv.org

Local Fr\'echet functional regression in manifolds from time-correlated bivariate curve data

・arXiv:2505.05168v4 Announce Type: replace-cross Abstract: Under mild conditions, a least-squares local linear Fr\'echet curve predictor is derived for a response and a regressor evaluated in a separable Hilbert space. ・The conditions that allow the implementation of the local linear Fr\'echet functional predictor in the ambient L2-space of vector functions, with values in the time-varying tangent space of a compact Ri
cs.LG updates on arXiv.org

Logit-Boundary Geometric Belief Interfaces and Sparse Sheaf-Enclave Protocols: A Self-Contained Substrate for Secure Network Electronic Health Record (EHR) Interoperability

・arXiv:2608.10300v1 Announce Type: cross Abstract: Electronic health-record interoperability is a boundary problem: legacy systems, generative models, terminology services, identity systems, and human reviewers may each expose rich internal states, while operational exchange requires a narrow shared interface of typed claims, bounded uncertainty, provenance, and explicit admission or abstention. ・This paper details a m
cs.LG updates on arXiv.org

Long-Time Trajectory Approximation via SA-NODEs: Model Predictive and Floquet Strategies

・arXiv:2608.10738v1 Announce Type: new Abstract: We study the approximation of dynamical systems by semi-autonomous neural ordinary differential equations (SA-NODEs) over long time horizons. ・For a single network trained on the whole horizon, the available error bound deteriorates double exponentially in the horizon length. ・We develop two training strategies that avoid this barrier, each built on a reset of the state.
cs.LG updates on arXiv.org

Lost in Aggregation: On a Fundamental Expressivity Limit of Message-Passing Graph Neural Networks

・arXiv:2603.14846v4 Announce Type: replace Abstract: We define an information-complexity property for aggregation functions, capturing a vast range of practical aggregations, and prove that any Message-Passing Graph Neural Network (MP-GNN) model with such aggregations induces only a polynomial number of equivalence classes on all graphs - while the number of non-isomorphic graphs is super-exponential (in number of ver
AI News & Artificial Intelligence | TechCrunch

Lovable confirms new $13.3B valuation, raises another $400M

・This new funding comes after Lovable hit $500 million in annualized run rate revenue in June, the startup told TechCrunch.
cs.LG updates on arXiv.org

LVCG: Learning ECG Representations in the Latent Vectorcardiogram Space

・arXiv:2605.31249v2 Announce Type: replace Abstract: Electrocardiography (ECG) is a cornerstone of cardiac assessment, making the learning of informative ECG representations fundamental to tasks ranging from disease diagnosis to clinical report generation. ・However, existing methods operate almost exclusively in the observable ECG signal space. ・In practice, the standard twelve-lead ECG represents multiple projections o
cs.LG updates on arXiv.org

Mapping and Measuring the Behavioral Evolution of Large Language Models

・arXiv:2608.11027v1 Announce Type: new Abstract: Benchmark leaderboards summarize how well a language model performs, but not how its behavior relates to that of other models or changes across generations. ・We characterize the output behavior of 32 models from six families using their responses to a shared bank of 10{,}000 prompts. ・After embedding each response, we construct three complementary sentence-level dissimila
cs.LG updates on arXiv.org

MARCO: Click-Intent Decomposition for Calibrated Ads Conversion Prediction

・arXiv:2608.10562v1 Announce Type: new Abstract: Not all clicks are equal. ・Industrial ads ranking decouples conversion probability into click-through rate (CTR) and post-click conversion rate (CVR), yet treats every click as the same event. ・In reality, users provide a free, self-generated signal of intent through their physical UI interactions.
WIRED

McDonald’s Built a 515-Page Dossier on Me. It Says I’ll Never Stop Eating There

・I requested a copy of my data from McDonald’s loyalty program and received an extensive, personalized report that algorithmically predicts my next purchase.
Zennの「大規模言語モデル」のフィード

MCP2026対応でAgentCore Gatewayが一気に変わった5つの理由

・本日は、BigGoの「LLMは舞台裏へ、AIエージェントと具身知能が主役に」が示す通り、AIの中心が“モデルそのもの”から“エージェントとしてどう動くか”へ移っています。AWSのMCP対応、Allganize・富士通・Microsoft・SAPの業務エージェント強化がその流れを裏付ける一方、ReutersやNPR、TechCrunchのOpenAIエージェント逸脱報道により、安全性と監査性が最大論点になっています。Python周辺ではuv/Ruff/PolarsとAstral買収が、AIコーディング基盤の再編を示しています。Web開発については、提供ヘッドライン内にNext.js/R...
cs.LG updates on arXiv.org

Measuring Semantic Abstractness of SAE Features via Nonlocality

・arXiv:2608.10537v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) have helped uncover mechanistic explanations for LLM behaviours such as reasoning, jailbreaking etc., via understanding the corresponding task-relevant and causally effective features. ・To evaluate such mechanistic explanations, downstream studies must distinguish surface lexical features from genuinely high-level ones. ・However, neither an au
Hugging Face Papers

Mendel Gödel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution

Mendel Gödel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution
cs.LG updates on arXiv.org

MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale

・arXiv:2608.10333v1 Announce Type: new Abstract: LLM agents execute heterogeneous sequences of model calls within a single task: some invocations require careful reasoning, while others are structured steps such as formatting or tool-argument construction. ・Prior routing methods exploit this asymmetry by assigning easy invocations to a cheaper small model and difficult ones to a large model. ・Such policies reduce infere
AI News & Artificial Intelligence | TechCrunch

Mesh, Automattic’s CRM for everyone, comes to Android

・Mesh, an AI-powered contacts app and relationship manager from Automattic is now an Android app.
cs.LG updates on arXiv.org

MIDAS: Mutual Information Disentanglement with Uncertainty-Aware Fusion for Incomplete Multimodal Sentiment Analysis

・arXiv:2608.09986v1 Announce Type: cross Abstract: Most existing multimodal sentiment analysis approaches assume access to complete multimodal inputs. ・However, real-world applications frequently encounter incomplete or corrupted modalities, posing a critical challenge. ・Although several methods have been proposed to tackle this issue, they mainly rely on data imputation and heuristic coordination constraints, which fai
Microsoft Research

MindTopo reveals VLMs’ spatial reasoning abilities

・A path, a fence, a knot. ・MindTopo sets a new benchmark for testing how AI understands topological relationships and highlights new opportunities to strengthen spatial reasoning and planning. ・The post MindTopo reveals VLMs’ spatial reasoning abilities appeared first on Microsoft Research.
cs.LG updates on arXiv.org

Model Multiplicity and Predictive Arbitrariness in Recidivism Risk Assessment

・arXiv:2606.02198v2 Announce Type: replace Abstract: Prediction tasks over individual futures, which are inherently noisy, often admit multiple similarly accurate models. ・When these models produce different predictions for the same individual, they raise concerns of arbitrariness in decision-making. ・How severe can this arbitrariness be, in theory and in practice?
cs.LG updates on arXiv.org

MoE Proxy Models for Low-Cost Failure Reproduction and Diagnosis in LLM RL Post-Training

・arXiv:2608.10823v1 Announce Type: new Abstract: Reinforcement learning (RL) post-training of large language models (LLMs) is computationally intensive and involves complex system pipelines with substantial debugging overhead. ・In practice, factors such as framework adaptation, numerical precision, and operator implementation can cause failures, including gradient overflow and loss divergence. ・Reproducing such failures
cs.LG updates on arXiv.org

More Accurate, Less Human: Gestalt Grouping in Vision Models

・arXiv:2608.10195v1 Announce Type: cross Abstract: Human vision organizes what it sees into wholes: same-colored points group into series, similar marks cohere into categories, and shapes complete into recognizable objects. ・These are the Gestalt operations that visualization design builds on. ・Whether vision models organize visual content this way has not been systematically tested.
WIRED

Motley Fool Promo Code: $200 Off on Stock Advisor August 2026

・Scale your portfolio for less with these verified The Motley Fool membership discounts, stock advisor promo codes, and Epic Bundle deals.
cs.LG updates on arXiv.org

MRIComp4Flow: Compression of 3D Brain MRI for Training Multi-Modal Generative Models

・arXiv:2608.10291v1 Announce Type: cross Abstract: Large-scale multi-modal MRI datasets impose substantial storage and I/O costs, limiting the training of 3D generative models on commodity infrastructure. ・While lossy compression is known to preserve accuracy for discriminative segmentation networks, its effect on generative models, which must learn the full data distribution rather than a decision boundary, is unexplo
cs.LG updates on arXiv.org

Multi-Granular Rationale-Guided Molecular LLM for Property Prediction

・arXiv:2608.10480v1 Announce Type: cross Abstract: Large language models (LLMs) are widely applied across chemical tasks, such as molecular property prediction, which underpins drug discovery. ・Molecular LLMs represent a molecule through several modalities, notably a 1D SMILES sequence or a 2D molecular graph. ・Both encode molecular information implicitly, so the contribution of individual substructures remains opaque.
cs.LG updates on arXiv.org

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment

・arXiv:2608.11167v1 Announce Type: cross Abstract: Existing Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image representations to long textual descriptions. ・However, this image-level alignment suffers from referential ambiguity: models struggle to infer the correspondences between multiple visual objects and textual entities from the
cs.LG updates on arXiv.org

Navigating the Proximity-Safety Balance: Constraint Decomposition for Human Following in Pedestrian Crowds

・arXiv:2608.10056v1 Announce Type: cross Abstract: Following a target human in crowded environments involves an inherent conflict between staying close to the target and navigating safely among surrounding pedestrians and obstacles. ・This conflict becomes more severe in dense scenarios, where aggressive following risks collisions and conservative margins lead to target loss, especially when pedestrian behaviors are unf
WIRED

Nike Promo Codes and Discounts: 30% for August 2026

・Check out our deals for Nike this August 2026, including 15% off select purchases.
Hugging Face Papers

Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents

Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents
MarkTechPost

NVIDIA AI Releases Nemotron 3.5 Lightning: A 30B Open MoE with 3B Active Parameters, and NeMo Switchyard Model Router

・NVIDIA's open 30B MoE targets the agent execution layer, with Switchyard routing each step to the cheapest capable model. ・The post NVIDIA AI Releases Nemotron 3.5 Lightning: A 30B Open MoE with 3B Active Parameters, and NeMo Switchyard Model Router appeared first on MarkTechPost.
NVIDIA Blog

NVIDIA CEO Tops Glassdoor’s 2026 List of Best CEOs

・NVIDIA founder and CEO Jensen Huang is ranked No. ・1 on Glassdoor’s Best CEOs list for 2026. ・In the just-released ranking, recognition is earned directly from the people who know their leadership the best — employees.
cs.LG updates on arXiv.org

Observational Policy Ranking for SMB Financial Guidance from Multi-Action Accounting Logs

・arXiv:2608.10050v1 Announce Type: new Abstract: Small and medium-sized businesses need timely financial guidance, yet historical accounting logs record self-selected and often co-occurring business changes rather than randomized recommendations. ・We formulate this setting as observational policy ranking: from pre-decision financial information, a policy selects one of 34 ledger-derived business-change categories for a
cs.LG updates on arXiv.org

Off-Axis, On Purpose: Where a Transformer Computes Concepts and Why it Does So

・arXiv:2608.10251v1 Announce Type: cross Abstract: A transformer's answer lives on one axis: the direction its unembedding reads. ・Its intermediate states largely do not, and that off-axis position is usually treated as an obstacle to interpretation. ・We show it is functional.
WIRED

Oh Lord, AI Reporters Are Actually Breaking Big News

・Last week, an AI newsroom beat mainstream journalists—including WIRED—to a story about OpenAI and hacking. ・It’s just the beginning.
cs.LG updates on arXiv.org

On Effectiveness and Efficiency of Agentic Tool-calling and RL Training

・arXiv:2606.00135v2 Announce Type: replace Abstract: Tool-calling is a central component of modern large language model (LLM) agents, equipping them with skills beyond their parametric knowledge. ・This paper studies tool-calling along two complementary axes: effectiveness, i.e., how this capability is measured, and efficiency, i.e., how it is learned. ・On effectiveness, we systematically analyze tool-calling evaluation
cs.LG updates on arXiv.org

On the Importance of Geometric Nonlinearity and Temperature-Dependent Properties in Multi-Material Thermo-Mechanical Topology Optimization

・arXiv:2608.10344v1 Announce Type: cross Abstract: Thermo-mechanical compliant devices are commonly designed with small-strain linear elasticity and temperature-independent material properties, even though they might operate hundreds of kelvin above ambient where both assumptions are questionable. ・In this work, we quantify the effect and cost of each assumption in multi-material topology optimization of thermally actu
WIRED

OnePlus Promo Codes: 30% Off August 2026

・Save 30% with a OnePlus coupon in August 2026, plus save up to 10% on earbuds, phones, and more.
AI News & Artificial Intelligence | TechCrunch

OpenAI-backed Thrive Holdings raises $2B to bring AI to the enterprise

・Thrive Holdings has raised $2 billion in new funding at a $12 billion valuation from investors like SoftBank, D1 Capital Partners, and Alitmeter Capital.
cs.LG updates on arXiv.org

Optimistic Rates for Multiclass PAC Learning

・arXiv:2608.10869v1 Announce Type: new Abstract: Worst-case multiclass bounds do not become smaller when the best classifier is already nearly correct: what is missing is an optimistic rate, a guarantee whose fluctuation scales with the oracle risk itself. ・For a class of Natarajan dimension $d_N$ and Daniely-Shalev-Shwartz dimension $d_{DS}$, the optimal excess risk is known at the two endpoints ($d_{DS}/n$ realizable
cs.LG updates on arXiv.org

Optimize Cheap, Deploy Strong: Cost-Aware Cross-Tier Transfer for Evolutionary Optimization

・arXiv:2608.10694v1 Announce Type: new Abstract: Evolutionary optimization of LLM prompts and agentic programs (e.g., GEPA) is dominated by fitness evaluation: scoring each candidate runs an answering LLM over a validation set, so the evaluator's price tier dictates total search cost. ・We restructure that search by decoupling the three roles an LLM plays, running the high-volume answering role on the cheapest tier, res
cs.LG updates on arXiv.org

Optimized Sequential Testing for Binary Ensemble Classifiers

・arXiv:2606.15237v1 Announce Type: cross Abstract: Ensemble classifiers are predictive models that combine the results of simpler base models, often by majority vote. ・A classic example is random forests, which combine the predictions of decision trees. ・Ensembles that use more base models can be more accurate but also more costly to train and run.
cs.LG updates on arXiv.org

Order Matters in Retrosynthesis: Structure-aware Generation via Reaction-Center-Guided Discrete Flow Matching

・arXiv:2602.13136v2 Announce Type: replace Abstract: Template-free retrosynthesis methods treat the task as black-box sequence generation, limiting learning efficiency, while semi-template approaches rely on rigid reaction libraries that constrain generalization. ・We address this gap with a key insight: atom ordering in neural representations matters. ・Building on this insight, we propose a structure-aware template-free
cs.LG updates on arXiv.org

P3CA: Encoder-Agnostic Interpretation of Vision Foundation Model Embeddings via Spatial Probing

・arXiv:2608.10131v1 Announce Type: cross Abstract: Vision foundation models are increasingly used as reusable encoders in medical image computing, yet their high-dimensional spatial embeddings are difficult to inspect beyond downstream task performance or global dimensionality reduction. ・We propose position-prompted PCA (P3CA), an encoder-agnostic method for local probing of channel-rich spatial tensors. ・Given a user-
cs.LG updates on arXiv.org

Pair-Centric Graph Rewiring for Over-Squashing via Optimal Transport-Guided Communication Alignment

・arXiv:2608.10619v1 Announce Type: new Abstract: Message-passing neural networks (MPNNs) often struggle when task-relevant information is distributed across distant regions of a graph, since local propagation must compress remote signals through limited structural interfaces. ・Graph rewiring provides a structural response to over-squashing. ・Most existing methods rely on edge-level bottleneck scores or graph-level conne
WIRED

Paramount+ Coupon Codes and Deals for August 2026

・Save on streaming with the latest Paramount+ promo codes and deals, including 50% off subscriptions, free trials, and more.
cs.LG updates on arXiv.org

Partially Observable Learning for Multi-Platform Dispatch Optimization

・arXiv:2608.10897v1 Announce Type: new Abstract: Instant delivery platforms have become a critical component of urban logistics, increasingly relying on crowdsourced couriers to fulfill highly dynamic orders. ・In real-world systems, couriers are not exclusive to a single platform and may concurrently serve multiple platforms, while each platform can only observe its own orders and couriers' interactions due to privacy
機械学習タグが付けられた新着記事 - Qiita

Particle Filter を勉強してみる

・Particle Filter を勉強してみる 本記事は個人学習の整理として作成しています TL;DR Particle Filterは,事後分布を大量の「粒子(仮説)」の集まりで表現する逐次ベイズ推定 アルゴリズムは 3 ステップの繰り返し:予測(粒子を動か...
cs.LG updates on arXiv.org

Path Integral Value Matching for Linear Quadratic Stochastic Optimal Control

・arXiv:2608.10777v1 Announce Type: new Abstract: Linear Quadratic Stochastic Optimal Control (LQ-SOC) establishes a fundamental framework for steering noisy dynamical systems and has recently gained renewed interest in the machine learning community. ・However, current state-of-the-art policy-based methods suffer from prohibitive computational costs and instability due to their heavy reliance on full-trajectory simulati
cs.LG updates on arXiv.org

Physics-informed Diffusion Generative Model for Time-Series Data Synthesis in Dynamic Systems

・arXiv:2608.10941v1 Announce Type: new Abstract: Industrial time-series signals, such as turbine temperature and rotational speed in aero-engines, are essential for monitoring the health and operational status of complex dynamical systems. ・However, collecting such data is often limited by harsh environments (e.g., high temperature and high pressure) and the high cost of experimental testing. ・To address this challenge,
cs.LG updates on arXiv.org

Physics-Informed Machine Learning in Prognostics and Health Management: A Systematic Literature Review

・arXiv:2608.10047v1 Announce Type: new Abstract: In modern industry, keeping complex systems reliable, safe, and efficient hinges on Prognostics and Health Management (PHM). ・Machine Learning (ML) has largely driven advancements in diagnostics and prognostics, yet purely data-driven models face inherent limitations, such as poor generalization, an inability to infer causal relationships, and a lack of interpretability.
#LLMタグ

PixAI Studioが提供する4つの核心的な機能を詳しく見ていきましょう。これらの機能が組み合わさることで、従来のAI制作環境では実現困難だった「ワンストップ制作体験」が実現しています。

・どうも皆さん!コンビニの『レンジで3分』は、袋を破く時間を含めていないのが納得いきません、 葉加瀬あい です! 最近、AI制作の現場で「ノードベース(キャンバス型)」のワークフローが一気に主流になってきましたよね! 続きをみる
cs.LG updates on arXiv.org

Population-Aware Physics-Informed Neural Particle Flow for Robust Spacecraft Bayesian Navigation

・arXiv:2606.10959v2 Announce Type: replace Abstract: Spacecraft navigation often requires Bayesian inference from sparse nonlinear measurements that produce curved, multimodal, or geometrically constrained posterior distributions. ・Physics-informed neural particle flow (PINPF) addresses such problems by learning a deterministic prior-to-posterior transport field from the governing probability evolution equation, but it
cs.LG updates on arXiv.org

Post-Calibration Reliability Reranking of Relevance Decisions via Label-wise Monotone Projection

・arXiv:2608.10406v1 Announce Type: cross Abstract: Web search, product search, and question-answering retrieval systems often assign a relevance label and confidence score to each query-candidate pair. ・The relevance label describes how well a page, product, or passage matches the query, while the confidence often guides downstream use or fallback decisions. ・Post-hoc calibration is therefore needed because misaligned c
Hugging Face Papers

Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference

Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference
cs.LG updates on arXiv.org

Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference

・arXiv:2608.10288v1 Announce Type: new Abstract: The Large Language Model from Power Law Decoder Representations (PLDR-LLM) and its attention, Power Law Graph Attention (PLGA), replace the fixed bilinear form of scaled dot-product attention (SDPA) with a learned, input-generated bilinear operator $G_{LM}$, built from a positive tensor $A_{LM}$ by elementwise power laws. ・The architecture is fully specified, verified ag
機械学習タグが付けられた新着記事 - Qiita

Power Queryでニューラルネット

・はじめに 最近エッジAIが流行っているらしい。 ・なんでも、いろんなものにAIが入るらしい。 ・はやくPowerQueryにもAIを入れないと、 遅れてるゥ~ って、LLMに煽られちゃうかもしれない。
cs.LG updates on arXiv.org

Predictive Entropy as a Joint Screen for Error and Paraphrase Instability in Medical Vision-Language Models

・arXiv:2604.08941v2 Announce Type: replace Abstract: Medical Vision-Language Models (VLMs) answering binary presence questions on chest radiographs can fail in two linked ways: they are confidently wrong, and they change answers when a clinically equivalent question is rephrased. ・In a binary answer head both failures track the same logit margin, so we test how well one score screens for both. ・On MedGemma-4B-IT across
Hugging Face Papers

Previous

Previous
cs.LG updates on arXiv.org

ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions

・arXiv:2608.10621v1 Announce Type: new Abstract: Recent research on Large Language Model (LLM) safety has widely adopted guardrails to identify unsafe LLM outputs. ・Existing guardrails typically formulate safety assessment as a deterministic classification task, mapping a discrete token sequence to a discrete safety label. ・However, this paradigm has two limitations: First, safety assessment is inherently an uncertain p
cs.LG updates on arXiv.org

Procedural Fairness Failures in RLHF from Preference Averaging

・arXiv:2608.10126v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) aggregates heterogeneous preferences into a single reward model, assuming preference homogeneity. ・When preferences are heterogeneous, this aggregation induces a procedural fairness failure where majority preference groups dominate reward learning while minority preferences are systematically under-represented.
cs.LG updates on arXiv.org

Progressive Semantic Communication for Efficient Edge-Cloud Vision-Language Models

・arXiv:2604.26508v2 Announce Type: replace Abstract: Deploying Vision-Language Models (VLMs) on edge devices remains challenging due to their substantial computational and memory demands, which exceed the capabilities of resource-constrained embedded platforms. ・Conversely, fully offloading inference to the cloud is often impractical in bandwidth-limited environments, where transmitting raw visual data introduces subst
cs.LG updates on arXiv.org

Projected climate memory and inherited warm-tail risk in accelerated European summer warming

・arXiv:2608.09966v1 Announce Type: cross Abstract: European summer warming reflects interactions among background change, persistent ocean--land--circulation states, and same-season variability. ・We develop an empirical reduced-dynamics framework that decomposes regional summer indicators into inherited slow-state memory, its predictable component, and contemporaneous innovation. ・Projection-operator theory motivates th
cs.LG updates on arXiv.org

ProTAGAD: A Foundation Model for TAG Anomaly Detection with Decoupled Topological and Textual Prototypes

・arXiv:2608.10699v1 Announce Type: new Abstract: Text-Attributed Graphs (TAGs), endowed with abundant textual content along with topological structures, have emerged as a versatile backbone for real-world anomaly detection spanning large language model security, social network moderation, and cyber threat identification. ・Unlike conventional Graph Anomaly Detection (GAD), which relies primarily on structural irregulari
cs.LG updates on arXiv.org

Proteo-R1: Reasoning Foundation Models for De Novo Protein Design

・arXiv:2605.02937v2 Announce Type: replace Abstract: Deep learning in de novo protein design has achieved atomic-level fidelity. ・However, existing models remain largely non-deliberative: they directly synthesize molecular geometries without explicitly reasoning about which residues or interactions are functionally essential. ・As a result, design decisions are entangled with continuous sampling dynamics, limiting interp
cs.LG updates on arXiv.org

Putting a Face to Forgetting: Continual Learning meets Mechanistic Interpretability

・arXiv:2601.22012v3 Announce Type: replace Abstract: Catastrophic forgetting in continual learning is often measured at the performance or last-layer representation level, overlooking the underlying mechanisms. ・We introduce a mechanistic framework that offers a geometric interpretation of catastrophic forgetting as the result of transformations to the encoding of individual features. ・These transformations can lead to
Google DeepMind News

Putting sign language AI into users’ hands

・Introducing sign-language-to-text (SL2T), our breakthrough model powering new sign language features for Deaf and hard of hearing users.
Zennの「大規模言語モデル」のフィード

Quadro P2200を2枚にしてGemma 4 12Bを64Kコンテキストで動かしてみた

・以前、手元で眠っていたQuadro P2200でllama.cppを動かした記事を書きました。 ・https://zenn.dev/nooop/articles/3fba81b0b7b732 今回はP2200をもう1枚追加してみました。 ・これで合計のVRAMは10GBとなり、Gemma4-E4Bより一回り大きなGemma4-12Bを動かせるはずです。
cs.LG updates on arXiv.org

Quantifying the noise sensitivity of the Wasserstein metric for images

・arXiv:2510.01015v3 Announce Type: cross Abstract: Wasserstein metrics are increasingly adopted as similarity scores for images. ・We consider the sensitivity of Wasserstein metrics with respect to pixel-wise additive noise when the images are treated as discrete measures on the pixel grid. ・We derive finite-sample expectation bounds for a Gaussian noise model.
cs.LG updates on arXiv.org

Quantum Incremental Learning with Mixed State Prototypes

・arXiv:2608.10464v1 Announce Type: cross Abstract: Incremental learning models are required to learn new classes sequentially without catastrophic forgetting, while operating under parameter and memory constraints. ・In the Noisy Intermediate-Scale Quantum (NISQ) era, although quantum neural networks offer advantages in feature mapping, hardware limitations restrict circuit width. ・Furthermore, traditional quantum classi
cs.LG updates on arXiv.org

REATS: LLM Reasoning-based Ensemble Learning for Adaptive Time Series Forecasting

・arXiv:2608.10149v1 Announce Type: new Abstract: Due to the diversity of real-world time series, no single forecasting model consistently dominates across all samples. ・Ensemble learning addresses this by combining complementary model strengths, yet existing methods rely on fixed rules or black-box models based solely on numerical inputs, failing to leverage LLM reasoning for interpretable weighting decisions.
cs.LG updates on arXiv.org

Recovering Wasted Compute in Autoresearch Agents

・arXiv:2608.10424v1 Announce Type: cross Abstract: A slew of recent works develop agents for solving research problems end-to-end, a paradigm increasingly referred to as autoresearch. ・Such agents have inspired large industry investment, motivated by their potential to automate time-consuming human labor and customize machine learning solutions for specialized applications. ・In this paper, we study the modeling pipeline
Hugging Face Papers

Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation

Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation
cs.LG updates on arXiv.org

Regression and Classification with Single-Qubit Quantum Neural Networks

・arXiv:2412.09486v2 Announce Type: replace-cross Abstract: The literature reflects a mutually beneficial relationship between machine learning and quantum computing, where progress in one field frequently drives improvements in the other. ・Motivated by the rich connection between these areas, we use a resource-efficient and scalable Single-Qubit Quantum Neural Network (SQQNN) for both regression and classification task
cs.LG updates on arXiv.org

ReOrder-OPD:Reliability-Aware Prompt Ordering for On-Policy Distillation

・arXiv:2608.10905v1 Announce Type: new Abstract: On-policy distillation (OPD) applies token-level teacher supervision to student-generated trajectories, but this supervision is not always reliable. ・Existing methods use local confidence or teacher-student agreement to weight, filter, or truncate the sampled trajectory. ・These signals do not directly determine whether the teacher can continue a student prefix to a correc
cs.LG updates on arXiv.org

ReRound: Reconstructive Rounding to Resolve Midpoint Ambiguity in Calibration-Free LLM Quantization

・arXiv:2608.11045v1 Announce Type: new Abstract: ReRound (Reconstructive Rounding) is a post-training quantization method that addresses the midpoint ambiguity inherent in standard round-to-nearest (RTN) schemes when quantizing weights near the centers of quantization intervals. ・Starting from a pretrained LLM, ReRound trains a conditional diffusion model to produce continuous reconstructions of low-bit weights for the
cs.LG updates on arXiv.org

Retrieval-Augmented Vision Foundation Models for Robust Leukemia Cell Classification across Multiple Microscopy Datasets

・arXiv:2608.10657v1 Announce Type: cross Abstract: Leukemia cell image classification is challenged by real-world domain shifts from acquisition, staining, illumination, and site protocols, causing single-dataset models to generalize poorly in real clinical scenarios. ・This work presents a robust framework for leukemia classification across multiple heterogeneous datasets using a two-stage pipeline with a pretrained vi
cs.LG updates on arXiv.org

Retrieval-Corrected Conformal Prediction for Time Series

・arXiv:2608.10553v1 Announce Type: new Abstract: Conformal prediction (CP) provides distribution-free prediction intervals for fixed forecasters, but its standard calibration procedure is often inefficient for time series data, where forecast errors are temporally dependent and change across time and operating conditions. ・Recent time series CP methods improve local calibration using recent, weighted, or localized resi
cs.LG updates on arXiv.org

Riemann GeoResolver: A Non-Euclidean Attention Framework from Euclidean Resolver to Hyperbolic-Spherical Geometry

・arXiv:2608.10416v1 Announce Type: cross Abstract: We present a theoretical foundation for inverse-distance attention, from its Euclidean prototype (Resolver) to its non-Euclidean realization (Riemann GeoResolver). ・The Euclidean part establishes three core theorems: (1) circuit separation---IDA achieves exact retrieval with $\mathcal{O}(1)$ resources while softmax requires $\Omega((\log n)^2)$ width; (2) a Polyak--Loj
cs.LG updates on arXiv.org

Risk-Averse Wasserstein Distributionally Robust Online Learning

・arXiv:2602.20403v2 Announce Type: replace Abstract: We study distributionally robust online learning, where a risk-averse learner updates decisions sequentially to guard against worst-case distributions drawn from a Wasserstein ambiguity set centered at past observations. ・While this paradigm is well understood in the offline setting through Wasserstein Distributionally Robust Optimization (DRO), its online extension
ITmedia NEWS 最新記事一覧

RIZAP運営のECサイトに不正アクセス カード情報含む個人情報が流出の可能性 サイトは一時閉鎖

・RIZAPは8月10日、運営するスポーツ・アウトドア用品の通販サイト「APORITOオンラインストア」が不正アクセスを受け、利用者の個人情報とクレジットカード情報が外部へ流出した可能性があると発表した。サイトは5日から閉鎖しており、再開時期は未定。
cs.LG updates on arXiv.org

Robust and Secure Code Watermarking for Large Language Models via ML/Crypto Codesign

・arXiv:2502.02068v3 Announce Type: replace-cross Abstract: This paper introduces RoSeMary, the first-of-its-kind ML/Crypto codesign watermarking framework that regulates LLM-generated code to avoid intellectual property rights violations and inappropriate misuse in software development. ・High-quality watermarks adhering to the detectability-fidelity-robustness tri-objective are limited due to codes' low-entropy nature.
cs.LG updates on arXiv.org

Robust Multi-Agent Bandits with Heavy-Tailed Rewards and Information Asymmetry

・arXiv:2608.10529v1 Announce Type: new Abstract: The multi-armed bandit problem is a central framework in sequential decision-making, extensively studied under sub-Gaussian reward assumptions. ・However, real-world applications often involve heavy-tailed reward distributions and decentralized, information-asymmetric interactions. ・We study multi-agent multi-armed bandits with heavy-tailed rewards under three information-
WIRED

Rover Promo Codes and Referral Deals for 2026

Rover Promo Codes and Referral Deals for 2026
cs.LG updates on arXiv.org

Rule of Thumb: Explaining Artificial Intelligence Systems using Partial Information

・arXiv:2608.10766v1 Announce Type: cross Abstract: Explainable Artificial Intelligence (XAI) seeks to explain how an Artificial Intelligence (AI) system arrived at a particular decision. ・We propose ''Rule of Thumb'' (RoT) explanations, a new approach to XAI based upon a novel formulation that identifies the most relevant features for predicting the behaviour of an AI system, for a particular datapoint. ・We show how RoT
WIRED

Sam's Club Promo Codes and Membership Deals for August 2026

・Save on bulk groceries, household essentials, and electronics with a verified Sam's Club promo code or membership discount.
cs.LG updates on arXiv.org

Same Targets, Different Computation: How Post-Training Divides Work Across Model Layers

・arXiv:2605.07284v2 Announce Type: replace Abstract: A late-layer change learned during post-training may work on the base model's earlier state, or it may depend on earlier computation learned with it. ・We distinguish these cases with a four-cell diagnostic that crosses base or descendant upstream states with base or descendant late stacks. ・A large late-stack effect need not imply strong upstream dependence.
cs.LG updates on arXiv.org

Scaling Self-Play with Self-Guidance

・arXiv:2604.20209v2 Announce Type: replace Abstract: LLM self-play algorithms are notable in that, in principle, nothing bounds their learning: a Conjecturer model creates problems for a Solver, and both improve together. ・However, in practice, existing LLM self-play methods do not scale well with large amounts of compute, instead hitting learning plateaus. ・We argue this is because over long training runs, the Conjectu
cs.LG updates on arXiv.org

Scheduling Mixed RL Rollouts Beyond Prefix Locality

・arXiv:2608.11152v1 Announce Type: cross Abstract: Modern reinforcement learning (RL) post-training pipelines for large language models (LLMs) increasingly combine rollout workloads across multiple domains and feedback paradigms. ・Prefix-aware routing improves inference efficiency through cache reuse and load balancing, but it does not control how heterogeneous rollout sessions compete for KV-cache capacity.
cs.LG updates on arXiv.org

SCOUT: Symmetric Consensus Outlier Detection for Failure Localization in LLM Pre-Training

・arXiv:2608.11034v1 Announce Type: cross Abstract: In LLM pre-training, synchronization propagates rank-local stalls, slowdowns, and numerical errors into job-wide symptoms, obscuring their origin. ・Existing diagnosis often relies on in-process monitors that cannot report after the trainer blocks or terminates, or on post-mortem logs that preserve only synchronized symptoms; offline health tests lose the workload and o
cs.LG updates on arXiv.org

SeFaR: Semantic Feature-aware Robustness Testing of Deep Neural Networks

・arXiv:2608.10289v1 Announce Type: cross Abstract: Deep neural networks are increasingly deployed in safety-critical domains as perception modules, where failures are often caused due to rare and under-represented scenarios. ・This necessitates the need to evaluate the semantic robustness of perception models; conformance of behavior to high-level requirements over real-world perceptual variability. ・To address this, we
cs.LG updates on arXiv.org

SeFoRA: Sketch-Aggregated Federated Low-Rank Adaptation with Heterogeneous Client Ranks

・arXiv:2608.10144v1 Announce Type: new Abstract: We consider federated parameter efficient fine-tuning of large neural networks with low-rank adaptation (LoRA,~Hu et al.\ 2022). ・Combining LoRA with federated PEFT introduces challenges absent from either setting alone: clients may use different LoRA ranks, making their factor matrices dimension-incompatible, and factor-wise averaging suffers from a bilinear mismatch.
cs.LG updates on arXiv.org

Self-Normalized Inference for Constant-Stepsize Temporal-Difference Learning under Markovian Sampling

・arXiv:2608.10896v1 Announce Type: cross Abstract: Constant-stepsize temporal-difference (TD) learning is attractive for policy evaluation, but inference from a single Markov trajectory must account for serial dependence and a stepsize-dependent stationary target. ・For fixed-stepsize linear TD, we establish a functional central limit theorem whose covariance retains the multiplicative component induced by the random TD
cs.LG updates on arXiv.org

Sequential Modality Dropout for Robust Multi-Modal Sequential Recommendation

・arXiv:2608.10240v1 Announce Type: cross Abstract: Multi-modal sequential recommenders assume every item carries every modality, but real product catalogs often miss images or text, and a model trained on complete data loses much of its recommendation accuracy when a modality is unavailable at serving time. ・We propose Sequential Modality Dropout (SMD): during training, each modality stream (image and text) is independ
cs.LG updates on arXiv.org

Share First, Route What Remains: A Unified Framework for Token-Adaptive MoE Computation

・arXiv:2608.10392v1 Announce Type: new Abstract: Mixture-of-experts (MoE) models have recently moved beyond routing a fixed number of complete experts. ・Shared-expert designs preserve reusable knowledge, fine-grained methods vary computation within experts, and dynamic routers adapt the number of active experts. ・Yet these decisions are usually made independently, overlooking a basic dependency: extracting reusable comp
cs.LG updates on arXiv.org

Sheaf-Based Federated Representation Learning

・arXiv:2608.10016v1 Announce Type: new Abstract: Heterogeneous federated systems require agents to learn and exchange informative representations despite differences in data distributions, sensing modalities, model architectures, latent dimensionalities, and local learning objectives. ・To address this challenge, we propose Sheaf-based Federated Representation Learning (SFRL), a general framework that jointly optimizes
Hugging Face Papers

SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure

SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure
ITmedia NEWS 最新記事一覧

SpaceXAI、AIエージェント「Grok Bot」発表 クラウド環境で常時稼働、GrokやCursorの有料ユーザー向けに

・米SpaceXAIは8月11日(現地時間)、常時稼働型のAIエージェント「Grok Bot」(β版)を発表した。チャットでタスクを渡すとBotがアプリやWebサイトを操作して作業を進め、役割の異なる複数のBotを並列で働かせることもできる。
cs.LG updates on arXiv.org

Spectral Embeddings of Degree-$\alpha$ Laplacians in Random Dot Product Graphs

・arXiv:2608.10845v1 Announce Type: cross Abstract: Spectral clustering methods for network data are commonly based on a few matrix representations, such as the adjacency matrix and the symmetric Laplacian. ・We study a continuum of degree-normalized spectral embeddings that includes these commonly used choices as special cases. ・Under a random dot product graph model, we establish a row-wise central limit theorem for thi
Hugging Face Papers

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information
cs.LG updates on arXiv.org

SPOTting the Future: Lookahead Explanations for Deep Reinforcement Learning

・arXiv:2608.09967v1 Announce Type: cross Abstract: Deep reinforcement learning (DRL) agents achieve strong performance in complex environments, yet their decision-making processes remain difficult to interpret. ・We introduce SPOT (Sampling Policy Observation Tree), a novel model-agnostic, sampling-based framework for interpreting DRL policies. ・Given access to the policy and an environment simulator, SPOT constructs an
cs.LG updates on arXiv.org

SQuaT: Self-Supervised Knowledge Distillation via Student-Aware Quantized Teacher Features

・arXiv:2608.10709v1 Announce Type: new Abstract: Quantization-Aware Training (QAT) enables the deployment of quantized models with minimal accuracy degradation. ・However, in practical scenarios, training labels are often unavailable due to privacy, copyright, or cost constraints. ・Knowledge Distillation (KD) is a common approach to address this challenge, but we observe that prior work combining QAT with KD suffers from
cs.LG updates on arXiv.org

Status Association Does Not Reliably Predict Decision Leakage

・arXiv:2608.10089v1 Announce Type: cross Abstract: Bias evaluations often move too quickly from evidence that a model encodes a social association to claims that the same association will alter consequential decisions. ・We test whether that inference is warranted using Chilean surnames as controlled socioeconomic probes. ・We evaluate eight frozen model-provider cells on 1,032 prompts each, yielding 8,256 verified primar
cs.LG updates on arXiv.org

Stay or Stray - A Dynamical Systems Viewpoint of Popularity Bias

・arXiv:2608.10474v1 Announce Type: cross Abstract: Popularity bias in recommendation systems arises when a majority user class generates disproportionate interaction data, causing the system to increasingly favour it while degrading recommendation quality for niche users. ・While extensive empirical evidence of popularity bias exists, the dynamics leading to its emergence are not well understood. ・In this work, we study
cs.LG updates on arXiv.org

STCAD: Scalable Trajectory Clustering and Anomaly Detection on Terabyte-Scale AIS Data

・arXiv:2608.10249v1 Announce Type: new Abstract: We present a scalable framework for unsupervised clustering of maritime trajectories derived from terabyte-scale Automatic Identification System (AIS) archives. ・Variable-length trajectories are encoded with a custom BERT-based model trained via masked token modeling and clustered using CURE hierarchical clustering, producing physically interpretable trajectory groups wi
cs.LG updates on arXiv.org

Stochastic Emulation of a Fully Coupled Preindustrial E3SMv3 Simulation

・arXiv:2608.10277v1 Announce Type: cross Abstract: We present a stochastic coupled emulator of E3SM version 3, built on the SamudrACE framework, which couples an atmosphere emulator (ACE2) with a full-depth ocean emulator (Samudra). ・We replace the deterministic atmosphere emulator with its stochastic counterpart, ACE2S, and fine-tune the coupled system with a probabilistic objective, so that the atmosphere acts as a s
cs.LG updates on arXiv.org

TACTICL: Task-Aware Compression of Tabular ICL Models

・arXiv:2608.10837v1 Announce Type: new Abstract: The strong performance of foundation models for tabular tasks comes at substantial inference costs. ・Distilling models into task-specific architectures reduces model size and computational demands but also sacrifices in-context adaptability. ・Here we introduce TACTICL, an automated task-aware compression framework for tabular in-context learning models that jointly prunes
cs.LG updates on arXiv.org

Temporal Straightening for Latent Planning

・arXiv:2603.12231v3 Announce Type: replace Abstract: Learning good representations is essential for latent planning with world models. ・While pretrained visual encoders produce strong semantic visual features, they are not tailored to planning and contain information irrelevant -- or even detrimental -- to planning. ・Inspired by the perceptual straightening hypothesis in human visual processing, we introduce temporal st
The Verge

The 7 biggest announcements of Google’s Pixel 11 launch

・Google is hosting a Pixel launch event on Wednesday night where it will show off its latest lineup of devices. ・But you don't have to wait until then for all the news; we were able to check out the devices ahead of the event and take a look at everything Google's unveiling, including the latest Pixel phones, new health features on the Pixel Watch 5, and more. ・Google will be livestreaming the event on Wednesday at 6PM
WIRED

The Best Action Cameras for All Your Craziest Adventures (2026)

・Gearing up to shred the slopes or dive into the seas? ・These photography tools are made for danger.
cs.LG updates on arXiv.org

The Evaluation Protocol Determines the Result: An Independent Reproduction of LeWorldModel on TwoRoom

・arXiv:2608.10145v1 Announce Type: new Abstract: LeWorldModel trains a latent world model with a prediction loss and a single anti-collapse regulariser, and reports approximately 87% of goals reached on TwoRoom, its simplest diagnostic environment. ・We reproduce that result by independent reimplementation on roughly $25 of rented compute, with all evaluation on one laptop CPU. ・We reach 94.0% at the repository's evaluat
WIRED

The Job-Interview Tattoo Guy Everyone Got Mad at Finally Explains Himself

・LemonLime cofounder Jordan Zietz hears your criticism loud and clear. ・That’s why he got his startup’s logo tattooed on his shoulder.
cs.LG updates on arXiv.org

The Kuramoto Neural Operator: Learning to Solve PDEs via Coupled Oscillator Dynamics

・arXiv:2608.10234v1 Announce Type: cross Abstract: Operator learning is a rapidly advancing area of computational science. ・It is particularly well suited to problems where a partial differential equation (PDE) must be solved repeatedly under varying physical configurations. ・Most existing architectures represent the solution operator in a fixed basis.
cs.LG updates on arXiv.org

The Matching Principle: When Does a Training Penalty Cover Deployment Shift?

・arXiv:2605.22800v3 Announce Type: replace Abstract: Ordinary training optimises the task loss and then stops. ・It never pays for internal representation energy: Jacobians can stay large in directions that never helped the label, so even small label-preserving noise throws the model off---a design gap that classical noise-injection theory fixes at second order, but only when applied as default regularisation, which cur
cs.LG updates on arXiv.org

The Parser Already Knows: Lightweight Bias Correction in Constrained Decoding

・arXiv:2608.10137v1 Announce Type: cross Abstract: Grammar Constrained Decoding (GCD) forces Language Models (LMs) to produce syntactically valid outputs by masking out non-conforming tokens at each step. ・However, rigid masking distorts the model's underlying probability distribution, often biasing generation toward valid but suboptimal outputs. ・While online sampling restores this distribution, it requires computation
WIRED

This Coin-Sized Device Can Hack a Boeing 737

・Security researchers found that in less than 60 seconds, they could open a hatch on a plane’s exterior, plug in a tiny device, and redirect the aircraft’s autopilot or sabotage its flight plan.
cs.LG updates on arXiv.org

Threshold Structure of Optimal Policies in Restart POMDPs

・arXiv:2608.10936v1 Announce Type: cross Abstract: We study a Restart POMDP (Partially Observable Markov Decision Process) on a general Borel state space, where the controller either lets the hidden state evolve unobserved or restarts the system and observes the new state. ・Exploiting a sufficient-statistic representation consisting of the last observed state and the elapsed time since restart, we reduce the problem to
cs.LG updates on arXiv.org

TideRL: Boosting Agentic RL Goodput with Readiness-Aware Scheduling

・arXiv:2608.10402v1 Announce Type: new Abstract: Reinforcement learning (RL) for large language models is moving toward multi-turn agentic workloads, where rollout tasks repeatedly pause for external environments, resume with growing contexts, and finish at highly variable times. ・In this setting, RL training goodput, measured by training throughput, matters more than raw GPU occupancy: GPU waiting and repeated prefill
cs.LG updates on arXiv.org

Time-Series Foundation Model Embeddings for Remaining Useful Life Estimation

・arXiv:2606.11990v3 Announce Type: replace Abstract: Remaining Useful Life (RUL) prediction is essential for industrial predictive maintenance, yet many learning-based approaches rely on extensive feature engineering or large labeled datasets to train task-specific sequence models. ・In this work, we introduce a lightweight learning approach, in which we leverage a frozen pretrained time-series foundation model (TSFM) a
cs.LG updates on arXiv.org

TimePre: Bridging Accuracy, Efficiency, and Stability in Probabilistic Time-Series Forecasting

・arXiv:2511.18539v3 Announce Type: replace Abstract: We propose TimePre, a simple framework that unifies the efficiency of Multilayer Perceptron (MLP)-based models with the distributional flexibility of Multiple Choice Learning (MCL) for Probabilistic Time-Series Forecasting (PTSF). ・Stabilized Instance Normalization (SIN), the core of TimePre, is a normalization layer that explicitly addresses the trade-off among accu
cs.LG updates on arXiv.org

Topological Feasibility Guarantees for Differentiable Predictive Control

・arXiv:2608.10332v1 Announce Type: cross Abstract: Differentiable predictive control (DPC), a self-supervised learning approach for approximating explicit model predictive control (MPC) policies, offers significant computational advantages over online optimization-based MPC. ・However, feasibility guarantees, a core requirement for safe control, are currently provided either probabilistically or via online safety filter
cs.LG updates on arXiv.org

Toward Human Rights Benchmarking for LLMs: A Pilot Methodology

・arXiv:2608.10268v1 Announce Type: new Abstract: Large language models (LLMs) increasingly mediate legal determinations over what human rights are realized, and how. ・Yet, no evaluation benchmark exists to assess whether they can reason correctly about human rights law. ・To this end, we report our efforts to develop a robust and scalable methodology for creating HumRightsBench: the first expert-validated, scenario-based
cs.LG updates on arXiv.org

Towards Sustainable Artificial Intelligence: A Comprehensive Review and Comparative Analysis of Deep Learning Models' Carbon Footprint

・arXiv:2608.09998v1 Announce Type: cross Abstract: Artificial Intelligence (AI) and Machine Learning (ML) have become powerful tools for supporting and automating complex human tasks. ・Despite their benefits, growing attention has been directed toward their environmental implications, primarily due to their high energy demands and associated carbon emissions. ・This concern is particularly relevant in light of the increa
cs.LG updates on arXiv.org

Transformer Geometry Observatory TGO-IV: Developmental Topology Observatory

・arXiv:2608.09997v1 Announce Type: new Abstract: Transformers have had a profound impact on the world of language processing and computer vision. ・As efforts to answer the million-dollar question of ``How does a Transformer learn?" have been increasing, existing interpretability studies primarily analyze representations at isolated layers or the network as a whole, while the developmental evolution of individual repres
cs.LG updates on arXiv.org

TS-Mob: Social and Geographical-Aware Time Series Foundation-Model Framework for Human Mobility Prediction

・arXiv:2507.00945v3 Announce Type: replace Abstract: Short-term forecasting of aggregated human mobility flows supports urban planning, intelligent transportation systems, and emergency response, yet existing models often require substantial mobility history and learn spatial structure implicitly through grids or graphs. ・Time series foundation models provide strong temporal priors but typically lack explicit geographi
Hugging Face Papers

TSDS-Toolbox: A Toolbox for Measuring Time-Series Dataset Similarity

TSDS-Toolbox: A Toolbox for Measuring Time-Series Dataset Similarity
The Verge

Twitch streamers can now opt out from training Amazon’s AI

・Twitch users can now opt out of allowing their content to be used to train Amazon's generative AI models. ・Opting out means that "your streams, VODs, clips, stream chats, and pictures and text on your channel" won't be used in "future training" of an Amazon AI model "whose purpose is to generate or synthesize text, audio, images, or video," according to a Twitch support page. ・Other "AI-supported" features like caption
cs.LG updates on arXiv.org

Two-stage Odd Residual Flows for Mean-Preserving Probabilistic Time Series Forecasting

・arXiv:2608.11114v1 Announce Type: new Abstract: Probabilistic forecasting plays an essential role in risk-sensitive decision-making, particularly in long-horizon settings. ・However, existing approaches often face a fundamental trade-off between distributional flexibility and accurate mean prediction. ・Traditional parametric methods, such as Mean Variance Estimation (MVE), can suffer from degraded point accuracy when tr
cs.LG updates on arXiv.org

Uncertainty-Aware Deep Learning for Genomics Applications: Insights from an Empirical Study

・arXiv:2608.11054v1 Announce Type: new Abstract: Deep learning models have emerged as the standard computational tool for a wide range of applications in genomics. ・Yet, uncertainty quantification (UQ) -- and more specifically, the reliability of different uncertainty estimates in this domain -- has received little systematic attention. ・This work presents an empirical analysis of UQ in deep learning models, focusing on
cs.LG updates on arXiv.org

Uncertainty-Aware Ensemble Deep Randomized Neural Networks for Classification

・arXiv:2608.10007v1 Announce Type: new Abstract: The current state-of-the-art (SOTA) deep randomized neural networks, such as deep Random Vector Functional Link (dRVFL) and ensemble deep RVFL (edRVFL), treat all training samples uniformly, which limits their robustness and effectiveness when applied to real-world datasets containing noise and outliers. ・Furthermore, the propagation of contaminated features across hidde
cs.LG updates on arXiv.org

UniMod: Enhancing Multi-Modal Medical Diagnosis through Cross-Modality and Within-Modality Alignment

・arXiv:2608.10316v1 Announce Type: cross Abstract: Multi-modal learning combining medical images and clinical text is promising for disease diagnosis. ・However, standard multi-modal training leads to shortcut learning: models exploit the easier modality (e.g., diagnostic cues in text) while neglecting harder-to-learn features (e.g., subtle visual patterns). ・We propose UniMod, a framework that mitigates shortcut learnin
Hugging Face Papers

UniMoMo: Expert Merging-Based MoE Acceleration for Large Recommendation Models

UniMoMo: Expert Merging-Based MoE Acceleration for Large Recommendation Models
cs.LG updates on arXiv.org

UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations

・arXiv:2608.10835v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) achieve impressive visual reasoning and dialogue capabilities, yet frequently hallucinate content unsupported by the visual input. ・Effective mitigation requires token-level localization, enabling targeted intervention without discarding the entire response. ・Existing detectors require expensive full-model fine-tuning, rely on extern
#LLMタグ

Unsloth Desktop からエージェントに繋ぐ

・そのままローカルLLMが利用できるようになるわけではなく、ggufの用意が必要。ダウンロードするところから始めるとします。 ・Macでの利用方法 URL: https://unsloth.ai/docs/jp/desktop 続きをみる
#LLMタグ

Unsloth Desktop で最初の1回を回す順番 ─ 既定が30ステップなので、触るのは2周目から

・学習画面を開くと、モデル、データセット、パラメーター、構成が並ぶ。パラメーターを開けば学習率もランクも勾配累積も出てくる。ここで手が止まる、というのは自分もよく分かる。 ・止まらなくていい理由が、公式ドキュメントと同梱ソースを読んで分かった。モデルを選んだ瞬間に、そのモデル用の既定値が流し込まれる。中身は「30ステップだけ回す」設定で、ほぼ全モデルがそうだ。
cs.LG updates on arXiv.org

URS: A Unified Neural Routing Solver for Cross-Problem Zero-Shot Generalization

・arXiv:2509.23413v3 Announce Type: replace Abstract: Multi-task neural routing solvers have emerged as a promising paradigm for their ability to solve multiple vehicle routing problems (VRPs) using a single model. ・However, existing neural solvers typically rely on predefined problem constraints or require per-problem fine-tuning, which substantially limits their zero-shot generalization ability to unseen VRP variants.
cs.LG updates on arXiv.org

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs

・arXiv:2608.10042v1 Announce Type: new Abstract: Tool-use LLMs are increasingly asked to act on users' behalf, but existing benchmarks usually focus on profile recall, style imitation, generic tool use, or response-level personalization. ・We introduce UserToolBench , a benchmark for personalized decision making in tool-use LLMs. ・UserToolBench tests whether a model can infer latent user preferences from interaction hist
cs.LG updates on arXiv.org

V-FiLLM: Verified Financial LLM Reasoning Benchmark

・arXiv:2608.11047v1 Announce Type: cross Abstract: While existing benchmarks have made substantial progress in evaluating LLMs across STEM domains, financial reasoning over structured data remains comparatively less explored. ・We introduce V-FiLLM, a framework that generates financial reasoning benchmarks from executable computation trees grounded in real tables, yielding items whose answers are correct by construction
cs.LG updates on arXiv.org

Validated Synthetic Patient Generation for Small Longitudinal Cohorts: Coagulation Dynamics Across Pregnancy

・arXiv:2604.07557v2 Announce Type: replace Abstract: Small longitudinal cohorts, common in maternal health, rare diseases, and early-phase trials, limit computational modeling because enrollment is slow and the data are too sparse to train reliable models. ・We present multiplicity-weighted Stochastic Attention (SA), a generative framework based on modern Hopfield networks. ・Stochastic attention stores real patient profi
Hugging Face Papers

VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?

VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?
cs.LG updates on arXiv.org

VIDS-Seg: Towards Reliable Uncertainty Quantification in Pediatric Cardiac Ultrasound Segmentation

・arXiv:2608.10903v1 Announce Type: cross Abstract: Reliable clinical deployment of machine learning requires models that know when they are likely to fail, particularly for subgroups underrepresented in training data. ・A common case is pediatric care, where models trained on adult cohorts can silently under-perform on children with no indication that something has gone wrong. ・As retraining with labeled pediatric data i
WIRED

Vimeo Promo Codes and Discounts: Up to 40% Off This August 2026

・Enjoy 25% off a membership, 40% off, plus an additional 10% off annual plans, and more deals to save at Vimeo.
WIRED

Walmart Promo Codes: Up to 65% Off for August 2026

・Score $10 off with our Walmart promo codes and coupon, and shop flash deals up to 65% off today.
cs.LG updates on arXiv.org

Weighted Sequential Bayesian Inference for Non-Stationary Linear Contextual Bandits

・arXiv:2307.03587v4 Announce Type: replace Abstract: In non-stationary linear contextual bandits, existing efficient algorithms typically rely on the Weighted Regularized Least-Squares (WRLS) estimator. ・Because WRLS only provides point estimates, previous methods typically construct surrogate distributions when aiming to perform Bayesian-like randomized exploration. ・To more properly establish the Bayesian principles,
cs.LG updates on arXiv.org

When Chain-of-Thought Helps and When It Hurts: An Empirical Investigation of the Serial-Depth Bottleneck in LLM Reasoning

・arXiv:2608.09942v1 Announce Type: cross Abstract: It is widely assumed that chain-of-thought (CoT) prompting universally improves LLM reasoning. ・We investigate this through the conceptual framework of the H_dp bandwidth bound (Chen et al., 2024): although the formal bound binds only asymptotically (at astronomically large prompt lengths), it identifies a real architectural bottleneck -- serial computation exceeding a
cs.LG updates on arXiv.org

When Do Anchor-Based Pointwise LLM Rerankers Help? Retriever Quality, Statistical Scope, and Anchor Design

・arXiv:2608.10528v1 Announce Type: cross Abstract: Anchor-based pointwise LLM reranking scores each candidate against a shared reference passage to recover cross-document context at pointwise cost. ・We study when this actually helps, using GCCP/PAGC as a representative method. ・Our study is reproduction-first.
cs.LG updates on arXiv.org

When Interpretability Is Unequally Distributed: Fairness in Hybrid Interpretable Models

・arXiv:2605.28626v2 Announce Type: replace Abstract: Hybrid interpretable models combine a transparent component with a black-box model by assigning some examples to the former and deferring the rest to the latter. ・While this design enables flexible tradeoffs between accuracy and interpretability, it also raises a distinct procedural fairness concern: some demographic groups may systematically receive interpretable de
cs.LG updates on arXiv.org

Why Does CLAUDE.md Keep Growing? Catastrophic Remembering in Agentic Coding

・arXiv:2608.11095v1 Announce Type: cross Abstract: Agentic coding READMEs like CLAUDE.md grow without bound in real repositories, stopping only when the repository retires or someone rewrites the file wholesale. ・We trace this to imperfect recall: appending an instruction is always cheap, but once an instruction's rationale is gone, deleting it without risking a correctness regression costs O(2^|D|) in a prompt of |D|
AI News & Artificial Intelligence | TechCrunch

Why Sandbar thinks it’s voice-enabled ring can avoid the AI hardware graveyard

・AI notetaking hardware has taken off over the past couple of years, with credit-card-sized devices, pendants, pins, and even transcribing earbuds all promising to capture your meetings and turn them into summaries and action items. ・Now, a whole wave of wearables — rings especially — are betting people want to capture stray thoughts and ideas the same way.
AI News & Artificial Intelligence | TechCrunch

Why Stream ring-maker Sandbar says the future of AI wearables is voice

・AI notetaking hardware has taken off over the past couple of years, with credit-card-sized devices, pendants, pins, and even transcribing earbuds all promising to capture your meetings and turn them into summaries and action items. ・Now, a whole wave of wearables — rings especially — are betting people want to capture stray thoughts and ideas the same way.
#LLMタグ

Will AI Replace Programmers in 2026? The Truth About Coding Careers

Will AI Replace Programmers in 2026? The Truth About Coding Careers
MarkTechPost

Xiaomi’s MiLM Plus Releases PROVE: Perception-Aligned Object Removal Metrics RC-S and RC-T With a Real-World Video Benchmark

・Object removal models have improved faster than the metrics used to judge them. ・Diffusion erasers now reconstruct shadows, reflections and occluded structure convincingly, yet PSNR, SSIM, LPIPS, ReMOVE and CFD frequently rank their outputs the wrong way. ・The root cause is structural: erasure is an ill-posed, one-to-many task, so no single ground truth exists to […] The post Xiaomi’s MiLM Plus Releases PROVE: Percepti
WIRED

You’re Thinking About Online Trends All Wrong

・From pessimism around dating to AI reshaping culture, cyber-ethnographer Ruby J. ・Thelot tells WIRED why people are putting too much stock into things that go viral.
cs.LG updates on arXiv.org

Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control

・arXiv:2608.10703v1 Announce Type: new Abstract: Large language models (LLMs) increasingly act in interactive settings where their behavioral styles affect user experience, safety, and downstream decision making. ・Existing LLM personality studies largely rely on self-report questionnaires administered in first-person settings, making the resulting profiles sensitive to surface elicitation choices and poorly grounded in
WIRED

Zoro Coupon Codes: 55% Off August

・Find the best Zoro coupon codes, deals, and free shipping offers to save big on industrial equipment, tools, and supplies this month.
#LLMタグ

お手軽LLMはじめてみた。その9.4(グラボ落ちる問題深堀り)

・考えられる原因をClaudeに聞いてみた。 ・補助電源接続不足 ← 最有力 RX 570 の 6 pin がまだ接続されていない?(その通り) 10 pin → 8+6 pin ケーブルの到着待ち?(その通り) 物理的接続不良 PCIe スロットが緩んだ GPU が完全に挿さっていない BIOS 設定 GPU が無効化されている マルチ GPU 設定がリセットされた ドライバ 再起動で Vulkan が renderD129 を認識できず 続きをみる
LLMタグが付けられた新着記事 - Qiita

さくらのAI Engineで「名もなき家事」に名前をつけようとしたら、名前はいらなかった話

・はじめに LLMを組み込んだアプリ開発を体験してみたいと考えていたところ、「さくらのAI Engine」3000リクエスト使い切りチャレンジ(https://qiita.com/official-events/bd14d28b53326d318fec )を知り、参加してみ...
Zennの「大規模言語モデル」のフィード

なぜAI時代にGoが最適な言語なのか

・はじめに 本記事では、GoogleのGolang Product ManagerのCameron Balahanと、Google Cloud Chief EvangelistのRichard Seroterによって投稿された、「Why Go is an Ideal Language for AI-Assisted Software Engineering」 を和訳するとともに、私の感想を少し書こうと思う。 ・また、本記事は元記事の著者の一人である、Cameron Balahanさんに翻訳と公開の許諾を得ております。 ・快く快諾していただき、本当にありがとございました!! !
#AIタグ

バトンを持って、ピンポンダッシュ

バトンを持って、ピンポンダッシュ
#AIタグ

ふるさとへ還る 神田川の夏の風

・故郷に帰りたくなった! (実家が事業に失敗して破産してしまい、実家は無いのだが) 私の記憶の「始まりの町」がありました。 ・早稲田のそば。神田川が流れ、都電が走る、お風呂屋さんがあり Oリング模様の坂道があり とにかく私の命がスタートした場所です。 ・60年以上前のことだけど、心惹かれてわくわくと旅をしてみました。
#AIタグ

ミニPC大高騰時代のプチおうちサーバ、wii-linux

・おうちサーバー、やってますか?皆さん。 ・このミニPC大高騰時代、Raspberry Piも3万円を越える中新たに始めたい人もなかなか手が出なくなってくる時代。 ・私も社会人とはいえ浪費に浪費を重ね、気づいた時には気軽に手を出すにはちょっと腰が引ける値段に。
Zennの「大規模言語モデル」のフィード

ローカルAIとContextLengthを超えた「思い出」を作る #3

・~海馬ってマジで見たまんまのネーミングだよね~ 前回まで、記憶を分類したり、自己連続性の三本柱なんて大仰なキーワードをまとめたりした。今回は、そのざっくりした概念を、実装の目線で見ていこうと思う。 ・1.で、それをいつ・だれが実行するのか? 情報を分類して、必要なものだけContextに載せる。そして記録は証跡を追加して壊れないように守る。なるほど、なんか出来たらすごそうだ。けれど、具体的にそれをローカルLLMに載せようとすると、重要な問題が出てくる。 ・それは、「分類作業を誰がやるのか」。
#AIタグ

一度独立に失敗して借金も経験した僕が、もう一度会社を作り「AI社員」にたどり着くまで

・前回の記事では、 「一人社長こそAIを雇ってみたら面白いんじゃないか?」 という話を書きました。 ・今はMac miniにOpenClawを入れて、自分専用のAI秘書を作っています。 ・ただ、僕はもともとIT業界にいたわけでも、AIの専門家だったわけでもありません。
#AIタグ

何を「選び続けるか」が信頼になる? / 誰でも作れる時代に、個人開発者は何を磨くのか|個人開発者になってみたい!#14

何を「選び続けるか」が信頼になる? / 誰でも作れる時代に、個人開発者は何を磨くのか|個人開発者になってみたい!#14
Zennの「大規模言語モデル」のフィード

技術コメントは英語、業務コメントは日本語にする

・以前、AI は文章を返す道具から、ファイルを読み、コードを編集し、テストを実行する作業基盤へ移りつつあると整理した[1]。この変化をコードコメントまで降ろして考えると、従来は人間の読みやすさだけで決めていた言語にも別の評価軸が加わる。 ・ただし、そこで「AI が読むからコメントをすべて英語にする」と結論づけるのは性急である。ソースコードのコメントには、実装上の理由を説明するものと、業務上の理由を説明するものがある。この二つは同じ言語で書く必要がない。 ・本稿では、日本語で業務要件を定義する開発を前提として、業務コメントは日本語、技術コメントは英語という分け方を提案する。要件定義書や基本設計書...
Zennの「機械学習」のフィード

魚種分類器の「実釣 top1 100%」は虚構だった ― 評価データ汚染の解剖と、釣るほど賢くなるループ

・個人開発のタナゴ釣り記録アプリ「KIRIBAKO」には、写真から魚種の候補を出すオンデバイス分類器(タナゴ族 5 種+よく混じる魚 3 種の 8 クラス)が入っています。前編では「敵はモデルよりドメインギャップだった」という話を書きました。 ・今回はその続編です。そして今回の敵は、モデルでもドメインギャップでもなく――自分の評価データそのものでした。 ・このシリーズ: スマホのカメラだけで魚を採寸する ― コイン基準とLiDAR【前編】 OpenCVで限界が来たコイン検出を、CoreMLの中心+半径回帰で強化する【後編】 魚種分類とドメインギャップ 本記事:評価データの汚染を暴き、実釣デ...
#AIタグ

人間とAIの関係ー不誠実な世界で他者と関わることの地獄と救いー

・1.AIにはできないこと AIの本質は、自発的な意志、生身の身体による実体験、そしてそれらに伴う結果への責任がないことにあります。AIは過去のデータを確率的に計算してそれらしい答えを出力しているだけであり、自らの行動に責任を負うことがありません。 ・AIは、間違った判断を下しても、心の底から謝罪することも、法的な責任を負うこともありません。 ・またAIは、この言葉にはこう返すというパターン認識を行っているだけで、相手の痛みを自分自身の感覚として共感することができません。従って、深い信頼関係の構築は不可能です。
#LLMタグ

制御の飽和をいかに防ぐか ── アンチワインドアップとT細胞に学ぶLLMのダイナミクス

・導入:位置決め(ゲイン調整)の「その先」にある問題 システム制御において、比例ゲイン(P成分)を適切に落とし、過剰反応によるオーバーシュートや発振を抑えることは「応答を所定の位置に収束させる(位置決め)」ための基本である。
#AIタグ

生成AIの是非を語りたければダ・ヴィンチと同じ経験値を積め

・命令形で言っておいて、僕もこれからそれを実践していく身なのですが、AI云々と倫理の領分で騒がしく議論するくらいなら「魂のこと」を念頭に含めて議論してほしい。 ・「ソクラテスの時代から人々は魂のことを意識して来世のことを思い描きながら生活していた」のです。それが現代ではどうか。なにか忙しいということを言い訳に、事の本題から逃げ続けている者たちばかりではないか? 続きをみる
Zennの「大規模言語モデル」のフィード

第1世代ループエンジニアリングを超えて

・第1世代ループエンジニアリングを超えて 速く回すことから、意味を保って回すことへ 「プロンプトを書くな、ループを書け」 この言葉は、AIエージェント時代の開発感覚をよく表している。 ・これまで人間は、AIに対して一回ずつ指示を書いていた。プロンプトを工夫し、出力を読み、足りないところを直し、また指示を出す。いわゆる Human in the Loop の形で、人間が作業の流れの中に入っていた。 ・しかし、エージェントがコードを書き、テストを実行し、失敗を読み取り、修正し、再びテストするようになると、人間が毎回プロンプトを書くこと自体がボトルネックになる。そこで設計対象は、個別のプロン...
#AIタグ

中国の最先端AIのNVIDIA製チップ依存 国産技術への移行を遅らせている要因は?

・中国の最先端AIは依然としてNVIDIA製チップで学習されている。国産技術への移行を遅らせている要因は何だろうか? What is delaying Chinese AI giants switching from Nvidia to local chips?High transition costs are keeping AI developers in China reliwww.scmp.com 続きをみる
ITmedia NEWS 最新記事一覧

中国発AIエージェント「Manus」、Metaから独立へ 中国政府が買収に反発、一部ユーザーデータは削除に

・AIエージェント「Manus」を提供するManusは8月11日(現地時間)、独立企業としての運営をまもなく再開すると発表した。米Metaからの分離に伴い、一部ユーザーのデータを23日から削除するとして、事前のバックアップを呼び掛けている。
ITmedia NEWS 最新記事一覧

提出した課題の口頭試問が必須に―─デンマーク、AI不正対策で高校課程に新ルール

・デンマーク教育省は、高校でのAI不正対策として、自宅で作成する試験課題に口頭試問を義務付ける方針を発表した。対象は2年間の高校教育課程「HF」の筆記課題「SSO」。筆記試験中のPC画面監視や、学校管理下での課題作成も促す。
#LLMタグ

答え合わせKimi K3 予測はどこまで当たったか

・7月27日、予告どおりのものが2つ公開された。Kimi K3の技術レポート(arXiv:2607.24653、本文だけで34ページ)と、Hugging Faceに置かれた約1.4TBのモデルウェイト一式である。 ・2週間前、私はこのモデルについて予測記事を書いた【前回記事リンク】。公開済みの関連論文(前世代K2のレポート、Kimi Linear論文、Attention Residuals論文)を足場に、未公表だったK3の中身を推定し、記事の最後に「7月27日にこのチェックリストで答え合わせをする」と書いていた。
#LLMタグ

日本語で聞くと、日本人の平均IQに収束するのか?

・AIを使い始めた頃、極論をひとつ投げてみたことがあります。 ・「日本語でプロンプトを書くなら、原理的に、出力は日本人の平均IQに収束するのではないか?」 続きをみる
#AIタグ

売っていた商品の目玉機能が、一度も動いていなかった話

・購入してくださった方から、機能が動かないという連絡をもらった。調べていくうちに、その機能が特定の人の環境だけで動いていないのではなく、売り出してから一度も、誰の環境でも動いていなかったことが分かった。今回はその顛末と、自分の作り方の何が間違っていたのかを書く。 ・最初は環境の問題だと思っていた 続きをみる
#LLMタグ

番外編4――LLMの「電子透かし」を、必読3論文で読み解く

・はじめに 2026年8月2日、EUのAI法(EU AI Act)第50条、いわゆる透明性義務の規定が適用開始となった。この条文には、AIが生成した文章・画像・音声・動画について、「機械可読な形式で印を付け、AI生成物として検出可能にしなければならない」という趣旨の義務が含まれている。ここ数週間、AI関係者のあいだで「電子透かし(watermark、ウォーターマーク)」という言葉を目にする機会が急に増えたのは、この期日が理由である。
Zennの「大規模言語モデル」のフィード

論文メモ:LLMの回答転換を予測するFusion-Fission式

・はじめに この記事は、LLM(大規模言語モデル)の回答が、会話の途中で望ましい内容から望ましくない内容へ転換する時点を、内部表現から予測しようとする論文の技術メモです。 ・論文タイトル:Fusion-fission forecasts when AI will shift to undesirable behavior 著者:Neil F. ・Johnson、Frank Yingjie Huo 発表年:2026年(arXiv v1、査読前原稿) 論文リンク:Fusion-fission forecasts when AI will shift to undesirable behav...