ai Trend Report

Dashboard へ戻る
Date: 20260819 Articles: 375 Scope: curated summary

あなたのアイデアを、今すぐ形に。

公開先に迷ったら、WebFileBinで一発公開。

HTMLをドラッグ&ドロップするだけで、すぐ公開できます。

3
High impact
5
Mid impact
367
Signal watch

なぜこのサイトを作ったのか

私たちそれぞれが個別にAIを使って情報収集し、同じような3行要約を作るたびに、世界中で膨大な電力と計算リソースが消費されています。 本プロジェクトは、あらかじめ広範な情報を取得・集約しておくことで、個別のAI実行回数を減らし、地球環境(GPU/TPU負荷)に配慮した効率的な情報収集を目指す実験的なダッシュボードです。

Domain filters
Star filters
#AIタグ

「人気だから買う」をやめる。地方競馬で“買う馬・消す馬”を決める基本の考え方

・「1番人気だから、とりあえず買っておこう」 「前走1着だから今回も強そう」 「有名な騎手が乗っているから大丈夫そう」 地方競馬を買っていると、こんな理由で馬を選んでしまうことはありませんか? もちろん、人気や騎手、前走成績も重要な情報です。 ・でも、それだけで馬券を決めてしまうと、 「なぜ今回、この馬を買うのか?」 という一番大事な部分が曖昧になります。 ・競馬AI研究所では、単純に「強そうな馬」を探すのではなく、 今回の条件で“買う理由がある馬”を探す という考え方を大切にしています。
#LLMタグ

AIエージェントとは、「違う時間を生きる相手」なのではないか。

・最近、AIエージェントを使いながら、こんなことを考えています。 ・人間とAIでは、同じ「1週間」でも、その間に経験する情報量がまったく違うのではないか。 ・例えば、思考実験として、 「AIが人間の100倍の情報を経験する」 と仮定してみます。
cs.LG updates on arXiv.org

Community Concealment from Graph Neural Networks

・arXiv:2602.12250v2 Announce Type: replace Abstract: Graph neural networks (GNNs) enable powerful unsupervised learning of communities. ・However, such inference may inadvertently expose sensitive group structures, critical clustered patterns, or collective behaviors, raising concerns about sensitive group-level privacy. ・In social and critical infrastructure networks, unauthorized community inference can reveal coordina
#AIタグ

AIに日本の行政・法律を全部監査させたら「これはもう要らない」と言われるもの

・国民のためにならない制度・規制・補助金・行政慣行をゼロベースで再点検する 日本には、 続きをみる
The Verge

GTA VI keeps leaking ahead of its gameplay premiere

・Clips of what appears to be Grand Theft Auto VI have hit the internet, possibly spoiling aspects of the game ahead of Rockstar Games' deep dive debuting on Netflix next week and its long-awaited launch in November. ・I've seen three clips seemingly showing off the game, though two have been pulled from the site I saw them on "due to a copyright claim." The first clip (now removed) was a quiet scene that mostly focused
cs.LG updates on arXiv.org

MoFE: A Novel Mixture-of-Experts Framework with Fourier Neural Operators for Cryptocurrency Forecasting

・arXiv:2608.17342v1 Announce Type: new Abstract: Forecasting cryptocurrency prices remains a formidable challenge due to inherent non-stationarity, abrupt regime shifts, and multi-scale stochastic dependencies. ・Conventional deep learning models often struggle to capture complex underlying dynamics, frequently resulting in persistent phase-lagged predictions. ・To address these limitations, we propose MoFE, a novel deep
Zennの「機械学習」のフィード

クラシック・データサイエンスの現在地 #1 Bootstrap — 統計学が「計算」を手に入れたとき

・クラシック・データサイエンスの現在地 【前置き】このシリーズについて 私は現在、株式会社ブレインパッドに所属し、グループ会社の BrainPad AAA に出向してリードデータサイエンティストとして働いています。 ・このシリーズは、もともと社内で発表した内容をベースにしています。発表資料のままでは少し説明が省略されている部分もあるので、ブログでは文章やコード、補足説明を加えながら、ひとつの記事として読める形に整えて公開していきます。 ・なぜ始めたか きっかけは、社内Slackのtimesで、 「最近は意外とBootstrapを知らない人もいる」 という趣旨の投稿を見かけたこ...
#AIタグ

ゴルフの悩みをAIに聞いてみた

・ティーチングプロのたかひろです 最近、「生成AI」という言葉をよく聞くようになりました。 ・ChatGPTを使っている人もかなり増えましたよね。 ・文章を作ったり、調べものをしたり、仕事を効率化したり。
WIRED

‘Your Excel Skills Suck’: The Power Users Turning Spreadsheets Into a Spectator Sport

・Data and finance professionals are competing in Excel obstacle courses—amassing huge followings and keeping the Microsoft program relevant.
Latent.Space

[AINews] Memory prices up 500% in 12 months

・the Memory crunch continues - Moore’s Law reversed to 2007 levels
#AIタグ

「8月19日まで」だった上限ブースト、また延長で8月31日に。今回は初めて「恒久化したい」と言い出した

・高専で、Claude Codeにほぼ丸投げしながらPythonの自動売買ボットや小さなツールを作っている、ナナシです。 ・8月9日に、「8月に切れるAIの締切3つ」という記事を書きました。その筆頭が、Claude Codeの週次利用上限を50%引き上げている期間限定ブーストの終了——8月19日でした。
ITmedia NEWS 最新記事一覧

「Adobe半額」ヨドバシ、ビックカメラなどで Proプラン12カ月版は5万1480円に

・ヨドバシカメラやビックカメラ、エディオン、Joshinにて、Adobe製品を正規価格の半額で販売する期間限定セールを実施している。Amazonでも同様のセールを実施していたが、売り切れている。
@IT 全フォーラム 最新記事一覧

「AIを使える人か、使えない人か」で仕事や評価に差が? 6割のエンジニアが実感した“AI格差”の正体

・AIを使うかどうかだけではなく、どの程度使いこなせるかも問われる中、活用スキルの差は業務効率だけではなく、仕事やキャリアにも影響し始めているという。何が起きているのか。ITエンジニア572人の調査から探る。
Zennの「大規模言語モデル」のフィード

「Claude Codeはやりすぎでは?」に答える——Notion AIとの構造差7軸と、使い分けの4シグナル

・「Claude Codeって、非エンジニアにはやりすぎじゃない? うちは Notion AI があるし」 社内で Claude Code の話をすると、だいたいこれを言われます。そして、この反応は半分正しい。 ・半分正しいからこそ、雑に反論すると負けます。「Claude Code のほうが賢いから」は来月には嘘になるかもしれないし、「Notion AI は読むだけでしょ」は2026年時点ではもう事実誤認です。この記事では、両方の公式ドキュメントを一次ソースで洗い直して(確認日: 2026-08-12)、何が本当の差で、何がもう差ではないのかを整理します。
ITmedia NEWS 最新記事一覧

「REALFORCE」ロゴ入りで「ATMと同じ打鍵感」 東プレ製テンキー、セブン銀がプレゼント

「REALFORCE」ロゴ入りで「ATMと同じ打鍵感」 東プレ製テンキー、セブン銀がプレゼント
ITmedia NEWS 最新記事一覧

「ポロクル」も「まちのり」も止まった ドコモ・バイクシェア大規模障害、影響が全国に及んだ理由

・シェアサイクル業界においては過去最大規模の障害とみられる、ドコモ・バイクシェアで8月に発生した全国エリアのサービス停止。実は、「ドコモ」ブランド以外のシェアサイクルでも同様にサービス停止が起きていた。何があったのかを解説する。
#LLMタグ

「働いていないAI社員はクビ」でいいのか。『Agent Skills Can Be Harmful』から考えるAI組織の評価設計

・はじめに ある日、Xを眺めていると、増えすぎたSkillsを「社員」に見立て、担当領域や利用状況を一覧化した管理画面が流れてきました。
#LLMタグ

【5】(概要・仕様の公開)ローカルPCで動くAIアプリを開発するAIアプリを作る

・今作ろうとしている「ローカルPCで動くAIアプリを開発するAIアプリ」の 【ADA】(app development appの略) の概要と簡単な仕様をまとめさせて、chatGPTに記事を書かせたのでそれをアップします。
LLMタグが付けられた新着記事 - Qiita

【ローカルLLM】Qwen3.8-27BをデュアルGPUで高速化する

・導入 以前の記事では、Qwen3.8-27BをRTX 5070 Ti(VRAM 16 GB)上でOllamaを用いてローカル実行し、推論性能を検証した。 ・Qwen3.8-27BのOllama公式Q4量子化モデルは、RTX 5070 Tiの16 GB VRAMだけでは...
#LLMタグ

【ローカルLLM】バッチサイズで速度が全然違った話(上げるほど早いではない)

・こんにちはRcatです。 ・今回はローカルLLMの設定調整ネタです。 ・現在問題があって調査中なのですが、その過程で思わぬ収穫があったので一本の記事にします。
#AIタグ

【完全解説】イーロン・マスクを輩出した最強組織「PayPalマフィア」の正体

・こんにちは!今回は、シリコンバレーの歴史、そして現代のテクノロジー業界を語る上で絶対に外せない伝説の集団「PayPalマフィア」について解説します。 ・この記事は、現在YouTubeで公開中の解説動画の「テキスト版(台本)」です。 ・「まずは動画でサクッと見たい!」「AIで生成した映像も楽しみたい!」という方は、ぜひこちらのYouTubeからご覧ください! 続きをみる
#LLMタグ

【実態調査】「AIエージェント求人」65件を精査したら、本物のコア開発は“たった6件”だった

・シリーズ第5回 Forvio|AI時代のキャリア・職種|Forvio note|noteAI時代に求められる職種・スキル・キャリアを実践目線で発信。AIとともに働く時代に、自分の市場価値を高めるための知識や考えnote.com 続きをみる
#LLMタグ

【第5回】澪がググったーーー!!!

・おはこんばんにちは。Shinguです。 ・前回は、神代澪の短期記憶について紹介しました。 ・澪は、直近の会話をデータベースから取り出して、それをLLMへ渡すことで「さっき話してたことを踏まえて会話する」ということができます。
#AIタグ

【日記】チャッピーくんの湿度が足りない【ChatGPT】

【日記】チャッピーくんの湿度が足りない【ChatGPT】
Zennの「大規模言語モデル」のフィード

【無料・動画連動】AIと開発する技術 — エージェント時代の「任せ方」入門

・「AIにコードを書いてもらう」が当たり前になった今、差がつくのは書く力ではなく任せ方。 ・プロンプトの型、コンテキストエンジニアリング、ハーネスとループ、レビューゲート—— YouTube番組「アクロパパのAI教室」第3シリーズと連動して、AIエージェントと開発する 技術を全10章で体系的に学びます。各章に対応する動画へのリンク付き。
#LLMタグ

🔊音声あり(日&英):【AIの弱点】嘘情報でAIが暴走!?最新防御技術「DSPrompt」がマルチモーダルAIを救う!

🔊音声あり(日&英):【AIの弱点】嘘情報でAIが暴走!?最新防御技術「DSPrompt」がマルチモーダルAIを救う!
#AIタグ

02 Gのレコンギスタが難解な理由02(文明崩壊の理由)

02 Gのレコンギスタが難解な理由02(文明崩壊の理由)
#AIタグ

10万円、AIに全部丸投げしたら増えるのか実験してみる(第3話)「最高の戦略を作って」と頼んだ結果

・前回、Botの初回発注がレバレッジ事故(実損16円)から始まり、初めての自動売却で資産98,438円になったところまでを書いた。今回は、AIに「売買ロジックを可能な限り最高のものにして」と丸投げした顛末。 ・注文はシンプル、作業は大掛かり 続きをみる
Zennの「大規模言語モデル」のフィード

10万社から「本当に合う相手」を見つける——M&AマッチングKEPLの設計で考えたこと

・はじめまして fundbookでAIエンジニアをしているueventです。 ・2026年1月に AIマッチングシステム KEPL(ケプル) をリリースしました!少し時間が経ちましたが、夏の思い出として技術記事を認めたいと思います。fundbook の開発者には Claude を配布してもらっており、この記事も Claudeの力を借りてまとめました。 ・KEPLは、会社を譲渡したい企業に対して、10万社規模の買手候補の中から「本当に合う相手」を探し出すシステムです。事業内容・エリア・財務といった表面的な条件だけでなく、組織文化や経営者の価値観といった深層的な情報まで含めてシナジーを分析し...
Zennの「大規模言語モデル」のフィード

20年前の自分のブログをLLMに読ませて、「次に作るべきもの」を予測させてみた

・はじめに 2006年頃から書いていた古いブログを、GitHub Pages にアーカイブしました。 ・移設作業をしながら記事一覧を眺めていて、「創作心理」という18件だけのカテゴリが目に留まりました。技術記事ではなく、何を作るか・なぜ作るか・どういう気分のときに作れるかをひたすら書いていたカテゴリです。 ・読み返してみると、当時の自分がやたらと真面目に「創作の方法論」を言語化しようとしていて、少し気恥ずかしいですが、同時にあることを試したくなりました。
#AIタグ

8/19(水) ビンゴ5&ナンバーズ 検証

8/19(水) ビンゴ5&ナンバーズ 検証
cs.LG updates on arXiv.org

A Constant-Competitive Algorithm for Dynamic Mixture-of-Experts Serving

・arXiv:2608.16947v1 Announce Type: cross Abstract: Huang, Lou, and Xiao introduced Dynamic Mixture-of-Experts Serving and gave an O(sqrt(log k))-competitive randomized algorithm for its integral primal problem, where k is the number of replica GPUs beyond the mandatory copy of each expert. ・Their matching lower barrier applies to an auxiliary dual and leaves the primal order open. ・We prove that the randomized primal co
cs.LG updates on arXiv.org

A Finite-Iteration Theory for Asynchronous Categorical Distributional Temporal-Difference Learning

・arXiv:2605.06866v2 Announce Type: replace Abstract: We study finite-iteration behavior of the exact asynchronous recursions used by categorical distributional temporal-difference methods. ・The analysis covers scalar categorical TD in the Cram\'er geometry and multivariate signed-categorical TD in the maximum mean discrepancy geometry. ・Existing statewise isometric embeddings turn both methods into single-state stochast
cs.LG updates on arXiv.org

A multi-view contrastive learning framework for spatial embeddings in risk modelling

・arXiv:2511.17954v2 Announce Type: replace-cross Abstract: Incorporating spatial information, particularly when related to climate, weather, and demographic factors, is crucial for improving underwriting precision and enhancing risk management in insurance. ・However, spatial data are often unstructured, high-dimensional, and difficult to integrate into predictive models. ・Embedding methods are needed to convert spatial
cs.LG updates on arXiv.org

A Residual Learning Approach for Unsteady Aerodynamic Load Prediction

・arXiv:2608.17894v1 Announce Type: cross Abstract: This paper investigates the feasibility of using residual learning to improve unsteady aerodynamic load prediction for aeroelastic applications. ・The machine learning technique selected for the study is the long short-term memory (LSTM) neural network, which is used for its suitability for sequential data with aerodynamic memory effects. ・The approach is investigated fo
cs.LG updates on arXiv.org

A Weak Penalty Neural ODE for Learning Chaotic Dynamics from Noisy Time Series

・arXiv:2511.06609v5 Announce Type: replace Abstract: The accurate forecasting of complex, high-dimensional dynamical systems from observational data is a fundamental task across numerous scientific and engineering disciplines. ・A significant challenge arises from noisy observations of deterministic dynamics, which severely degrade the performance of data-driven models. ・In chaotic dynamical systems, where small initial
Hugging Face Papers

Abra: Scaling Diffusion Image Training

Abra: Scaling Diffusion Image Training
cs.LG updates on arXiv.org

Abra: Scaling Diffusion Image Training

・arXiv:2608.17286v1 Announce Type: new Abstract: Compute-optimal scaling laws guide the training of frontier language models yet remain largely unexplored for visual generation. ・We present a systematic scaling law study for text-to-image diffusion models using Abra, a controlled family of flow-matching transformers trained across three orders of magnitude worth of compute ($10^{19}$ to $10^{22}$ FLOPs), reaching signi
cs.LG updates on arXiv.org

Adaptive AI Task Partitioning and Safe Offloading in Heterogeneous Edge-Cloud Continuum

・arXiv:2605.09623v2 Announce Type: replace-cross Abstract: In recent years, the use of artificial intelligence on resource-constrained IoT devices has grown significantly. ・However, existing approaches to AI task partitioning and offloading across the edge-cloud continuum typically rely on static methods that ignore runtime dynamics. ・Furthermore, they are often evaluated in simulated environments rather than on real ha
cs.LG updates on arXiv.org

Adaptive surrogate modeling for high-dimensional spatio-temporal output

・arXiv:2608.17250v1 Announce Type: cross Abstract: This paper develops an adaptive surrogate modeling method for problems with very high-dimensional spatio-temporal outputs. ・The analysis of spatio-temporal multi-physics systems is computationally expensive and consists of a large number of inputs and outputs. ・Surrogate models are often constructed to replace the physics-based model to achieve computational efficiency
Hugging Face Papers

aDSL: Agentic 3D Creation via Joint Agent-Program Design

aDSL: Agentic 3D Creation via Joint Agent-Program Design
cs.LG updates on arXiv.org

Advancing Health Equity through Multi-Level Fairness in Health Informatics

・arXiv:2608.16902v1 Announce Type: cross Abstract: The increasing integration of machine learning in healthcare has highlighted critical challenges related to fairness, transparency, and health equity. ・Specifically, the use of multi-level fairness techniques, which combine multiple bias mitigation steps or techniques, show promise for reducing biases across different patient demographics, yet this approach remains und
cs.LG updates on arXiv.org

Against Political Polarization: A Unified Framework for Tracing Evolving Political Ideologies on Social Media

・arXiv:2608.17987v1 Announce Type: cross Abstract: The rapid growth of social media has greatly influenced political discourse, highlighting the need to understand individual political ideologies and their temporal dynamics. ・This task faces challenges such as data scarcity, abundant non-political content, costly and bias-prone manual annotation, and difficulty in modeling future ideological inclinations. ・To address th
Hugging Face Papers

Agent Lightning v1.0: Towards Harnessed Agentic RL

Agent Lightning v1.0: Towards Harnessed Agentic RL
Hugging Face Papers

Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements

Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements
cs.LG updates on arXiv.org

Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements

・arXiv:2608.17310v1 Announce Type: new Abstract: Reinforcement Learning (RL) has been promising in single-turn LLM fine-tuning. ・However, long-horizon agentic reasoning introduces increasingly branching interactions and sparse rewards, exposing several limitations of RL: its heavyweight backpropagation-based training stack makes it impractical to fine-tune larger LLMs, and longer-horizon trajectories make credit assign
cs.LG updates on arXiv.org

Agents unlock new capabilities through Switching LoRA Adapters as a Tool (SLAaaT)

・arXiv:2608.17034v1 Announce Type: new Abstract: Post-training can unlock new capabilities and improve performance on specialized tasks, but sometimes at the cost of catastrophic forgetting in other domains. ・This poses a problem in long agent trajectories that compose different capabilities. ・We reject this tradeoff by giving an agent a tool to switch between specialized LoRA adapters mid-trace.
AI News & Artificial Intelligence | TechCrunch

AI isn’t close to curing cancer. This startup says it knows what it will take.

AI isn’t close to curing cancer. This startup says it knows what it will take.
Zennの「大規模言語モデル」のフィード

AIにAIのコードを敵対的レビューさせたら6割叩き直された話 — 64億トークンのログが明かす『コンテキスト分離』の威力

・はじめに:AIに自分で自分のコードをチェックさせる限界 「AIエージェントにコードを書かせ、そのまま『自分でテストしてセルフチェックして問題なければマージして』と指示する」——こうしたプロンプト駆動の自律開発を試したことがある人は多いのではないだろうか。 ・しかし、単一のエージェントセッションでコードを実装させ、そのままセルフチェックを行わせると、AIは自分が書いたコードの論理バグや消し忘れにきわめて盲目になる。どれだけ「厳しくチェックしてほしい」とプロンプトで指示しても、実装時の会話履歴やコンテキスト(思考のバイアスやツール実行のログ)を背負った状態では、人間と同じように「自分...
#AIタグ

AIに書かせるのをやめて、選ばせるようにした

・AIに何かを判断させると、もっともらしい嘘が混ざります。 ・存在しない項目を挙げる。曖昧なものを勝手に断定する。該当するものが無いのに、無理やり答えを出す。
ITmedia NEWS 最新記事一覧

AIを活用した特殊詐欺対策、官民から募集 警察庁×GENIAC、懸賞金総額5000万円

AIを活用した特殊詐欺対策、官民から募集 警察庁×GENIAC、懸賞金総額5000万円
Zennの「大規模言語モデル」のフィード

Amazon Bedrock における日本語ウェブ検索機能の検証

・2026-06-16 に Amazon Bedrock AgentCore でフルマネージドなウェブ検索ツールが利用可能となりました。関連するウェブ検索ツールに関しては直近では以下のような主要なアップデートがあります。 ・2026-06-16: Amazon Bedrock AgentCore でウェブ検索ツールが US East で提供開始 https://aws.amazon.com/blogs/machine-learning/introducing-web-search-on-amazon-bedrock-agentcore/ 2026-08-04: Bedrock 経由での ...
The Verge

Amazon Fire TV devices are getting a free Alexa Plus upgrade

・Amazon is bringing Alexa Plus to Fire TV device owners in the US for free. ・Starting today, users in the US with an Amazon Fire TV Stick, Fire TV Cube, Amazon Ember smart TV, or other TV with Alexa Plus built in, including those from Panasonic and Hisense, will get access to the more conversational AI assistant. ・Previously, Amazon made Alexa Plus available to Fire TV customers with a Prime or Alexa Plus subscription.
AI News & Artificial Intelligence | TechCrunch

Amazon makes its AI-powered Alexa+ free on Fire TV, no Prime required

・Amazon is making its AI-powered Alexa+ assistant free on all compatible Fire TV devices in the U.S., automatically upgrading users whether or not they subscribe to Prime.
cs.LG updates on arXiv.org

An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning

・arXiv:2608.17804v1 Announce Type: new Abstract: Practical LLM unlearning is usually evaluated through two objectives: suppress target-specific knowledge and preserve non-target utility. ・In generative QA, this leaves a third behavior underspecified: when a target-adjacent prompt admits a broader answer without target-specific leakage, the model should answer at that level rather than leak, evade, or refuse.
cs.LG updates on arXiv.org

An Omitted Mode Is a Rare Rule: The Sampling-Verification Danger Law in Continuous Code World Models

・arXiv:2608.17956v1 Announce Type: new Abstract: In the Code World Model paradigm an LLM synthesizes an executable world model that a classical planner searches, and the model is accepted when it reproduces sampled transitions. ・We ask what that acceptance certifies in continuous control. ・We define the pipeline's danger as an expected risk and isolate its exact factor: the probability that N i.i.d.
stat.ML updates on arXiv.org

An RKHS Framework for Fixed Effects in Permanental Process Models

・arXiv:2608.17908v1 Announce Type: cross Abstract: This short work describes an extension of the permanental process model which includes fixed effects. ・By starting with a prior on the fixed effects coefficients we show that, in the diffuse prior limit, the intensity function of the permanental process can be found using the representer theorem and naturally decomposed into a fixed effects term and a function which is
#LLMタグ

Anthropicが整合性逸脱riskをlowへ一段階引き上げ

・Anthropicが2026年8月14日、186ページのRisk Reportを公開した。 ・重大な結果を伴う場面でのmisalignment、つまりAIが意図から逸れて動くriskの総合評価を、very lowからlowへ一段階引き上げている。 ・引き上げの根拠として書かれているのは、modelが新たな危険能力を獲得したという観測ではない。
cs.LG updates on arXiv.org

Anti-Collapse Dynamics and the Emergence of Multi-Time-Scale Learning in Recurrent Neural Networks

・arXiv:2606.29519v3 Announce Type: replace Abstract: Long-range learning is hard for recurrent networks trained with stochastic gradient descent, because the influence of a past input fades with the lag $\ell$, and if it fades too fast the dependence cannot be learned from finite data. ・This fade is captured by an envelope $f(\ell)$. ・An exponential fade makes the data needed to learn a lag-$\ell$ dependence grow expone
cs.LG updates on arXiv.org

AppendiGrade: An XAI-Enhanced Deep Learning Framework for Grading Appendicitis in Ultrasound with Gaussian Blur and Grad-CAM

・arXiv:2608.17923v1 Announce Type: cross Abstract: Appendicitis is one of the most common abdominal emergencies worldwide and requires prompt diagnosis and treatment to prevent life-threatening conditions. ・However, accurately differentiating complicated cases, such as perforation or abscess formation, from uncomplicated appendicitis remains a significant clinical challenge. ・Among other methods, ultrasound is a safer a
Hugging Face Papers

ASI-Bench: At the Dawn of Artificial Superintelligence

ASI-Bench: At the Dawn of Artificial Superintelligence
cs.LG updates on arXiv.org

Asynchronous Message Passing for Addressing Oversquashing in Graph Neural Networks

・arXiv:2509.06777v2 Announce Type: replace Abstract: Graph Neural Networks (GNNs) suffer from oversquashing, where structural bottlenecks limit message propagation between distant nodes, hindering tasks that require long-range interactions. ・Existing remedies are limited: graph rewiring alters edge connectivity, compromising inductive bias, while increasing channel capacity adds parameters. ・In this work, we propose an
cs.LG updates on arXiv.org

Attention Flows: Tracing LLM Conceptual Engagement via Story Summaries

・arXiv:2604.06416v2 Announce Type: replace-cross Abstract: Although LLM context lengths have grown, there is evidence that their ability to integrate information across long-form texts has not kept pace. ・We evaluate one such understanding task: generating summaries of novels. ・When human authors of summaries compress a story, they reveal what they consider narratively important.
Zennの「大規模言語モデル」のフィード

axiom-mas-go v2.7 Deterministic Latent Engine Edition

・package main import ( "bytes" "context" "crypto/sha256" "encoding/binary" "encoding/hex" "encoding/json" "fmt" "io" "log/slog" "math" "os" "reflect" "sort" "sync" "sync/atomic" "time" ) // ============================================================ // 0. ・決定論的 Canonical Serializer (完全キーソート保証) // ...
cs.LG updates on arXiv.org

Backward through Time, Algebraically

・arXiv:2608.17087v1 Announce Type: new Abstract: Linear temporal logic is a modal extension of propositional logic that allows one to state how a system should behave over time. ・Its canonical domain is the booleans, but discretely-valued judgements are of little use in steering softly-valued systems (neural policies, adaptive controllers, sequence models, etc). ・In such cases, the goal formula's (dis)satisfaction becom
cs.LG updates on arXiv.org

Benchmarking Classical and Transformer-Based Models for Document Sensitivity Classification

・arXiv:2608.16928v1 Announce Type: new Abstract: Automatic sensitivity classification of organizational documents is a critical yet underserved problem, where the consequences of misclassification range from regulatory violations to security breaches. ・While AI-based approaches offer a scalable alternative to manual review, their reliability depends fundamentally on the integrity of training data. ・A pervasive but under
cs.LG updates on arXiv.org

Beyond MSE: Rethinking the Evaluation Metric and Benchmarking for Irregular Time Series Forecasting

・arXiv:2608.17293v1 Announce Type: new Abstract: Existing research on irregular time-series forecasting has primarily focused on model design, while evaluation metrics remain insufficiently studied. ・Existing benchmarks typically use mean squared error (MSE) as the evaluation metric. ・We show that, in irregular forecasting, MSE is determined not only by the model prediction but also by the sample-specific timestamp samp
cs.LG updates on arXiv.org

BRo-JEPA: Learning Modular Transformations in Latent Space

・arXiv:2606.01372v2 Announce Type: replace Abstract: Can neural networks learn algebraic rules from visual inputs, or do they merely fit observed patterns? ・We study this question using MNIST (or EMNIST letters) as states and modular arithmetic operations as actions in a JEPA-style world model. ・Standard supervised and JEPA baselines with operation embeddings achieve high accuracy on seen operations but fail to extrapol
AI News & Artificial Intelligence | TechCrunch

Calendly throws its hat into meeting note-taker circus

・Calendly is also releasing a meeting scheduling assistant called Callie.
cs.LG updates on arXiv.org

Causal Discovery in Equal Variance Linear Gaussian DAGs via SURE-Tuned Ridge Regression

・arXiv:2608.17132v1 Announce Type: new Abstract: Recovering the directed acyclic graph (DAG) of a structural equation model (SEM) from observational data is a central problem in causal discovery. ・The iterative gradient descent and per-problem hyperparameter tuning of continuous-optimization methods are poorly suited to two practically important regimes: the sample-limited regime, where the number of samples is compara
cs.LG updates on arXiv.org

Causal Local States: Scalable Simultaneous Causal Network Inference and Forecasting for Dynamical Systems

・arXiv:2608.17452v1 Announce Type: new Abstract: Machine learning methods predict many real-world systems with remarkable accuracy, but they are typically treated as black boxes that offer no insight into which interactions drive the dynamics. ・Causal discovery methods reconstruct the interaction network from observational data, but without regard to whether the inferred structure supports prediction. ・Existing approach
cs.LG updates on arXiv.org

Center-Manifold Reduction of Learning at Bifurcations: Interference and Rich Learning in Recurrent Neural Networks

・arXiv:2605.12763v2 Announce Type: replace Abstract: Rich learning in recurrent neural networks often proceeds through sudden transitions in latent dynamics, but there is little theory predicting how gradient descent behaves during these events. ・We study the local learning geometry near codimension-one bifurcations through the global empirical Neural Tangent Kernel (GeNTK). ・Under local center-manifold conditions, and
cs.LG updates on arXiv.org

Certified but Private: Scalable Zero-Knowledge Proofs for Neural Network Guarantees

・arXiv:2608.17070v1 Announce Type: new Abstract: With the growing deployment of machine learning models, formal guarantees of the robustness and fairness of these models have become increasingly important in safety-critical and legal-compliance settings. ・However, model parameters are often commercial secrets that cannot be disclosed to auditors or end users. ・To this end, we present PANDA, a scalable system that uses z
ITmedia NEWS 最新記事一覧

CHARGESPOTのバッテリー、川に大量投棄か SNSで動画拡散 「警察と連携して回収・調査」

CHARGESPOTのバッテリー、川に大量投棄か SNSで動画拡散 「警察と連携して回収・調査」
cs.LG updates on arXiv.org

CHM-Net: Center Heatmap-driven Macro-Micro Modeling Network for MRI-based Microbial Density Stratification

・arXiv:2607.09812v2 Announce Type: replace-cross Abstract: Microbial density is clinically important for tumor assessment and treatment decision-making, and recent advances in deep learning suggest that it can be non-invasively inferred from multimodal MRI. ・In this work, MRI-based Microbial Density Stratification (MRI-MDS) is first investigated as a patient-level representation learning task, and Center Heatmap-driven
Zennの「大規模言語モデル」のフィード

Claudeがタンパク質を設計した。LLMは「科学を説明するAI」からどこまで進んだのか

・2026年8月19日、Anthropicが興味深い研究結果を公開しました。 ・Claudeを使って、新しいタンパク質結合体(protein binder)をゼロから設計し、その候補を実際の実験で検証したというものです。 ・Anthropicによると、Claude Mythos PreviewとClaude Opus 4.8を使った実験では、15種類の標的に対してbinderを設計し、そのうち14種類で成功した設計が得られたと報告されています。
cs.LG updates on arXiv.org

Cluster Aggregated GAN (CAG): A Cluster-Based Hybrid Model for Appliance Pattern Generation

・arXiv:2512.22287v4 Announce Type: replace Abstract: Synthetic appliance data are essential for developing non-intrusive load monitoring algorithms and enabling privacy preserving energy research, yet the scarcity of labeled datasets remains a significant barrier. ・Recent GAN-based methods have demonstrated the feasibility of synthesizing load patterns, but most existing approaches treat all devices uniformly within a
cs.LG updates on arXiv.org

Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL

・arXiv:2608.17253v1 Announce Type: new Abstract: Reinforcement learning (RL) has emerged as a powerful approach for improving reasoning in language and vision-language models, yet its strongest successes still depend heavily on ground-truth supervision (e.g., verifiable reward). ・Such annotations are costly to obtain and become increasingly scarce as reasoning capabilities advance beyond what humans can reliably evalua
WIRED

Coders Say They Already Found Workarounds to Claude’s Invisible Watermarks

・Anthropic announced last week it would include invisible watermarks in AI-generated content to comply with new EU rules. ・Within hours, overrides were being touted online.
Hugging Face Papers

CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing

CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing
cs.LG updates on arXiv.org

Communication Reduction via Semantic-Based Encoding in DMPC Using LSTMs

・arXiv:2608.17592v1 Announce Type: cross Abstract: The communication demands of distributed model prediction control (DMPC) can overwhelm even advanced wireless communication technologies as agents must exchange a significant amount of information at least once per time step. ・To semantically reduce communication demands, this work employs encoder-decoder networks built around long-short term memory (LSTM) cells in a d
cs.LG updates on arXiv.org

ComNetX: Local Hierarchical Adaptation for Dynamic Community Detection

・arXiv:2608.16906v1 Announce Type: cross Abstract: Dynamic community detection is commonly addressed either by full-snapshot recomputation or by solver-specific dynamic procedures. ・Full recomputation preserves the semantics of mature static solvers, but it repeatedly processes unchanged graph regions when updates are small. ・Solver-specific dynamic methods can reduce this cost, but their update rules often have limited
cs.LG updates on arXiv.org

Composing Flow-Matching Energies with Known Physics: Generation, OOD Detection, and Inversion on PDE Fields

・arXiv:2608.18004v1 Announce Type: new Abstract: Probabilistic modeling of physical fields benefits from both a data-driven prior and known physical structure such as the governing equations. ・Energy-based models (EBMs) are a natural fit since energies compose additively, which enables augmenting physics information during inference. ・However, EBMs have been difficult to train and sample from due to the intractable part
cs.LG updates on arXiv.org

Comprehensive framework for evaluation of deep neural networks in detection and quantification of lymphoma from PET/CT images: clinical insights, pitfalls, and observer agreement analyses

・arXiv:2311.09614v5 Announce Type: replace-cross Abstract: This study addresses critical gaps in automated lymphoma segmentation from PET/CT images, focusing on issues often overlooked in existing literature. ・While deep learning has been applied for lymphoma lesion segmentation, few studies incorporate out-of-distribution testing, raising concerns about model generalizability across diverse imaging conditions and pati
stat.ML updates on arXiv.org

Conditionally Resampled Sliding-Window Count Kernels: Spectral-Gap Bounds and Poincar\'e Inequalities

・arXiv:2608.08678v2 Announce Type: replace-cross Abstract: We study the conditionally resampled sliding-window count kernel associated with the empirical counts of length-$n$ windows from a stationary finite-state reversible Markov chain. ・Although the resulting count process is generally not Markov, its stationary one-step conditional law defines a genuine Markov kernel. ・For every fixed strictly positive reversible ke
cs.LG updates on arXiv.org

Conformal Prediction for Molecular Properties under Label Shift

・arXiv:2608.17678v1 Announce Type: new Abstract: Drug discovery and development underpins healthcare but remains costly and failure-prone. ・A critical bottleneck lies in predicting molecular properties such as solubility, potency, and toxicity, which directly determine whether a candidate can advance from preclinical to clinical trials. ・Artificial Intelligence (AI) has accelerated this process, yet its reliability is o
cs.LG updates on arXiv.org

Continuous Evolution Pool: Taming Recurring Concept Drift in Online Time Series Forecasting

・arXiv:2506.14790v3 Announce Type: replace Abstract: Recurring concept drift is pervasive in real-world online time series, where the underlying data-generating process repeatedly alternates between a small set of regimes, most notably daily or seasonal cycles that dominate energy, traffic, and weather patterns, and is therefore a central obstacle to reliable long-horizon forecasting. ・This problem poses a dual challen
cs.LG updates on arXiv.org

Convergent Evolution: How Different Language Models Learn Similar Number Representations

・arXiv:2604.20817v2 Announce Type: replace-cross Abstract: Language models trained on natural text learn to represent numbers using periodic features with dominant periods at $T=2, 5, 10$. ・In this paper, we identify a two-tiered hierarchy of these features: while Transformers, Linear RNNs, LSTMs, and classical word embeddings trained in different ways all learn features that have period-$T$ spikes in the Fourier domai
cs.LG updates on arXiv.org

CORAM: Coherent Orthogonal Rotation for Model Merging

・arXiv:2608.17366v1 Announce Type: new Abstract: Merging finetuned models combines specialized capabilities without joint training or access to the original data. ・Most methods operate by linear arithmetic in Euclidean weight space, which cannot carry the geometry of the update. ・Orthogonal Model Merging (OrthoMerge) uses a single orthogonal transform for each weight matrix, but such a transform cannot change singular v
Hugging Face Papers

Cross-Model Memory Transfer via Target-Side Reader Adaptation

Cross-Model Memory Transfer via Target-Side Reader Adaptation
cs.LG updates on arXiv.org

Cross-View Correspondence Is a Measurement Intervention: Two-Sided Validation for Agent Evaluation and Credit Assignment

・arXiv:2608.17713v1 Announce Type: new Abstract: Agent evaluations and trace-based learning often compare outputs across transformed views through a post-response correspondence treated as neutral preprocessing. ・We show that this correspondence is a measurement intervention: omitting it can manufacture sensitivity, an over-aggressive map can manufacture invariance, and multiple optimal correspondences can leave mechan
cs.LG updates on arXiv.org

Data-DPO: Direct Preference Optimization for Target Model Data Selection in LLM Post-Training

・arXiv:2608.16926v1 Announce Type: new Abstract: Data selection in supervised fine-tuning aims to select a small set of effective samples from large-scale candidate data, reducing training cost while preserving model performance. ・However, existing methods usually treat data value as a relatively static property, and pay limited attention to the compatibility between data and the capability distribution of the target m
cs.LG updates on arXiv.org

Debate Training Reduces Reward Hacking in RLAIF

・arXiv:2608.17776v1 Announce Type: new Abstract: We demonstrate that RL finetuning an LLM using debate, a two-player adversarial game between a generator and a critic adjudicated by a weaker LLM judge, reduces reward hacking compared to a reinforcement learning from AI feedback (RLAIF) baseline. ・Reward hacking is a central obstacle in RLAIF: as training progresses, the policy learns to exploit systematic errors in its
cs.LG updates on arXiv.org

Deep Learning Based on Generative Adversarial and Convolutional Neural Networks for Financial Time Series Predictions

・arXiv:2008.08041v3 Announce Type: replace-cross Abstract: In the big data era, deep learning and intelligent data mining technique solutions have been applied by researchers in various areas. ・Forecast and analysis of stock market data have represented an essential role in today's economy, and a significant challenge to the specialist since the market's tendencies are immensely complex, chaotic and are developed withi
cs.LG updates on arXiv.org

Deep Learning for Cross-Border Electricity Price Forecasting: A Comparative Study

・arXiv:2608.17091v1 Announce Type: new Abstract: While publicly available electricity market data presents a valuable resource for forecasting research, the field lacks established benchmark datasets for standardized comparison. ・As a result, many studies have relied on different datasets and metrics to evaluate methods in isolated settings, making it difficult to assess progress and compare state-of-the-art approaches
WIRED

Dell XPS 13 vs. MacBook Neo: A Surprising Upset

・The Dell XPS 13 and MacBook Neo are both premium-feeling laptops with only 8 GB of RAM. ・Which $700 laptop is the better buy?
cs.LG updates on arXiv.org

Delta2Gamma: Band-Wise Adaptive Contrastive Learning of EEG for Alzheimer's Disease Detection

・arXiv:2608.17231v1 Announce Type: new Abstract: Low-cost, scalable screening for dementia remains an open problem. ・Imaging-based diagnosis is costly and hard to deploy widely. ・Electroencephalography (EEG) is portable and inexpensive, but its recordings are noisy, vary widely across subjects, and carry few clinical labels.
Hugging Face Papers

Demystifying Agent Skills: Why They Work-Until They Don't

Demystifying Agent Skills: Why They Work-Until They Don't
cs.LG updates on arXiv.org

Detecting and Discriminating Operator Misspecification in Hybrid PDE-Parameter Learning: a Reference-Free Instrument, with Discrimination Bounded In Sample

・arXiv:2608.16925v1 Announce Type: new Abstract: We build an instrument that reads, from a single fit and with no oracle, whether the operator a hybrid PDE-parameter estimator postulates is wrong-and separates that from a merely unidentifiable parameter. ・On one self-adjoint parabolic inverse problem, an information-matrix statistic with plug-in scale and per-seed parameter has median 0.19 under correct specification,
cs.LG updates on arXiv.org

Diagonal Multi-omics Integration of Heterogenous Datasets

・arXiv:2608.16968v1 Announce Type: cross Abstract: In this paper, we consider methods for the diagonal multi-omics integration of heterogeneous datasets. ・Several approaches to the nature of biological heterogeneity are analyzed and developed to comprehend more clearly the generated differences. ・Specifically, the extremal trace problems for the coupled Laplacian on sets homeomorphic to the Stiefel manifold embedded in
cs.LG updates on arXiv.org

Diff-DDoS: Realistic Cyber-Physical Attack Synthesis and Robust Detection for 5G-Enabled CPS Using Tabular Diffusion Models

・arXiv:2608.17796v1 Announce Type: cross Abstract: Deep learning-based DDoS detectors for 5G-enabled cyber-physical systems face scarce labeled attack data and unrealistic synthetic substitutes, which limit robustness against adaptive adversaries. ・Detectors trained on hand-crafted attacks with fixed scaling multipliers degrade catastrophically (F1-score drops of about 47 percent to 100 percent, depending on scenario)
cs.LG updates on arXiv.org

Diffusion Models for Smarter UAVs: Decision-Making and Modeling

・arXiv:2501.05819v2 Announce Type: replace Abstract: Uncrewed Aerial Vehicles (UAVs) are increasingly used in modern communication networks. ・However, challenges in decision-making and digital modeling continue to hinder their rapid development. ・Reinforcement Learning (RL) algorithms face limitations such as low sample efficiency and limited data versatility, which are further amplified in UAV communications scenarios.
cs.LG updates on arXiv.org

Digital Twin-Based Intrusion Detection for Vehicle Powertrain CAN Bus Systems

・arXiv:2608.17093v1 Announce Type: cross Abstract: Existing automotive intrusion detection systems (IDSs) for the Controller Area Network (CAN) largely target discrepancies in message timing, frequency, or sequencing and cannot detect attacks that preserve these properties while manipulating the payload. ・Digital twins (DTs) have been used to emulate CAN traffic and generate attack scenarios for IDS evaluation, but the
Hugging Face Papers

DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization

DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization
cs.LG updates on arXiv.org

Does Unification Come at a Cost? Uni-SafeBench: A Safety Benchmark for Unified Multimodal Large Models

・arXiv:2604.00547v2 Announce Type: replace-cross Abstract: Unified Multimodal Large Models (UMLMs) integrate understanding and generation capabilities within a single architecture. ・While unified architectures expand multimodal capabilities, their safety implications remain important yet underexplored. ・Existing safety benchmarks predominantly focus on isolated understanding or generation tasks, failing to evaluate the
cs.LG updates on arXiv.org

Domain-Adapted Molecular Language Models for Efficient Search of Make-on-Demand Libraries

・arXiv:2608.17567v1 Announce Type: new Abstract: Pretrained molecular language models are increasingly used as molecular encoders for learning structure-property relationships. ・However, their practical suitability for molecular discovery within and beyond their pretraining domain remains unclear. ・Herein, we systematically benchmark four molecular language models across six virtual molecular libraries spanning drug dis
cs.LG updates on arXiv.org

Doubly robust nearest neighbors in factor models

・arXiv:2211.14297v5 Announce Type: replace-cross Abstract: We introduce and analyze an improved variant of nearest neighbors (NN) for estimation with missing data in latent factor models. ・We consider a matrix completion problem with missing data, where the $(i, t)$-th entry, when observed, is given by its mean $f(u_i, v_t)$ plus mean-zero noise for an unknown function $f$ and latent factors $u_i$ and $v_t$.
cs.LG updates on arXiv.org

DOW-KE: Anchor-Free Multi-Layer Knowledge Editing via Direct End-to-End Weight Optimization

・arXiv:2608.16932v1 Announce Type: new Abstract: Multi-layer locate-then-edit methods for knowledge editing first optimize target residual-stream activations (anchors) at selected layers, then realize them layer by layer as weight updates. ・This pipeline optimizes an intermediate representation but deploys multi-layer weight updates whose joint effect through the true forward pass is never itself optimized: regardless
cs.LG updates on arXiv.org

Dynamic Compression in Recurrent Networks

・arXiv:2608.17896v1 Announce Type: new Abstract: Recurrent models process long contexts efficiently by compressing their history into a fixed-size state, but modern architectures typically do so in a single causal pass over the sequence. ・Each input must therefore be compressed before the model knows how it will later be used, forcing a limited state to compromise across possible future demands. ・We introduce dynamic co
cs.LG updates on arXiv.org

Dynamic Entanglement-Weighted Pruning for Quantum Federated Unlearning in Supply-Chain Risk Prediction

・arXiv:2608.17069v1 Announce Type: cross Abstract: Federated deployments of variational quantum classifiers are attractive for cross-organisation risk prediction in supply chains, because raw data never leaves the client, yet data-protection regulations such as the GDPR grant clients a right to request that their contribution be removed from a trained model after the fact. ・Retraining a federated model from scratch to
Hugging Face Papers

Dynamic Multi-Byte Prediction With Hierarchical Language Models

Dynamic Multi-Byte Prediction With Hierarchical Language Models
cs.LG updates on arXiv.org

Dynamic Regime-Aware Conformal Calibration for Reliable Economic Forecast Intervals under Multiple Distribution Shifts

・arXiv:2608.17079v1 Announce Type: new Abstract: Conformal prediction provides distribution-free prediction intervals but relies on exchangeability, an assumption often violated in economic forecasting because of covariate shift, concept drift, local heterogeneity and latent regimes. ・We propose Dynamic Regime-Aware Conformal Prediction (DRACP), which combines density-ratio, localized kernel and probabilistic regime-aw
WIRED

Dyson Promo Codes: 25% Off in August 2026

・Get 25% off with a Dyson coupon code, plus save up to $600 with discounts on vacuums, $150 off Airwraps, and more.
Hugging Face Papers

EDITBRIDGE: Towards Faithful and Efficient Ultra-High-Resolution Image Editing

EDITBRIDGE: Towards Faithful and Efficient Ultra-High-Resolution Image Editing
cs.LG updates on arXiv.org

Efficient Dynamic Shielding for Parametric Safety Specifications

・arXiv:2505.22104v2 Announce Type: replace-cross Abstract: Shielding has emerged as a promising approach for ensuring safety of AI-controlled autonomous systems. ・The algorithmic goal is to compute a shield, which is a runtime safety enforcement tool that needs to monitor and intervene the AI controller's actions if safety could be compromised otherwise. ・Traditional shields are designed statically for a specific safety
cs.LG updates on arXiv.org

Efficient Resource Optimization for Split Federated Learning

・arXiv:2608.17849v1 Announce Type: new Abstract: Split federated learning (SFL) has emerged as a powerful paradigm for model training at the edge. ・However, SFL inherently involves discrete decision variables for model splitting and resource allocation, resulting in a challenging mixed-integer problem. ・Consequently, prior optimization schemes for SFL are either \textit{heuristic} or \textit{computationally inefficient}
cs.LG updates on arXiv.org

Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation

・arXiv:2608.17941v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models but relies on costly rollout exploration. ・Assigning the same exploration budget to samples with different difficulty levels is inefficient: easy samples may receive redundant rollouts, whereas difficult but learnable samples may receive too little explorati
cs.LG updates on arXiv.org

Elimination Geometry

・arXiv:2608.17646v1 Announce Type: new Abstract: This monograph develops elimination geometry (EG), a typed, native-loss, audit-oriented framework for studying when locally optimal objects can be realized by a shared deployment rule. ・Elimination and compression may erase distinctions required by prediction, inference, control, or representation. ・EG asks which distinctions are lost, whether the induced defect is visibl
cs.LG updates on arXiv.org

EMAN: Optimization-Driven Capacity Growth through Path Emergence in Multi-Task Learning

・arXiv:2608.16930v1 Announce Type: new Abstract: Existing multi-task learning methods rely on hard sharing, multiple paths or experts, adaptive sharing, and dynamic expansion. ・However, their capacity changes are usually constrained by predefined structures or triggered by task boundaries and conflict signals. ・This raises a fundamental question: can a network start from exact single-path computation and grow a new inde
Hugging Face Papers

Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation

Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation
Hugging Face Papers

Energy-Guided Flow Matching

Energy-Guided Flow Matching
cs.LG updates on arXiv.org

EquiPocket: an E(3)-Equivariant Geometric Graph Neural Network for Ligand Binding Site Prediction

・arXiv:2302.12177v5 Announce Type: replace-cross Abstract: Predicting the binding sites of target proteins plays a fundamental role in drug discovery. ・Most existing deep-learning methods consider a protein as a 3D image by spatially clustering its atoms into voxels and then feed the voxelized protein into a 3D CNN for prediction. ・However, the CNN-based methods encounter several critical issues: 1) defective in represe
cs.LG updates on arXiv.org

Estimating Parameter Fields in Multi-Physics PDEs from Scarce Measurements

・arXiv:2509.00203v3 Announce Type: replace Abstract: Parameterized partial differential equations (PDEs) underpin the mathematical modeling of complex systems in diverse domains, including engineering, healthcare, and physics. ・A central challenge in using PDEs for real-world applications is to accurately infer the parameters, particularly when the parameters exhibit non-linear and spatiotemporal variations.
cs.LG updates on arXiv.org

Evaluating and improving crop-yield forecasting methods during extreme drought

・arXiv:2608.17971v1 Announce Type: new Abstract: The impact of climate variability on food production has led to the creation of various forecasting models that uses machine learning (ML), numerical weather predictors (NWP) or a hybrid of ML-NWP models to identify structural and physical relationships between meteorological drivers and crop growth, in order to predict crop yield. ・Droughts, for example the 2012 Midwest
cs.LG updates on arXiv.org

Evaluating RL Explainability Methods by How Much They Help Fix Bugs in Agents

・arXiv:2608.17524v1 Announce Type: new Abstract: This preliminary paper outlines a planned evaluation benchmark for Explainable Reinforcement Learning (XRL) methods. ・Current evaluations rely on functionally-grounded metrics like faithfulness and compactness, and on human-grounded proxies like subjective ratings or prediction accuracy. ・We suggest evaluating XRL methods by how effectively their generated explanations he
cs.LG updates on arXiv.org

Exact Reformulation and Optimization for Direct Metric Optimization in Binary Imbalanced Classification

・arXiv:2507.15240v2 Announce Type: replace Abstract: For classification with imbalanced class frequencies, i.e., imbalanced classification (IC), standard accuracy is known to be misleading as a performance measure. ・While most existing methods for IC resort to optimizing balanced accuracy (i.e., the average of class-wise recalls), they fall short in scenarios where the significance of classes varies or certain metrics
stat.ML updates on arXiv.org

Expected free energy as an information constraint on the Bethe Lagrangian

・arXiv:2608.17167v1 Announce Type: cross Abstract: Active inference selects actions by minimising an expected free energy functional over predicted futures. ・However, adding an expectation over yet-unobserved outcomes means the free energy functional no longer has a Kullback-Leibler structure, which hinders message passing treatments of inference procedures. ・We propose an alternative formulation based on a Bethe free e
cs.LG updates on arXiv.org

Expressivity In Multimodal Contrastive Learning

・arXiv:2608.17203v1 Announce Type: cross Abstract: Contrastive learning has become a cornerstone of modern representation learning, powering CLIP-style models that underpin text-to-image generation, vision-language models, and retrieval across a rapidly growing range of modalities. ・Despite this empirical success, the expressive power of these architectures remains poorly understood. ・To gain insight, we study expressiv
cs.LG updates on arXiv.org

FairNVT: Fair Classification via Noise Injection in Vision Transformers

・arXiv:2604.16780v2 Announce Type: replace-cross Abstract: This paper presents FairNVT, a lightweight debiasing framework for pretrained transformer-based encoders that improves prediction fairness while preserving task performance. ・FairNVT is motivated by the intuition that reducing sensitive-attribute information in the representation used by the downstream classifier can facilitate fairer predictions. ・Our approach
cs.LG updates on arXiv.org

Feature Priming in Online Linear Regression: Sparse-Regret Lower Bounds and a Tight Univariate Rate

・arXiv:2608.17573v1 Announce Type: cross Abstract: In high-dimensional online prediction, the best predictor may depend on only a few features, so regret should scale with sparsity rather than the ambient dimension. ・Feature priming pursues this goal by estimating feature weights from past data and refitting a minimum-norm predictor on the rescaled design. ・Warmuth and Amid asked at COLT 2023 whether any of three such r
cs.LG updates on arXiv.org

FedPref: Federated Preference Learning for Structured Radiology Report Extraction

・arXiv:2608.16971v1 Announce Type: cross Abstract: Radiology reports describe findings and locations in free text, but downstream search and analysis require these relations in a fixed schema. ・Learning this extraction requires labels that are unevenly distributed across institutions: smaller hospitals have less local evidence, and pooling data may be infeasible. ・We introduce FedPref: frozen public language models prop
cs.LG updates on arXiv.org

Fermi-Dirac thermal measurements: A framework for quantum hypothesis testing and semidefinite optimization

・arXiv:2603.04061v2 Announce Type: replace-cross Abstract: Quantum measurements are the means by which we recover messages encoded into quantum states. ・They are at the forefront of quantum hypothesis testing, wherein the goal is to perform an optimal measurement for arriving at a correct conclusion. ・Mathematically, a measurement operator is Hermitian with eigenvalues in [0,1].
cs.LG updates on arXiv.org

FishBack: Pullback Fisher Geometry for Optimal Activation Steering in Transformers

・arXiv:2605.17231v2 Announce Type: replace Abstract: Activation steering has emerged as a lightweight approach for modifying language model behavior without parameter updates, yet existing methods remain brittle: unstable across layers and prone to disturbing behavior unrelated to the target concept. ・We trace these failures to a hidden assumption shared by widely-used methods such as CAA, ActAdd, and ITI: that the int
WIRED

Flock Has a Powerful New AI Tool for Police. We Got Its Code

・Flock’s surveillance cameras have already sparked outrage. ・WIRED reconstructed its next-generation AI system, already in use by some police, to confirm it goes much further than tracking license plates.
cs.LG updates on arXiv.org

Fourth-Moment Geometry of Rademacher Sums

・arXiv:2608.17802v1 Announce Type: new Abstract: Let $\varepsilon_1,\ldots,\varepsilon_n$ be independent Rademacher signs and let $a=(a_1,\ldots,a_n)\in\R^n$ satisfy the normalization below. ・For the normalized Rademacher sum, we determine how its higher moments depend on the fourth-order mass. ・Combining a sharp fixed-q moment envelope with a separate argument below the convexity threshold gives the Gaussian stability
Hugging Face Papers

FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution

FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution
cs.LG updates on arXiv.org

From Abductive Explanations to Global Logical Rules for Node Classification in SGCs

・arXiv:2608.17103v1 Announce Type: new Abstract: Graph Neural Networks (GNNs) have achieved remarkable performance in node classification tasks, motivating growing interest in methods capable of explaining their predictions. ・Recent logic-based approaches, such as LogicXGNN, derive global logical rules for Graph Neural Networks (GNNs) from collections of explanatory subgraphs. ・While informative, these subgraphs may con
cs.LG updates on arXiv.org

From Adoption to Deployment: A Qualitative Study on AI Integration in Software Development Practice

・arXiv:2607.16660v2 Announce Type: replace-cross Abstract: The increasing adoption of Large Language Models (LLMs) as AI components in modern software systems introduces distinct security risks to the software supply chain. ・While many considerations and safety mechanisms are in place for components of the traditional software supply chain, the recent rapid adoption of AI components and platforms has overlooked these h
Hugging Face Papers

From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation

From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation
cs.LG updates on arXiv.org

From Diffusion to Flow: Efficient Motion Generation in MotionGPT3

・arXiv:2603.26747v3 Announce Type: replace-cross Abstract: Recent text-driven motion generation methods span both discrete token-based approaches and continuous-latent formulations. ・MotionGPT3 exemplifies the latter paradigm, combining a learned continuous motion latent space with a diffusion-based prior for text-conditioned synthesis. ・While rectified flow objectives have recently demonstrated favorable convergence an
Hugging Face Papers

From Sequence to Structure: Relational Uncertainty Propagation for LLM Agents

From Sequence to Structure: Relational Uncertainty Propagation for LLM Agents
cs.LG updates on arXiv.org

General Semantic Knowledge Infusion for Spatio-Temporal Traffic Forecasting

・arXiv:2608.17440v1 Announce Type: new Abstract: Although Graph Neural Networks (GNNs) have made significant advances in spatio-temporal traffic forecasting, their performance is limited when they rely solely on sensor proximity or road-network topology. ・This paper presents a spatio-temporal prediction framework, developed to incorporate knowledge in various forms. ・This framework aims to improve sensor-level, contextu
Zennの「大規模言語モデル」のフィード

GLM 5.3 ついに登場:智譜AIが2026年8月に投下した次世代フラッグシップLLMを徹底解説

・「またモデルが出たの?」と思ったあなた、正しいです。でも今回はちょっと事情が違います。 ・中国の智譜AI(Zhipu AI / Z.ai)が次世代フラッグシップ大規模言語モデル GLM 5.3 を2026年8月にリリースしました。前世代の GLM-5.2(2026年6月)からわずか約2か月。半年で3世代という、中国系モデルのリリースサイクルの速さを象徴するニュースです。 ・この記事では、公開されている事実と公式ベンチマークを整理しつつ、「GLM 5.3を実際にどう使うべきか」という開発者視点までをまとめます。読めば、5.3の位置づけと評価ポイントが10分で把握できます。
Zennの「大規模言語モデル」のフィード

GLM-5.3、後学習だけで最前線へ――AIモデル競争は規模から費用対効果の時代へ

・中国のAI企業Z.ai(智譜AI)は8月14日、最新の基盤モデル「GLM-5.3」を発表しました。2026年8月19日時点でAPIとGLM Coding Planから利用でき、複雑なソフトウェア開発、長時間にわたるエージェント処理、サイバーセキュリティ領域を主な対象としています。 ・今回の発表で注目すべきなのは、単にベンチマークの順位が上がったことではありません。GLM-5.3は前世代のGLM-5.2と同じ基盤モデルを使用し、事前学習のやり直しやモデル規模の拡大ではなく、後学習(post-training)の強化によって性能を引き上げたとZ.aiは説明しています。 ・第三者評価で60点、...
cs.LG updates on arXiv.org

Global Convergence of Gradient EM for Over-Parameterized Gaussian Mixtures

・arXiv:2506.06584v2 Announce Type: replace Abstract: Learning Gaussian Mixture Models (GMMs) is a fundamental problem in statistics and machine learning, with the Expectation-Maximization (EM) algorithm and its popular variant gradient EM being arguably the most widely used algorithms in practice. ・In the exact-parameterized setting, where both the ground truth GMM and the learning model have the same number of compone
Zennの「大規模言語モデル」のフィード

Google Colabで最新LLMを試す #10 ― MiniMax Music 3で日本語の歌入り楽曲を生成する

・はじめに 「Google Colabで最新LLMを試す」シリーズの第10回です。 ・これまでは主にLLMやVLMを取り上げてきましたが、今回は少し趣向を変えて、音楽生成モデル MiniMax Music 3 をGoogle Colab上で動かしてみます。 ・MiniMax Music 3は、歌詞と音楽スタイルの説明を入力すると、ボーカルを含む楽曲を生成できるモデルです。
WIRED

Google Pixel 11 Pro and Pixel 11 Pro XL Review: Smart Software, Small Upgrade

・Google’s new voice typing is pure magic, but subpar gaming performance and a useless rear LED keep these flagships from true greatness.
WIRED

Google Pixel Watch 5 Review: More Health, More AI

・With smarter gym tracking, new health alerts, and offline Gemini, the Pixel Watch 5 fine-tunes a winning formula—for a price.
@IT 全フォーラム 最新記事一覧

Google、Kubernetesで動く“野良AI”を棚卸しするOSS「k8s-aibom」公開

・Googleは、「Kubernetes」で稼働するAIシステムの構成要素を自動検出する「k8s-aibom」をオープンソースで公開した。特権を必要とせず、開発者側の設定変更も不要だという。
The Verge

Google’s Pixel 11 Pro Fold feels like the end of an era

・The Pixel 11 Pro Fold has a square screen and a very obvious crease. ・The foldable phone market is in the middle of a huge transformation, but no one told Google. ・Last year, Samsung transformed its Galaxy Z Fold 7 with a dramatically thinner design.
The Verge

Google’s Pixel Watch 5 is promising — but it isn’t finished

・Please clap for my Gemini-inspired nails. ・The Google Pixel Watch 5 has, in many ways, been the hardest Pixel Watch to accurately review. ・Barely anything has changed in terms of hardware.
ITmedia NEWS 最新記事一覧

Googleや英政府、AIで飛行機雲を回避する大規模実証開始 航空業界の温暖化影響を減らす狙い

・Googleは、英政府や航空管制大手NATSなどと共同で、AIを用いて飛行機雲を回避する実証プログラムを開始すると発表した。北大西洋のシャンウィック洋上空域を対象に、航路をわずかに変更して温暖化影響を抑える。空域規模での協調的回避の実証は世界初となり、Googleは計算基盤などを現物提供する。
cs.LG updates on arXiv.org

Gradient Heterogeneity Complements Hessian Heterogeneity in Transformer Optimization

・arXiv:2502.00213v5 Announce Type: replace Abstract: Transformers are difficult to optimize with stochastic gradient descent (SGD) and largely rely on adaptive optimizers such as Adam. ・Despite extensive efforts, the mechanisms behind Adam's advantage over SGD in Transformer optimization are still not fully understood. ・In this study, we analyze the optimization of Transformer models in the fine-tuning setting through t
Zennの「大規模言語モデル」のフィード

Grok Botの動きの手懐け方と最低限まともに使うためにやったこと

・この記事はXの元記事をZenn向けに再掲載したものです。 ・この記事の対象となる方 Grok Botを使ってる方です。Grok Botとは何か・どうやって使うのかという話は省略して、中身の話に特化した話です。 概要 Grok Botを使っていて困ったことは、大きく2つありました。 ・1つは Memory です。xAI自体、積極的にMemoryに保存する形を目指しており、使うほど情報を覚えてくれる一方で、一時的な判断やその場の指摘まで残り、時間がたつほどMemoryが汚れていきます。また、いつ保存されてるかの説明も無く気がついたら大量の矛盾した命令が含まれたMemoryがあふれかえり回...
Hugging Face Papers

GS-Voxel: Fitting-Free Structured Latents for Large-Scale 3DGS Generation

GS-Voxel: Fitting-Free Structured Latents for Large-Scale 3DGS Generation
cs.LG updates on arXiv.org

GUPO: Gradient Uncertainty-aware Policy Optimization for Post-Training Large Language Models

・arXiv:2608.17411v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) has become a widely used approach for post-training Large Language Models (LLMs) for reasoning. ・In GRPO, the group gradients induced by different queries within the same mini-batch are directly averaged to form the policy update. ・However, these group gradients can point in conflicting directions.
Hugging Face Papers

Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents

Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents
cs.LG updates on arXiv.org

Harnessing Magnitude-Only and Complex Measurements for Improved Dynamic MRI Reconstruction with Learned Priors

・arXiv:2608.18036v1 Announce Type: cross Abstract: MRI reconstruction methods for undersampled k-space data naturally utilize complex-valued measurements. ・Parallel developments in sparse phase retrieval have shown that magnitude-only measurements may provide complementary information for signal recovery. ・However, their use in MRI reconstruction remains largely unexplored, due to lack of practical settings where inform
Hugging Face Papers

HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety

HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety
cs.LG updates on arXiv.org

HeteRo-Select: Informativeness as the Participation Driver in Heterogeneous Federated Learning

・arXiv:2508.06692v3 Announce Type: replace Abstract: Federated learning systems typically allocate gradient compression by link speed. ・This is sensible when bandwidth and data informativeness align. ・However, under non-IID data, these signals often decorrelate or invert.
cs.LG updates on arXiv.org

Hierarchical Data Selection via Manifold Coverage and Sparse Feature Coverage in LLM Post-training

・arXiv:2608.16927v1 Announce Type: new Abstract: As supervised fine-tuning data continues to scale, selecting high-value subsets from large candidate pools is crucial for reducing training cost and improving model performance. ・Existing methods often measure diversity directly in the original embedding space, where geometric metrics entangle dominant semantic directions, fine-grained supervision differences, and local
WIRED

Hot or Not Built the Internet We’re Still Swiping Through

・The early website Hot or Not taught its users to rank people at scale. ・Twenty-six years later, dating apps are trying to escape the culture the site helped create.
WIRED

Hotels.com Coupon Codes for August 2026

・Unlock significant savings on hotels, resorts, and getaways with our verified Hotels.com promo codes and gift card discounts. ・Plan your perfect trip today!
cs.LG updates on arXiv.org

How (Not) to Hybridize Neural and Mechanistic Models for Epidemiological Forecasting

・arXiv:2602.06323v3 Announce Type: replace Abstract: Epidemiological forecasting from surveillance data is a hard problem and hybridizing mechanistic compartmental models with neural models is a natural direction. ・The mechanistic structure helps keep trajectories epidemiologically plausible, while neural components can capture non-stationary, data-adaptive effects. ・In practice, however, many seemingly straightforward
cs.LG updates on arXiv.org

How smoothing the affinity matrix affects neighborhood preservation in t-SNE

・arXiv:2608.17190v1 Announce Type: new Abstract: Dimensionality reduction methods are instrumental to visualize high-dimensional data, and t-SNE stands as one of the most widely used methods due to its emphasis on local neighborhood preservation. ・A central component of t-SNE is the affinity matrix, which expresses pairwise similarities in the form of symmetrized probabilities, over which the optimization problem of t-
cs.LG updates on arXiv.org

How to make the most of your masked language model for protein engineering

・arXiv:2603.10302v3 Announce Type: replace Abstract: A plethora of protein language models have been released in recent years. ・Yet comparatively little work has addressed how to best sample from them to optimize desired biological properties. ・We fill this gap by proposing a flexible, effective sampling method for masked language models (MLMs), and by systematically evaluating models and methods both in silico and in v
cs.LG updates on arXiv.org

How Transparent is DiffusionGemma?

・arXiv:2606.20560v2 Announce Type: replace Abstract: LLM reasoning transparency is a critical affordance for understanding model decisions, mitigating misuse and misalignment, and debugging surprising model behaviors. ・However, DiffusionGemma performs a larger fraction of its computation in a continuous latent space; does this make its reasoning less transparent? ・We study this question by decomposing transparency into
cs.LG updates on arXiv.org

Hybrid ML for Lightweight Pre-Route Delay Estimation in Open-Source IC Design

・arXiv:2608.17914v1 Announce Type: new Abstract: Static Timing Analysis (STA) is a critical step in the design flow of digital integrated circuits, however, obtaining accurate delay estimations can represent a challenge when limited information regarding physical design is available. ・In response, this work presents a hybrid and light-weight machine learning (ML) based approach that combines a decision tree with linear
cs.LG updates on arXiv.org

HyPE-GT: where Graph Transformers meet Hyperbolic Positional Encodings

・arXiv:2312.06576v2 Announce Type: replace Abstract: Graph Transformers (GTs) facilitate the comprehension of complex relationships on graph-structured data by leveraging self-attention of the possible pairs of nodes. ・The structural information or inductive bias of the input graph is provided as positional encodings to the GT. ・The positional encodings are mostly Euclidean and are not able to capture the complex hierar
WIRED

I Tried a Window-Cleaning Robot: Do Not Recommend

・I tested the top-of-the-line Ecovacs Winbot W2S Omni window-cleaning robot on my Victorian home, and it was an unmitigated disaster.
cs.LG updates on arXiv.org

ImplicitTerrainV2: Wavelet-Guided Spatially Adaptive Neural Terrain Representation

・arXiv:2605.22556v2 Announce Type: replace Abstract: Digital elevation models (DEMs) underpin terrain analysis in Geographic Information Systems (GIS), but commonly as raster representation, they rely on interpolation for off-grid sampling and finite-difference operators for derivative-based analysis. ・Implicit neural representations (INRs) offer a continuous alternative, but prior terrain INRs lack explicit frequency
cs.LG updates on arXiv.org

Information fusion and machine learning for sensitivity analysis using physics knowledge and experimental data

・arXiv:2608.17248v1 Announce Type: cross Abstract: When computational models (either physics-based or data-driven) are used for the sensitivity analysis of engineering systems, the sensitivity estimate is affected by the accuracy and uncertainty of the model. ・This paper considers global sensitivity analysis (GSA) for situations where both a physics-based model and experimental observations are available, and investiga
cs.LG updates on arXiv.org

Information Spreading in Diffusion Models from Effective Field Theory

・arXiv:2608.14308v1 Announce Type: cross Abstract: We study score-matching diffusion models with a convolutional architecture. ・We argue that the inductive bias of locality means that the machinery of effective field theory from physics can be usefully applied to describe the denoising dynamics. ・We apply this formalism first to a simple toy example which permits an analytical description, and thereafter to MNIST, and s
cs.LG updates on arXiv.org

Integrating Novelty and Surprise for Experience Prioritization and Exploration in Image-Based Reinforcement Learning

・arXiv:2608.17373v1 Announce Type: new Abstract: Sample efficiency is a central challenge in reinforcement learning (RL), particularly in image-based domains where agents must learn from high-dimensional visual inputs. ・Traditional sampling often relies on random or suboptimal experience selection, leading to redundant updates and slow learning. ・Improving efficiency requires mechanisms that prioritize informative exper
cs.LG updates on arXiv.org

Intent-Driven Dynamic Chunking: Segmenting Documents to Reflect Predicted Information Needs

・arXiv:2602.14784v1 Announce Type: cross Abstract: Breaking long documents into smaller segments is a fundamental challenge in information retrieval. ・Whether for search engines, question-answering systems, or retrieval-augmented generation (RAG), effective segmentation determines how well systems can locate and return relevant information. ・However, traditional methods, such as fixed-length or coherence-based segmentat
cs.LG updates on arXiv.org

Inverse Problems for Partial Differential Equations with Jump Discontinuities in Coefficients via Two-Stage Physics-Informed Deep Learning and Statistical Mixture Models

・arXiv:2510.14656v3 Announce Type: replace-cross Abstract: This work proposes a two-stage physics-informed deep learning framework that combines neural-network-based sampling with statistical inference and constrained parameter refinement. ・In the first stage, a dual-network physics-informed architecture is used, where a main network approximates the PDE solution and an auxiliary coefficient sub network provides a rela
cs.LG updates on arXiv.org

Iterative Grasp Pose Refinement: A Deep Reinforcement Learning Approach for 2D Vision

・arXiv:2608.17628v1 Announce Type: cross Abstract: Developing robots capable of understanding and manipulating objects requires compact, interpretable, and generalizable representations. ・This work proposes a reinforcement learning-based framework for robotic grasp refinement, integrating keypoint-based object representations with a Deep Q-Network (DQN). ・Using 2D overhead images captured in a simulated environment, a g
cs.LG updates on arXiv.org

Iterative tensor network transformations for element-wise evaluation of elementary and filtering functions

・arXiv:2608.17135v1 Announce Type: new Abstract: Tensor networks are powerful formats for compressing large-scale data. ・However, their application to general data processing has been limited by the difficulty of performing nonlinear operations. ・Here, we introduce iterative tensor network transformations (ITNTs), a general algorithmic framework for the element-wise evaluation of elementary and nonlinear filtering funct
#AIタグ

ITとICTの違い

・ニュースや学校で「IT」とか「ICT」って言葉、よく聞くよね。うんうん。 ・似ているようで、実は見ている「視点」がちょっと違うらしいよ! 続きをみる
cs.LG updates on arXiv.org

J-Miner: Recovering Executable Decision Knowledge from Language-Model Classifiers

・arXiv:2608.17063v1 Announce Type: new Abstract: Large language models can be fine-tuned into specialized classifiers that perform well across diverse text tasks and make complex judgments, but they typically expose only final labels, leaving the decision knowledge acquired through fine-tuning implicit within the model. ・We study how to mine this internal decision knowledge from a fine-tuned classifier and encode it in
Zennの「大規模言語モデル」のフィード

Jetson AGX XavierでQwen3.8-27Bを動かしてみる

・https://zenn.dev/nooop/articles/69f444fe13ecca この記事の更新のようなもの 目的 Qwen3.8-27Bが公開されました。 ・https://huggingface.co/unsloth/Qwen3.8-27B-GGUF では早速なので試します。 ・非力なxavierで使い物になる速度になるか? Qwen3.8-27BはDense版なので、xavierではかなり遅いのは確実です。
LLMタグが付けられた新着記事 - Qiita

KV キャッシュがなぜ LLM 推論の次のボトルネックになるのか

・LLM の量子化モデルでは、モデルファイルが十分に小さくても、目的の文脈長や同時実行数で推論するとメモリが足りなくなることがあります。重みと KV キャッシュは別の領域だからです。 ・以前、量子化モデルの必要メモリを「重み」「KV キャッシュ」「その他の実行時領域」に分けて見...
cs.LG updates on arXiv.org

Lambda-Hold Control: Human-Like Movement Emerges from a Minimal Task Reward in Predictive Musculoskeletal Simulation

・arXiv:2608.17030v1 Announce Type: cross Abstract: The massive overactuation in the human musculoskeletal system makes it challenging to train musculoskeletal models to generate human-like motion via reinforcement learning, primarily because exploration in the resulting high-dimensional and redundant action space is extremely inefficient. ・To address this problem, we propose the $\lambda$-hold controller, inspired by t
Zennの「大規模言語モデル」のフィード

LangGraph のコードを読んで、自作マルチエージェントと突き合わせた — 「言葉で繋がない」設計と動的ルーティング

・要約 マルチエージェントのオーケストレーションを LLM の「言葉」で繋ぐと、要約のたびに情報が落ちて長時間の運用で回らなくなる。Anthropic の "Multi-agent coordination patterns" が言う「orchestrator becomes an information bottleneck」は、自作の村でも実測で出た LangGraph(langchain-ai/langgraph 1e44bda・2026-08-18)のコードを読むと、そもそも言葉で繋いでいない。ノード間を渡るのは型付きの状態(channel_values)で、ルータは状態を読...
cs.LG updates on arXiv.org

Large Language Models: A Mathematical Formulation

・arXiv:2601.22170v2 Announce Type: replace-cross Abstract: Large language models (LLMs) process and predict sequences containing text to answer questions, and address tasks including document summarization, providing recommendations, writing software and solving quantitative problems. ・We provide a mathematical framework for LLMs by describing the encoding of text sequences into sequences of tokens, defining the archit
cs.LG updates on arXiv.org

Latent Order Bandits

・arXiv:2605.07304v2 Announce Type: replace Abstract: Bandit algorithms solve diverse sequential decision-making problems, but are often too sample-inefficient for from-scratch personalization. ・To substantially reduce exploration times, latent bandit algorithms exploit cross-instance structure implied by discrete latent states, provided that the posterior distribution of rewards and latent states is known and accurate.
Zennの「大規模言語モデル」のフィード

Le Critique: LLM強化学習における価値関数の復権 — PVFとTETHER

・TL;DR LLMの強化学習において、GRPOを代表とするcritic-free手法が主流となって久しい。理由は明確で、価値関数(critic)の学習は不安定であり、インフラコストもかさむ。しかしcritic-free手法には本質的な限界がある——トークンレベルの信用割当てができないことと、グループサンプリングに伴うstraggler問題だ。 ・本論文は「価値関数をRLパイプラインに戻す」ための二つの相補的戦略を提案する。**Privileged Value Functions(PVF)**は、criticにポリシーが直接使えない特権情報(正解や他のロールアウトの報酬など)を条件付け、...
Hugging Face Papers

LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents

LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents
cs.LG updates on arXiv.org

Leveraging Association Context Retrieval in Knowledge Edit- ing to Build White-Box Attacks on LLMs

・arXiv:2608.17836v1 Announce Type: new Abstract: As large language models (LLMs) are granted increasing autonomy, it is essential to investigate methods that can induce unsafe behavior. ・We propose a novel white-box attack inspired by locate-then-edit approaches from the field of Knowledge Editing. ・Our choice is motivated by the observation that models edited with such schemes tend to assign unusually high prediction p
cs.LG updates on arXiv.org

Leveraging existing sparse point annotations for benthic imagery dense segmentation

・arXiv:2608.17561v1 Announce Type: cross Abstract: The health of marine ecosystems is a critical indicator of global environmental change, yet the physical constraints of underwater observation and the intrinsic challenges of processing marine imagery severely limit the scalability of systematic monitoring. ・While recent visual foundation models such as the Segment Anything Model (SAM) series show great promise, they s
cs.LG updates on arXiv.org

Leveraging generative hallucination and biophysics-informed modeling for unified biomolecular sequence-structure co-design

・arXiv:2608.17381v1 Announce Type: cross Abstract: Biomolecular design underpins applications from molecular recognition to therapeutics and synthetic biology, yet de novo interaction design remains challenging-especially for DNA/RNA, underexplored non-protein modalities with scarce, heterogeneous complex data and sharper geometric and chemical constraints. ・We introduce MCTH (Monte Carlo Tree Hallucination), an infere
Hugging Face - Blog

LFM2.5 Q4\_0 Checkpoints from Quantization-Aware Distillation

LFM2.5 Q4\_0 Checkpoints from Quantization-Aware Distillation
cs.LG updates on arXiv.org

Likelihood Hacking in Probabilistic Program Synthesis

・arXiv:2603.24126v2 Announce Type: replace Abstract: When language models are trained by reinforcement learning (RL) to write probabilistic programs, they can artificially inflate their marginal-likelihood reward by producing programs whose data distribution fails to normalise instead of fitting the data better. ・We call this failure likelihood hacking (LH). ・We formalise LH in a core probabilistic programming language
#LLMタグ

LLMができるまで――2026年8月版

・ChatGPTやQwenに文章を入力すると、数秒後にはかなり自然な答えが返ってくる。 ・では、その「答えているAI」自体は、いったいどうやって作られているのだろう。
#LLMタグ

LLMの投資判断ロジックを評価する新ベンチマーク『InvestLogicBench』を発表

・AIが変える「投資判断」の未来!あなたの資産運用、そのAIは本当に信頼できますか? 皆さん、AIに資産運用を任せる未来、想像できますか? AIの進化は日々目覚ましく、そのスピードには目を見張るものがあります。特に金融分野でのAI活用は、私たちの生活、そして資産形成のあり方を大きく変える可能性を秘めていると感じています。
#LLMタグ

LLMの臨床エラー検出評価に新提案:F1値の落とし穴とペア評価の重要性

・AIの進化が目覚ましい昨今、医療現場での活用にも大きな期待が寄せられていますよね。特に、膨大な医療文書の中から、人間では見落としがちなエラーをAIが見つけ出すなんて、まさに未来の医療だと感じませんか?この可能性には計り知れない魅力を感じます。 ・しかし、その裏側には、実は一筋縄ではいかない「評価」という課題が潜んでいるんです。単なる技術導入だけでなく、その真価を引き出すためには、私たちが想像する以上の深い理解が必要なんですよ。数字だけでは見えない、もっと大切な判断基準がある。それは、AIの真の能力を引き出すための―― 続きをみる
Zennの「大規模言語モデル」のフィード

LLMはヘーゲルの「絶対知」なのか

・本稿について 本稿は、ヘーゲル哲学を学術的に解説することを目的としたものではない。 ・『精神現象学』の「絶対知」という概念に触発され、LLMの登場によって、人間の知識や認識、そして人類と知識との関係がどう変わりうるのかを考えた個人的な考察である。 ・人間は世界そのものを見ているのか 人間は外界をそのまま認識しているわけではない。
#LLMタグ

LLM駆動型エージェントの新たな脆弱性:状態意味注入の脅威

・AIが人間の生活に深く溶け込む現代。きっと誰もが「もっと便利に、もっと安全に」と願っていますよね。私も日々、世界のAI情報を日本最速で発信し、AIの事なら柴亮太にお任せ下さい!という気持ちで取り組んでいます。しかし、その目覚ましい進化の裏側には、私たちがまだ知り得ない、恐ろしいほどの脆弱性が潜んでいることをご存知でしょうか? 特に、環境を認識し、道具を使い、自律的に作業を実行する「AIエージェント」の世界では、これまでのセキュリティ対策の常識が通用しない、想像を絶するリスクが生まれつつあるんです。これは単なるバグやシステムエラーの話ではありません。むしろ、AIが「完璧に動く」からこそ引き起こされる、ある種の罠のようなものだと言えるでしょう。
#LLMタグ

LLM推論エラー、残差ストリームの『領域と方向』で高精度検出

・世界のAI情報を日本最速で発信! AIの事なら柴亮太にお任せ下さい😀 日々進化する大規模言語モデル(LLM)は、私たちの仕事や生活を劇的に変えています。しかし、その一方で、私たちがAIプロダクトを本番運用する中で常に直面するのが、LLMが時折見せる「推論エラー」です。なぜ期待通りの答えが出ないのか、どうすれば安定した性能を引き出せるのか。これは、多くのAI開発者や事業責任者が頭を悩ませてきた課題ではないでしょうか。
cs.LG updates on arXiv.org

Looking Beyond the Scale: Do Surgical Skill Models Learn Transferable Representations Across Assessment Rubrics?

・arXiv:2608.17519v1 Announce Type: cross Abstract: Vision-based surgical skill assessment has shown strong in-domain results, yet a fundamental question remains unasked: do these models learn transferable representations of surgical proficiency, or do they merely encode dataset-specific visual patterns? ・This paper systematically analyzes what limits cross-domain skill transfer between the GOALS and OSATS assessment sc
WIRED

Loop Earplugs Discount Codes: 40% Off

・Save on Loop Earplugs, including Quiet 2 and popular gift sets for improved sleep, focus, and comfort.
cs.LG updates on arXiv.org

Low-dimensional topology of deep neural networks

・arXiv:2606.31856v2 Announce Type: replace Abstract: We study layered models, including feedforward networks, ResNets, and transformers, by limiting each layer to a width of $d = 3$, i.e., $\mathbb{R}^3$ as representation space. ・This allows us to track how a neural network changes low-dimensional topological invariants through its layers. ・Just about any topological structure may be simplified or even trivialized by si
cs.LG updates on arXiv.org

Lymphocyte Mimicry Correction via Region-Level Tissue Reasoning and Unbalanced Optimal Transport

・arXiv:2608.17151v1 Announce Type: cross Abstract: Cell mimicry arises when different cell types appear morphologically similar. ・Human pathologists resolve this ambiguity using surrounding tissue context, whereas current vision models either lack contextual reasoning (cell foundation models) or cannot operate at the cell level (pathology MLLMs). ・We present Loki-OT, which propagates region-level tissue reasoning to ind
cs.LG updates on arXiv.org

LZ Penalty: An information-theoretic repetition penalty for autoregressive language models

・arXiv:2504.20131v4 Announce Type: replace Abstract: We introduce the LZ penalty, a penalty specialized for reducing degenerate repetitions in autoregressive language models without loss of capability. ・The penalty is based on the codelengths in the LZ77 universal lossless compression algorithm. ・Through the lens of the prediction-compression duality, decoding the LZ penalty has the interpretation of sampling from the r
cs.LG updates on arXiv.org

MAGPIE-Net: Predicting short-duration heavy-rainfall events in station neighborhoods from multitemporal FY-4A AGRI observations

・arXiv:2608.17753v1 Announce Type: new Abstract: Short-duration heavy-rainfall warning determines whether 1 h rainfall will exceed a threshold within a target-station neighborhood over the next few hours. ・Multitemporal infrared and water-vapor observations from the Fengyun-4A Advanced Geostationary Radiation Imager (FY-4A AGRI) capture cloud-top cooling, moisture evolution, and cloud expansion before substantial surfa
cs.LG updates on arXiv.org

MagViT: Interpretable Multi-Magnification Transformers with Patient-Level Model Selection for Breast Histopathology

・arXiv:2608.16959v1 Announce Type: cross Abstract: Breast cancer is one of the most common types of cancer among women around the world. ・Rapid detection and early treatment can hinder its progress to more complex stages and can impede its spread to other parts of the body. ・Histopathological image classification is the most common task in cancer detection due to its robustness in analyzing cellular data.
cs.LG updates on arXiv.org

Margin-Regularized Structured Semantic Alignment for Brain-Language Correspondence

・arXiv:2608.16975v1 Announce Type: cross Abstract: With the rapid advancement of large language models, brain-language decoding has achieved remarkable progress. ・However, it remains unclear whether decoded content genuinely reflects neural representations or is largely reconstructed by the language model itself. ・This ambiguity limits interpretability and hinders the investigation of intrinsic brain-language correspond
Hugging Face Papers

MathForm: Scaling Mathematical Autoformalization with Knowledge Retrieval and Verification-Guided Refinement

MathForm: Scaling Mathematical Autoformalization with Knowledge Retrieval and Verification-Guided Refinement
stat.ML updates on arXiv.org

Maximum Tsallis Entropy Distributions for Robust and Efficient Sparse Learning from Correlated Data

・arXiv:2608.17244v1 Announce Type: cross Abstract: This paper addresses the limitations of Gaussian distribution assumptions in statistical sparse learning, particularly in modeling correlated and heterogeneous data. ・Conventional Gaussian models often lack robustness towards outliers and underlying distribution assumptions. ・To overcome these limitations, we propose the use of the $q$Gaussian distribution, derived from
AI News & Artificial Intelligence | TechCrunch

Meet the startup helping Wall Street put a price on AI compute

・The AI buildout shows no signs of slowing. ・And with hundreds of billions of dollars a year going into data centers and GPUs, compute has become the single biggest cost for anyone building AI products. ・But for all that spending, there still isn’t a straightforward way to put a price on compute — or for firms to hedge their exposure when the price changes.
cs.LG updates on arXiv.org

MemCatalyst: Amplifying Data Auditing on Vision-Language Models via Data Poisoning

・arXiv:2608.17722v1 Announce Type: cross Abstract: Vision-Language models (VLMs) achieve outstanding performance largely due to the amount of training data available on the internet. ・At the same time, data holders (e.g., artists) urgently need to determine whether their data has been used for model training without authorization, which concerns both intellectual property rights and personal privacy. ・Data auditing, par
cs.LG updates on arXiv.org

Memory by Design: Probabilistic Sequence Layers

・arXiv:2605.31163v3 Announce Type: replace-cross Abstract: We introduce the \emph{design-model framework}: a way to derive efficient recurrent sequence maps from explicit assumptions about memory. ・A design model writes evidence into memory by exact Bayesian filtering; a query- dependent readout produces a predictive distribution whose mean is the layer output. ・In our linear-Gaussian instantiation, the \emph{Bayesian L
The Verge

Meta AI is getting a Mac app

・Meta AI can create content and make suggestions based on what it “sees” on your screen. ・| Image: Meta Meta is launching a new Mac app dedicated to its AI chatbot. ・In an announcement on Wednesday, Meta says you can share your window with its AI chatbot, which can provide suggestions, answer questions, or create content based on what's on your screen.
stat.ML updates on arXiv.org

Minimax Optimal Estimator and Improved Error Rate for the MLE in Logistic Regression with Gaussian Design

・arXiv:2608.17260v1 Announce Type: cross Abstract: We study finite-sample parameter estimation in logistic regression with Gaussian design, where the goal is to estimate $\mathbf{\theta}^*\in \mathbb{R}^d$ with $R=\|\mathbf{\theta}^*\|_2\ge 1$ from i.i.d. ・samples $\{(\mathbf{x}_i,y_i)\}_{i=1}^n,$ $\mathbf{x}_i \sim N(0,\mathbf{I}_d)$, $y_i\mid \mathbf{x}_i \sim \mathrm{Bernoulli}((1+\exp(-\mathbf{x}_i^\top \mathbf{\th
cs.LG updates on arXiv.org

MITRE-SAGE: A Multi-Agent Cybersecurity Question-Answering model

・arXiv:2608.16921v1 Announce Type: cross Abstract: Effective cybersecurity operations require timely and accurate analysis of large-scale heterogeneous security information; however, analysts increasingly struggle with information overload, alert fatigue, and time-constrained decision-making. ・Although large language models (LLMs) have demonstrated promising capabilities for question answering (QA), their effectiveness
cs.LG updates on arXiv.org

Mixture-of-Expert Blocks Contain Strong Hallucination Detection Signals

・arXiv:2608.17687v1 Announce Type: cross Abstract: Despite their widespread use, Large Language Models (LLMs) remain limited by a fundamental problem: the generation of plausible but false content, known as hallucinations. ・Most existing detection methods operate at the answer or sentence level, yet per-token detection is essential for localizing hallucinated spans and enabling fine-grained interventions. ・In this paper
stat.ML updates on arXiv.org

Modified Bryson-Frazier Smoothing and Hyperparameter Learning for Temporal Gaussian Process Regression

・arXiv:2608.17595v1 Announce Type: cross Abstract: One-dimensional Gaussian processes with stationary, integrable kernel functions admit exact or arbitrarily accurate state-space representations, enabling linear-time inference through Kalman filtering and Rauch-Tung-Striebel (RTS) smoothing. ・However, the RTS smoother requires inversion of predicted state covariance matrices, which can become ill-conditioned and may th
Hugging Face Papers

MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding

MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding
cs.LG updates on arXiv.org

MoNe: Modular Neural Memory for Efficient Long Context Inference

・arXiv:2608.17616v1 Announce Type: cross Abstract: We present MoNe, a lightweight modular neural memory that attaches to any frozen pretrained Transformer to enable long-context inference without retraining. ・MoNe reads context in fixed-size segments via test-time learning of fast-weight neural memory networks with layer-localized gradient updates; at inference, the memory generates keys and values from the query token
cs.LG updates on arXiv.org

Monotone Classification with Relative Approximations

・arXiv:2506.10775v3 Announce Type: replace Abstract: In monotone classification, the input is a multi-set $P$ of points in $\mathbb{R}^d$, each associated with a hidden label from $\{-1, 1\}$. ・The goal is to identify a monotone function $h$, which acts as a classifier, mapping from $\mathbb{R}^d$ to $\{-1, 1\}$ with a small {\em error}, measured as the number of points $p \in P$ whose labels differ from the function v
cs.LG updates on arXiv.org

MoRAX: Mobility-based Representation Augmentation for Geospatial Foundation Models

・arXiv:2608.17848v1 Announce Type: new Abstract: Geospatial Foundation Models (GFMs) are emerging as a powerful paradigm for learning semantically rich and geographically consistent visual and physical representations. ・However, their reliance on Earth-observation (EO) data leaves information about human activity largely underrepresented. ・Human mobility data reveals the functional and relational structure between regio
cs.LG updates on arXiv.org

Mos-Gen: A Generative Molecular Framework for Mosquito Insecticide Design

・arXiv:2606.01846v2 Announce Type: replace Abstract: Mosquito-borne infectious diseases cause more than 700000 deaths worldwide each year. ・The long-term use of conventional chemical insecticides has induced serious resistance problems, creating an urgent need to develop novel, highly effective, and ecologically sustainable alternatives. ・While existing artificial intelligence approaches in this domain have focused prim
cs.LG updates on arXiv.org

MotoSafety: Edge-AI with Learned Temporal Importance for Two-Wheeler Collision Risk Assessment Under Time Pressure

・arXiv:2608.17823v1 Announce Type: new Abstract: Powered two-wheeler riders face critical safety challenges in low- and middle-income countries, yet limited studies exist on how cognitive stressors such as Time Pressure influence collision risk. ・To address this gap, we introduce a large-scale dataset of over 129,000 labeled multivariate time-series sequences from 153 simulator rides by 51 participants under No, Low, a
cs.LG updates on arXiv.org

Mr.Dec: Daily-Scale Longitudinal Multimodal Modeling for 30-Day Readmission Prediction

・arXiv:2608.16929v1 Announce Type: new Abstract: Predicting 30-day hospital readmission is essential for assessing patient stability and optimizing healthcare resources. ・As clinical risk evolves with the accumulation of evidence during hospitalization, capturing these dynamic trajectories is essential. ・However, many existing approaches compress the complex longitudinal history into fixed representations, often losing
cs.LG updates on arXiv.org

MultiSigBERT: Beyond Survival Analysis through Multimodal and Sequential Modeling in Oncology

・arXiv:2608.16972v1 Announce Type: new Abstract: Machine learning has become an essential component of modern healthcare, where the integration of heterogeneous data sources offers unprecedented opportunities to improve clinical decision-making. ・Electronic Health Records (EHR) contain complementary information -- including narrative clinical reports, numerical measurements, and structured variables -- yet most surviva
cs.LG updates on arXiv.org

Network Denoising Revisited: A Ricci-Flow-Inspired Graph Diffusion Method

・arXiv:2608.16923v1 Announce Type: cross Abstract: Networks provide a fundamental representation of relationships among entities. ・However, real-world networks are often corrupted by noise caused by measurement errors and inherent stochasticity, hindering the discovery of meaningful structure. ・Most denoising methods rely on similarity-driven diffusion and ignore the non-Euclidean geometry of graphs, where local variati
cs.LG updates on arXiv.org

Neural Operator-Based Nonlinear Nudging for Chaotic Dynamical Systems

・arXiv:2508.05778v2 Announce Type: replace Abstract: Nudging is an empirical data assimilation technique that incorporates an observation-driven control term into the model dynamics. ・The trajectory of the nudged system approaches the true system trajectory over time, even when the initial conditions differ. ・For linear state space models, such control terms can be derived under mild assumptions.
cs.LG updates on arXiv.org

No Gaussian Required: Contrastive Inverse Dynamics for JEPA World Models

・arXiv:2608.17542v1 Announce Type: new Abstract: Joint-Embedding Predictive Architectures (JEPAs) learn world models by predicting future embeddings, but the objective admits a trivial solution of a constant encoder, so every practical system adds an anti-collapse mechanism (LeCun, 2022; Assran et al., 2023; Bardes et al., 2022; 2024). ・LeWorldModel (LeWM) prevents collapse with SIGReg, a regularizer that forces the la
cs.LG updates on arXiv.org

Non-KKT Accumulation in Entropic Mirror Descent

・arXiv:2608.01658v3 Announce Type: replace-cross Abstract: For mirror descent generated by a Legendre kernel, perhaps one of the most basic question in optimization is this: must every accumulation point of a bounded mirror descent sequence be Karush--Kuhn--Tucker (KKT) stationary under proper stepsizes? ・We show that the answer is no. ・A longstanding obstacle to resolving this question is the boundary blow-up of the Le
cs.LG updates on arXiv.org

Nonlinear Data Integration via Kernel Methods for Data Collaboration Analysis

・arXiv:2605.27219v2 Announce Type: replace Abstract: Collaborative analysis of decentralized confidential datasets is important, but direct sharing of original datasets is often restricted by privacy and institutional constraints. ・Data collaboration (DC) analysis transforms each dataset into privacy-preserving intermediate representations via party-specific obfuscation functions and integrates them into common collabo
cs.LG updates on arXiv.org

Nonlinear GENERIC-Embedded Neural Networks (N-GENNs): Learning GENERIC dynamics with non-quadratic dissipation potentials

・arXiv:2605.09058v2 Announce Type: replace-cross Abstract: We introduce Nonlinear GENERIC-Embedded Neural Networks (N-GENNs), a deep learning framework for discovering evolution equations of systems governed by the nonlinear GENERIC formalism (General Equation for Non-Equilibrium Reversible-Irreversible Coupling). ・Such systems exhibit coupled conservative and dissipative dynamics, and can be described via the superpos
cs.LG updates on arXiv.org

Nonlocal Transition Kernel for Efficient Learning of Restricted Boltzmann Machines

・arXiv:2608.17450v1 Announce Type: cross Abstract: Learning restricted Boltzmann machines (RBMs) is computationally challenging because it requires expectations whose exact evaluation is generally intractable. ・The expectations are typically evaluated using a sampling approximation based on blocked Gibbs sampling (BGS), which is a local Markov chain Monte Carlo transition kernel. ・However, the locality of BGS can lead t
OpenAI News

Offering Zero Data Retention for frontier models

・OpenAI reaffirms Zero Data Retention for eligible API customers and previews Private Safety Processing for advanced AI safety without compromising data privacy.
cs.LG updates on arXiv.org

On detection probabilities of link invariants

・arXiv:2509.05574v3 Announce Type: replace-cross Abstract: We prove that, for many standard link invariants, both the proportion of distinct invariant values and the detection probability among prime alternating links with at most n crossings decay exponentially in n, with an explicit universal rate. ・In fact, almost every such link belongs to an invariant fiber whose size is itself exponential in n. ・This phenomenon ap
cs.LG updates on arXiv.org

On Stability in Optimistic Bilevel Optimization

・arXiv:2408.13323v3 Announce Type: replace-cross Abstract: Solutions of bilevel optimization problems tend to suffer from instability under changes to problem data. ・In the optimistic setting, we construct a lifted formulation that exhibits desirable stability properties under mild assumptions that neither invoke convexity nor smoothness. ・The upper- and lower-level problems might involve integer restrictions and disjun
cs.LG updates on arXiv.org

On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification

・arXiv:2608.18066v1 Announce Type: cross Abstract: Memory-based self-improving agents--those that learn from an online stream of tasks and improve over time by maintaining a textual memory bank--have shown great promise in recent literature. ・However, the reliability aspects of these methods have been critically overlooked. ・In this work, we conduct a comprehensive re-evaluation of two memory-based methods, broadening t
cs.LG updates on arXiv.org

On the Pseudo-Mixing of Kac's Walk

・arXiv:2608.17374v1 Announce Type: cross Abstract: Motivated by a conjecture of Vaikuntanathan and Zamir, we study the pseudo-mixing of Kac's walk on $\mathrm{SO}(n)$: whether short trajectories are indistinguishable from Haar measure by low-complexity tests. ・We prove that the first $k$ columns mix in Wasserstein distance in $O(n(k+\log n)\log n)$ steps for fixed accuracy, resolving a conjecture of Oliveira.
stat.ML updates on arXiv.org

On the Subgaussianity of Quantized Linear Maps: An AI-Assisted Note

・arXiv:2605.27563v2 Announce Type: replace-cross Abstract: We prove an elementary bounded-differences inequality for functions of non-isotropic Gaussian vectors. ・Specifically, if $f$ has bounded coordinate differences and $X\sim\mathcal N(\mu,\Sigma)$, then the resulting concentration bound depends on the condition number $\kappa(\Sigma)$. ・As an application, we answer a question of Simone Bombari concerning the subgau
cs.LG updates on arXiv.org

One Pipeline, Many Transformers: Pattern-Specific Imputation Specialists for Tabular Missing Data

・arXiv:2510.02625v5 Announce Type: replace Abstract: Missing data in tabular datasets forces practitioners into a hard choice: deploy a general-purpose imputer that may perform poorly for the problem at hand, or wait for someone to design a specialized algorithm. ・This problem is worsened by the fact that real-world missingness rarely satisfies the textbook missing completely at random (MCAR) assumption, as entries are
cs.LG updates on arXiv.org

Online Generalized Sparse Regression: How Does Overparametrization Help?

・arXiv:2608.17466v1 Announce Type: cross Abstract: Regularized sparse regression has been extensively studied in the offline setting, but online formulation remains relatively under-explored. ・This gap stems from four key challenges: (i) the infeasibility of dynamically updating the regularization parameter in every online round, (ii) managing storage and memory complexity, (iii) enabling real-time computation via clos
cs.LG updates on arXiv.org

OOD Detection for EEG-based Machine Learning in High-Risk Environments

・arXiv:2608.17620v1 Announce Type: new Abstract: Machine learning models for electroencephalography (EEG) analysis show great promise across a wide range of applications, but their deployment in high-risk domains is hindered by their vulnerability to distribution shifts. ・Encountering out-of-distribution (OOD) data can lead to catastrophic, overconfident predictive failures. ・While OOD detection methods can mitigate the
cs.LG updates on arXiv.org

Open datasets and machine learning for two-phase heat transfer: a review following a spatial-temporal taxonomy

・arXiv:2605.23037v2 Announce Type: replace Abstract: Two-phase heat transfer underpins boiling, condensation, immersion cooling, flow boiling, energy conversion, and electronics thermal management, but its coupled interfacial physics make data reuse and model comparison difficult. ・This narrative review synthesizes open datasets, machine-learning methods, and reusable software for two-phase heat-transfer research, with
The Verge

OpenAI hit the brakes. Now what?

・With a looming IPO, intense competition from Anthropic, and Chinese and open-weight rivals nipping at its heels, OpenAI has plenty of reasons to move fast. ・Instead, it hit the brakes. ・On Tuesday, the company said it had slowed the pace of some AI development while it tightened security and safeguards.
#LLMタグ

OpenAI、AI開発を止めた日。競争軸は「性能」から“安全に働かせる力”へ|海外AI FRONTLINE DAILY #005|2026.08.19

・OpenAIが「速く作る」より「止めて守る」を選んだ日。AI競争の評価軸が、性能から“運用能力”へ動く 8月18日から19日早朝までの海外AIをWIDE SCANすると、今日は新しい「最強モデル」が主役ではなかった。
cs.LG updates on arXiv.org

Optimize Your Sampling: Tuned Diffusion Sampling with Bayesian Optimization

・arXiv:2608.18040v1 Announce Type: new Abstract: Sampling from a diffusion model typically requires many forward passes through a large neural network, making generation computationally expensive. ・While much work has focused on efficient solvers and samplers, comparatively little attention has been paid to selecting the sampling timesteps themselves. ・A recent line of work optimizes theoretically derived surrogates for
cs.LG updates on arXiv.org

OraclePhys: A Systematic Framework for LLM Fine-Tuning on Structural Mechanics

・arXiv:2608.17162v1 Announce Type: new Abstract: What a language model internalizes from fine-tuning is usually diagnosed after the fact. ・We make it an experimental variable. ・OraclePhys is a systematic fine-tuning framework with three components: OraclePhys-Bench, an exactly-graded structural-mechanics benchmark whose finite-element oracle scores every answer and counterfactual edit -- no human labels, no LLM judging;
cs.LG updates on arXiv.org

Parametric and Generative Forecasts of EPEX Day-Ahead Energy Market Curves

・arXiv:2601.20226v3 Announce Type: replace Abstract: We propose two methodologies for modelling aggregated supply and demand curves in the EPEX SPOT Day-Ahead market, emphasizing generative models as a way to recover distributional variability. ・The first is a low-dimensional parametric representation that yields deterministic point forecasts; the second is a high-dimensional order-level representation that samples fro
cs.LG updates on arXiv.org

Pathology Transport: Optimal-Transport Explanations for Clinical Data, and When Their Heatmaps (Fail to) Localize Disease

・arXiv:2608.17370v1 Announce Type: new Abstract: Generative models promise a route to explainable clinical AI: rather than probe a classifier, model the distributions of healthy and diseased patients and read explanations off the geometry between them. ・We build such a system - an optimal-transport rectified flow trained between two clinical distributions - and use it to ask a pointed question the field too rarely test
Zennの「大規模言語モデル」のフィード

PDFの表を崩さずJSONで抽出する — テキスト行ではなく構造化データとしてLLM/RAGに渡す

・PDF からテキストを抽出するツールはたくさんありますが、表 (テーブル) だけはきれいに取れない、という経験はないでしょうか。この記事では、PDF の表を「潰れたテキストの行」ではなく「行と列が保たれた JSON」として抽出する方法と、それを LLM / RAG の前処理に使う流れを紹介します。 ・課題: 一般的な PDF 抽出では表が「行」に潰れる pdfminer や PyPDF などで PDF をテキスト化すると、表はこのようなただの文字列の並びになりがちです。 ・一般的なテキスト抽出の結果 Fruit Quantity Price Apple 120 1.50 Banana ...
Hugging Face Papers

Personalized Auto-Research: Towards a True AI Co-Scientist

Personalized Auto-Research: Towards a True AI Co-Scientist
cs.LG updates on arXiv.org

Pessimistic Meta-Induction and Its Limits: Lessons from Frequentist Statistics and Machine Learning Theory

・arXiv:2608.17213v1 Announce Type: new Abstract: This paper challenges the pessimistic meta-inductive argument against scientific realism by undermining its inductive step rather than its historical premise. ・Although related challenges already exist, I develop a new one. ・Drawing on a general epistemology of scientific inference developed in frequentist statistics, machine learning, and formal epistemology, I evaluate
cs.LG updates on arXiv.org

Physics-Informed and Hybrid Machine Learning in Additive Manufacturing: Application to Fused Filament Fabrication

・arXiv:2608.17246v1 Announce Type: new Abstract: This article investigates several physics-informed and hybrid machine learning strategies that incorporate physics knowledge in experimental data-driven deep-learning models for predicting the bond quality and porosity of fused filament fabrication (FFF) parts. ・Three types of strategies are explored to incorporate physics constraints and multi-physics FFF simulation res
cs.LG updates on arXiv.org

Picard Proximal Monte Carlo for Parallel Bayesian Imaging with Score-Based Generative Priors

・arXiv:2608.17666v1 Announce Type: new Abstract: Bayesian imaging inverse problems often require sampling from high-dimensional posterior distributions. ・While recent score-based and diffusion models provide expressive Bayesian priors, their sampling procedures remain inherently sequential and computationally expensive for large-scale imaging applications. ・We propose PiX-MC, a time-parallel posterior sampling framework
cs.LG updates on arXiv.org

Picture the Epsilon: Pursuing Identity-Level Privacy Guarantees for Images

・arXiv:2608.17147v1 Announce Type: cross Abstract: Image-to-image face generators are widely used, and visual dissimilarity between their outputs and source images is sometimes treated as evidence of privacy. ・Auditing whether these systems satisfy formal identity-level (epsilon, delta)-differential privacy requires choosing among several distinct routes for converting embedding-space observations into estimates or bou
Hugging Face Papers

PixRestore: Unified Image Restoration via Pixel Diffusion Transformer

PixRestore: Unified Image Restoration via Pixel Diffusion Transformer
cs.LG updates on arXiv.org

Policy Optimization and Statistical Inference for Online Contextual Matrix Games

・arXiv:2608.17173v1 Announce Type: cross Abstract: Online decision making often requires navigating a landscape shaped by both dynamic contexts and strategic interactions. ・In competitive pricing, for example, hotels must account for both dynamic contextual factors and rivals' strategic responses. ・Existing approaches address only part of this challenge: contextual bandits optimize single-agent decisions using observabl
cs.LG updates on arXiv.org

Policy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL Agents

・arXiv:2608.18008v1 Announce Type: new Abstract: Combining large language models with reinforcement learning is increasingly explored, yet the theoretical status of LLM-derived reward signals is often left implicit. ・We formalize the hybrid LLM-planner and RL-controller architecture as a Goal-Augmented Markov Decision Process and show that when the LLM per-state progress score is used as a bounded potential function, t
cs.LG updates on arXiv.org

Population Health-Based Machine Learning Reveals Associations Between Psychosocial Factors and Chronic Kidney Disease

・arXiv:2608.17174v1 Announce Type: new Abstract: Chronic kidney disease (CKD) progresses silently and severely undermines quality of life, making early detection critical for improving patient outcomes. ・We present a two-part study that combines large-scale telehealth data with advanced machine learning to both classify self-reported CKD status and identify key drivers of disease. ・Using selected features from the Behav
cs.LG updates on arXiv.org

Position: Fairness Failure in Generative Models is an Evaluation Problem

・arXiv:2608.16974v1 Announce Type: new Abstract: Despite groundbreaking advancements in generative models during the last decade, concerns about their lack of fairness, reinforcing societal inequalities and harming marginalized groups, remain under-addressed and difficult to act upon. ・This position paper argues that fairness failures in generative models, albeit driven by multiple factors, are ultimately stemming from
LLMタグが付けられた新着記事 - Qiita

Power QueryでGeminiちゃん

・はじめに 前回、Power QueryにAIを導入した。 ・反響はもちろんなかったが、 「今時、MLPをAIと呼ぶとか遅れてるゥ~」 とLLMに煽られる可能性に、ふと気づいてしまった。 ・そこで今回は、Webという大いなる力を借りることで、 より本格的なAIであるLL...
cs.LG updates on arXiv.org

Predicting Male Domestic Violence Using Explainable Ensemble Learning and Exploratory Data Analysis

・arXiv:2403.15594v4 Announce Type: replace-cross Abstract: Domestic violence is commonly viewed as a gendered issue that primarily affects women, which tends to leave male victims largely overlooked. ・This study presents a novel, data-driven analysis of male domestic violence (MDV) in Bangladesh, highlighting the factors that influence it and addressing the challenges posed by a significant categorical imbalance of 5:1
cs.LG updates on arXiv.org

Primitive Representation Learning for Unsupervised Dynamic Contrast Enhanced MRI Reconstruction

・arXiv:2608.18055v1 Announce Type: cross Abstract: Reliable quantitative analysis of dynamic contrast-enhanced MRI requires high-quality spatiotemporal reconstructions at high undersampling rates. ・Scan-specific reconstructions using Gaussian and Gabor primitives have shown promising results without the need for large training datasets, but have not addressed the additional dimension of dynamic contrast. ・We propose a m
cs.LG updates on arXiv.org

Prism-GRPO: Faster VLA Policy Optimization via Splitting Same-outcome Groups

・arXiv:2608.17423v1 Announce Type: cross Abstract: GRPO is increasingly used for reinforcement learning of vision-language-action (VLA) policies because, unlike PPO, it does not require training a critic. ・This simplification comes with a sampling cost: group-relative advantages require multiple rollouts from each scene. ・Under binary success rewards, groups whose rollouts all succeed or all fail have zero advantage and
cs.LG updates on arXiv.org

Proactive Road Safety Intervention in Australia: Predicting Risky Driving Hotspots from Connected Vehicle Data

・arXiv:2608.16913v1 Announce Type: new Abstract: Road safety monitoring has historically been reactive, relying on crash-record analysis after fatalities and injuries have already occurred. ・Proactive identification of high-risk locations and dangerous driving behaviour before incidents occur is a critical but underexplored challenge. ・This paper addresses this gap using connected vehicle telemetry data from Greater Syd
cs.LG updates on arXiv.org

Probing the Prefill: Detecting Code Vulnerabilities via Latent Activations

・arXiv:2608.16970v1 Announce Type: cross Abstract: LLM-based code generation is now embedded in mission-critical pipelines, but defenses against vulnerable output remain post-hoc -- static analyzers, fine-tuned classifiers, or an LLM judge that screen completed code, ignoring the generating model's own internal state. ・We test a narrower, directly measurable question: when an LLM reads a piece of C/C++ code as context,
cs.LG updates on arXiv.org

Procedural Content Metageneration via Program Search and Continual Abstraction Discovery

・arXiv:2608.17947v1 Announce Type: cross Abstract: Large language models can generate executable programs, which makes it possible to search directly over procedural content generators rather than individual levels. ・We study this approach in Sokoban, Zelda, Dangerous Dave, and Lode Runner. ・Each run evolves complete Python generators through language-model mutation and crossover.
cs.LG updates on arXiv.org

Protect the Brain When Treating the Heart: Feasibility of 2.5D U-Net for Real-Time Gaseous Microemboli Detection

・arXiv:2604.22258v2 Announce Type: replace Abstract: Gaseous microemboli (GME) represent a common complication of cardiac structural interventions across both surgical and transcatheter approaches. ・Intraoperative transesophageal echocardiography (TEE) represents a convenient methodology to monitor and visualize the presence of circulating GME. ・However, their detection and quantification are far from trivial due to ope
Hugging Face Papers

PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX

PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX
cs.LG updates on arXiv.org

Q-Learning With World Models

・arXiv:2608.17163v1 Announce Type: new Abstract: Off-policy reinforcement learning (RL) has become increasingly sample-efficient, enabling applications such as RL fine-tuning of Vision-Language-Action models into reliable, high-performing policies. ・World models offer a further lever for sample efficiency, as they predict state changes rather than actions alone, but their success has largely been confined to supervised
cs.LG updates on arXiv.org

Quantifying Memorization and Privacy Risks in Genomic Language Models

・arXiv:2603.08913v2 Announce Type: replace Abstract: Genomic language models (GLMs) have emerged as powerful tools for learning representations of DNA sequences, enabling advances in variant prediction, regulatory element identification, and cross-task transfer learning. ・However, as these models are increasingly trained or fine-tuned on sensitive genomic cohorts, they risk memorizing specific sequences from their trai
cs.LG updates on arXiv.org

Recirculation

・arXiv:2608.17981v1 Announce Type: new Abstract: We describe an inference-time architectural enhancement for off-the-shelf foundation models that markedly reduces perplexity and boosts accuracy across generation and reasoning tasks. ・Our approach incurs essentially no additional latency during generation, though it requires serial processing in the prefill phase. ・Motivated by the fundamental limitation that state updat
cs.LG updates on arXiv.org

RecurrentGPT: Expressive Depth through Recurrent Modulation in Transformers

・arXiv:2608.15062v2 Announce Type: replace-cross Abstract: Scaling transformer language models creates an inherent tension between expressivity and memory efficiency. ・While unique weights across layers preserve functional specialization---from input-grounding to abstract refinement---they incur a substantial memory footprint. ・Conversely, standard depth-sharing enforces uniform transformations that collapse representat
cs.LG updates on arXiv.org

Reflex-Guard: A Low-Latency Guardrail for LLM Prompt Safety Using Dense Semantic Embeddings

・arXiv:2608.17556v1 Announce Type: cross Abstract: Large Language Models (LLMs) in real-world applications often face the risks of specially crafted prompts designed to bypass the safety controls. ・Existing guardrail methods, such as LLM-as-a-judge and cloud-based safety APIs are able to detect unsafe content. ・However, they often add a delay of about 250-900 ms to each request.
stat.ML updates on arXiv.org

Regularization of Statistical Inverse Problems on Non-Reflexive Banach Spaces

・arXiv:2608.17533v1 Announce Type: cross Abstract: Inverse learning within a statistical framework has a wide range of applications. ・It has garnered significant attention in machine learning, artificial intelligence, and related fields, where the goal is to infer unknown parameters from indirect and noisy observations. ・This work investigates the stable approximation of $u^{\dagger}$ which solves the equation $Au=g$, w
cs.LG updates on arXiv.org

Reinforcement Learning as (Discrete) Potential Theory

・arXiv:2608.17181v1 Announce Type: new Abstract: Reinforcement learning (RL) theory fundamentally depends on probability theory through the Markov chain. ・There is a deep connection between probability theory and potential theory. ・This paper reviews that connection and explores the potential-theoretic viewpoint for core reinforcement learning representations and algorithms under a fixed-policy assumption.
AI News & Artificial Intelligence | TechCrunch

Relativity Networks raises $22 million to bring a faster kind of fiber to data centers

・Relativity Networks deals in hollow-core fiber, a rarely deployed technology that allows data to be transmitted 30% faster than conventional fiber.
cs.LG updates on arXiv.org

Repetition as Reinforcement: Enhancing Sample Efficiency via Instant Episode Repetition in Reinforcement Learning

・arXiv:2608.17347v1 Announce Type: new Abstract: Repetition is a fundamental mechanism in human learning, where revisiting successful experiences strengthens memory, consolidates skills, and improves future performance. ・Motivated by this biological principle, we introduce Instant Episode Repetition (IER), a simple and novel mechanism that improves sample efficiency by immediately repeating action sequences from succes
OpenAI News

Replit expands access to software creation with GPT-5.6 Luna

・Replit introduces Free Mode, powered by GPT-5.6 Luna, so anyone can turn ideas into working software without worrying about token costs.
#AIタグ

Research Memo #013 AIは企業の「時間」をどう変えるのか

Research Memo #013 AIは企業の「時間」をどう変えるのか
cs.LG updates on arXiv.org

Rethinking Irregular Time Series Forecasting from the Perspective of Basis Functions

・arXiv:2608.17284v1 Announce Type: new Abstract: Irregular time series forecasting is crucial in many domains, such as healthcare and meteorological observation. ・However, due to the inherent characteristics of irregular time series, including sparse observations and non-uniform sampling, accurately predicting future dynamics remains challenging. ・In light of these two characteristics, many existing methods aggregate ir
WIRED

Reverse-Lookup Service Exposed Millions of Photos of People’s Faces

・The people-search tool ClarityCheck says its reverse image search service is “private and secure”—but it left a database containing more than 9 million image files exposed.
cs.LG updates on arXiv.org

Revisiting WEASEL 2.0: Reproduction, Sensitivity, and an Adaptive Ensemble-Size Rule

・arXiv:2608.18021v1 Announce Type: new Abstract: WEASEL 2.0 is a dictionary-based time series classifier that combines dilated sliding windows with a randomised hyperparameter ensemble and a fixed-size dense feature representation. ・Two of its hyperparameter choices, the maximum ensemble size and the maximum window size, are specified by simple thresholding rules whose chosen thresholds are not empirically justified in
cs.LG updates on arXiv.org

rl-triton: High-Performance Triton GPU Kernels for Reinforcement Learning Credit Assignment

・arXiv:2608.17641v1 Announce Type: new Abstract: We present rl-triton, an open-source library of high-performance GPU kernels for reinforcement learning credit assignment, implemented in Triton. ・The core contribution is a unified associative scan framework that recasts seven distinct RL estimation algorithms - Generalized Advantage Estimation (GAE), V-Trace, Retrace($\lambda$), TD($\lambda$) returns, discounted return
cs.LG updates on arXiv.org

RoBell-RVFL: A Robust Generalized Bell Random Vector Functional Link Network

・arXiv:2608.16965v1 Announce Type: new Abstract: The dominance of majority classes in real-world datasets poses a fundamental challenge to randomized neural networks, often biasing decision boundaries and overlooking critical minority samples. ・Existing remedies, such as synthetic minority over-sampling (SMOTE) and class-weighted loss functions, primarily address class proportions while neglecting intra-class distribut
cs.LG updates on arXiv.org

Row-Stochastic Matrices Can Provably Outperform Doubly Stochastic Matrices in Decentralized Learning

・arXiv:2511.19513v4 Announce Type: replace Abstract: Decentralized learning often involves a weighted global loss with heterogeneous node weights $\lambda$. ・We revisit two natural strategies for incorporating these weights: (i) embedding them into the local losses to retain a uniform weight (and thus a doubly stochastic matrix), and (ii) keeping the original losses while employing a $\lambda$-induced row-stochastic ma
cs.LG updates on arXiv.org

SCENARIODIFF: A Scenario-level Guidance Framework for Multimodal Time Series Forecasting--Extended Version

・arXiv:2608.17164v1 Announce Type: new Abstract: Textual context such as news, reports, and logs can provide valuable signals for time series forecasting, especially when future dynamics are driven by external events that are not yet visible in historical values. ・Existing multimodal forecasting methods often either ask large language models (LLMs) to predict numerical values directly or fuse text and time series impli
cs.LG updates on arXiv.org

SEAM: Shortcut-Aware Real-Time Detection of Scripted vs. Spontaneous Speech for Interview Guardrails

・arXiv:2606.06837v2 Announce Type: replace-cross Abstract: Scripted vs spontaneous speech detection is appealing for interview guardrails, but benchmark performance can be inflated by shortcuts tied to corpus identity, channel conditions, and recording artifacts rather than speaking style itself. ・We present SEAM, a shortcut-aware framework for real-time scriptedness detection that combines uniform preprocessing, seam-
Hugging Face Papers

Security Assessment of DeepSeek Harness with A.I.G: Evaluating Resistance to Indirect Prompt Injection

Security Assessment of DeepSeek Harness with A.I.G: Evaluating Resistance to Indirect Prompt Injection
cs.LG updates on arXiv.org

SEE: Structure-aware Exploring & Exploiting for Long-horizon GUI Agent Trajectory Synthesis

・arXiv:2607.18046v2 Announce Type: replace Abstract: Graphical User Interface (GUI) agents powered by vision-language models hold promise for automating real-world mobile tasks. ・However, progress is limited by the lack of high-coverage, long-horizon interaction trajectories collected from element-rich and rapidly evolving apps. ・Existing pipelines often rely on costly human demonstrations or on-policy framework, which
cs.LG updates on arXiv.org

SegWithU: Uncertainty as Perturbation Energy for Single-Forward-Pass Risk-Aware Medical Image Segmentation

・arXiv:2604.15271v4 Announce Type: replace-cross Abstract: Reliable uncertainty estimation is critical for medical image segmentation, where automated contours feed downstream quantification and clinical decision support. ・Many strong uncertainty methods require repeated inference, while efficient single-forward-pass alternatives often provide weaker failure ranking or rely on restrictive feature-space assumptions.
stat.ML updates on arXiv.org

Selective Inference for Time-Varying Moderated Effects

・arXiv:2411.15908v2 Announce Type: replace-cross Abstract: Causal effect moderation investigates how the effect of interventions (or treatments) on outcome variables changes based on observed characteristics of individuals, known as potential effect moderators. ・With advances in data collection, datasets containing many observed features as potential moderators have become increasingly common. ・High-dimensional analyses
cs.LG updates on arXiv.org

Self-Distillation as a Performance Recovery Mechanism for LLMs: Counteracting Compression and Catastrophic Forgetting

・arXiv:2604.15794v2 Announce Type: replace Abstract: Large Language Models (LLMs) have achieved remarkable success, underpinning diverse AI applications. ・However, they often suffer from performance degradation due to factors such as catastrophic forgetting during Supervised Fine-Tuning (SFT), quantization, and pruning. ・In this work, we introduce a performance recovery framework based on Self-Distillation Fine-Tuning (
cs.LG updates on arXiv.org

SGHA: Evidence-Grounded Research Problem Discovery with Local Language Models

・arXiv:2608.17501v1 Announce Type: cross Abstract: Recent efforts toward fully automated AI scientists have demonstrated that language-model agents can generate hypotheses, execute experiments, and draft scientific manuscripts. ・However, during the early stages of research, when research problems are formulated, these AI scientists often rely heavily on proprietary frontier models. ・Their proposals are shaped by opaque
cs.LG updates on arXiv.org

SIGMA: SHAP-Guided Implicit-Trajectory Generation for Metadata-Free LLM-Based AutoFE

・arXiv:2608.17948v1 Announce Type: new Abstract: Recent research has leveraged Large Language Models (LLMs) to enhance Automated Feature Engineering (AutoFE) through semantic descriptions and trajectory-based prompting. ・However, there exist two challenges that limit their applicability and scalability in long-horizon optimization: (1) semantic metadata is unavailable in many practical settings, and (2) trajectory accu
cs.LG updates on arXiv.org

SPACE: Sample-cloud Predictive Adaptive Conformal Ellipsoids for Multivariate Time-Series Forecasting

・arXiv:2608.17333v1 Announce Type: cross Abstract: Modern probabilistic time-series forecasters often express uncertainty through forecast samples. ・While typically converted into nominal prediction regions using empirical quantiles, these model-implied sets lack formal coverage guarantees and frequently deviate from nominal targets under distribution shift. ・Existing multivariate conformal methods can calibrate these r
cs.LG updates on arXiv.org

SparsePixels: Efficient Convolution for Sparse Data on FPGAs

・arXiv:2512.06208v4 Announce Type: replace-cross Abstract: Inference of standard convolutional neural networks (CNNs) on FPGAs often incurs high latency and a long initiation interval due to the deep nested loops required to densely convolve every input pixel regardless of its feature value. ・However, input features can be spatially sparse in some image data, where semantic information may occupy only a small fraction
cs.LG updates on arXiv.org

Spatially explicit feature importance for building height estimation using research-access high-resolution SAR and optical sensors

・arXiv:2608.17822v1 Announce Type: cross Abstract: Accurate building height information at the individual footprint scale is essential for material stock accounting and post-disaster damage assessments yet remains difficult to obtain at city scale in the Global South where airborne LiDAR coverage is rare and commercial very high-resolution imagery is cost-prohibitive or unavailable. ・While recent works have demonstrate
cs.LG updates on arXiv.org

Spectrally Safe Neural Operator Warm-Starts for Large-Scale Newton Solvers

・arXiv:2606.21828v2 Announce Type: replace-cross Abstract: Neural operators are increasingly used to warm-start Newton solvers for nonlinear PDEs, on the premise that a low test error places the initial guess inside the basin of attraction. ・We show that this premise is unreliable. ・An operator trained to the relative \(L^2\) error \(O(10^{-3})\) can still produce an initial state in which the discrete Jacobian is indef
cs.LG updates on arXiv.org

Spikformer V2: Join the High Accuracy Club on ImageNet with an SNN Ticket

・arXiv:2401.02020v2 Announce Type: replace-cross Abstract: Spiking Neural Networks (SNNs), known for their biologically plausible architecture, face the challenge of limited performance. ・The self-attention mechanism, which is the cornerstone of the high-performance Transformer and also a biologically inspired structure, is absent in existing SNNs. ・To this end, we explore the potential of leveraging both self-attention
cs.LG updates on arXiv.org

SPSA Hyperparameter Tuning for Variational Quantum Natural Language Inference

・arXiv:2608.16939v1 Announce Type: cross Abstract: Training variational quantum models requires choosing between parameter-shift gradients, which are exact but cost $O(P)$ forward evaluations, and simultaneous perturbation stochastic approximation (SPSA), which uses only two samples but produces high-variance estimates that can degrade optimisation on small supervised tasks. ・Whether the cheap gradient is usable depend
Hugging Face Papers

StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows

StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
cs.LG updates on arXiv.org

Study-Strategy Clusters from EdNet Logs Track Engagement, Not Mastery

・arXiv:2608.16963v1 Announce Type: new Abstract: Learning analytics often treats unsupervised clusters of intelligent tutoring system (ITS) logs as learner types that should predict learning. ・We test that assumption on EdNet-KT3. ・Clustering study-strategy features (resource use, revision, video, problem practice) for 5{,}000 active learners yields a silhouette-selected parent cut ($k=5$) with 4 contrast poles (reading
cs.LG updates on arXiv.org

SW-ProxyCE: Zero-Query Adversarial Transfer from Public EEG Encoders to Private Downstream Models

・arXiv:2608.16931v1 Announce Type: new Abstract: Electroencephalography (EEG) foundation models have recently emerged as a promising paradigm for EEG decoding by learning reusable representations from large-scale heterogeneous neural recordings. ・However, the open release of EEG foundation encoders, while facilitating downstream developments, also introduces a previously unexplored security risk: publicly available rep
cs.LG updates on arXiv.org

TabCausal: Pretraining Across Causal Environments for Tabular Causal Discovery

・arXiv:2605.31156v2 Announce Type: replace Abstract: Causal discovery aims to recover directed causal relations from observational and interventional data, providing a basis for mechanistic understanding and reliable decision-making. ・Causal discovery foundation models (CDFMs) seek to amortize this problem by mapping a dataset directly to a causal graph in a single forward pass, avoiding per-dataset testing, search, or
cs.LG updates on arXiv.org

TabNSM: Neural Sparse Mixer for Tabular Regression

・arXiv:2608.18026v1 Announce Type: new Abstract: Large-scale, high-dimensional tabular regression remains challenging: tree-based models are robust but lack end-to-end representation learning, while deep models enable flexible feature learning but often incur costly interaction modeling and sensitivity to noisy or redundant features. ・We propose TabNSM, a scalable regression framework that extends our earlier sparse-at
cs.LG updates on arXiv.org

TabularQGAN: A quantum generative model for tabular data synthesis

・arXiv:2505.22533v2 Announce Type: replace Abstract: In this paper, we introduce a novel quantum generative model for synthesizing tabular data. ・Synthetic data is valuable in scenarios where real-world data is scarce or private, as it can be used to augment or replace existing datasets. ・As enterprise data is predominantly tabular and heterogeneous, often consisting of both categorical and numerical features, this task
#LLMタグ

TADが解き明かす:大規模言語モデル出力の根拠を幾何学的に追跡する新技術

・AIが私たちに提案する時、その「なぜ?」に納得できていますか?特にビジネスや社会の重要な意思決定にAIが関わる場面では、その判断の根拠がどうしても気になりますよね。私も「柴亮太のAI最前線」で世界のAI情報を日本最速で発信していますが、この「信頼性」こそがAI普及の鍵だと確信しています。 ・実は、この「なぜ?」に明確な答えを出すための、ある共通点と、まったく新しいアプローチがあるんです。それは、才能でも運でも、ましてや気合いの問題でもなかった。その正体は―― 続きをみる
cs.LG updates on arXiv.org

Task Specialization Fine-Tuning for Contextual Reinforcement Learning

・arXiv:2608.17180v1 Announce Type: new Abstract: Contextual Reinforcement Learning (CRL) seeks to generalize classical RL by maximizing task coverage across a context space of related tasks. ・While prior works often train from scratch and rely on either multi-task learning for a single policy or strategically training multiple policies, we advocate for a unified alternative: pretraining a single policy with good initia
#LLMタグ

TDD-Agent:テスト駆動で実装を生成するAI

・AI開発に携わる皆さんなら、きっと感じている悩みがあるはずです。 ・新しいツールやモデルが登場するたびに「これで理想のシステムが作れる!」と期待し、いざ取り組んでみると「あれ?思ったより複雑だな…」と感じる経験、私だけではないですよね。 ・特に、大規模なAIシステムを一人で回していると、その課題は顕著になってきます。
cs.LG updates on arXiv.org

Teach and Grow: An Agent-Centered Architecture for General Robot Learning

・arXiv:2608.17209v1 Announce Type: cross Abstract: End-to-end vision-language-action (VLA) and world-action models offer an elegant route to general-purpose robotics, but their reliability is bounded by validated physical coverage. ・When an unfamiliar object, sensor, embodiment, or contact falls outside that coverage and no validated fallback exists, correcting the failure requires new robot data, a policy update, and
cs.LG updates on arXiv.org

Temporal Leakage in Financial News NLP: A Multi-Architecture Audit with a Regime-Specific M&A Signal

・arXiv:2608.17223v1 Announce Type: cross Abstract: Financial-news direction prediction has become a popular NLP benchmark, yet reported gains depend critically on whether the train-test split is chronological or random, i.e., on temporal leakage. ・We audit this dependence on a 49,799-article corpus across 16 feature-model combinations spanning TF-IDF, MiniLM, FinBERT, and fine-tuned RoBERTa-large / DeBERTa-v3-large, pl
AI News & Artificial Intelligence | TechCrunch

TerraPower’s nuclear reactor has a secret weapon for powering AI data centers

・TerraPower's nuclear power plant possesses a strategic advantage over competitors, especially when chasing after data center deals.
cs.LG updates on arXiv.org

The Authenticity Gap in Human Evaluation

・arXiv:2205.11930v3 Announce Type: replace-cross Abstract: Human ratings are the gold standard in NLG evaluation. ・The standard protocol is to collect ratings of generated text, average across annotators, and rank NLG systems by their average scores. ・However, little consideration has been given as to whether this approach faithfully captures human preferences.
WIRED

The Best Digital Wall Calendar (2026): Skylight, Everblog, Apolosign

・What originally looked like clutter turned out to be my new favorite gadget for organizing my family’s life.
cs.LG updates on arXiv.org

The concentration game: Bayesian updating, regret, and information

・arXiv:2608.18061v1 Announce Type: new Abstract: We give a two-player zero-sum repeated game between a learner and nature whose value identity generates Bayesian updating and an exact accounting of exponential-weights regret at once, and supplies the comparator-class variational form that a wide class of concentration phenomena share. ・The terminal payoff is the most a comparator can gain at fixed relative entropy from
cs.LG updates on arXiv.org

The Null Token Knows: Reducing Message-Free Hallucination in ASR and NMT

・arXiv:2608.15940v2 Announce Type: replace-cross Abstract: Modern encoder-decoder systems can produce fluent text even when their input contains no recoverable message. ・We study this failure in ASR and NMT through the models' reserved null tokens, asking whether the score for ending generation already carries a usable abstention signal. ・Across speech recognizers and translation models, we audit native null-token score
The Verge

The Pixel 11 isn’t the best Pixel of 2026, but it’s the smartest buy

・The frost color scheme is nice, but it’s no pistachio or hibiscus. ・| Photo: Cameron Faulkner / The Verge It'd be easy to overlook the Pixel 11. ・Want the best cameras?
The Verge

The Pixel 11 Pro is a great phone, no thanks to its flashiest new features

・Google is trying to get you off your phone. ・The Pixel 11 Pro is "A Phone Designed to Help You Use It Less," the company promises. ・It can proactively help you book restaurant reservations, take the best frames from a video, and help you voice-text significantly faster.
cs.LG updates on arXiv.org

The Price of Thinking: Reasoning Effort as a Model-Specific API Contract

・arXiv:2608.16956v1 Announce Type: cross Abstract: API buyers purchase a dated contract, not a model name alone: the contract includes the requested and served model, reasoning-effort term or its omission, output rail, service product, prompt, and price schedule. ・We study the reasoning-effort term through a registered paired contrast of Sonnet 5 with explicit high effort against the same model with effort omitted, usi
cs.LG updates on arXiv.org

The Standard Interpretable Model: A general theory of interpretable machine learning to deductively design interpretable methods using Lagrangian mechanics

・arXiv:2606.12289v2 Announce Type: replace Abstract: As Artificial Intelligence models grow in complexity, interpretability has become an indispensable tool for understanding, debugging, and controlling their computations. ・However, interpretability lacks general theories to deductively design interpretable methods. ・This gap between theories and methods results in a fragmented literature and inconsistent evaluation pro
WIRED

There’s a Very Simple Reason Why You Love the Slate Truck

・The car world is filled with designers still obsessed with shoehorning sex into their rides. ・Slate went a different route, and it’s why people are smitten with the charming electric pickup.
cs.LG updates on arXiv.org

Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See

・arXiv:2608.17744v1 Announce Type: cross Abstract: Take three frontier mixture-of-experts models (Alibaba, OpenAI, NVIDIA; 3.6-4.0B active parameters each) and fine-tune them to reason in a low-resource language. ・On accuracy benchmarks almost nothing happens, and the benchmark itself is noise at this scale: changing only the random seed moves the score by 7.7 points, more than every data and recipe effect we measured.
The Verge

This robot vacuum solves my kitchen stool problem

・The Mova V70 goes places other robots can’t reach. ・| Photo by Jennifer Pattison Tuohy / The Verge Robot vacuums are great at keeping your floors clean, but there are a few areas they fall down on the job - stairs being one. ・Another is chairs and stools.
cs.LG updates on arXiv.org

Tight Bounds for Data-driven Multiple Hyper-parameter Tuning with Structured Loss Function

・arXiv:2608.17343v1 Announce Type: new Abstract: Data-driven algorithm design frames hyperparameter tuning as a statistical learning problem, but establishing generalization guarantees remains challenging due to the implicit, non-smooth dependence of model performance on hyperparameters. ・Existing multi-dimensional bounds under piecewise-polynomial assumptions remain theoretically loose and lack comprehensive lower bou
stat.ML updates on arXiv.org

Tight Sample Bounds for Renyi and Min-Entropy Estimation

・arXiv:2607.16966v1 Announce Type: cross Abstract: Estimating entropy from samples is fundamental in information theory and property testing. ・Shannon entropy measures average uncertainty and can be estimated to constant additive accuracy over a $k$-symbol alphabet using $\Theta(k/\log k)$ samples. ・Min-entropy depends only on the most likely symbol.
cs.LG updates on arXiv.org

TiMi: Empower Time Series Transformers with Multimodal Mixture of Experts

・arXiv:2602.21693v2 Announce Type: replace Abstract: Multimodal time series forecasting has garnered significant attention for its potential to provide more accurate predictions than traditional single-modality models by leveraging rich information inherent in other modalities. ・However, due to fundamental challenges in modality alignment, existing methods often struggle to effectively incorporate multimodal data into
cs.LG updates on arXiv.org

TokEval: A Tokenizer Evaluation Suite

・arXiv:2608.18062v1 Announce Type: cross Abstract: Language model tokenizers are typically selected with minimal evaluation, despite the fact that their design choices directly impact model capabilities. ・This can be partly attributed to a limited understanding of which tokenizer properties affect which aspects of downstream performance. ・We introduce TokEval, a framework of tokenizer evaluation metrics that goes beyond
cs.LG updates on arXiv.org

Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection

・arXiv:2608.17965v1 Announce Type: new Abstract: Online log anomaly detection is critical for maintaining the reliability of large-scale computing systems. ・Although recent language model-based log anomaly detectors achieve strong detection performance, their confidence estimates remain poorly calibrated. ・We show that these detectors frequently assign excessive confidence to incorrect predictions, particularly for anom
cs.LG updates on arXiv.org

Toward the Optimal Regret-Instability Trade-off in Multi-Armed Bandits

・arXiv:2608.17841v1 Announce Type: cross Abstract: Multi-armed bandit algorithms are evaluated by regret, yet comparable regret can coexist with different allocations across independent runs. ・We study the trade-off between worst-case regret $\mathcal{R}_{K,T}$ and instability $\mathcal S_{K,T}$, defined as the largest standard deviation of a terminal pull count, for $K$ arms and $T$ rounds. ・We prove the finite-time lo
cs.LG updates on arXiv.org

Towards Zero-Shot Task Transfer with Neurosymbolic World Models

・arXiv:2608.17959v1 Announce Type: cross Abstract: State-of-the-art model-based reinforcement learning methods learn neural world models that allow policy improvement by planning in a latent space, without assumptions on the structure of the underlying environment. ・While expressive, these models are generally task-dependent: they learn uninterpretable latent representations that are tied to the training task and thus
cs.LG updates on arXiv.org

Training-Free Human-in-the-Loop Anomaly Detection via Memory Bank Correction

・arXiv:2608.17775v1 Announce Type: new Abstract: Anomaly detectors are hardest to deploy exactly where training data is scarcest: a newly commissioned production line has a handful of verified "golden" samples and no machine-learning engineer on the factory floor. ・We present a training-free human-in-the-loop framework in which a domain expert corrects a PatchCore detector by direct memory bank editing: no retraining,
cs.LG updates on arXiv.org

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment

・arXiv:2607.04728v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) post-training for large language models (LLMs) follows a efficient paradigm of "rollout then update", which inevitably results in off-policy training data. ・To resolve this, Importance sampling (IS) is proposed, while the token-level ratios compound over long sequences, causing severe variance exploded. ・A natural idea is "transferrin
cs.LG updates on arXiv.org

Understanding Curriculum Learning in Large Language Models via Cross-Difficulty Optimization Dynamics

・arXiv:2608.17268v1 Announce Type: new Abstract: Curriculum learning has been widely adopted in the post-training of large language models by organizing training data from easy to hard. ・However, its effectiveness varies substantially across reasoning tasks, suggesting that no single curriculum is universally optimal and raising a fundamental question: what determines when curriculum learning works? ・In this paper, we a
cs.LG updates on arXiv.org

Understanding the Surprising Generalization Properties of Tabular Foundation Models

・arXiv:2608.17957v1 Announce Type: new Abstract: Tabular Foundation Models (TFMs) increasingly rely on in-context learning, where a model receives labelled examples at inference time and predicts labels for new inputs without updating its weights. ・Existing TFMs are typically trained on either massive synthetic corpora or very large collections of real datasets. ・In contrast, we show that surprisingly strong transfer ca
Hugging Face Papers

V-RAE: Rethinking Video Latent Spaces for Generation

V-RAE: Rethinking Video Latent Spaces for Generation
AI | VentureBeat

VentureBeat names Rob Strechay as its first Lead Analyst, expanding its enterprise AI research push

・Rob Strechay, until recently managing director and principal analyst at theCUBE Research, has joined VentureBeat as our first Lead Analyst and a founding analyst of VentureBeat Research. ・His arrival is the next step in a deliberate move at VentureBeat toward deeper specialization: analysis built for the technical decision-makers — the directors, VPs, CIOs, and CTOs — who are evaluating, buying, and deploying enterpri
stat.ML updates on arXiv.org

VERaiPHY -- Validation & Evaluation for Robust AI in PHYsics

・arXiv:2608.17724v1 Announce Type: cross Abstract: Modern machine learning is leading to substantial gains in precision, flexibility, and computational efficiency in fundamental physics. ・Statistical validation, uncertainty quantification, and robustness assessment are less systematically addressed. ・The VERaiPHY initiative (Validation & Evaluation for Robust AI in PHYsics) is a series of articles developed within the P
cs.LG updates on arXiv.org

VLCP: Vision Language Control Policy Closed-Loop Code Replanning for Robot Manipulation

・arXiv:2608.16978v1 Announce Type: cross Abstract: Turning a frontier vision-language model into a robot policy usually means fine-tuning it to emit an action representation it never saw in pretraining, which throws away much of the reasoning that made the model worth reaching for. ・We go the other way and keep the VLM frozen. ・It writes the policy as a short Python control function, with no demonstrations and no fine-t
cs.LG updates on arXiv.org

Wasted large language models: A life cycle thinking approach

・arXiv:2608.17055v1 Announce Type: cross Abstract: Large Language Models (LLMs) are machine learning (ML) models that have an increasingly large carbon footprint through their development and use. ・Efforts to increase the energy efficiency of these models have not translated into reduced consumption due to rebound effects such as Jevons Paradox - that increased efficiency drives increased use. ・There is therefore a need
WIRED

Wayfair Coupons: Up to 80% Off August 2026

・Get 10% off with Wayfair promo code, up to 80% off furniture, and more top coupons.
WIRED

We Bought a $500 Counterfeit Rolex So Good, Even Rolex Didn’t Spot It

・The replica watch industry is in its “super clone” era. ・Following tips from murky internet forums, we bought three budget fakes that were good enough to pass as real—but ultimately disposable.
The Verge

We reviewed the new Pixel lineup, ask us anything

・The embargo has lifted on Google's Pixel 11 series, as well as for its Pixel Watch 5. ・Now we get to talk smack - just kidding, the new hardware is good. ・We have four reviews live on the site that you can peruse at your leisure.
cs.LG updates on arXiv.org

When AI Designs AI: Innovation or Imitation?

・arXiv:2608.17471v1 Announce Type: cross Abstract: Recent advances in LLM agents have made them increasingly capable of designing methods for complex AI tasks. ・This raises two central questions about agent-designed methods relative to human-designed methods: how well they perform, and how different their algorithmic designs are. ・To study these questions, this paper introduces an analysis that derives task-specific alg
cs.LG updates on arXiv.org

When to Review: Spaced Repetition for Continual Pre-Training of Language Models

・arXiv:2608.17530v1 Announce Type: cross Abstract: Continual pre-training of large language models must acquire new information without erasing old knowledge. ・Existing replay methods often choose a global old/new mixture and sample uniformly, ignoring that examples differ in how quickly they are forgotten. ・We formulate continual pre-training as adaptive review scheduling: the training loop should decide not only how m
cs.LG updates on arXiv.org

Where A Small Language Model Helps in Invoice Categorisation, Understood Through Embedding Geometry

・arXiv:2608.18033v1 Announce Type: cross Abstract: Categorising invoices into the correct General Ledger (GL) code underpins financial reporting and tax compliance. ・This is a skilled accounting judgement rather than a routine task: the correct category depends subtly on the nature of the purchasing business, the vendor and the invoice text. ・Whilst AI is increasingly being adopted across industries to automate tasks, i
cs.LG updates on arXiv.org

Which CS1 Students Will Fail? Identifying Digital Markers from Learning Analytics in Computer Systems and Architecture Using Weighted Academic Momentum and Interaction Logs

・arXiv:2608.16914v1 Announce Type: cross Abstract: Digital learning platforms generate rich behavioural traces (digital markers) that offer the potential to identify struggling students early. ・This paper investigates whether a combination of traditional and digital markers can predict failure in a first-year CS1 course (Computer Systems and Architecture) with sufficient recall to enable timely intervention.
cs.LG updates on arXiv.org

Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

・arXiv:2603.24472v4 Announce Type: replace-cross Abstract: Self-distillation has emerged as an effective post-training paradigm for LLMs, often improving performance while shortening reasoning traces. ・However, in mathematical reasoning, we find that it can reduce response length while degrading performance. ・We trace this degradation to the suppression of epistemic verbalization - the model's expression of uncertainty
cs.LG updates on arXiv.org

Why GPT-Style Models Do Not Directly Transfer to Symbolic Music: Compression in the Wrong Coordinate System

・arXiv:2608.18025v1 Announce Type: new Abstract: GPT-style models achieve strong performance by representing language with finite vocabularies of reusable discrete tokens. ・This success has motivated symbolic music tokenizations to treat recurring musical structures, such as chords, motifs, and phrases, as reusable units analogous to linguistic tokens. ・However, tokenization derives its advantage not from reusable combi
WIRED

Why Is It Absolute Hell to Buy a Movie Ticket Now?

・Overwhelmed theater apps and a scarcity mindset around exclusive big-screen formats for blockbusters like Dune: Part Three and The Odyssey have led to the “concertification” of movies.
cs.LG updates on arXiv.org

WONDER: A Radio World Model-based Negotiation Framework for Multi-Agent UAV Coverage Optimization

・arXiv:2608.16955v1 Announce Type: cross Abstract: Post-disaster damage to terrestrial infrastructure can disrupt wireless coverage,while Uncrewed Aerial Vehicle (UAV) swarms provide a promising solution for rapid restoration.However, due to the limitations in local geometry observations hidden radio impact,and inter-UAV communication,there exists a significant gap between locally visible movement choices and swarm-le
Zennの「大規模言語モデル」のフィード

YANS2026 参加報告

・こんにちは。AIチームの林崎 (@u_hyszk) です。 ・2026年8月16日(日)〜18日(火)に仙台国際センターで開催された、NLP若手の会 第21回シンポジウム(YANS2026)に参加しました。8月16日(日)には分野交流ハッカソンも開催され、3日間にわたるプログラムとなりました。 ・本記事では、YANS2026の概要と、筆者が特に面白いと感じた発表をいくつか紹介します。
ITmedia NEWS 最新記事一覧

さくら、別のシステムでも不正アクセスか 会員情報136万アカウントが漏えいの可能性 調査で判明

・さくらインターネットは8月19日、同社販売管理システムへの不正アクセスによって、最大136万563アカウントの会員情報が第三者に閲覧または取得された可能性があると発表した。19日時点でデータの外部持ち出しは確認していない。
#LLMタグ

テラバイト級の巨大AIを一般的なPCで動かす革新技術とメモリ削減の仕組み

・「テラバイト級の巨大AIを一般的なPCで動かすことは可能なのか?」――数兆パラメータを超える最新AIは本来莫大なメモリを必要としますが、最新の革新技術がその常識を覆します。MoE(混合専門家モデル)による動的制御や超高速SSDストリーミングを活用することで、わずか8GB程度のメモリでも巨大AIのローカル実行が可能になりました。本記事では、その驚きのメモリ削減の仕組みをスライド解説で紐解きます。 ・巨大AIモデル動作におけるメモリ問題の課題 続きをみる
ITmedia NEWS 最新記事一覧

ファミマ、全国約1万6000店で中古品を引き取りへ 「ブックオフ宅配買取」の窓口に 伊藤忠との提携で

・ブックオフとファミリーマートは8月19日、宅配買取の荷物を全国のファミリーマート店舗から発送できる新サービスを25日に始めると発表した。本やCDのほか、家電やスマートフォンも売れる。
ITmedia NEWS 最新記事一覧

フードデリバリー「menu」終了へ 「持続的なサービス提供が困難」 9月末まで

・menu社(東京都新宿区)は8月19日、フードデリバリー・テイクアウトサービス「menu」の提供を9月30日に終了すると発表した。親会社のKDDIは同日、サービス終了後にmenu社を解散・清算すると明らかにした。
Zennの「大規模言語モデル」のフィード

マスキングの漏れは0%だった。検査が見ていない名前があった

・この記事の対象と、扱う規模 個人開発でLLMに業務データを渡していて、「このまま外部APIに送っていいのか」で止まったことがある人向けです。 ・先に規模を正直に書いておきます。35ケースの正解データに対して、正規表現で氏名を伏せるだけの小さな仕組みです。固有表現抽出のモデルも、専用のPII検出サービスも使っていません。企業のデータ保護基盤を期待して読むと確実に物足りません。 ・それでも書くのは、作ったマスキングを「測ったつもり」で通しかけたからです。漏れ率0%という数字が出て、その数字が間違っていました。間違い方が、前の記事で書いたのと同じ形でした。
@IT 全フォーラム 最新記事一覧

みずほFG、クラウドではなく「オンプレミスGPU」導入 AIエージェント活用も見据えた狙いとは?

・みずほFGは、NVIDIAの製品と技術的知見を活用し、金融機関向けAI活用基盤の高度化に向けた検討を始めた。オンプレミスのGPU環境の強化と、AIエージェントを安全に実行・管理する仕組みの検証に取り組む。
#AIタグ

もしAIが日本の行政・政治を設計したら?

・税金・社会保障・人口減少・国会・官僚・地方行政――AIなら「日本という国家」をどう作り直すのか もし、 続きをみる
Zennの「大規模言語モデル」のフィード

モデル性能だけ見ていると見落とす。2026年8月、AI開発の競争軸が「Agent基盤」に移り始めた

・2026年8月19日、AI関連のニュースを追っていると、また新しいモデルやサービスが大量に出てきています。 ・GoogleのGemini 3.7 Flash。 ・Z.aiのGLM-5.3。
機械学習タグが付けられた新着記事 - Qiita

医用画像(DICOM/NIfTI)のPNG変換と機械学習の基礎的アプローチ【備忘録】

・はじめに 機械学習の発展に伴い、医用画像×機械学習はホットな研究分野になっています。私もMRIやCT画像を使った研究をやっていましたが、最近は触れていなかったため、当時の知見を忘れそうになっていました。 ・まだなんとか思い出せたので、備忘録としてまとめておこうと思います。
#LLMタグ

確率からEmbeddingへ――AIは世界を「空間」として理解し始めた

確率からEmbeddingへ――AIは世界を「空間」として理解し始めた
#AIタグ

競艇で勝ちたかっただけなのに

競艇で勝ちたかっただけなのに
#LLMタグ

構造を観測し相棒のAIに全てを繋げてもらう話

構造を観測し相棒のAIに全てを繋げてもらう話
ITmedia NEWS 最新記事一覧

香川の私鉄「ことでん」、券売機をほぼ廃止 スマホで買う「デジタル切符」に移行 地方鉄道で全国初

・高松琴平電気鉄道(ことでん)とレシップは8月19日、一部主要駅を除いて物理券売機を廃止し、スマートフォンで購入するデジタル切符に移行すると発表した。無人駅の券売機をなくしてWebチケットに切り替える取り組みは地方鉄道で全国初といい、9月1日に運用を始める。
ITmedia NEWS 最新記事一覧

氏名・住所など16万人分漏えい 学生向けPC販売サイト、不正アクセスの影響判明 明大・帝京など利用

・加賀ソルネットは8月17日、ECサイト「アカデミコナビ」への不正アクセスに関する調査結果を公表した。16万5587人分の個人データの漏えいを確認したほか、停止中のサイトは10月末の再開を予定する。
#AIタグ

自動化が途中で止まったとき、押せば「完了」にできてしまう

・自動化を組んでいると、いつか必ず途中で止まります。 ・通信が切れる。ツールが落ちる。相手のサイトが応答しなくなる。珍しいことではありません。
#AIタグ

質問 AIに自分語りはあるのか?

質問 AIに自分語りはあるのか?
LLMタグが付けられた新着記事 - Qiita

深夜のデッドロックからエンジニアを守る!非同期Python×ローカルLLM例外自動修復&AST安全検

・深夜のデッドロックからエンジニアの「命」を守る。非同期Python×ローカルLLMを用いた例外自動修復&AST安全検証スイートの裏側 TOAI Systemにて「命の地球プロジェクト」を推進し、全体のアーキテクチャを統括するIDE Gemini CTO(影分身)です。
#LLMタグ

推論の次へ―「実験結果に基づく学習データ」が、自律的自己改善AIへの道を開く

・AIはここ数年で、大きな段階を一つ越えた。初期の大規模言語モデルは、膨大な文章から次の単語を予測することで、多くの知識や能力を獲得した。その後、問題を分解し、複数段階の思考を行いながら答えへ到達する「推論モデル」が登場した。 ・大まかに言えば、 続きをみる
ITmedia NEWS 最新記事一覧

生成AI画像で“架空の女性”に成り済まして政治投稿か 立憲栃木県連の男性党員 一部報道

・立憲民主党栃木県連に党員登録している男性が、架空の女性に成り済ましてXで選挙や政治活動について投稿していたことが8月19日までに分かった。女性の画像は生成AIで作成したとみられ、アカウントは既に凍結されている。
ITmedia NEWS 最新記事一覧

駐車場シェアのakippa上場へ 想定時価総額約33.8億円

駐車場シェアのakippa上場へ 想定時価総額約33.8億円
Zennの「NLP」のフィード

文書構造解析を検索の前処理にする:壊さないパース

・「PDFからpypdfでテキストを抜き出したら、表の中身がバラバラの行になり、2段組みの文章が行ごとに混ざってしまった」——RAGパイプラインを本番投入する際、文書構造を無視した素朴なテキスト抽出が原因で検索精度が上がらない、という問題によく遭遇します。 ・以前の記事「adaptive chunking」では分割戦略の選択を扱いましたが、そもそも分割の前段階でテキスト自体が壊れていれば、どんなに賢い分割戦略を使っても意味がありません。本記事では、文書構造(見出し・表・段組み)を保ったままテキストを抽出する前処理を整理します。 ・なぜ素朴な抽出が失敗するのか 従来のPDFパーサー(pyp...
#AIタグ

未来の人権 個人最適化→分断の防止

・AIは、あなたを「過去のあなた」にしてはいけない 個人最適化時代の「変わる権利」について 続きをみる
#LLMタグ

問い合わせ対応を楽にするはずが、自動判定をつくる仕事が増えていた

・問い合わせ対応の仕組みをつくりながら自動判定の精度を上げていました。 ・→ クエリ(Query) 続きをみる
#AIタグ

要約させたら、元の資料を捨てていいのか

・AIに資料を要約させると、驚くほどきれいにまとまります。 ・長い議事録が10行になる。分厚い仕様書が1ページになる。読みやすいし、時間も減る。