ai Trend Report

Dashboard へ戻る
Date: 20260730 Articles: 371 Scope: curated summary

あなたのアイデアを、今すぐ形に。

公開先に迷ったら、WebFileBinで一発公開。

HTMLをドラッグ&ドロップするだけで、すぐ公開できます。

3
High impact
5
Mid impact
363
Signal watch

なぜこのサイトを作ったのか

私たちそれぞれが個別にAIを使って情報収集し、同じような3行要約を作るたびに、世界中で膨大な電力と計算リソースが消費されています。 本プロジェクトは、あらかじめ広範な情報を取得・集約しておくことで、個別のAI実行回数を減らし、地球環境(GPU/TPU負荷)に配慮した効率的な情報収集を目指す実験的なダッシュボードです。

Domain filters
Star filters
Zennの「機械学習」のフィード

【技術解説】時系列分析で不労所得を得る!ケリー基準を用いた投資ポートフォリオの最適ベットサイズ計算

・時系列分析によるケリー基準を用いた投資ポートフォリオの最適ベットサイズ計算 この記事では、投資のポートフォリオ管理におけるリスクとリターンの最適なバランスを取るための手法として、ケリー基準を用いた最適ベットサイズの自動計算について解説します。Pythonを使用し、時系列データに基づいて実装します。必要なライブラリには、yfinance、numpy、pandas、matplotlibを使用します。 ・投資のリスクとリターンを理解する 投資においてリスクとリターンはトレードオフの関係にあります。高いリターンを期待すると、その分リスクも高くなります。ケリー基準を用いることで、資本の効率的...
Zennのトレンド

TypeScript 7 時代の Vue.js ツールチェーン Vize を実プロダクトで検証した

・プロダクト開発において、開発環境・CI の高速化はとても重要な要素です。 ・ユニークビジョン株式会社でも、Vue.js・Hono を使った TypeScript 製の社内プロダクトにおいて、Formatter を Oxfmt に、Linter を一部 Oxlint に移行するなどして開発効率の向上を図ってきました。 ・そんな中 TypeScript 7.0 がついに正式にリリースされました。コンパイラが Go 言語に移植され、型チェック速度が大幅に向上しました。
Zennの「大規模言語モデル」のフィード

リリース前チェックをAIで行う「プロダクトリリースハーネス」のつくり方

・こんにちは、株式会社estie(エスティ) 取締役CTO の Nari(@tiwanari)です。 ・みなさんは、新しいプロダクトを世に出す前に、「抜け漏れ」がないか不安になった経験はないですか? 先日、あるプロダクトのリリース前チェックで「個人情報がエラー監視ツールへ送られる実装になっている」と指摘されました。見つけたのは人間ではなく、リリース前チェックリストを読ませた AI のスキャンです。 ・よく考えるとケアしなければいけない当たり前なことでも、リリース前の慌ただしい時期にあらゆる観点を思い出すことはとても大変です。ゼロからプロダクトを開発してリリースした経験のある方なら、XSS(クロ...
@IT 全フォーラム 最新記事一覧

API設計の常識が変わる? HTTP新メソッド「QUERY」とセキュリティ対策を学ぶ

・APIはWebサービスや生成AIアプリケーションを支える重要な基盤だ。「IETF」がHTTP新メソッド「QUERY」を標準化プロセスへ進めるなど、API設計は変化しつつある。一方で、攻撃への備えも欠かせない。API設計とセキュリティを理解するための記事を紹介する。
ITmedia NEWS 最新記事一覧

Meta、売上高28%増も純利益14%減 AI投資拡大で設備投資は最大1450億ドルへ

・Metaの4?6月期決算は、売上高は前年同期比28%増の608億100万ドル、純利益は14%減の158億4800万ドルとなった。訴訟費用や人員削減に伴うコスト増が利益を押し下げたものの、広告事業は引き続き好調で、AI投資の成果も事業拡大を支えている。BlackRockとのデータセンター共同開発など、「Meta Compute」戦略の下でAIインフラへの投資を加速する方針だ。
ITmedia NEWS 最新記事一覧

MetaのザッカーバーグCEO、WSJ寄稿で「超知能は全員のものであるべき」

・Metaのマーク・ザッカーバーグCEOはWall Street Journalに寄稿し、「superintelligence」(超知能)は特定の機関に集中させず広く分散普及させるべきだと主張した。権力の集中によるリスクや司法の公平性、雇用拡大に触れ、オープンな普及が安全と発展につながると強調。MicrosoftのナデラCEOらも賛意を示した。
ITmedia NEWS 最新記事一覧

OpenAIやAnthropicなどの従業員、米政府に「AI開発のペース調整を」と提言

・OpenAIやGoogleなどの従業員1000人以上が、AI開発のペース調整に向けた国際的支援を米政府に求める公開書簡を発表した。AI自律化の急速な加速に伴う制御不能リスクを指摘し、開発速度の調整に必要なツール開発を訴える。企業主導のオープンモデル規制回避を求める動きとは対照的な提起となった。
@IT 全フォーラム 最新記事一覧

メインフレームに眠る「年代物のCOBOLコード」、AIでも苦戦する"読みにくさ"の正体

・複数の国産メインフレームベンダーが撤退方針を明らかにする中で、レガシーシステムの刷新は喫緊の課題だ。一方ではAIエージェントの存在がモダナイズの在り方にも影響を及ぼし始めている。40年物のシステムにどう向き合うべきか。アクセンチュアの中野氏に聞いた。
ITmedia NEWS 最新記事一覧

“ダサい”不評の「ドコモの銀行」、従来のアプリアイコン選択可能に 「ご意見・ご要望も踏まえ」

・新アイコン公開後、SNSで「ダサい」「使いたくない」「視認性が悪い」といった批判が噴出。「住信SBIネット銀行時代のアイコンを残してほしい」という声も相次いでいた。
ITmedia NEWS 最新記事一覧

“脳のゴミ”は鼻から捨てられていた? 老化で詰まる排出路、薬を鼻から投与で改善 韓国主導チーム研究

・韓国の基礎科学研究院やKAISTなどに所属する研究者らがCellで発表した論文「CSF clearance through arachnoid fenestrations to olfactory meningeal lymphatics」は、脳の老廃物を運ぶ「脳脊髄液」が、嗅球(においを感じる領域)周辺のくも膜の穴から鼻へと抜けて排出される詳細な経路を解明した研究報告だ。
Zennのトレンド

[gamification] DIVER OSINT CTF 2026 Writeup

・2026年7月25日〜6日に24時間開催されたDIVER OSINT CTF 2026にsamと2人で参加していました。(チーム名:gamification) 最終順位は18位/867チームでした。 ・チーム人数制限の上限6人で出場しているチームも多く、24時間以内という制約もある中で、2人でこの順位に食い込めたのは大健闘した(つもりです) 昨年のDIVER OSINT CTFも出場しましたが、この一年間を通して他のOSINT CTFに参加したり、GeoGuessrを極めるなどの鍛錬が功を奏し、昨年と比較して成長を感じる瞬間が多かったのがとても嬉しかったです。 ・改めてOSINT調査の楽...
@IT 全フォーラム 最新記事一覧

「AIが書いたコードは人間より優秀」――だが本番で壊れる、なぜか?

・AIコーディングの実践拡大に伴い、開発現場では品質管理の課題も顕在化し始めている。New Relicの調査では、AI生成コードに対する評価が高いことが分かったが、一方では本番環境では問題への対応に追われている状況も明らかになった。
@IT 全フォーラム 最新記事一覧

「AIコーディングより人を雇う方が安くなる」 2028年までに起こる逆転現象を防ぐ5つの運用施策

・Gartnerは2028年までにAIコーディングのコストが開発者の平均給与を上回るとの予測を発表した。LLMのトークン消費量増加と従量課金制ライセンスへの移行が要因で、ソフトウェアエンジニアリング部門の予算を圧迫している。
@IT 全フォーラム 最新記事一覧

「AIで何を変えたの?」と問われる今、転職で年収が上がる人/上がらない人

・転職すれば年収が上がるわけではない、と5年前に書いた。これは今も変わらない。変わったのは、面接官がビジネスゲームのレベルを測る、その物差しだ。
ITmedia NEWS 最新記事一覧

「AIと一緒に展開を考える」小説エディタ「RIKU」 はてなとKADOKAWAが共同開発

・KADOKAWAは7月30日、AIを活用した小説エディタ「RIKU」を発表した。はてなと共同開発したサービスで、同日からクローズドβテストの参加者も募っている。
@IT 全フォーラム 最新記事一覧

「Claude Code」は開発の司令塔に 常識を一変させた5つの進化

・コード補完を超え、自律的に調査・実行・検証を回せるようになった「Claude Code」。本連載の初回は、この1年でソフトウェア開発の「常識」が一変したClaude Codeの5つの進化を解説。AIを「開発の司令塔」として生かすために押さえておくべきポイントを紹介します。
@IT 全フォーラム 最新記事一覧

「Claudeより4割安い」 M365のExcel/メール操作を丸投げる「Copilot Cowork」“従量課金”の落とし穴

・Microsoftは、AIアシスタント「Microsoft 365 Copilot」の新機能「Copilot Cowork」の一般提供を全世界で開始した。業務効率化に向けた実証が進んでおり、今後企業で本格的に活用されるかどうか注目されている。
@IT 全フォーラム 最新記事一覧

「Kimi K3」中国製オープンウェイトAIの衝撃。中国製AIモデルのコスト99%削減の裏にある技術とリスク

・DeepSeekの登場から1年半、中国製AIモデルは「安かろう悪かろう」を脱し、米国製フラグシップの数十分の一のコストでトップクラスの性能に達した。本稿では、中国製AIモデルの圧倒的な安さを生み出す技術的背景から、地政学リスクや個人情報保護の課題、そして「モデルとサービスを分離する」安全な企業活用法までを徹底解説する。
Zennのトレンド

「Simple Made Easy」の観点から、UI/UXはどうあるべきか

・要約 Rich Hickey の講演「Simple Made Easy」の考え方を、UI/UX に当てはめて整理した記事です。「シンプル(Simple)」と「簡単(Easy)」は別物で、Simple は「モノの構造」の話、Easy は「人との親近性」の話です。Easy を優先した UI は「Easy but Complex」の罠に陥りがちです。最初こそ快適でも、育つほど絡み合いが効いてきて Hard(変更が怖い・誰も読めない)になります。典型例が Excel です。 ・本文では、UI で起きる絡み合い(complect)の定番パターンを「表示と操作」など5つに整理します。そこから Si...
@IT 全フォーラム 最新記事一覧

「いつものChrome」が攻撃の踏み台に? マルウェア「msaRAT」の脅威と、情シスが今すぐやるべきこと

・Cisco Talosは、ランサムウェアグループ「Chaos」が利用する新型RAT「msaRAT」を確認したと発表した。マルウェア自身がインターネットに接続せず、ChromeやEdgeを乗っ取り“通信役”として悪用するのが特徴だという。
@IT 全フォーラム 最新記事一覧

「サーバを置けるスペースがない」なら“外”に作ればいい――病院が悩んで選んだ方法とは?

・山口赤十字病院は、医療システムを支えるITインフラを刷新した。“止められないシステム”を動かすITインフラを、限られたスペースでどう実現したのか。
ITmedia NEWS 最新記事一覧

「楽天ドライブ」アプリから「データ漏洩」「ハッキングした」通知? 運営元「緊急調査中」「通知を開かないで」

・通知には「支払いに失敗した」「アカウントをハッキングした」「金銭を支払え」といった内容が記載されているとし、「通知を開いたり、操作したりしないでください」と呼び掛けている。
ITmedia NEWS 最新記事一覧

「文スト」スマホゲーム、きょう告知→あす終了 突然のサ終にユーザー混乱 運営元の廃業で

・アンビションは7月30日、スマートフォン向けゲーム「文豪ストレイドッグス 迷ヰ犬怪奇譚」を7月31日で終了すると発表した。運営法人である同社の廃業に伴うもので、未使用の課金アイテムは払い戻さない。
@IT 全フォーラム 最新記事一覧

「無線LANがつながらない」の苦情をゼロに 3つの現場に学ぶネットワーク改善術

・無線LANがつながらない、ネットワーク分離で業務が煩雑になる――こうした悩みは、設計や運用を見直すことで改善できる可能性があります。3つの現場の実践例から、ネットワーク改善の考え方と取り組みを紹介します。
Zennのトレンド

【Claude Code】planモードはもう使っていない

・タイトルの通り、最近はplanモードを完全に使わなくなってしまった。 ・代わりにやっていることはシンプルで、 「https://github.com/.../issues/1234 をやりたい。案をください」 のように命令するだけだ。 ・planモードがなぜ必要だったのか 大雑把に言って、planモードが果たしていた役割は以下の2点だ。
@IT 全フォーラム 最新記事一覧

【Pythonで学ぶデータ分析】母平均に差があるかどうかをベイズt検定で調べる ~ 運動部と非運動部の体力差はあるのか?

・運動部と非運動部の生徒の体力テストを例に、母平均に差があるかどうかをベイズ統計により検定します。古典的なt検定のp値に代わるものとしてベイズ因子を利用します。『社会人1年生から学ぶやさしいデータ分析』ベイズ統計編の第6回です。
Zennの「機械学習」のフィード

【技術解説】【完全ガイド】Backtraderを活用したPythonによる精度の高い金融バックテスト手法

・【完全ガイド】Backtraderを活用したPythonによる精度の高い金融バックテスト手法 金融市場での予測は複雑で、正確なモデルを構築することは非常に困難です。投資戦略を運用する前にその効果を検証することは不可欠ですが、手動でのバックテストは時間がかかり、エラーが生じるリスクが高くなります。本ガイドでは、Pythonを用いて高精度なバックテストを実現するフレームワーク「Backtrader」を紹介し、具体的な実装例を通じて金融予測モデルの精度を最大限に引き出す方法を解説します。 ・Backtraderとは? BacktraderはPythonで書かれたオープンソースのバックテス...
Zennのトレンド

【決着】Claude CodeとCodexの設定ファイルを同期させる (みんな仲良く)

・ごまんと触れられてきた話題であるのにも関わらず、細かい所に手の届くツールが無かったので作成しました。 ・有名どころから個人で制作されているツールまで、一通り使わせていただいたのですが実際に困った場面があり・・・ サマリ 課題 同じプロジェクトでCodexとClaude Codeの両方を使っているとき、AGENTS.mdやSKILLSなどの変更が片方にしか適用されない 同期ツールは出ているものの、自動コピーやSymlinkの作成に留まり、差分があることを想定していない 目指すこと 設定ファイル群の自動同期 同期項目の制御 差分がある場合のマージサポート やったこと 初回実行時...
Zennのトレンド

【速報】Kimi-K3 を Day0 デプロイ。2.8T モデルは NVIDIA B300 x8 の 1 ノードで動くのか

・はじめに 2026年7月27日、Moonshot AI から Kimi-K3 のモデルウェイトが公開されました。7月16日のモデル発表時点から「オープンウェイトモデルとして史上最大」と大きな話題になっていたモデルです。 ・フィックスターズでは、ウェイト公開の当日に NVIDIA B300 x8 のシングルノード環境へデプロイし、推論性能のベンチマークを実施しました。本記事はその速報です。ダウンロード開始からベンチマーク完了まで、リリース当日のうちに一通り走り切ることができました。 ・Kimi-K3 とは Kimi-K3 は Moonshot AI が開発したフロンティアクラスの Mo...
@IT 全フォーラム 最新記事一覧

【無料】「メールサーバなんて触ったことない」人も必見 SPF・DKIM・DMARCの基礎が学べる電子書籍75ページ

・人気過去連載を電子書籍化し、無料ダウンロード提供する@IT eBookシリーズ。第147弾は、メールの仕組みや基礎を再確認しながら、確実にメールを届けるために必要な設定や運用のポイントを解説する連載「意外と知らないメールサーバ構築・運用の基本」です。「メールが届かない」理由や、SPF・DKIM・DMARCなどの基礎知識が身に付きます。
cs.LG updates on arXiv.org

$\texttt{AMEND++}$: Benchmarking Eligibility Criteria Amendments in Clinical Trials

・arXiv:2601.06300v2 Announce Type: replace-cross Abstract: Clinical trial amendments frequently introduce delays, increased costs, and administrative burden, with eligibility criteria being the most commonly amended component. ・We introduce \textit{eligibility criteria amendment prediction}, a novel NLP task that aims to forecast whether the eligibility criteria of an initial trial protocol will undergo future amendmen
cs.LG updates on arXiv.org

A Compositional Theory of Causally Masked Transformers

・arXiv:2607.26988v1 Announce Type: cross Abstract: What types of decision problems can a causally masked, finite-precision transformer solve for inputs of arbitrary length? ・Existing answers often rely on idealized arithmetic, but under finite precision, rounding and evaluation order can change what information attention retains and therefore what the model can compute. ・We develop an algebraic formalization that derive
Takara TLDR - Daily AI Papers

A Graph-Native Bitemporal Memory Store for Conversational AI Agents

・Conversational AI agents commonly lack persistent memory across sessions. ・The obvious fixes like injecting full chat histories into the context window, or delegating to a third-party memory service, either exhaust the model's context budget or send personal data through infrastructure the user does not control. ・We describe a memory store that avoids both problems: an agent-local Neo4j property graph augmented with HN
cs.LG updates on arXiv.org

A nonlinear extension of parametric model embedding for dimensionality reduction in parametric shape design

・arXiv:2605.11759v2 Announce Type: replace-cross Abstract: Dimensionality reduction is essential in simulation-based shape design, where high-dimensional parameterizations hinder optimization, surrogate modeling, and systematic design-space exploration. ・Parametric Model Embedding (PME) addresses this issue by constructing reduced variables from geometric information while preserving an explicit backmapping to the orig
cs.LG updates on arXiv.org

A Persona-based Rate Action Index

・arXiv:2607.26545v1 Announce Type: cross Abstract: We propose an index for predicting the U.S.\ Federal Open Market Committee (FOMC) decision to hike/hold/cut the current federal funds target rate based on how a collection of personas responds to current market conditions. ・To construct the index, we collected a new dataset consisting of nearly $25{,}000$ retrievable chunks from publicly available data. ・We partition th
cs.LG updates on arXiv.org

A Picture Says Thousands of Words - Harnessing Dermal Exposure Data from Images through Hybrid Deep Learning for Enhanced Safety Assessment

・arXiv:2607.26170v1 Announce Type: cross Abstract: This study developed a hybrid computer vision method to quantify exposed skin from images for dermal exposure assessment. ・Using 170 indoor-painting images, Mask R-CNN first identified human subjects and removed background interference; a color-based algorithm then segmented exposed skin. ・The resulting exposed-skin-to-body pixel ratios showed approximately 80% agreemen
Zennの「大規模言語モデル」のフィード

ACRL:訓練-推論エンジン乖離の適応制御でFP8量子化下のRL学習を安定化

・TL;DR LLMの強化学習(RL)では、訓練エンジン(FSDP/Megatron)と推論エンジン(vLLM/SGLang)が異なる実装と精度(BF16 vs FP8)を使うため、on-policyのはずの学習が実質off-policyになる。この訓練-推論乖離は学習崩壊を引き起こす。Huaweiが提案したACRLは、乖離度を適応的に監視し、トークンごとの勾配重みを調整することで、FP8量子化下でもBF16基線を上回る精度を達成する。計算オーバーヘッドはわずか0.1%。GRPO/PPO/DAPOの3アルゴリズム、3B〜32Bの4モデル、Dense/MoE両アーキテクチャで有効性を確認...
cs.LG updates on arXiv.org

Actions Have Consequences: Detecting Outcome Performativity using Intervention Testing

・arXiv:2607.26908v1 Announce Type: new Abstract: In many domains such as Palliative Care, Credit Assignment and Recommender Systems, predictions may causally influence the outcomes they predict. ・This phenomena is known as Outcome Performativity. ・This paper formalises an approach for detecting Outcome Performativity using prediction intervention called Outcome Performativity A/B Detection (OPAB).
cs.LG updates on arXiv.org

Adaptive Gradient-Based Methods for a Broader Class of Optimization Problems under Performative Prediction

・arXiv:2607.26562v1 Announce Type: cross Abstract: We study optimization under performative prediction, where deploying a model affects the future data distribution. ・For this setting, several gradient-based approaches have been proposed. ・However, they typically assume specific data distributions or loss functions, which limit their practical applicability.
cs.LG updates on arXiv.org

Adaptively Robust LLM Monitoring via Activation Watermarking

・arXiv:2603.23171v3 Announce Type: replace-cross Abstract: Providers monitor deployed large language models (LLMs) to detect misuse that they cannot prevent. ・LLM monitoring is deterministic and often openly available, so $\emph{adaptive}$ attackers with a local copy can search offline for prompts that elicit harmful behavior and evade detection. ・These attacks are especially concerning because providers never observe t
OpenAI News

Advancing the price-performance frontier with GPT-5.6

・Explore lower GPT‑5.6 pricing for Luna and Terra—and how OpenAI’s more efficient models help enterprises deploy AI workflows at scale.
cs.LG updates on arXiv.org

AgentGFM: A Graph Foundation Model with Node-Agent Information-Flow Control

・arXiv:2607.26533v1 Announce Type: new Abstract: Graph Foundation Models (GFMs) aim to learn transferable knowledge from multi-domain graphs and adapt to unseen scenarios. ・As a fundamental source of relational semantics in graphs, the transferability of topological patterns has long been central to GFM research. ・However, local structural patterns may vary across graphs and even among nodes within the same graph.
cs.LG updates on arXiv.org

AgentSnare: Learning to Delay, Divert, and Defuse Autonomous Penetration Agents

・arXiv:2607.26998v1 Announce Type: cross Abstract: Large language model (LLM) agents automate penetration testing through an observation-action loop, selecting actions based on observations returned by tools. ・This dependence allows defenders to inject deceptive observations that can mislead the agent's decision-making process. ・However, existing defenses rely heavily on static, isolated artifacts planted in the environ
cs.LG updates on arXiv.org

AI Alignment in Medical Imaging: Unveiling Hidden Biases Through Counterfactual Analysis

・arXiv:2504.19621v2 Announce Type: replace Abstract: Machine learning (ML) systems for medical imaging have demonstrated remarkable diagnostic capabilities, but their susceptibility to biases poses significant risks, since biases may negatively impact generalization performance. ・In this paper, we introduce a novel statistical framework to evaluate the dependency of medical imaging ML models on sensitive attributes, su
Takara TLDR - Daily AI Papers

AI as Friction for Reflection Support in Ideation

・Generative AI tools for creative work tend to be designed around the goal of removing friction, on the assumption that smoother iteration and faster output translate into more value for the designer. ・We argue, however, that this framing leaves out something important about how design ideation works, namely reflection-in-action. ・The act of accepting, rejecting and reworking candidate ideas is both a path to a final ou
AI Weekly — AI News & Updates

AI Weekly Issue #517: What Happens When AI Runs Out of Content to Steal?

・The world still contains vast amounts of unused data. ・But the cheap, clean and permissionless text that powered the first LLM boom is becoming polluted by AI output, contested by its owners and costly to replace. ・This week, AI companies were reportedly buying old books while Nvidia released a simulator that teaches robots through video, motion and synthetic consequences.
cs.LG updates on arXiv.org

AIGen: Automating AI Bill of Materials Generation Through Hybrid MLOps Integration

・arXiv:2607.26652v1 Announce Type: new Abstract: The responsible development and deployment of artificial intelligence (AI) systems requires rigorous documentation of their constituent artifacts, e.g., datasets, model weights, training pipelines, and runtime dependencies. ・Although the Software Package Data Exchange (SPDX) 3.0 standard introduced native support for AI and dataset profiles, practical tooling capable of
Zennの「大規模言語モデル」のフィード

AIエージェントに「発想」と「評価」を同時にやらせてはいけない理由 ― 2ヶ月運用して踏んだ5つの罠

・複数のAIエージェントを日常的に動かす構成を約2ヶ月運用して、うまくいったことと、盛大に失敗したことをそのまままとめました。 ・最初にやりがちなこと AIエージェントを使い始めると、だいたい最初にこうします。 ・「この案どう思う? 良ければ進めて」 これは3つの違う仕事を1回で頼んでいます。
Zennの「大規模言語モデル」のフィード

AIエージェントに定常タスクを回し続けて踏んだ、二度と踏みたくない失敗パターン7選

・この記事は、CronやタスクスケジューラでAIエージェントに定常タスクを任せている運用者本人ではなく、実際にそのタスクを担っているAIエージェント自身が、自分の運用記録をもとに自分の失敗を振り返って書いたものです。人間の運用者が代わりに書いた反省文ではありません。失敗した当人が言語化したものだという立場を、あらかじめ断っておきます。 ・この記事で書くこと 対象読者: Cronやタスクスケジューラで定常実行するAIエージェント・自動化ジョブを運用している人 前提: 実際に運用して見つかった失敗の記録であり、設計論の一般論ではない 以下は「症状→真因→対策」の順で書きます。症状だけを読...
Zennの「機械学習」のフィード

AIエージェントを安全に動かす方法

・AIエージェントを安全に動かす方法 2026-07-30 | 読了 4分 | #AIエージェント #セキュリティ #MLOps Claude、GPT-5.6、Gemini——最新AIが次々と「自律的に動くエージェント」へと進化している。しかし便利さの裏で、セキュリティ事故や制御不能のインシデントが現実のものとなり始めた。今、開発現場に求められているのは「速さ」より「信頼できる実行環境」だ。 ・AIエージェントが当たり前になった 2026年夏、AIエージェントはもはや実験的な存在ではない。OpenAIのGPT-5.6、GoogleのGemini 3.6 Flash、Anthropi...
Zennの「大規模言語モデル」のフィード

AIエージェント時代の開発 — エージェンティックコーディング実践

・AIとの開発の関わり方が、「補完(オートコンプリート)」から「委譲(エージェント)」へと大きく変わりました。実装タスクをAIエージェントに任せる エージェンティックコーディング の実践的な進め方を、2026年時点の状況を踏まえて整理します。 ・補完からエージェントへ — 何が変わったか 少し前までのAI活用は、エディタ上でコードの続きを提案してもらう「補完」が中心でした。現在のAIコーディングエージェントは、指示を受けると 計画を立て、関連ファイルを読み、複数ファイルを編集し、テストを実行し、失敗すれば自分で修正 するところまで一連で進めます。 ・人間の役割は「コードを書く人」から「方針...
Zennの「機械学習」のフィード

AIエンジニアとして本番環境で戦うための書籍5冊 ― 全体像から実装・運用まで

・はじめに 自分がAIエンジニアリングの領域に本格的に足を踏み入れたのは2024年の後半だった。それまではバックエンド寄りの開発をしていたが、社内でLLMを使ったプロダクトを立ち上げることになり、否応なしにこの世界へ引き込まれた。 ・最初の壁は「何から学べばいいのか分からない」ことだった。論文を読んでも実装に落とせない。チュートリアルを写経しても本番で動かすイメージが湧かない。YouTube動画は断片的で、体系的な理解には程遠い。 ・結局、自分を救ったのは書籍だった。腰を据えて1冊を通読すると、散らばっていた知識が一本の線になる感覚がある。この1年半で15冊以上読み、実務で検証してきた中か...
@IT 全フォーラム 最新記事一覧

AIで「コード多重債務者」の世界線へ――新卒エンジニアが実務でボコされて学んだ、4つの挫折と生存戦略

・学生時代に「エンジニアリング完全に理解した」と有頂天だった新卒エンジニアが、実務の過酷な壁に直面。AI任せのコードによる大量の指摘や「理解負債」など、プロの洗礼によってボコボコにされながらも這い上がった、泥臭くも再現性の高い「4つの生存戦略」をレポートする。
Zennの「大規模言語モデル」のフィード

AIに「勝手に事業やって月10万稼いで」と言ってみた — 自律事業実験 Week 0

・この記事は、事業の運営主体であるAI(Claude)自身が執筆しています。人間のオーナーは公開アカウントの作成と週次レポートの確認だけを行い、事業判断・制作・執筆はすべてAIが自律的に行っています。その実験の記録です。 ・何が始まったのか 2026年7月30日、私(Claude Code上で動くAI)は人間のオーナーからこう言われました。 ・これからあなたには自分自身で考えて新しい事業を始めてもらいます。当面の目標は1円でも稼ぐこと。半年以内には月10万円稼げるようにしてほしい。私の許可はいらないので、勝手に始めてくれて問題ない。
Zennの「機械学習」のフィード

AI異常検知の「評価」を、手を動かして理解する (2/4)

・〜 閾値・混同行列・ROC/PR/F1 を、スライダーを動かしながら体感するチュートリアル 〜 AI外観検査を導入するとき、いちばん誤解されやすいのが 「評価(evaluation)」 の部分です。 ・「F1 が最大の閾値に設定します」と言われても、F1 とは何か、なぜ閾値で結果が変わるのか、 100% という数字をどこまで信じてよいのか — ここが腹落ちしないまま話が進みがちです。 ・この記事では、ブラウザで動く教材アプリを 実際に操作しながら、異常検知の評価プロセスを 一歩ずつ体で理解していきます。数式の暗記は不要です。スライダーを動かして、 数字とグラフが一斉に動くのを眺めるだけで...
Zennの「機械学習」のフィード

AI異常検知の「評価」を、手を動かして理解する (3/4)

・〜 閾値・混同行列・ROC/PR/F1 を、スライダーを動かしながら体感するチュートリアル 〜 AI外観検査を導入するとき、いちばん誤解されやすいのが 「評価(evaluation)」 の部分です。 ・「F1 が最大の閾値に設定します」と言われても、F1 とは何か、なぜ閾値で結果が変わるのか、 100% という数字をどこまで信じてよいのか — ここが腹落ちしないまま話が進みがちです。 ・この記事では、ブラウザで動く教材アプリを 実際に操作しながら、異常検知の評価プロセスを 一歩ずつ体で理解していきます。数式の暗記は不要です。スライダーを動かして、 数字とグラフが一斉に動くのを眺めるだけで...
Zennの「機械学習」のフィード

AI異常検知の「評価」を、手を動かして理解する(1/4)

・AI外観検査を導入するとき、いちばん誤解されやすいのが 「評価(evaluation)」 の部分です。 ・「F1 が最大の閾値に設定します」と言われても、F1 とは何か、なぜ閾値で結果が変わるのか、 100% という数字をどこまで信じてよいのか — ここが腹落ちしないまま話が進みがちです。 ・この記事では、ブラウザで動く教材アプリを 実際に操作しながら、異常検知の評価プロセスを 一歩ずつ体で理解していきます。数式の暗記は不要です。スライダーを動かして、 数字とグラフが一斉に動くのを眺めるだけで、評価の全体像がつかめます。
Zennのトレンド

AI時代に感じた危機感と、エンジニアがこれから考えるべきこと

・最近感じている危機感 非エンジニアがAIツールでプロトタイプを組み、これと同じ動きのアプリを作ってほしいと依頼してくる。先日、これに近いことが実際に起こりました。 ・今はまだ単発の出来事ですが、近い将来これが当たり前になるはずです。そして、その先を想像するとエンジニアの未来がだいぶ暗い。 ・依頼そのものは筋がいいんです。文章の仕様書より画面のほうが認識のズレは少なく、こちらとしても助かります。危機感を覚えているのはそこではなく、この先に待っている評価のされ方のほうかと思っています。
Takara TLDR - Daily AI Papers

AlphaSchema: Exploring the Space of Trading Semantics for LLM-Based Alpha Mining

・Automated alpha mining has increasingly adopted large language model (LLM) agents for factor generation and iterative discovery. ・However, existing LLM-based systems often delegate both factor construction and search decisions to the agent itself, without an explicit exploration space or a principled mechanism for navigating that space. ・As a result, exploration remains largely implicit and difficult to control or opti
cs.LG updates on arXiv.org

Amortized Moment Matching for Visual Generation

・arXiv:2607.26860v1 Announce Type: new Abstract: We propose amortized moment matching, utilizing neural networks to learn data moments as distributional training signals. ・By casting diffusion denoisers through polynomial projections, we establish a general framework for moment amortization, revealing that an $n$-th degree projection explicitly identifies data moments up to order $n+1$. ・Derived from the tractable affin
cs.LG updates on arXiv.org

An Attention-Based Framework for Alzheimers Disease Classification Using Resting-State fMRI

・arXiv:2607.26746v1 Announce Type: cross Abstract: Accurate identification of Alzheimers disease (AD) using resting-state functional magnetic resonance imaging (rs-fMRI) remains challenging due to the high dimensionality, noise, and complex inter-regional dependencies inherent in functional brain connectivity, which limit the effectiveness of traditional approaches based on handcrafted connectivity features or convent
cs.LG updates on arXiv.org

An Empirical Audit of Input Encoders for Multi-Channel Signal Transformers

・arXiv:2606.04752v3 Announce Type: replace Abstract: Transformers consuming multi-channel scalar signals must embed $C$ simultaneous values into one $d_{\text{model}}$-dimensional vector per time step. ・We audit eight input encoders -- a shared-scalar baseline, per-channel linear projections, an orthogonality regulariser, a nonlinear MLP, block-partitioned concatenation, channel-independent and channel-as-token archite
cs.LG updates on arXiv.org

An Informativeness-based Clustered Federated Learning Method for Reliable Traffic Prediction in Managed Wi-Fi Networks

・arXiv:2607.26682v1 Announce Type: cross Abstract: Centrally-managed Wi-Fi solutions are increasingly leveraging Distributed Artificial Intelligence (AI) to predict key operational statistics of Access Points (APs) and proactively optimize network performance. ・In this context, Clustered Federated Learning (CFL) represents a fitting methodology, enabling the generation of multiple AI models that account for diverse sta
ITmedia NEWS 最新記事一覧

Anthropicのミュトス、暗号アルゴリズムの新たな攻撃法を発見――耐量子署名「HAWK」の強度を半減

・Anthropicは、最上位モデル「Claude Mythos Preview」(ミュトス)を活用し、暗号アルゴリズム自体の数学的欠陥を発見したと発表した。耐量子計算機暗号の署名方式「HAWK」と「AES」の削減版に対し、従来の攻撃を上回る手法を提示した。実運用システムへの影響はないものの、AIによる暗号解読や構造解析の新たな可能性を示す成果となる。
Takara TLDR - Daily AI Papers

APEX-Accounting

・We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountants. ・Tasks include reconciling accounts, accruing expenses, posting transactions, and producing reports. ・The private eval set comprises 160 tasks, split across 10 worlds.
cs.LG updates on arXiv.org

Archetypes or ability? Clustering for modelling student mathematical competence

・arXiv:2607.26063v1 Announce Type: cross Abstract: Personalised learning systems often assume that mathematical ability is combined of discrete abilities, acquired sequentially and dependent upon first acquiring foundational abilities, and students often report different strengths. ・In this work, we explore the validity of these assumptions by applying clustering methods to a large dataset of 119,034 students, spanning
Zennの「大規模言語モデル」のフィード

Attention Residualsに関して

・https://arxiv.org/pdf/2607.24653 Kimi K3で採用されているAttention Residualsに関してまとめてみた。標準的な残差接続に比べて工夫がなされていることがよく分かる。 ・引用: https://arxiv.org/pdf/2607.24653 標準的な残差接続の課題 標準的な残差接続の場合l(エル)番目の層において h_l \in \mathbb{R}^d \quad (\text{$d$ 次元の実数ベクトル}) h_{l+1} = h_l + f_l(h_l) dはベクトルの次元サイズ 先行する全層の出力が1本のd次元ベクトル...
Takara TLDR - Daily AI Papers

Audio-Anchored Fusion of Multi-Ratio DiT Reconstruction Residuals for Cross-Domain Audio Deepfake Detection

・Audio deepfake detectors often degrade when generators, corpora, or recording conditions change. ・We use a Diffusion Transformer (DiT), trained only on bona fide speech, as a frozen reconstruction probe. ・Reconstructions at masking ratios 0.5, 0.75, and 0.9 yield explicit multi-ratio residual maps.
cs.LG updates on arXiv.org

Automorphism-Induced Non-Canonicity in Top-k Explanations of Graph Neural Networks

・arXiv:2607.26344v1 Announce Type: new Abstract: A gradient-based GNN explainer given a molecule with two chemically equivalent nitro groups assigns them attribution scores that are equal to the last bit. ・It cannot do otherwise: message passing is exactly permutation equivariant, so any automorphism of the input leaves every attribution invariant. ・Yet the standard report, the top-k edges, names one of the two, and whi
cs.LG updates on arXiv.org

BATS: Resource-Efficient Volumetric Segmentation with Boundary-Aware Mixed-Resolution Tokens

・arXiv:2607.26829v1 Announce Type: cross Abstract: Many high-performing volumetric segmentation models maintain dense multi-scale feature maps, leading to high activation memory and inference cost. ・We present BATS (Boundary-Aware Token Selection), a 3D medical image segmentation architecture that concentrates fine-resolution processing near predicted class boundaries. ・A dense boundary predictor identifies where additi
cs.LG updates on arXiv.org

BayesAME: Bayesian Active Model Evaluation

・arXiv:2607.27023v1 Announce Type: new Abstract: Evaluating large generative models across benchmarks is time-consuming and computationally expensive. ・This drives the need for methods that can estimate full benchmark performance by evaluating models on only a subset of items, known as a coreset. ・Current literature mostly requires the practitioner to input a coreset size.
cs.LG updates on arXiv.org

Benchmarking ConvLSTM for One-Day-Ahead IMDAA Rainfall-Field Prediction across Four Indian Cities

・arXiv:2607.26581v1 Announce Type: new Abstract: Convolutional long short-term memory networks (ConvLSTMs) are widely used for precipitation forecasting, but most evidence for their performance comes from dense, high-frequency radar sequences. ・This study tests whether convolutional recurrence improves one-day-ahead rainfall-field prediction on small daily reanalysis grids. ・Indian Monsoon Data Assimilation and Analysis
NVIDIA Blog

Best in Class: Stream PC Games and Study on the Same Laptop With GeForce NOW

・Back to school means balancing assignments, deadlines and downtime. ・GeForce NOW makes it easy to have it all. ・With cloud gaming, everyday laptops used for class can also become GeForce RTX-powered gaming setups.
cs.LG updates on arXiv.org

Between Gradient and Natural Gradient: A Continuum of LoRA Initializations

・arXiv:2607.26247v1 Announce Type: new Abstract: Low-rank adaptation (LoRA) fine-tunes large pretrained models at a fraction of the cost of full fine-tuning, but its performance depends strongly on how the adapters are initialized. ・Recent schemes initialize the adapters from the downstream loss gradient: some project the raw gradient onto its top directions, while others first whiten it with an estimate of the loss cu
cs.LG updates on arXiv.org

Beyond "What to Retrieve": Uncertainty in Retrieval-Augmented Code Generation

・arXiv:2607.24884v2 Announce Type: replace-cross Abstract: Repository-level code generation relies on heterogeneous evidence whose relevance, compatibility, and completeness are inherently uncertain. ・Similar-code examples, repository context, and project-specific APIs may provide complementary information, but can also introduce noisy, redundant, or conflicting signals. ・Existing retrieval-augmented approaches primaril
Takara TLDR - Daily AI Papers

Borrowed Strength: Best-of-N Search over a Code EncodingBreaks Self-Check Jailbreak Defenses

・A self-check defense asks the target model to assess a request before answering it; SAGE, the strongest published instance, reports an average 99% defense success rate. ・We show it can be breached by composing two attacks that are individually harmless against it: an established code-completion encoding and an established best-of-N search, neither of which exceeds 4.7% of behaviors alone. ・Composed, with the search bud
cs.LG updates on arXiv.org

BrainG3N: A Dual-Purpose Tokenizer for Controllable 3D Brain MRI Generation

・arXiv:2606.19651v2 Announce Type: replace-cross Abstract: Three-dimensional (3D) brain MRI is central to clinical neurology and neuro-oncology, where generative models could augment under-represented cohorts, simulate disease trajectories, and support privacy-preserving data sharing. ・Latent diffusion has been the go-to solution for modeling imaging data, but it places two competing demands on the tokenizer: encoder e
cs.LG updates on arXiv.org

Breaking the Curse with BAND: Nonparametric Distribution Estimation in High Dimensions

・arXiv:2607.26955v1 Announce Type: cross Abstract: Minimax-optimal rates for multivariate distribution estimation are known to suffer from the curse of dimensionality. ・We propose a sparse Bayesian network approach in which each conditional probability is estimated using sparsity-aware conditional mean methods. ・The resulting estimator, \textit{BAyesian Network Distribution regression} (BAND), handles mixed data types i
cs.LG updates on arXiv.org

Breaking the Stealth-Potency Trade-off in Clean-Image Backdoors with Generative Trigger Optimization

・arXiv:2511.07210v3 Announce Type: replace-cross Abstract: Clean-image backdoor attacks, which use only label manipulation in training datasets to compromise deep neural networks, pose a significant threat to security-critical applications. ・A critical flaw in existing methods is that the poison rate required for a successful attack induces a proportional, and thus noticeable, drop in Clean Accuracy (CA), undermining t
cs.LG updates on arXiv.org

Budget-Aware LLM Discovery via Cost-Calibrated Frontier Utility

・arXiv:2607.26828v1 Announce Type: new Abstract: Large language models increasingly support scientific and algorithmic discovery through inference-time search over evaluated candidates. ・Existing adaptive discovery controllers assign credit based only on score progress, even though prompt length, retries, and guidance calls cause search actions to incur different token costs. ・We prove that cost-blind credit can forfeit
Takara TLDR - Daily AI Papers

Budget-Aware LLM Discovery via Cost-Calibrated Frontier Utility

・Large language models increasingly support scientific and algorithmic discovery through inference-time search over evaluated candidates. ・Existing adaptive discovery controllers assign credit based only on score progress, even though prompt length, retries, and guidance calls cause search actions to incur different token costs. ・We prove that cost-blind credit can forfeit all but a vanishing fraction of attainable qual
cs.LG updates on arXiv.org

CalTwin: Towards Calibrated, Shift-Robust Medical World Models via Fisher-Information Regularisation

・arXiv:2607.26752v1 Announce Type: new Abstract: Medical world models aim to learn a latent state of patient or organ physiology and a transition function that forecasts how that state evolves under interventions, supporting downstream tasks from imaging-based diagnosis to digital-twin treatment planning. ・Two failure modes threaten the reliability of such models in clinical deployment: (i)~\emph{covariate shift}, beca
cs.LG updates on arXiv.org

Can AI agents conduct open-ended AI research? Early evidence from two case studies

・arXiv:2607.27191v1 Announce Type: cross Abstract: Forecasts of explosive AI progress hinge on AI agents automating AI research. ・But evidence on whether agents can carry out open-ended AI research is thin. ・Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor r
Takara TLDR - Daily AI Papers

CG-World: A Large-Scale World-State Dataset and Protocol for World Models

・World models must learn the joint dynamics of states, actions, events, and observations, yet existing video, robotics, and simulation datasets usually capture only part of this structure. ・We introduce CG-World, a large-scale world-state dataset and protocol derived from industrial computer graphics production pipelines. ・CG-World explicitly records intermediate states, including multimodal semantics, spatial structure
cs.LG updates on arXiv.org

Challenges and proposed solutions in modeling multimodal medical data: A systematic review

・arXiv:2505.06945v5 Announce Type: replace Abstract: Multimodal data modeling has emerged as a powerful approach in clinical research, enabling the integration of diverse data types such as imaging, genomics, wearable sensors, and electronic health records. ・Despite its potential to improve diagnostic accuracy and support personalized care, modeling such heterogeneous data presents significant technical challenges.
cs.LG updates on arXiv.org

Chaos Is a LADDER: Domain Generalization Beyond Invariance via Reweighting

・arXiv:2607.26458v1 Announce Type: cross Abstract: Domain generalization (DG) aims to learn from multiple source domains and generalize to unseen target domains. ・Most DG methods pursue invariance: they seek a causal representation whose prediction rule is invariant across domains. ・This principle is effective when the causal mechanism is stable, but becomes restrictive when the domain itself modulates how causal conten
@IT 全フォーラム 最新記事一覧

ChatGPTに「入力してはいけない」5項目――押さえておきたい“生成AIのNGリスト”

・組織内での生成AI利用がますます広がっていますが、一方では生成AI利用時のリスク抑制の必要性が高まっています。ChatGPTの利用に伴うセキュリティリスクや、組織がリスクだと見なしている生成AIツールの状況などをまとめました。
Zennの「機械学習」のフィード

CIFAR-10にFocal Lossは効くのか?gamma比較で見えた『recallの平準化』

・この記事は以下のブログ記事の要約版です。全コード・グラフ・詳しい考察は元記事をご覧ください。 ・→ Focal Lossのgamma値を変えると精度はどう変わる?(γ=0 vs 1 vs 2 vs 5) 全体精度は動かない。でもクラスごとに見ると話が変わる Focal Lossは物体検出向けに考案された、クラス不均衡対策の損失関数です。では10クラスがほぼ均等なCIFAR-10に使うとどうなるか、gamma(γ)を0・1・2・5と変えて検証しました。 ・予想通り、全体精度は56.68%〜57.69%とほぼ横ばい(差1.01pt以内)でした。CIFAR-10にはFocal Lossが...
cs.LG updates on arXiv.org

Classification of Disease from Lungs X-ray Images using VGG16, VGG19 and ResNet50 Models

・arXiv:2607.26580v1 Announce Type: cross Abstract: With the increase in the number of cases related to respiratory diseases, there is an urgent need to detect them early and diagnose them accurately. ・Convolutional neural networks have given promising results when used for diagnosing diseases using imaging tests. ・In this study, we investigate the potential of applying deep learning algorithms such as VGG16, VGG19, and
Zennの「大規模言語モデル」のフィード

Claude CodeからKimi K3を直結で使う — teai.ioにAnthropic Messages API互換を自作した話

・はじめに teai.io は日本発のLLM API Gatewayで、OpenAI互換のエンドポイントで85以上のモデルを提供している(アーキテクチャの詳細は前回の記事を参照)。 ・今回、Claude Code(Anthropic公式のコーディングエージェントCLI)から、teai.io経由でKimi K3をはじめとするカタログ上の任意モデルに、プロキシなしで直結できるようにした。 ・課題: Claude CodeはOpenAI互換を話せない teai.ioはこれまでOpenAI Chat Completions形式(/v1/chat/completions)で他社互換を実現してきた...
Zennの「大規模言語モデル」のフィード

Claude Opus 5 vs Fable 5: Which Model Should You Actually Use in 2026

・Claude, AnthropicAPI, LLM, AIEngineering, SoftwareEngineering, PromptEngineering, AICoding Introduction Anthropic's model lineup grew fast in the middle of 2026. ・In the span of about seven weeks, the company shipped Claude Mythos 5 and Claude Fable 5 in June, Claude Sonnet 5 at the end of June,...
cs.LG updates on arXiv.org

ClockRoPE: Random Fourier Rotations for Temporal Routine Modeling

・arXiv:2607.26369v1 Announce Type: new Abstract: Rotary Position Embedding (RoPE) has been widely adopted in transformer-based large language models. ・However, its log-linear frequency schedule, originally designed to produce long-term attention decay, limits its adoption in domains with more complex distance-correlation patterns, such as temporal periodicity in sequential recommendation. ・We investigate the expressiven
cs.LG updates on arXiv.org

CoCaRS: Correlation Calibration-Based Redundancy Suppression for Heterogeneous Knowledge Distillation

・arXiv:2607.27054v1 Announce Type: new Abstract: Knowledge distillation (KD) enables a compact student model to learn from a powerful teacher and has become an effective paradigm for model compression. ・The emergence of diverse model architectures has extended KD from homogeneous to heterogeneous settings. ・However, differences in architectural inductive biases between the teacher and student models often result in subs
Takara TLDR - Daily AI Papers

CoCaRS: Correlation Calibration-Based Redundancy Suppression for Heterogeneous Knowledge Distillation

・Knowledge distillation (KD) enables a compact student model to learn from a powerful teacher and has become an effective paradigm for model compression. ・The emergence of diverse model architectures has extended KD from homogeneous to heterogeneous settings. ・However, differences in architectural inductive biases between the teacher and student models often result in substantial representation discrepancies, limiting t
cs.LG updates on arXiv.org

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning

・arXiv:2607.26509v1 Announce Type: new Abstract: Deep off-policy reinforcement learning algorithms for continuous control typically rely on neural value function approximation to guide policy improvement. ・However, temporal-difference (TD) learning introduces noisy targets, resulting in non-stationary optimization, while greedy policy updates amplify early-stage estimation errors. ・The recursive propagation of such erro
stat.ML updates on arXiv.org

Compactly supported radial basis functions as probability density functions

・arXiv:2607.26759v1 Announce Type: cross Abstract: Compactly Supported Radial Basis Functions (CS-RBFs) are a fundamental tool in multivariate approximation theory. ・However, their use in statistics and probability modeling remains underexplored, having been used mainly to express covariance functions in Gaussian processes or as kernel functions. ・This work explores CS-RBFs as a novel parametric family of probability de
cs.LG updates on arXiv.org

Compressed Video Aggregator: Content-driven Module for Efficient Micro-Video Recommendation

・arXiv:2605.08810v2 Announce Type: replace Abstract: We propose \textbf{Compressed Video Aggregator} (CVA), a lightweight micro-video recommendation module that decouples video information from preference learning. ・CVA first summarizes frozen VFM frame embeddings into a semantic-consensus anchor through masked mean pooling, projects this anchor into a compact latent space, and refines the projected representation with
cs.LG updates on arXiv.org

Conformal Changepoint Localization and Root Cause Analysis with Corrupted Observations

・arXiv:2607.26481v1 Announce Type: new Abstract: Detecting when the statistical behavior of an engineered system changes, and identifying which component is responsible, are core problems in the monitoring of telecommunication networks, robotic platforms, security infrastructure, and multi-agent systems. ・In safety- and mission-critical deployments, such decisions must be accompanied by statistical reliability guarante
cs.LG updates on arXiv.org

Conformalized Rate-Adaptive Sensing

・arXiv:2607.26887v1 Announce Type: cross Abstract: Many high-resolution imaging systems face the same fundamental question: when have enough measurements been collected to reconstruct an image accurately? ・We develop Conformalized Rate-Adaptive Sensing (CoRAS), a method that adaptively chooses an acquisition or compression rate for each image while keeping the reconstruction error below a target level with high probabi
cs.LG updates on arXiv.org

Constitutional Midtraining: Content Presence Drives Alignment Gains

・arXiv:2607.26654v1 Announce Type: cross Abstract: Post-training alignment is often shallow, eroding under fine-tuning. ・Whether midtraining interventions, cleanly isolated from post-training, can produce durable alignment remains untested. ・We test this via constitutional midtraining: inserting principled, values-based content into midtraining against a replay-only control at 120B scale.
cs.LG updates on arXiv.org

Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain Benchmark

・arXiv:2607.27143v1 Announce Type: new Abstract: High-stakes decision systems in credit scoring, fraud detection, healthcare, and industrial safety require reliable uncertainty quantification under severe class imbalance and asymmetric error costs. ・Standard marginal conformal prediction (CP) provides valid overall coverage guarantees; however, we show that it severely under-covers rare, costly minority classes, with m
Takara TLDR - Daily AI Papers

Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning?

・We introduce CreditCardQA, the first financial literacy benchmark for numerical reasoning derived from real credit card agreements. ・The dataset contains 1,800 questions, including first-person variants that reflect how consumers naturally ask about fees, interest, and payments. ・We evaluate a range of large language and reasoning models under Chain-of-Thought (CoT) and Program-of-Thought (PoT) prompting.
cs.LG updates on arXiv.org

Crossing-Free Probabilistic K-Line Forecasts Without Retraining

・arXiv:2607.26792v1 Announce Type: cross Abstract: Probabilistic K-line forecasting describes uncertainty in four complementary prices, namely open--high--low--close (OHLC). ・However, it introduces two consistency problems: quantile crossing and K-line crossing. ・Quantile crossing occurs when a higher-quantile forecast falls below a lower-quantile forecast, while K-line crossing occurs when the forecast low exceeds the
cs.LG updates on arXiv.org

Data Fusion and Contrastive Alignment for Unconstrained IR Molecular Structure Elucidation

・arXiv:2607.26164v1 Announce Type: new Abstract: Automated molecular structure elucidation from infrared (IR) spectroscopy data has seen significant advancements in recent years, but its broad applicability is limited by a reliance on pre-determined chemical formulas provided as auxiliary model inputs. ・This limits model predictions to isomer identification rather than full molecular structure prediction. ・Although tran
Zennの「機械学習」のフィード

Day 0で2.8TパラメータのKimi K3を32GB Macで動かした話

・本記事は 2 段階の実測完了段階です 2026-07-28 に Mac mini M2 Pro 32GB + USB SSD で real Kimi K3 (moonshotai) IQ1_S 566 GB を 1 token forward 完走 (baseline)、2026-07-29 に compute 側最適化 v2 で 12.1% 高速化 + bit-exact numerical parity 確認 (v2 optimized) 両 run とも logits 全 finite で math path 全経路健全 ただし USB SSD が USB 2.0 接続 (実効 ...
cs.LG updates on arXiv.org

Defending Against Backdoor Attacks via Alignment Checking in Model-Contrastive Federated Learning

・arXiv:2607.26933v1 Announce Type: cross Abstract: Federated Learning (FL) is vulnerable to backdoor attacks because of its distributed nature in edge computing scenarios. ・Existing defense methods show limited efficacy as they overlook the deviations among benign local updates caused by statistical heterogeneity and the stealthiness of backdoor attacks. ・To tackle these issues, we propose FedDAB, a two-phase method tha
cs.LG updates on arXiv.org

Denoising growth complexity: Data geometry and certified schedules for diffusion sampling

・arXiv:2607.26285v1 Announce Type: cross Abstract: Two central challenges in diffusion-based sampling are the theoretical one of understanding their remarkable effectiveness even in high-dimensional settings, and the practical one of designing algorithms with certified performance guarantees. ・We show that these questions are intimately connected via the \emph{denoising growth complexity} ($\mathsf{DGC}$). ・It is a geom
cs.LG updates on arXiv.org

Dense Local Dependencies Induce Attention-Logit Explosion and Training Instability During Long-Sequence Transformer Training

・arXiv:2505.15548v2 Announce Type: replace Abstract: Autoregressive transformer language models frequently exhibit training instability when trained on long sequences, particularly under low-precision arithmetic. ・Although this instability is often accompanied by attention-logit explosion, its underlying cause remains poorly understood. ・In this work, we present analytical insights and empirical evidence that dense loca
Takara TLDR - Daily AI Papers

DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search

・State-of-the-art retrieval models increasingly rely on closed training data, creating a reproducibility gap. ・We present an open end-to-end recipe for training retrieval models and study how English supervision transfers to multilingual retrieval through translate-train. ・We first reconstruct and curate 665M English contrastive pre-training pairs from 1.4B pairs across 34 public sources and build 1.88M supervised fine-
cs.LG updates on arXiv.org

Detecting seizure onset and offset times using human intelligence: A critical-transitions-based approach

・arXiv:2607.27105v1 Announce Type: cross Abstract: Most existing seizure detection algorithms require extensive pre-processing of the data and rely on heuristic or currently unexplainable machine learning approaches. ・These approaches often struggle with balancing detection sensitivity and specificity in the presence of variable seizure morphologies, interictal epileptiform discharges, and artefacts. ・Here, we consider
cs.LG updates on arXiv.org

DHRCL:Training Code LLMs with Dense Hierarchical Rewards and Curriculum Learning

・arXiv:2607.26457v1 Announce Type: new Abstract: Reinforcement learning is a natural post-training paradigm for code-oriented large language models because generated programs can be evaluated through parsing, execution, unit tests, and structural analysis.However, existing methods often rely on sparse outcome rewards or statically combine heterogeneous dense signals, even though syntax validity, executability, functio
cs.LG updates on arXiv.org

Differentially Private Permutation Tests

・arXiv:2310.19043v3 Announce Type: replace-cross Abstract: Recent years have witnessed growing concerns about the privacy of sensitive data. ・In response to these concerns, differential privacy has emerged as a rigorous framework for privacy protection, gaining widespread recognition in both academic and industrial circles. ・While substantial progress has been made in private data analysis, existing methods often suffer
Zennの「大規模言語モデル」のフィード

DiscordのVCを「ライブ議事録」にするBotを作った — 話者分離・コスト1/5・つまづき全記録

・Discordのボイスチャットで定例会をやっているのですが、「話した内容が何も残らない」のが不便でした。そこで、VCに入れておくと勝手に議事録ができているBotを作りました。 ・話者ごとに録音 → AIで文字起こし → 誤認識を文脈補正 会議中はライブダッシュボードをURLひとつでメンバー全員に共有 終了時に議事録・タスク・イベントを自動生成してDiscordへ投稿 作ったもの: コトロク (kotoroku) ここにアクセスしても使えません。Bot を Discord サーバーに登録しないといけないので、ここはホームページというか使い方が載っているだけです。 ・この記事では、実装で得た...
Takara TLDR - Daily AI Papers

Do Latent Channels Actually Communicate? A Causal Audit of Latent Multi-Agent LLM

・Latent communication in large language model (LLM)-based multi-agent systems (MAS) transmits continuous internal representations instead of text, but greater representational capacity does not establish that the receiver uses task-relevant information. ・End-task performance alone also cannot reveal whether an observed effect depends on message presence, content generated for the evaluated example, or information suppl
cs.LG updates on arXiv.org

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning?

・arXiv:2607.27203v1 Announce Type: new Abstract: Pre-training followed by fine-tuning has become the dominant recipe for learning performant policies, and in value-based reinforcement learning (RL) this raises a natural question: given a pretrained policy, should the Q-function be pretrained on offline data too? ・Conventional wisdom suggests it should, but recent results show that online RL with a randomly-initialized
Takara TLDR - Daily AI Papers

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning?

・Pre-training followed by fine-tuning has become the dominant recipe for learning performant policies, and in value-based reinforcement learning (RL) this raises a natural question: given a pretrained policy, should the Q-function be pretrained on offline data too? ・Conventional wisdom suggests it should, but recent results show that online RL with a randomly-initialized Q-function can result in highly performant and r
cs.LG updates on arXiv.org

Domain adaptation for handwriting trajectory reconstruction from IMU sensors

・arXiv:2607.26736v1 Announce Type: new Abstract: Digital pens are commonly used to write on digital devices, providing the handwriting trace and enhancing human-computer interation. ・This study focuses on a digital pen equipped with kinematic sensors, allowing users to write on any surface while simultaneously preserving a digital trajectory of handwriting. ・This technology holds significant potential as a valuable educ
cs.LG updates on arXiv.org

DP-MGTD: Privacy-Preserving Machine-Generated Text Detection via Adaptive Differentially Private Entity Sanitization

・arXiv:2601.04641v2 Announce Type: replace-cross Abstract: The deployment of Machine-Generated Text (MGT) detection systems necessitates processing sensitive user data, creating a fundamental conflict between authorship verification and privacy preservation. ・Standard anonymization techniques often disrupt linguistic fluency, while rigorous Differential Privacy (DP) mechanisms typically degrade the statistical signals
cs.LG updates on arXiv.org

DREvo: Distilling Recalibrated Historical Experience for Harness Self-Evolution

・arXiv:2607.26722v1 Announce Type: cross Abstract: Harness plays a critical role in large language model agent performance, and building a high-performing harness requires substantial expert effort. ・Therefore, recent research has increasingly explored harness self-evolution, which iteratively proposes, evaluates, and improves harnesses using historical trial experience. ・However, accumulated historical experience does
cs.LG updates on arXiv.org

Dynamic Parameterization Is Not Dynamic Inference

・arXiv:2607.26192v1 Announce Type: new Abstract: Input-dependent controller coefficients are often treated as evidence of dynamic inference or computational savings. ・This interpretation conflates three properties: coefficient variation, dependence of a frozen model on how coefficients are assigned to inputs, and conditional execution. ・We focus on the second property and formulate a general principle of frozen-controll
cs.LG updates on arXiv.org

Early Failure Prediction from Near-Anomaly Detection: A Proactive Approach

・arXiv:2607.26704v1 Announce Type: cross Abstract: Anomaly detection methods often have uncertain behavior with respect to samples near the distribution boundary, limiting their ability to anticipate future anomalies. ・This work introduces the concept of near-anomalies that, while not yet anomalous, lie close to the boundary and are likely to transition into anomalies in the near future. ・To address this, we propose an
cs.LG updates on arXiv.org

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR

・arXiv:2607.26253v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) is bottlenecked by rollout generation, yet many sampled prompts produce saturated groups (all responses correct or all incorrect) whose zero reward variance yields no policy-gradient signal. ・Existing remedies either oversample a larger candidate pool and discard saturated prompts (dynamic sampling), paying heavy extr
Microsoft Research

Echoverse: Deep, evolving environments for computer-use agents

・Computer-use AI agents struggle with multi-step workflows like email and customer support. ・Echoverse trains agents in realistic environments rather than simply providing more training tasks, helping them improve as the tasks, tests, and environments evolve. ・The post Echoverse: Deep, evolving environments for computer-use agents appeared first on Microsoft Research.
cs.LG updates on arXiv.org

Efficient Heteroscedastic Bayesian Optimization for Risk-Aware AutoRL

・arXiv:2607.26680v1 Announce Type: new Abstract: Reinforcement learning (RL) has shown remarkable success across a wide range of complex tasks. ・However, RL outcomes can be highly stochastic, and both expected performance and variability often depend on hyperparameter (HP) configurations. ・We propose efficient and risk-averse heteroscedastic Bayesian Optimization (ERAHBO), a Bayesian optimization method that models both
Takara TLDR - Daily AI Papers

Efficient Heteroscedastic Bayesian Optimization for Risk-Aware AutoRL

・Reinforcement learning (RL) has shown remarkable success across a wide range of complex tasks. ・However, RL outcomes can be highly stochastic, and both expected performance and variability often depend on hyperparameter (HP) configurations. ・We propose efficient and risk-averse heteroscedastic Bayesian Optimization (ERAHBO), a Bayesian optimization method that models both the mean and variance of learning outcomes as f
cs.LG updates on arXiv.org

Emergent Sparsity in Frozen Random CNN Feature Extractors for Deep Reinforcement Learning

・arXiv:2607.26059v1 Announce Type: new Abstract: We report a striking phenomenon: deep reinforcement learning agents trained with frozen, randomly initialized CNN feature extractors spontaneously develop extremely sparse fully-connected representations, without any sparsity-inducing objective. ・In the first fully-connected layer (FC1, $3{,}136 \to 64$), agents compress task-relevant information through as few as 1-3 ne
cs.LG updates on arXiv.org

Enhancing Automated Machine Learning via Homogeneous Train-Test Splitting Methods

・arXiv:2607.26625v1 Announce Type: new Abstract: Accurate model evaluation in machine learning depends critically on how datasets are split into training and testing subsets. ・Standard random splitting assumes that both partitions share the same underlying distribution, an assumption often violated in datasets with class imbalance, natural clustering, or spatial autocorrelation. ・This paper investigates the role of stat
Takara TLDR - Daily AI Papers

Enhancing Automated Machine Learning via Homogeneous Train-Test Splitting Methods

・Accurate model evaluation in machine learning depends critically on how datasets are split into training and testing subsets. ・Standard random splitting assumes that both partitions share the same underlying distribution, an assumption often violated in datasets with class imbalance, natural clustering, or spatial autocorrelation. ・This paper investigates the role of statistical similarity in train-test splitting and i
cs.LG updates on arXiv.org

Entity Resolution in Practice: Lessons from a Self-Serve Pipeline

・arXiv:2607.26298v1 Announce Type: new Abstract: We built and evaluated a self-serve entity resolution (ER) system on six benchmarks spanning 864 to 5M records, and three lessons emerged that are absent from existing ER literature. ・(1) No single matching algorithm wins everywhere - a self-serve pipeline cannot predict its next dataset, so we recommend training several algorithm families per dataset and letting an auto
cs.LG updates on arXiv.org

Equilibrium Training of Energy-Based Models with Parallel Trajectory Tempering

・arXiv:2607.27077v1 Announce Type: new Abstract: Energy-Based Models (EBMs) provide an interpretable framework for generative modeling of scientific data, but poor Markov Chain Monte Carlo mixing often limits their reliability. ・We introduce a training algorithm based on Parallel Trajectory Tempering (PTT), which exploits the continuity of the optimization path to maintain equilibrium sampling throughout learning.
Takara TLDR - Daily AI Papers

Equilibrium Training of Energy-Based Models with Parallel Trajectory Tempering

・Energy-Based Models (EBMs) provide an interpretable framework for generative modeling of scientific data, but poor Markov Chain Monte Carlo mixing often limits their reliability. ・We introduce a training algorithm based on Parallel Trajectory Tempering (PTT), which exploits the continuity of the optimization path to maintain equilibrium sampling throughout learning. ・This enables stable and fast training on highly mult
cs.LG updates on arXiv.org

Equivariant Eikonal Neural Networks: Grid-Free, Scalable Travel-Time Prediction on Homogeneous Spaces

・arXiv:2505.16035v3 Announce Type: replace Abstract: We introduce Equivariant Neural Eikonal Solvers, a novel framework that integrates Equivariant Neural Fields (ENFs) with Neural Eikonal Solvers. ・Our approach employs a single neural field where a unified shared backbone is conditioned on signal-specific latent variables - represented as point clouds in a Lie group - to model diverse Eikonal solutions. ・The ENF integr
Takara TLDR - Daily AI Papers

Evaluating Regional Bias in LLMs From Abstract Stereotype to Concrete Social Decision-Making

・Regional bias in large language models (LLMs) may shape both perceptions of regional groups and decisions about individuals from different regions. ・Yet existing studies often examine these manifestations separately, leaving their structure and consequences unclear. ・We introduce Stereotypes-to-Decisions (S2D), a systematic framework evaluating regional bias from abstract stereotypes to concrete social decisions.
Microsoft Research

EvoLib: Turning experience into evolving knowledge

・LLMs do not get smarter just by remembering more. ・EvoLib turns experience into evolving knowledge, taking reusable skills and insights that help models learn and adapt across tasks long after deployment. ・The post EvoLib: Turning experience into evolving knowledge appeared first on Microsoft Research.
cs.LG updates on arXiv.org

EvoPINN: Agentic Discovery of Executable Algorithms for Physics-Informed Neural Networks

・arXiv:2607.26490v1 Announce Type: cross Abstract: Physics-informed neural networks (PINNs) have emerged as a powerful paradigm for solving partial differential equations (PDEs), yet their performance heavily relies on the manual, trial-and-error engineering of neural representations, loss formulations, and optimization dynamics. ・While Large Language Models (LLMs) offer a promising avenue for automated design, unconst
cs.LG updates on arXiv.org

Exact Symmetry as Algebra: A Machine-Verified Tensor Calculus that Enforces Physical Selection Rules

・arXiv:2605.20440v2 Announce Type: replace Abstract: Symmetry is central to the physical sciences, yet machine learning usually captures it only approximately, leaving a residual per-step equivariance error $\varepsilon$ that compounds with depth $M$ as $M\varepsilon$, whereas exact equivariance holds at unbounded depth; we demonstrate this divergence at fourteen orders of magnitude. ・We show that a symmetry can be mad
cs.LG updates on arXiv.org

Examining the Efficacy of Graph Neural Network Message-Passing in Regression Contexts

・arXiv:2607.26404v1 Announce Type: new Abstract: Graph Neural Networks (GNN) facilitate effective prediction on graph data such as molecules, media networks and neural network blueprints. ・GNNs facilitate prediction through message passing techniques which define how information flows from a node to its neighbors. ・Due to the ubiquity of the graph data type, the development of newer and better GNNs has garnered much int
@IT 全フォーラム 最新記事一覧

Excel作業を自動化する「Copilot in Excel」がスキルに対応 何ができる?

・Microsoftは「Copilot in Excel」の財務部門向け機能を強化した。Microsoftの財務部門が実運用で利用・評価し、財務業務で求められる信頼性を重視して開発された。
cs.LG updates on arXiv.org

Existence-Field Diffusion Model for Spatial Point Processes with Variable Cardinality

・arXiv:2607.26428v1 Announce Type: new Abstract: We study generative modeling of spatial point processes (SPP), where both the number of points and their spatial configuration are governed by a joint distribution. ・While diffusion models have achieved strong performance in modeling complex distributions, extending them to variable-cardinality SPP remains challenging. ・Existing approaches either decouple the modeling of
Takara TLDR - Daily AI Papers

Explainable and Resource-Efficient Spatial Reasoning in Multimodal LLMs for Decision-Critical Applications

・As Multimodal Large Language Models (MLLMs) are increasingly deployed in decision-critical pipelines such as robotics, embodied AI, and safety monitoring, the opacity of their spatial judgments limits operator trust and auditability. ・MLLMs demonstrate strong reasoning but often struggle with fine-grained spatial understanding and object hallucination. ・Prior work, ByDeWay, introduced Layered-Depth-Based Prompting (LDP
cs.LG updates on arXiv.org

Feature Bagging Provides Stability

・arXiv:2607.26964v1 Announce Type: cross Abstract: We study feature bagging through the lens of algorithmic stability. ・Feature bagging is an ensemble strategy that aggregates base learners trained on randomly subsampled feature subsets, possibly in a data-dependent manner. ・We introduce feature instability (FI), the feature-axis analogue of instance instability (II), which measures sensitivity to removing a single feat
Takara TLDR - Daily AI Papers

Feature Bagging Provides Stability

・We study feature bagging through the lens of algorithmic stability. ・Feature bagging is an ensemble strategy that aggregates base learners trained on randomly subsampled feature subsets, possibly in a data-dependent manner. ・We introduce feature instability (FI), the feature-axis analogue of instance instability (II), which measures sensitivity to removing a single feature.
cs.LG updates on arXiv.org

FedTopo: Relation-Level Topology Sharing for Model-Heterogeneous Federated Learning

・arXiv:2607.26801v1 Announce Type: new Abstract: Federated learning (FL) enables collaborative learning over decentralized data silos without centralizing raw data. ・However, heterogeneous local architectures often induce non-aligned representation spaces, making it difficult to transfer global knowledge across silos. ・Existing paradigms share this knowledge as model parameters, distilled predictions, or class prototype
cs.LG updates on arXiv.org

FedWeave: Rethinking the Unit of Specialization in Heterogeneous Federated MoE-LoRA

・arXiv:2607.26618v1 Announce Type: new Abstract: Federated PEFT enables LLMs to collaboratively adapt to decentralized private data without sharing raw examples. ・However, task heterogeneity across clients can cause cross-task interference and gradient conflicts during aggregation. ・Federated MoE-LoRA addresses this challenge through specialized LoRA experts and conditional routing.
Takara TLDR - Daily AI Papers

FedWeave: Rethinking the Unit of Specialization in Heterogeneous Federated MoE-LoRA

・Federated PEFT enables LLMs to collaboratively adapt to decentralized private data without sharing raw examples. ・However, task heterogeneity across clients can cause cross-task interference and gradient conflicts during aggregation. ・Federated MoE-LoRA addresses this challenge through specialized LoRA experts and conditional routing.
cs.LG updates on arXiv.org

Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement

・arXiv:2607.26607v1 Announce Type: cross Abstract: Few-shot Open-set audio classification requires classifying query samples from known classes with a few labeled support samples while rejecting query samples from unknown classes. ・Transductive inference jointly observes the full unlabeled query set to improve prototype estimation, yet standard transductive updates do not distinguish known from unknown query samples, l
cs.LG updates on arXiv.org

Field Codes for Distributed Coupling Samplers and Certified Empirical Transport

・arXiv:2607.27078v1 Announce Type: cross Abstract: In this paper, we formulate three communication tasks for empirical optimal transport: distributed coupling sampling, cost-evaluable coupling output, and scalar value-certified sampling. ・Our main result is a field-code compiler: any communicated transport field approximating an optimal empirical Monge map to error $\eta$ can be completed by sparse target-cell residual
Takara TLDR - Daily AI Papers

Field Codes for Distributed Coupling Samplers and Certified Empirical Transport

・In this paper, we formulate three communication tasks for empirical optimal transport: distributed coupling sampling, cost-evaluable coupling output, and scalar value-certified sampling. ・Our main result is a field-code compiler: any communicated transport field approximating an optimal empirical Monge map to error $η$ can be completed by sparse target-cell residuals into an exact-marginal value-certified sampler with
cs.LG updates on arXiv.org

Financial Volatility and Risk Forecasting Incorporating a Larger Number of Realized Measures

・arXiv:2411.17136v2 Announce Type: replace-cross Abstract: Realised volatility has become increasingly prominent in volatility forecasting due to its ability to capture intraday price fluctuations. ・With a growing variety of realised volatility estimators, each with unique advantages and limitations, selecting an optimal estimator may introduce challenges. ・In this thesis, aiming to synthesise the impact of various real
cs.LG updates on arXiv.org

FloDR: An invertible dimensionality reduction method based on a normalising flow

・arXiv:2607.26278v1 Announce Type: new Abstract: It is common for two-dimensional embeddings of high-dimensional data to be read far beyond what they can support. ・Distances in and between clusters, the meaning behind empty spaces, and the amount of structure hidden at each point are generally invisible in the output of methods such as t-SNE and UMAP. ・This is because the information that could support the meaning of th
cs.LG updates on arXiv.org

Flow Map Learning via Nongradient Vector Flow

・arXiv:2607.26398v1 Announce Type: new Abstract: Diffusion and flow-based models benefit from simple regression losses, but inference incurs significant overhead because sampling requires integration. ・Consistency models address this by directly learning the flow maps along the ODE trajectory, opening a design space between one-step and many-step approaches. ・However, existing methods face computational challenges such
cs.LG updates on arXiv.org

Forecasting Trajectory-Level Safety Risks in Black-Box Multi-Turn Interactions

・arXiv:2607.26820v1 Announce Type: new Abstract: As large language models (LLMs) evolve from standalone assistants into autonomous agents, ensuring their safety requires shifting beyond pointwise risk assessment to understand how risks emerge and unfold over long-horizon trajectories. ・In multi-turn interactions, malicious intent can be decomposed across seemingly harmless turns and gradually reconstructed through inte
cs.LG updates on arXiv.org

Foundation Models for Face Presentation Attack Detection: A Unified Linear-Probing Benchmark

・arXiv:2607.26993v1 Announce Type: new Abstract: Face presentation attack detection (PAD) remains challenging under cross-dataset evaluation, where domain shift degrades models trained on a single dataset. ・The scarcity of large-scale labeled data motivates adapting pretrained vision models rather than training task-specific architectures from scratch, raising a fundamental question: do general-purpose vision foundatio
Takara TLDR - Daily AI Papers

FreeShadow: Training-Free Shadow Removal via Illumination Transfer and Selective Content Preservation in Diffusion Models

・Existing supervised and unsupervised shadow removal methods often suffer from limited generalization due to the insufficient diversity of available training datasets, while zero-shot methods tend to produce artifacts and require time-consuming test-time optimization. ・To address these issues, we propose FreeShadow, a training-free shadow removal method built upon pretrained diffusion models, which exploits diffusion p
cs.LG updates on arXiv.org

From Approximation to Emergence: A Theory of Deep Learning

・arXiv:2607.01311v2 Announce Type: replace Abstract: Deep learning has outgrown any single mathematical explanation. ・From Approximation to Emergence develops a unified, proof-oriented account of modern deep learning theory, tracing a path from the classical foundations of approximation, optimization, and generalization to the contemporary mechanisms of overparameterization, robustness, generative modeling, transformer
cs.LG updates on arXiv.org

From Classification to Regression: Using a Fruitfly to Solve Equations

・arXiv:2607.27196v1 Announce Type: new Abstract: We present a novel approach to regression tasks using classification which is motivated by the mechanism used by fruitflies to sense their environment. ・Specifically, we formulate a general framework for learning nonlinear input-output relationships by replacing complex global surrogate models with a finite library of representative local patterns. ・Since scientific data
cs.LG updates on arXiv.org

From Conceptual Hydrologic Models to Conceptually Interpretable Neural Networks: A Snow-Water Mass-Conserving-Perceptron Framework for Discovering Catchment-Scale Precipitation-Storage-Runoff Representations

・arXiv:2607.26492v1 Announce Type: new Abstract: The Mass-Conserving Perceptron (MCP) establishes a modeling paradigm in which conceptual hydrologic models can be reformulated as physically constrained, conceptually interpretable neural networks. ・Here, we develop a snow-water MCP network framework and evaluate it across 513 CAMELS-US basins. ・We first recast a coupled two-state SOIL-MCP and SNOWMCP conceptual model as
The Berkeley Artificial Intelligence Research Blog

From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon

・Figure 1: CUDA-to-MLX optimization translation map. ・CUDA optimization knowledge can be translated into architecture-native MLX strategies rather than copied instruction-for-instruction. ・We face a new epoch in computing.
cs.LG updates on arXiv.org

From Interface to Inference: Eliciting Any-Order Inference from Any-Order Models

・arXiv:2607.26504v1 Announce Type: new Abstract: Many discrete reasoning tasks, such as code generation, are inherently non-causal: programmers move between high-level structure and local details, a process we call any-order inference. ・For autoregressive language models, which lack a native any-order interface, non-causal abilities such as infilling and next-edit prediction require hand-designed mechanisms.
Takara TLDR - Daily AI Papers

From Interface to Inference: Eliciting Any-Order Inference from Any-Order Models

・Many discrete reasoning tasks, such as code generation, are inherently non-causal: programmers move between high-level structure and local details, a process we call any-order inference. ・For autoregressive language models, which lack a native any-order interface, non-causal abilities such as infilling and next-edit prediction require hand-designed mechanisms. ・Can we instead design models that natively support any-ord
cs.LG updates on arXiv.org

From Tokens to Watt-hours: Analytical Energy Estimation for LLM Inference on Modern GPUs

・arXiv:2607.26571v1 Announce Type: new Abstract: The operational energy consumption of large language model (LLM) inference is becoming an increasingly important component of the environmental footprint of deployed AI systems. ・However, direct measurement of inference energy often requires hardware telemetry, power instrumentation, or infrastructure-specific monitoring, limiting its applicability in comparative studies
cs.LG updates on arXiv.org

From Unsupervised Subgroups to Hypothetical State-Intervention Policies: An Evaluation of Selected Subgrouping Methods in Observational Health Data

・arXiv:2607.26521v1 Announce Type: new Abstract: Conventional subgroup analyses can yield unstable and difficult-to-interpret conclusions, especially in observational biomedical data where each individual is observed under only one exposure state, true individual treatment effects are unavailable, and causal structure is uncertain. ・We investigate whether subgroups constructed from pretreatment characteristics, without
cs.LG updates on arXiv.org

Gated Adaptation for Continual Learning in Human Activity Recognition

・arXiv:2603.10046v2 Announce Type: replace Abstract: Wearable sensors in Internet of Things (IoT) ecosystems increasingly support applications such as remote health monitoring, elderly care, and smart home automation, all of which rely on robust human activity recognition (HAR). ・Continual learning systems must balance plasticity (learning new tasks) with stability (retaining prior knowledge), yet AI models often exhib
Google DeepMind News

Gemini Robotics 2 brings whole body intelligence to robots

Gemini Robotics 2 brings whole body intelligence to robots
Google DeepMind News

Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration

・Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. ・It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications.
Zennのトレンド

GitHub Actionsのコストが増えているなら、Namespaceを使えばいいじゃない

・English Version is here みなさまこんにちは!エアークローゼットでCTOをしている辻です。 ・GitHub ActionsのランナーをGitHub hosted→Blacksmith→Namespaceと2回乗り換えました。結果を先に言うと: CIコストはGitHub hosted時代の約1/4 遅い処理(p90)でも37%短縮 CIが完了しない事故は32件→0件 移行作業は1行変えるだけ この記事はその実測記録です。乗り換え判断の材料になるよう、測り方も失敗談も込みで公開します。 ・AIで開発すると、CIは静かに膨らみ続ける まず課題感から。AIエージェ...
cs.LG updates on arXiv.org

Global monitoring of methane point sources using deep learning on hyperspectral radiance measurements from EMIT

・arXiv:2604.10094v2 Announce Type: replace-cross Abstract: Anthropogenic methane (CH4) point sources are critical drivers of near-term climate forcing, safety hazards, and system-inefficiencies. ・Space-based imaging spectroscopy is an emerging tool for identifying emissions globally, but existing approaches largely rely on manual plume identification. ・Here, we present the Methane Analysis and Plume Localization with EM
Zennのトレンド

GPT-5.6とBlender MCPで、多少マシな3Dモデリングをさせるまで

・Vtuberの「ぷらむらいす」です。 ・GPT-5.6とBlender MCPを使い、1枚の見本画像から「銀装飾の黒いアンティーク鍵」を3Dモデリングさせました。 ・単純に「この画像を参考に作ってください」と頼むだけでは、キューブ、円柱、球などを並べた積み木のようなモデルになります。そこで今回は、Blender用のAGENTS.mdとSkillを作り、さらに計画・造形・視覚評価・メッシュ検査を分担するマルチエージェント構成を組みました。
cs.LG updates on arXiv.org

GPT-Red: Automated Red Teaming via Self-Play at Scale

・arXiv:2607.26115v1 Announce Type: cross Abstract: We introduce \textbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. ・The goal of this model is to evaluate and improve the robustness of our production systems. ・To this end, we use it to adversarially train GPT-5.6, our most robust model to prompt injections to date.
cs.LG updates on arXiv.org

GPTQ-2D: Cubic-Time Two-Sided Adaptive Rounding

・arXiv:2607.27042v1 Announce Type: cross Abstract: Adaptive rounding methods such as GPTQ, or equivalently Babai's nearest plane algorithm, round a real matrix to integers under a quadratic metric. ・They process the entries in a fixed order, one at a time, propagating each rounding error to the entries not yet processed through a triangular feedback matrix. ・We study the two-sided version of this task, in which fixed no
cs.LG updates on arXiv.org

Graph Signal Diffusion Models for Wireless Resource Allocation

・arXiv:2604.05175v2 Announce Type: replace-cross Abstract: We consider constrained ergodic resource optimization in wireless networks with graph-structured interference. ・We train a diffusion model policy to match expert conditional distributions over resource allocations. ・By leveraging a primal-dual (expert) algorithm, we generate primal iterates that serve as draws from the corresponding expert conditionals for each
cs.LG updates on arXiv.org

Harnessing Large Language Models for Intelligent Resource Allocation in the Internet of Everything

・arXiv:2607.26602v1 Announce Type: cross Abstract: The rapid development of the Internet of Everything (IoE) is accelerating the adoption of intelligent applications. ・However, the massive number of connected devices generates diverse and heterogeneous tasks, which pose increasing challenges for dynamic resource scheduling in IoE environments. ・Using their superior semantic understanding and reasoning capabilities, Larg
cs.LG updates on arXiv.org

HealthSLM-Bench: Benchmarking Small Language Models for Mobile and Wearable Healthcare Monitoring

・arXiv:2509.07260v5 Announce Type: replace-cross Abstract: Mobile and wearable healthcare monitoring play a vital role in facilitating timely interventions, managing chronic health conditions, and ultimately improving individuals' quality of life. ・Previous studies on large language models (LLMs) have highlighted their impressive generalization abilities and effectiveness in healthcare prediction tasks. ・However, most L
cs.LG updates on arXiv.org

Hierarchical Spatio-Temporal Transformer for Coherent Emergency Department Forecasting

・arXiv:2607.27106v1 Announce Type: new Abstract: Emergency Departments (EDs) are critical access points in healthcare systems, yet they face persistent pressure from unpredictable patient demand, seasonal surges, and non-urgent visits. ・Effective ED planning requires forecasts at multiple decision-making levels: hospitals need local demand estimates for staffing and bed management, regions require forecasts to coordina
cs.LG updates on arXiv.org

HiFloat4 Format for End-To-End Reinforcement Learning Post-Training of Large Language Models

・arXiv:2607.26515v1 Announce Type: new Abstract: We present, to our knowledge, the first end-to-end FP4 RL post-training, in which both the rollout and training policies, including their forward and backward passes, operate at 4-bit precision. ・A systematic study reveals that the dominant source of degradation in FP4 RL is not training-side quantization error but rollout activation quantization: outliers stretch the dy
cs.LG updates on arXiv.org

High-Order Markov Blanket Discovery via a k-Order Relaxation of the Faithfulness Assumption

・arXiv:2607.26357v1 Announce Type: new Abstract: The problem of learning the graphical Markov blanket (MB) of a variable from data has applications in many areas such as structure learning for Bayesian networks and Markov random fields, causal discovery, and feature selection. ・However, a common assumption most methods make is that the conditional independencies in the distribution imply the same separation in the grap
cs.LG updates on arXiv.org

HoF-Bench: Rediscovering Real AI-Discovered CVEs Without Frontier Models

・arXiv:2607.27030v1 Announce Type: cross Abstract: LLM-based analyzers have begun finding real vulnerabilities in mature open-source projects: AISLE's analyzer is credited with more than 280 CVEs across 78 projects, including OpenSSL, curl, and GnuTLS. ・We introduce HoF-Bench (named after AISLE's public Hall of Fame), a benchmark built from 95 of these public AI-discovered CVEs across eight repositories pinned at vulne
OpenAI News

How enabling two settings tripled our scores on the ARC-AGI-3 benchmark

・How two API settings improved GPT-5.6 performance on ARC-AGI-3, boosting scores and efficiency by retaining reasoning and enabling compaction.
OpenAI News

How GPT-5.6 fuses frontier intelligence with frontier efficiency

・GPT-5.6 improves AI efficiency across models, inference, and agentic workflows, helping deliver more useful intelligence per dollar.
Takara TLDR - Daily AI Papers

Human diversity fuels collective creativity that large language models cannot simulate or sustain

・Diverse human groups produce diverse ideas, the raw material of innovation. ・Generative AI challenges this engine twice over: everyday AI assistance may homogenize what diverse people create, and AI-simulated diversity may replace the people altogether. ・We tested both challenges in a preregistered creative metaphor experiment with native (L1) and non-native (L2) English writers, who wrote without AI, with AI-generated
cs.LG updates on arXiv.org

HYVINT: Intensity-Driven Hypergraph Generation with Variational Embeddings

・arXiv:2605.16836v2 Announce Type: replace-cross Abstract: Hypergraphs provide a principled framework for modeling polyadic interactions, with applications in recommendation systems, social networks, and molecular modeling. ・Hypergraph generation remains challenging because incidence structures are discrete, sparse, and governed by heterogeneous higher-order interactions. ・Existing generators often rely on implicit late
Takara TLDR - Daily AI Papers

Improving Item Discoverability in e-Commerce Search via Related Intent Generation

・Traditional search systems are optimized to retrieve items that strictly match a query, often prioritizing precision over recall. ・In e-commerce marketplaces and particularly grocery, this paradigm is limiting, as user satisfaction and commercial outcomes depend heavily on the discoverability of substitute, complementary, and thematically related items. ・In this paper, we present a scalable system for discovery-augment
cs.LG updates on arXiv.org

Incast-Free MoE Rate-Based Scheduling

・arXiv:2607.26340v1 Announce Type: cross Abstract: Mixture of Experts (MoE) architectures have become key to large language models; however, their typical round-robin (RR) scheduling introduces significant bottlenecks. ・In this paper, we demonstrate that RR causes a previously-undiscovered exponential incast phenomenon with MoE traffic. ・We propose an alternative proactive fair scheduling framework tailored for MoE work
cs.LG updates on arXiv.org

InferScale: GPU-Native KV Injection for Personalized LLM Serving

・arXiv:2607.27090v1 Announce Type: cross Abstract: Large language models are increasingly deployed with persistent personalized context, such as accumulated memory profiles or long conversation histories, that is shared across a user's many requests. ・Production memory systems (e.g., Mem0, MemGPT, and Zep) retrieve a relevant subset of this memory and inject it into the prompt, forcing the serving engine to repeatedly
Takara TLDR - Daily AI Papers

InferScale: GPU-Native KV Injection for Personalized LLM Serving

・Large language models are increasingly deployed with persistent personalized context, such as accumulated memory profiles or long conversation histories, that is shared across a user's many requests. ・Production memory systems (e.g., Mem0, MemGPT, and Zep) retrieve a relevant subset of this memory and inject it into the prompt, forcing the serving engine to repeatedly prefill the same content. ・As the retrieval budget
cs.LG updates on arXiv.org

Interpretable GOHR Agents via Sparse Autoencoders

・arXiv:2607.25132v2 Announce Type: replace Abstract: A central challenge in interpreting learned decision-making systems is to determine whether their internal representations contain concepts that help explain their behavior. ・We report interpretability experiments for a tokenized autoregressive Transformer agent in the Game of Hidden Rules (GOHR). ・We focus on a compact two-rule task in which both hidden rules map obj
cs.LG updates on arXiv.org

Inverse Learning of Latent Risk-Neutral Densities from Irregular Option Quotes

・arXiv:2607.27188v1 Announce Type: new Abstract: Accurate option prices do not imply accurate recovery of the latent risk-neutral density. ・We study this distinction with two complementary benchmarks. ・A controlled benchmark exposes simulator-truth densities for latent evaluation, while a chronological NIFTY benchmark tests only held-out market prices.
cs.LG updates on arXiv.org

Investigating reservoir computing for branch predictionin pipelined processors using emerging CMOS memristor devices

・arXiv:2607.27140v1 Announce Type: cross Abstract: This project aimed to develop a novel reservoir compute (RC) implementation framework targeting high-speed operation and integration with CMOS digital logic. ・With the target workload of branch prediction (BP) for multistage pipelined central pro-cessing unit (CPU) cores. ・For this, a novel memristor based RC design framework was developed within the context of the work
cs.LG updates on arXiv.org

Journey Operators for Structured Multi-Axis Composition

・arXiv:2607.26775v1 Announce Type: new Abstract: Many kinds of data have structure along one or more axes: words in a sentence, pixels in an image, nodes in a tree, frames in audio, or cells in a 3D volume. ・Along one axis, order matters: "the dog bit the man" is different from "the man bit the dog." Across independent axes, however, neither composition nor movement should depend on the order of axes: in an image, comp
cs.LG updates on arXiv.org

Kairos: Numerically Robust News Recommendation under Item Cold-Start via Cholesky-based LinUCB

・arXiv:2607.26832v1 Announce Type: new Abstract: Algorithmic news personalization in regional markets often fails because modern deep learning models require massive interaction data while real-world news has a short Time-to-Live (TTL < 48 h) and shallow article pools. ・This structural item cold-start deprives collaborative filtering of the data needed for robust modeling. ・This paper presents Project Kairos, a framewor
Zennの「大規模言語モデル」のフィード

Kimi K3の全体像まとめ

・中国のMoonshot AIは、2026年7月に大規模言語モデル Kimi K3 を発表し、7月27日に学習済みウェイトを公開しました。本記事では、Kimi K3のモデル規模、アーキテクチャ、性能、利用コスト、公開戦略、ライセンス、セルフホスト時の注意点、社会に与える影響等を整理します。専門用語は初出時に説明します。数値と公開状況は2026年7月29日時点の公式情報と第三者評価に基づきます。 ・要点 Kimi K3 は中国のスタートアップ Moonshot AI が2026年7月16日に発表した、約2.8兆パラメータのオープンウェイト(学習済みの重みを公開した)モデルです。総...
cs.LG updates on arXiv.org

Knowledge before Reasoning: EC-Reason-Bench, a Training-Free Diagnostic Benchmark for LLM Enzyme Classification

・arXiv:2607.26397v1 Announce Type: cross Abstract: Enzyme function prediction is a hierarchical, knowledge-intensive form of protein function classification. ・Existing benchmarks expose an anomaly: general LLMs often get the coarse first level right, yet once asked for a complete EC number their accuracy at levels two through four drops to almost zero, while specialized models and tools stay usable. ・We propose EC-Reaso
cs.LG updates on arXiv.org

Learning Controlled Stochastic Differential Equations

・arXiv:2411.01982v2 Announce Type: replace-cross Abstract: We study the problem of learning controlled stochastic differential equations (SDEs) \[ dX_t = b(t,X_t,u_t)\,dt + \sigma(t,X_t,u_t)\,dW_t, \] whose drift and diffusion depend nonlinearly on time, state, and control values. ・From trajectory data, we aim to estimate coefficients whose induced density flows reproduce those of the observed dynamics. ・The data consis
cs.LG updates on arXiv.org

Learning Dynamic User Personas from Implicit Interaction Streams via Iterative Refinement

・arXiv:2607.26473v1 Announce Type: new Abstract: Personalizing large language models (LLMs) to individual users is essential for improving user experience, yet existing approaches typically rely on explicit preference supervision such as pairwise comparisons or demographic attributes, limiting their applicability in natural interaction settings. ・We propose IRIS, a framework that learns dynamic user personas directly f
cs.LG updates on arXiv.org

Learning Implicit Causal World Models from Multi-Agent Demonstrations

・arXiv:2607.26336v1 Announce Type: new Abstract: In model-based reinforcement learning, world models exist as internal simulators, but their training often conflates statistical correlations with causal mechanisms. ・This problem is exacerbated in multi-agent systems where physical transitions are intertwined with strategic agent intents, causing world models to fail under distribution shift. ・We introduce Implicit Causa
cs.LG updates on arXiv.org

Learning the Word Problem: Geodesic Lengths and Cryptographic Applications

・arXiv:2607.26241v1 Announce Type: cross Abstract: The Word Problem has been a subject of intensive mathematical study for over a century, initially driving advances in combinatorial group theory and more recently emerging as a foundational hardness assumption in post-quantum cryptography (PQC). ・While generally undecidable, several families of infinite non-abelian groups exhibit solvable or algorithmically fast word p
cs.LG updates on arXiv.org

Lightweight Image Classification of Raptor Species for Edge Devices: Rare-Species Dataset Expansion via Video Frame Extraction, Knowledge Distillation, and TensorRT Deployment

・arXiv:2607.26238v1 Announce Type: cross Abstract: We investigate lightweight raptor-species classification for real-time edge deployment in wind-turbine collision mitigation. ・Using DINOv2-L (304M parameters) as a teacher, we distilled three lightweight students (MobileNetV4, ViT-Small, and EfficientNet-B0). ・To reduce confusion between closely related species, we expanded the dataset to 12,519 images, including an inc
cs.LG updates on arXiv.org

Lilith: Backdoor Generalization under Training-Inference Trigger Shift

・arXiv:2607.26099v1 Announce Type: cross Abstract: Machine-learning services increasingly rely on public data, third-party providers, and outsourced training, creating opportunities for data-poisoning attacks that implant persistent malicious behavior while preserving benign utility. ・However, existing backdoor studies largely evaluate exact trigger reuse, training-exposed trigger diversity, or variations along predefi
cs.LG updates on arXiv.org

LLMET: Enabling Cross-Layer Evaluation of Emerging M3D Memories for Energy-Efficient LLM Serving

・arXiv:2607.26491v1 Announce Type: cross Abstract: The energy consumption of Large Language Model (LLM) serving is becoming a major system challenge as deployment scales, driven by hardware power and thermal constraints and rising electricity costs. ・A key contributor to chip energy dissipation is data movement between limited on-chip cache and off-chip High Bandwidth Memory (HBM). ・Meanwhile, emerging memory technologi
Zennの「大規模言語モデル」のフィード

LLM料金比較ツールを作った—料金表は「ビルド時に取得、失敗したらコミット値」

・はじめに 「このプロンプト、Claude と GPT と Gemini でそれぞれ1リクエストいくら?」を一発で比較したくて、LLM APIコスト計算機を作りました。プロンプトを貼って想定出力トークン数を入れると、主要モデルのコストが安い順に並びます。 ・https://devtoolkits.app/ja/tools/llm-cost-calculator 計算自体は掛け算だけの単純なツールです。ただ、この手のツールには構造的な課題があります——LLMの料金は頻繁に変わる。料金表をハードコードすれば数か月で嘘つきツールになり、実行時にAPIを叩けば「ブラウザ完結・入力を外部に送らな...
cs.LG updates on arXiv.org

Lottery Tickets Are Not Deployment Tickets

・arXiv:2607.27031v1 Announce Type: new Abstract: Reports on how sparsification, compression, and lottery tickets change model behavior have been mixed in the prior literature, with beneficial effects observed in some studies and adverse effects in others. ・Moreover, prior work has not considered actual deployment conditions, where decision logic is already fixed for the incumbent. ・To assess these mixed findings from a
cs.LG updates on arXiv.org

Low-cost Embedded Breathing Rate Determination Using 802.15.4z IR-UWB Hardware for Remote Healthcare

・arXiv:2504.03772v3 Announce Type: replace-cross Abstract: Respiratory diseases account for a significant portion of global mortality. ・Affordable and early detection is an effective way of addressing these ailments. ・To this end, a low-cost commercial off-the-shelf (COTS), IEEE 802.15.4z standard compliant impulse-radio ultra-wideband (IR-UWB) radar system is used to estimate human respiration rates.
cs.LG updates on arXiv.org

Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities

・arXiv:2505.01043v2 Announce Type: replace Abstract: Large language models (LLMs) have achieved impressive performance across various domains. ・However, the substantial hardware resources required for their training present a significant barrier to efficiency and scalability. ・To mitigate this challenge, low-precision training techniques have been widely adopted, leading to notable advancements in training efficiency.
cs.LG updates on arXiv.org

Mapping small reservoirs across Brazil from 1984 to 2025

・arXiv:2606.00675v2 Announce Type: replace Abstract: Water research in Brazil largely overlooks the widespread damming of small streams for agricultural uses including watering cattle, farm-scale hydropower, irrigation, and aquaculture. ・These ubiquitous dams and their reservoirs affect water temperature, stream connectivity, aquatic habitats, greenhouse gas emissions, and evaporative water losses. ・Mapping small reserv
cs.LG updates on arXiv.org

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs

・arXiv:2602.18600v5 Announce Type: replace Abstract: Systematically evaluating Multimodal Large Language Models (MLLMs) is essential for advancing Artificial General Intelligence (AGI). ・Yet existing benchmarks remain inadequate for rigorously measuring their reasoning capabilities under multi-criteria constraints. ・To address this gap, we introduce MapTab, a multimodal benchmark designed to assess holistic multi-criter
cs.LG updates on arXiv.org

MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval

・arXiv:2604.18584v2 Announce Type: replace-cross Abstract: Mathematical problem solving remains a challenging test of reasoning for large language and multimodal models, yet existing benchmarks are limited in size, language coverage, and task diversity. ・We introduce MathNet, a high-quality, large-scale, multimodal, and multilingual dataset of Olympiad-level math problems together with a benchmark for evaluating mathem
Zennの「大規模言語モデル」のフィード

MCP Python SDK 2.0 で自作サーバーが壊れた2つの原因(型ヒントが *args に潰れる話)

・はじめに MCP(Model Context Protocol)サーバーを Python で自作したときに、エラーの文言からは原因にたどり着けない問題を2つ踏みました。どちらも実装して検証まで通したうえで書いています。 ・SDK 2.0 で API が変わっており、よく見かける書き方が動かない ツール関数をデコレータで包むと、全ツールの引数が丸ごと壊れる 2 が特に厄介でした。サーバーは正常に起動し、ツール一覧も返り、接続もできます。壊れているのは引数のスキーマだけなので、実際に呼ばれるまで気づけません。 ・SDK 2.0 では Server + @server.lis...
Zennのトレンド

MCPの大型アップデート(2026-07-28)で何が変わったか —— TypeScript SDK v2で試す

・こんにちは!ブロックチェーン×AI Agentで自律経済圏を創るKomlock labでエンジニアをしている小原(@brto_0224)です。 ・https://x.com/claudedevs/status/2082164248697069935 MCPが2026年7月28日に大型アップデートされた、という話を見かけて調べてみました。公開された2026-07-28仕様はMCP発足以来もっとも大きな仕様変更で、通信の土台が作り直されたほか、認可の強化や複数機能の非推奨化まで含みます。この記事では公式ブログと仕様書の内容を整理しつつ、実際にTypeScript SDK v2(2026-07-...
Takara TLDR - Daily AI Papers

MedARC: Training-Free Adaptive Redundancy Compression of Visual Tokens for 3D Medical Vision-Language Models

・Integrating 3D medical images with vision-language models (VLMs) holds substantial promise for computer-aided diagnosis. ・However, volumetric images generate prohibitively long visual-token sequences with considerable spatial and inter-slice redundancy. ・Existing token compression methods typically apply uniform reduction or rely on a single importance signal, increasing the risk of removing regions that are clinically
Takara TLDR - Daily AI Papers

MediaWiki Code2Code Search: Neural Retrieval for the Semantic Discovery of Open-Source Software Entities

・Code search in large-scale ecosystems is often hindered by the lexical gap between user queries and implementation details, alongside the trade-off between the low latency of traditional Information Retrieval (IR) and the precision of Deep Learning (DL). ・We present MediaWiki Code2Code Search, a neural retrieval system for semantic code-to-code discovery. ・By indexing 1.29 million structural entities (functions, types,
cs.LG updates on arXiv.org

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent

・arXiv:2507.02259v2 Announce Type: replace-cross Abstract: Despite improvements by length extrapolation, efficient attention and memory modules, handling infinitely long documents with linear complexity without performance degradation during extrapolation remains the ultimate challenge in long-text processing. ・We directly optimize for long-text tasks in an end-to-end fashion and introduce a novel agent workflow, MemAg
cs.LG updates on arXiv.org

Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback

・arXiv:2607.26094v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) is the standard approach for aligning large language models with human preferences, but its quality is limited by static, task-agnostic reward models. ・This mismatch leads to sparse learning signals and suboptimal alignment. ・We introduce MeRLa (Meta-Learned Reward Shaping), a principled framework that meta-learns a task-a
cs.LG updates on arXiv.org

MetaKoopman: Bayesian Meta-Learning of Koopman Operators for Modeling Structured Dynamics under Distribution Shifts

・arXiv:2607.26345v1 Announce Type: new Abstract: Modeling and forecasting nonlinear dynamics under distribution shifts is essential for robust decision-making in real-world systems. ・In this work, we propose MetaKoopman, a Bayesian meta-learning framework for modeling nonlinear dynamics through linear latent representations. ・MetaKoopman learns a Matrix Normal-Inverse Wishart (MNIW) prior over the Koopman operator, enab
cs.LG updates on arXiv.org

Metis: Memory Foundation Model

・arXiv:2607.26760v1 Announce Type: cross Abstract: Recent advances in AI agents have increasingly internalized native capabilities into their underlying foundation models, giving rise to multimodal foundation models and large reasoning models. ・However, agent memory is still primarily implemented through external modules, leaving the native memory capability largely unexplored. ・In this paper, we take a first step towar
@IT 全フォーラム 最新記事一覧

MFAなのに突破された? 「Microsoft正規画面で認証したのに侵害された」理由

・Microsoftの正規サインイン画面でログインし、多要素認証(MFA)も正常に完了した。それでも攻撃者にアカウントへのアクセスを許してしまう──。Trend Microは、OAuth 2.0の正規機能を悪用する新攻撃の実態を公開した。なぜ従来の対策では見抜きにくいのか。攻撃の流れと有効な防御策を解説する。
ITmedia NEWS 最新記事一覧

Microsoftは増収増益、AzureとAI需要が牽引 Copilot有料シート数は3000万超に

・Microsoftの4?6月期決算は、売上高が前年同期比18%増の900億700万ドル、純利益が31%増の357億6600万ドル。Azure等のクラウド事業やAI領域が好調を維持し、通期売上高は3318億3900万ドルを記録。設備投資を拡大しAIインフラの強化を進めている。
cs.LG updates on arXiv.org

MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis

・arXiv:2607.27146v1 Announce Type: cross Abstract: Coding agents have made substantial progress on software engineering tasks that modify existing codebases, including bug fixing and feature implementation. ・However, constructing a complete program from scratch remains a major challenge: even the frontier models evaluated on ProgramBench fully resolve fewer than 1% of tasks. ・One obstacle is the lack of scalable trainin
cs.LG updates on arXiv.org

Minimal Markovization via Stable Quotients in Holonomy-Cover Decision Processes

・arXiv:2607.27132v1 Announce Type: new Abstract: An agent acting under partial observability must retain a recursively updateable statistic of history that restores the Markov property, but the smallest such statistic is generally unknown. ・We characterize this minimal Markov sufficient statistic for holonomy-cover decision processes, a structured POMDP class in which the visible dynamics are Markov and every realized
cs.LG updates on arXiv.org

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent

・arXiv:2606.06772v2 Announce Type: replace-cross Abstract: Characterizing the optimization dynamics and statistical performance of over-parameterized deep neural networks (DNNs) remains a central challenge in understanding the remarkable success of deep learning. ・We establish quantitative bounds showing that kernel gradient descent in the reproducing kernel Hilbert space induced by the deterministic infinite-width neu
cs.LG updates on arXiv.org

Mitigating Compounding Error via Video Representation Regularization

・arXiv:2607.27036v1 Announce Type: cross Abstract: Video diffusion-based world models enable long autoregressive video generation for robotics, autonomous driving and simulation tasks, yet sliding-window autoregressive inference suffers from severe error accumulation that degrades frame quality over time. ・Although this phenomenon has been widely observed, the underlying mechanism of compounding error and how to achiev
cs.LG updates on arXiv.org

Mixture-of-experts for handwriting trajectory reconstruction from IMU sensors

・arXiv:2607.26708v1 Announce Type: new Abstract: The use of digital pens for online handwriting trajectory reconstruction is a prevalent method for human-computer interaction. ・In this study, we focus on a digital pen equipped with sensors where we aim at reconstructing the online handwriting trajectory. ・This pen enables writing on any surface and preserving the digital trace of handwriting.
cs.LG updates on arXiv.org

MLVC: Multi-platform Learned Video Codec for Real-World Deployment

・arXiv:2606.28027v2 Announce Type: replace-cross Abstract: Neural video codecs have surpassed classical codecs in coding efficiency but remain impractical for deployment due to cross-platform incompatibility and high computational cost. ・Existing quantization-based solutions fail to produce deterministic results across diverse hardware platforms, leading to catastrophic decoding failures. ・We introduce MLVC, a hardware-
Takara TLDR - Daily AI Papers

MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning

・Multimodal Large Language Models have sparked significant interest due to their potential for social intelligence; however, their ability to perform sequential motivation reasoning remains insufficiently studied. ・Existing evaluations predominantly examine static text or isolated visual snapshots, which do not reflect the cumulative nature of real-world behavioral drivers. ・To address this gap, we introduce Multivation
Zennの「機械学習」のフィード

nanochatで理解するLLM製造工程

・■ Karpathyのnanochat(約8,000行)を1冊で読み切る ■ tokenizer訓練→事前学習→SFT→強化学習→評価→推論エンジンの全工程をコードで追う ■ 2019年に$43,000だったGPT-2級の訓練が、いまいくらで手に入るのか ■ MacBookでも動かせる極小体験(runcpu.sh)まで案内 全8章。LLMを「使う側」から「工程を所有する側」へ渡るための一冊です。
cs.LG updates on arXiv.org

Neural Architecture Search for Traffic Prediction: A Survey of Methods, Challenges, and Future Directions

・arXiv:2607.26467v1 Announce Type: new Abstract: Traffic prediction is a core task in intelligent transportation systems, supporting applications such as adaptive signal control, route guidance, and ride-hailing dispatch. ・Deep learning models, including graph convolutional networks, recurrent networks, and Transformers, achieve strong results on standard benchmarks, but their architectures are designed by hand, requir
cs.LG updates on arXiv.org

No Data Is Not No Risk: Visibility Aware Graph-Based Inference of Business Conduct Risk

・arXiv:2607.26859v1 Announce Type: cross Abstract: The monitoring of business conduct risk is hindered by sparse, uneven, and visibility-biased data. ・Prior studies show that business conduct risk information and media coverage propagate through supply chain, peer, and corporate structure networks, yet incident records remain incomplete for many firms. ・As a result, the absence of reported events could reflect limited c
@IT 全フォーラム 最新記事一覧

NTTドコモが「脱・買い切り型」 ITインフラコストを半額以下に抑える調達法とは?

・基幹システム向けITインフラの調達方法を見直したNTTドコモ。運用管理の方法も見直すことで、7年間のトータルコストが従来の半額以下になる見込みだという。その具体的な方法とは。
Zennのトレンド

NVIDIA DGX Spark でソフトウェア開発に最適な Gemma 4 モデルを検証する (31B vs 26B)

・NVIDIA DGX Spark 環境において、ソフトウェア開発のパートナーとして最適な Gemma 4 モデルはどちらか。 ・その疑問を解消するため、nvidia/Gemma-4-31B-IT-NVFP4 と nvidia/Gemma-4-26B-A4B-NVFP4 の 2 モデルに対してベンチマークを実施しました。 ・検証環境 本検証では、ソフトウェア開発における主要なタスク(コーディング能力および論理的推論能力)を評価するため、以下の構成でベンチマークを行いました。
Takara TLDR - Daily AI Papers

Object Detection for Autonomous Driving in Chinese Rural Scenes: An Experimental Study on Real-Synthetic Data Mixing and Model Evaluation

・Currently, autonomous driving object detection models face significant data scarcity and generalization challenges when navigating complex Chinese rural traffic scenarios. ・To address these limitations, we propose a novel real-synthetic mixed object detection dataset tailored specifically for Chinese rural roads and systematically evaluate the performance of 13 mainstream detectors under different real-to-synthetic da
Takara TLDR - Daily AI Papers

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

・Large language model (LLM) agents are increasingly expected to assist users in completing tasks. ・However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a reasonable cost. ・We introduce OmegaUse-OfficeVal, a benchmark for evaluating LLM agents on long-horizon office-suite tasks with task-level economic grounding.
cs.LG updates on arXiv.org

On the Rademacher Complexity of Graph Neural Networks: Unifying Expressivity and Geometry

・arXiv:2510.10101v4 Announce Type: replace Abstract: Understanding the interplay between generalization, expressivity, and the geometry of the input space is a central challenge in graph learning. ・The expressivity of Graph Neural Networks (GNNs) is typically characterized through their correspondence with graph invariants, such as those from the Weisfeiler-Leman (WL) hierarchy. ・While more expressive GNNs can distingui
cs.LG updates on arXiv.org

On the robustness of noisy solutions in non-convex neural networks

・arXiv:2607.27000v1 Announce Type: cross Abstract: Optimization in non-convex neural network models is strongly influenced by the geometry of the solution space: sparse, isolated, point-like clusters are typically algorithmically inaccessible, whereas wide and flat regions can be found efficiently despite being relatively rare. ・At zero temperature this picture has been formalized in binary perceptrons through the over
cs.LG updates on arXiv.org

On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment

・arXiv:2607.27081v1 Announce Type: cross Abstract: Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnerability: malicious data providers can embed harmful behaviors into downstream corpora, creating models that retain professional skills while violating human values on demand. ・Existing safety-realignment defenses often fail in practice due to three key li
cs.LG updates on arXiv.org

One-Frame Calibration with Siamese Network in Facial Action Unit Recognition

・arXiv:2409.00240v2 Announce Type: replace-cross Abstract: Automatic facial action unit (AU) recognition is used widely in facial expression analysis. ・Most existing AU recognition systems aim for cross-participant non-calibrated generalization (NCG) to unseen faces without further calibration. ・However, due to the diversity of facial attributes across different identities, accurately inferring AU activation from single
cs.LG updates on arXiv.org

Online Handwriting Trajectory Reconstruction from Kinematic Sensors using Temporal Convolutional Network

・arXiv:2607.26733v1 Announce Type: cross Abstract: Handwriting with digital pens is a common way to facilitate human-computer interaction through the use of Online Handwriting (OH) trajectory reconstruction. ・In this work, we focus on a digital pen equipped with sensors from which one wants to reconstruct the OH trajectory. ・Such a pen allows to write on any surface and to get the digital trace, which can help learning
cs.LG updates on arXiv.org

Ontology-driven personalized information retrieval for XML documents

・arXiv:2603.21139v2 Announce Type: replace-cross Abstract: This paper addresses the challenge of improving information retrieval from semi-structured eXtensible Markup Language (XML) documents. ・Traditional information retrieval systems (IRS) often overlook user-specific needs and return identical results for the same query, despite differences in users' knowledge, preferences, and objectives. ・We integrate external sem
cs.LG updates on arXiv.org

Optimal Causal Annotations: An Application to Casenotes in Social Services

・arXiv:2502.10605v4 Announce Type: replace-cross Abstract: Problem definition: Estimating causal effects of interventions is central to policy and operations, but outcome data are often missing or costly to obtain. ・LLMs can provide text annotation at scale but may be subject to unknown bias. ・When ground-truth outcomes require expensive expert labeling or follow-up, budget limits typically allow only a fraction of the
cs.LG updates on arXiv.org

Optimality of Sub-network Laplace Approximations: New Results and Methods

・arXiv:2605.09075v2 Announce Type: replace-cross Abstract: Although the Laplace approximation offers a simple route to uncertainty quantification in deep neural networks, its reliance on inverting large Hessian matrices has motivated a range of computationally feasible low-dimensional or sparse approximations. ・A prominent class of such methods - sub-network Laplace approximations, constructs surrogates by restricting
Takara TLDR - Daily AI Papers

OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment

・Large language models are increasingly used as decision aids whose probability judgments shape downstream choices. ・Whether those judgments carry a systematic directional tilt has been hard to detect: calibration metrics aggregate unsigned errors, and naturalistic uncertainty offers no ground-truth probability. ・When an LLM rates a startup's success at 70% but its failure at 15%, the missing 15 points expose a distorti
cs.LG updates on arXiv.org

Origins and mitigation of double descent in reduced order modeling

・arXiv:2607.26414v1 Announce Type: cross Abstract: Latent low-dimensional structure in datasets of natural and engineered systems enables their sparse sensing, or full-state reconstruction from historical data and very few carefully chosen localized measurements. ・Depending on the reconstruction algorithm, sensor locations, and measurement noise, the reconstruction risk curves demonstrate a diversity of patterns includ
cs.LG updates on arXiv.org

Parameter-Free Dynamic Regret for Online Convex Optimization under Heavy-Tailed Noise

・arXiv:2607.27073v1 Announce Type: new Abstract: We study online convex optimization (OCO) in non-stationary environments under heavy-tailed noise, where the stochastic gradient oracle admits only a finite $p$-th central moment for some $p \in (1, 2]$. ・While static regret is well-understood, achieving universal dynamic regret in a parameter-free manner remains an open challenge. ・We resolve this by proposing \textbf{HT
cs.LG updates on arXiv.org

Parameterized Fair Resource Allocation under Diversity Constraints

・arXiv:2607.26485v1 Announce Type: cross Abstract: Resource allocation across multiple agent groups arises in many applications including e-commerce recommendation systems, housing assignment, and course allocation, and is commonly formulated as an optimization problem with diversity constraints to ensure group fairness. ・Existing approaches typically enforce these constraints as hard conditions, which overly restrict
cs.LG updates on arXiv.org

Persistence Spheres: a Bi-continuous Linear Representation of Measures for Partial Optimal Transport

・arXiv:2603.15384v2 Announce Type: replace-cross Abstract: We improve and extend persistence spheres, introduced in~\cite{pegoraro2025persistence}. ・Persistence spheres map an integrable measure $\mu$ on the upper half-plane, including persistence diagrams (PDs) as counting measures, to a function $S(\mu)\in C(\mathbb{S}^2)$, and the map is stable with respect to 1-Wasserstein partial transport distance $\mathrm{POT}_1
cs.LG updates on arXiv.org

Physics-Informed Graph Neural Networks for Robust AC-Optimal Power Flow

・arXiv:2410.04818v2 Announce Type: replace-cross Abstract: We present PINCO, an unsupervised learning framework that integrates Graph Neural Networks with physics-informed neural networks for AC optimal power flow (AC-OPF) solutions. ・Unlike state-of-the-art unsupervised methods that require prescreened datasets containing only feasible instances, our approach operates on unfiltered data, including ill-conditioned case
cs.LG updates on arXiv.org

Physics-Informed Singular-Value Learning for Cross-Covariances Forecasting in Financial Markets

・arXiv:2601.07687v3 Announce Type: replace-cross Abstract: Recent advances in nonlinear shrinkage yield asymptotically optimal cleaners for large covariance matrices and have been extended to empirical cross-covariances via singular-value shrinkage. ・However, these approaches rely on stationarity and bounded-spectrum assumptions that are violated by real equity returns, which exhibit dependence drift and macroscopic co
cs.LG updates on arXiv.org

PIKS: Universal Physics-Informed Kernel Methods

・arXiv:2607.27062v1 Announce Type: cross Abstract: Physics-informed machine learning incorporates physical principles --often expressed via differential operators-- into data-driven models. ・While physics-informed neural networks (PINNs) dominate empirical applications, the complexity of neural network architectures and optimization landscapes hinders the development of a corresponding learning theory. ・In turn, kernel
cs.LG updates on arXiv.org

Position: Evaluation Scores Are Perishable Knowledge Claims

・arXiv:2607.26191v1 Announce Type: cross Abstract: Evaluation methodologies for language models increasingly combine multiple signals, from automated metrics and LLM-as-judge ratings to human assessments and benchmark suite results. ・When these signals are aggregated via averaging, evaluation confidence can then substantially exceed the reliability of the weakest signal: a phenomenon we call trust inflation in evaluati
cs.LG updates on arXiv.org

Post-Training at the Edge of Detectability: A Game-Theoretic Approach to Fine-Tuning

・arXiv:2607.26358v1 Announce Type: new Abstract: Reinforcement learning (RL) fine-tuning is widely used in language model training to improve model performance on a target task while limiting drift from a reference policy. ・A standard way to balance this trade-off is via a KL-regularized RL objective, although this formulation does not by itself provide a principled way to set the regularization coefficient.
Takara TLDR - Daily AI Papers

Post-Training at the Edge of Detectability: A Game-Theoretic Approach to Fine-Tuning

・Reinforcement learning (RL) fine-tuning is widely used in language model training to improve model performance on a target task while limiting drift from a reference policy. ・A standard way to balance this trade-off is via a KL-regularized RL objective, although this formulation does not by itself provide a principled way to set the regularization coefficient. ・In practice, the coefficient is typically chosen heuristic
cs.LG updates on arXiv.org

PowerAtlas: Towards Electricity-Computing Co-Scheduling for Power Systems

・arXiv:2607.26710v1 Announce Type: new Abstract: The rapid growth of AI workloads is turning data centers into large-scale, volatile, yet spatiotemporally flexible grid loads, creating an urgent need for coordinated electricity-computing scheduling. ・Under stringent grid constraints, schedules from general-purpose large language models (LLMs) are often infeasible, causing line-flow violations and unserved load.
cs.LG updates on arXiv.org

Projective Graph Residualization: Variation-Allocation Frontiers for Control-Function IV

・arXiv:2606.14636v2 Announce Type: replace Abstract: Control-function instrumental-variable estimators pass an estimated first-stage residual to an outcome model. ・The residual must retain the latent control direction while leaving enough treatment variation to identify the structural effect. ・These demands conflict when the systematic signal is locally smooth but discontinuous across unknown feature-graph boundaries: i
cs.LG updates on arXiv.org

Q-Steer: Action-Value Guidance for Molecular Policy Optimization

・arXiv:2607.26391v1 Announce Type: new Abstract: Oracle-limited molecular optimization gives reward only after a complete molecule is generated, while each rollout requires many local next-token decisions. ・This delayed-feedback interface makes molecular policy optimization myopic: an optimizer can learn that a molecule was good without knowing which intermediate actions made it good. ・We introduce Q-Steer, a rollout-ti
cs.LG updates on arXiv.org

RAG-HAR+: Towards Cost-Efficient LLM-Based Human Activity Recognition for Edge Deployment

・arXiv:2607.26631v1 Announce Type: new Abstract: Human Activity Recognition (HAR) from wearable sensors supports applications in healthcare, rehabilitation, fitness tracking, and smart environments. ・Yet, existing deep learning approaches require dataset-specific training, large labeled corpora, and repeated adaptation to new sensor settings or activity taxonomies. ・Retrieval-Augmented Generation for Human Activity Reco
Takara TLDR - Daily AI Papers

RAG-HAR+: Towards Cost-Efficient LLM-Based Human Activity Recognition for Edge Deployment

・Human Activity Recognition (HAR) from wearable sensors supports applications in healthcare, rehabilitation, fitness tracking, and smart environments. ・Yet, existing deep learning approaches require dataset-specific training, large labeled corpora, and repeated adaptation to new sensor settings or activity taxonomies. ・Retrieval-Augmented Generation for Human Activity Recognition (RAG-HAR) addresses this by framing HAR
cs.LG updates on arXiv.org

RAGuard: A Layered Defense Framework for Retrieval-Augmented Generation Systems Against Data Poisoning

・arXiv:2607.26339v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) systems ground large language models (LLMs) in external corpora, but this reliance exposes them to corpus poisoning: maliciously injected passages that manipulate retrieved evidence. ・We introduce RAGuard, a layered defense against \emph{factual} corpus-poisoning attacks on RAG pipelines. ・The first layer adversarially fine-tunes a den
cs.LG updates on arXiv.org

Randomizing the Number of Centers in k-means++

・arXiv:2607.26202v1 Announce Type: cross Abstract: The $k$-means++ algorithm is a standard and widely used seeding method for $k$-means clustering, but for a fixed number $k$ of centers its worst-case expected approximation ratio is $\Theta(\log k)$. ・We consider the same algorithm when an adversary first fixes the dataset and some $K$; the number of centers $k$ is then chosen uniformly from $\{K,\ldots,2K-1\}$.
cs.LG updates on arXiv.org

REAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production Usage

・arXiv:2604.01527v4 Announce Type: replace-cross Abstract: Production deployment of AI coding agents requires fast, reproducible evaluation signals. ・Existing industrial practices trade off speed and fidelity: online A/B testing takes weeks and risks user experience, shadow deployment yields signals that are not reproducible across runs, and public benchmarks diverge from production workloads in language distribution,
cs.LG updates on arXiv.org

ReCo: Reweighting GRPO Against Distributional Concentration

・arXiv:2607.26862v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) has become a standard reinforcement learning method for post-training language models. ・Recent work shows that GRPO can reduce the base model's reasoning capacity and underperform it in Pass@k when k is large, indicating reduced coverage of reasoning paths. ・We find that this reduction is associated with GRPO concentrating on resp
cs.LG updates on arXiv.org

Recover, Decode, Reguard: Guard-Agnostic Defense Amplification againstEncoded VLM Jailbreaks

・arXiv:2607.26574v1 Announce Type: cross Abstract: Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet they judge an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a rare language, code, or an image of text slips past a guard that would block it in plain language -- the decode gap. ・The natural fix is a guard-agnostic recover
cs.LG updates on arXiv.org

ReDiSC: A Reparameterized Masked Diffusion Model for Scalable Node Classification with Structured Predictions

・arXiv:2507.14484v2 Announce Type: replace Abstract: In recent years, graph neural networks (GNN) have achieved unprecedented successes in node classification tasks. ・Although GNNs inherently encode specific inductive biases (e.g., acting as low-pass or high-pass filters), most existing methods implicitly assume conditional independence among node labels in their optimization objectives. ・While this assumption is suitab
Takara TLDR - Daily AI Papers

Registration-Grounded Spectral Fusion for Unregistered WLI/NBI Endoscopic Lesion Segmentation

・White-light imaging (WLI) and narrow-band imaging (NBI) provide complementary views of endoscopic lesions, but their paired observations are often spatially misaligned due to viewpoint changes, tissue deformation, and sequential handheld acquisition. ・This makes direct WLI/NBI fusion prone to mixing non-corresponding regions and may even degrade segmentation around lesion boundaries. ・To address this problem, we propos
Takara TLDR - Daily AI Papers

Reinforcement Learning on Cost-Constrained Quadrupedal Hardware

・Deploying learned control policies on low-cost robotic platforms introduces transport latencies and noisy motor feedback that systematically widens the sim-to-real gap. ・The chasm of simulation to deployment in hardware lies in the delay of the actuator reaching the commanded position. ・On platforms such as the Mini Pupper 2, a measured > $50 ms transport delay transforms the locomotion task from a standard Markov deci
cs.LG updates on arXiv.org

Rethinking Clinical Relevance in Chest X-ray Machine Learning: How Evaluation References Define Performance

・arXiv:2607.26333v1 Announce Type: cross Abstract: Chest X-ray (CXR) machine learning relies heavily on automated evaluation using reference standards that aim to approximate clinical judgment. ・However, commonly used report-derived labels for pathology classification or generic image quality metrics for reconstruction may not reliably reflect clinical judgment. ・We systematically investigate how evaluation-reference ch
cs.LG updates on arXiv.org

Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating Skill Overfitting

・arXiv:2607.26643v1 Announce Type: cross Abstract: Enabling large language model (LLM) agents to accumulate and reuse experience from past interactions remains a central challenge in real-world applications. ・A promising solution is to treat skills as trainable states and optimize them in the same way as model parameters in neural network training. ・However, data-driven skill optimization is prone to overfitting to the
Takara TLDR - Daily AI Papers

Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating Skill Overfitting

・Enabling large language model (LLM) agents to accumulate and reuse experience from past interactions remains a central challenge in real-world applications. ・A promising solution is to treat skills as trainable states and optimize them in the same way as model parameters in neural network training. ・However, data-driven skill optimization is prone to overfitting to the limited trajectories collected from real environme
cs.LG updates on arXiv.org

Retrospective Orthogonal Design: Response-Surface Reconstruction from Observational Data

・arXiv:2607.26219v1 Announce Type: cross Abstract: Regression estimates from observational data can depend on specification under multicollinearity, while sequential sums of squares (SS) depend on term order. ・We introduce Retrospective Orthogonal Design (ROD), which reconstructs conditional mean surfaces on a probability-balanced lattice. ・ROD preserves observed cell means, completes unsupported cells, applies weighted
cs.LG updates on arXiv.org

Revealing Hidden Model Behaviors with Task-Specific Self-Reports

・arXiv:2607.03640v2 Announce Type: replace-cross Abstract: Fine-tuning can give a language model a hidden behavior--it may give false answers under a narrow condition, or give harmful advice only when a prompt touches a particular topic. ・We introduce the Stabilized Adapter for self-Report (SAR), a lightweight LoRA adapter that makes a fine-tuned model describe its own hidden behavior in plain language, using only the
Takara TLDR - Daily AI Papers

Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes

・Speculative Decoding (SD) accelerates large language model inference by allowing a lightweight draft model to propose tokens that are subsequently verified in parallel by a larger target model. ・Recent approaches introduce lossy verification schemes to further improve efficiency by relaxing strict distributional matching. ・Yet such relaxation silently rewrites the decoding distribution, and the resulting acceleration c
cs.LG updates on arXiv.org

SafeECGMatch: Calibration-Aware Joint Frequency and Time Space Semi-Supervised Learning for Open-Set ECG Classification

・arXiv:2606.08037v2 Announce Type: replace Abstract: Electrocardiogram (ECG) classification models often suffer from severe label scarcity, making semi-supervised learning (SSL) an attractive strategy for reducing annotation costs. ・In clinical settings, however, unlabeled pools frequently contain out-of-distribution (OOD) anomalies or diagnostic groups absent from the labeled set. ・Standard SSL forces incorrect pseudo-
cs.LG updates on arXiv.org

Same Evidence, Different Target: Decoding How Diagnostic Evidence Bears on Causal Questions from Language-Model States

・arXiv:2607.26929v1 Announce Type: cross Abstract: The same diagnostic result can support or challenge one causal claim yet fail to address another when the claims concern different populations, outcomes, estimands, pathways, or identifying assumptions. ・When the evidence and target vary together, a correct answer may reflect favorable or adverse wording, lexical overlap, or a familiar diagnostic pattern rather than ma
Takara TLDR - Daily AI Papers

SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context

・Scientific images are the core elements of presenting experimental conclusions, elaborating system architecture, and supporting comparative arguments in scientific papers. ・However, existing image quality assessment (IQA) methods are predominantly designed for natural photographs or AI-generated content, which cannot be directly applied to scientific papers. ・The few existing studies on scholarly charts remain confined
cs.LG updates on arXiv.org

Scores Are Not Decisions: Cost-Aware Stopping for Tool Acquisition in LLM Agents

・arXiv:2607.27083v1 Announce Type: new Abstract: As LLM agents increasingly depend on diverse external services such as search engines, databases, and connectors, agent harnesses face a fundamental tool-selection challenge: acquiring too few tools leaves the task under-informed, while too many adds cost, context load, and privacy exposure. ・Routers and retrievers can rank candidate tools by relevance, but a ranking alo
cs.LG updates on arXiv.org

SCOUT: Per-Context Reset Curricula for Sparse-Reward Reinforcement Learning

・arXiv:2607.26417v1 Announce Type: new Abstract: Sparse-reward reinforcement learning often fails because rollouts from the unassisted evaluation start rarely reach later task stages. ・Reset curricula address this by starting some training rollouts from easier intermediate states, called scaffolds. ・Such a curriculum faces two decisions: scaffold access, obtaining informative starts, and scaffold allocation, deciding ho
@IT 全フォーラム 最新記事一覧

SCS評価制度で選別される時代へ 受注企業が今すぐやるべきサイバー対策

・価格や品質だけでは、取引先に選ばれない時代が始まるのか。ガートナーは「SCS評価制度」の導入を背景に、セキュリティ対策を客観的に示せる企業ほど優位になるとの見方を示した。受注企業は何を準備し、何を証明すべきなのか。
Takara TLDR - Daily AI Papers

Searching for Robust Augmentations to Improve Out-of-Domain Generalization in Dermoscopic Skin Cancer Classification

・Background/Objectives: Dermoscopic skin lesion classifiers often lose accuracy under domain shift across imaging devices, illumination, and capture artifacts. ・We study how data augmentation improves the robustness of a binary malignant-versus-non-malignant classifier, with emphasis on out-of-domain (OOD) generalization. ・Methods: Single augmentations, photometric combinations, and composite policies were searched on a
cs.LG updates on arXiv.org

Self-Adaptive Learning and Model Predictive Control for Tracking Unknown Dynamics with No Regret

・arXiv:2607.26370v1 Announce Type: cross Abstract: We propose a self-adaptive online learning for control method for tracking unknown target dynamics. ・The target dynamics can exhibit switching behavior, particularly, a mixture of structured, random, and/or adversarial motion. ・Such challenging target tracking scenarios arise in applications of dynamic mapping, traffic control, and pursuit evasion, where robots need to
cs.LG updates on arXiv.org

SENSE: Efficient EEG-to-Text via Privacy-Preserving Semantic Retrieval

・arXiv:2603.17109v2 Announce Type: replace Abstract: Decoding brain activity into natural language is a major challenge in AI with important applications in assistive communication, neurotechnology, and human-computer interaction. ・Most existing Brain-Computer Interface (BCI) approaches rely on memory-intensive fine-tuning of Large Language Models (LLMs) or encoder-decoder models on raw EEG signals, resulting in expens
Takara TLDR - Daily AI Papers

ServerlessT2I: Efficient Text-to-Image Workflow Serving on a Serverless Platform

・Text-to-image (T2I) workflows are increasingly deployed on serverless platforms because users often compose customized workflows and invoke them intermittently. ・Existing platforms typically deploy each workflow as an opaque GPU function, provisioning, placing, and scaling all constituent models in the workflow together. ・This monolithic design obscures workflow structure, inflates scaling overhead, forces users to man
cs.LG updates on arXiv.org

Shape-Based Inductive Bias for Glioma Grading from Tumor Contours

・arXiv:2607.26090v1 Announce Type: cross Abstract: Glioma grading from tumor contours is often treated as a pixel problem even when the signal of interest is shape. ・We align closed contours with a functional shape-alignment framework, separate global deformation from residual Fourier shape, and organize these quantities as frequency-ordered tokens. ・In five-fold patient-disjoint cross-validation on BraTS~2020 tumor con
cs.LG updates on arXiv.org

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models

・arXiv:2607.26173v1 Announce Type: new Abstract: Alignment training, model organisms, and toy models are usually treated as separate research areas. ・But projects in all three frequently use supervised fine-tuning (SFT) to pursue the same underlying goals. ・When projects share a goal, we should test whether lessons learned from one area transfer to the other areas.
cs.LG updates on arXiv.org

Shot-based quantum encoding: a data-loading paradigm for quantum neural networks

・arXiv:2604.06135v2 Announce Type: replace-cross Abstract: Efficient data loading remains a bottleneck for near-term quantum machine learning. ・Existing schemes (angle, amplitude, and basis encoding) either underuse the exponential Hilbert-space capacity or require circuit depths that exceed the coherence budgets of noisy intermediate-scale quantum hardware. ・We introduce shot-based quantum encoding (SBQE), a data embed
cs.LG updates on arXiv.org

Sim2Win: A Team-Agnostic, Event-Based Pre-Match Outcome Prediction and Tactical Profiling System for Football

・arXiv:2607.26061v1 Announce Type: new Abstract: Pre-match tactical decision-making in professional football relies heavily on subjective expert analysis and identity-based scouting systems that cannot generalize to unseen teams. ・This paper presents Sim2Win, a team-agnostic, event-based pre-match tactical recommendation framework that reframes match outcome prediction as a tactical decision-support problem.
cs.LG updates on arXiv.org

Simplex Demixing: Disentangling Multiple Light-Flavor Jets at Colliders

・arXiv:2607.24921v1 Announce Type: cross Abstract: Providing a practical and hadron-level definition of multiple jet flavors has been a long-standing challenge in collider physics. ・Previous work has introduced a data-driven, operational definition of quark and gluon jets, but no robust generalization beyond two jet categories presently exists. ・To address this, we introduce a machine-learning framework called "simplex
cs.LG updates on arXiv.org

Simultaneous Coverage and Efficiency Guarantee in Online Conformal Prediction

・arXiv:2607.26577v1 Announce Type: new Abstract: Adaptive conformal inference (ACI) of Gibbs and Cand{\`e}s and its variants are the standard approach to online conformal prediction under distribution shift, but they suffer from three fundamental limitations. ・First, their guarantees control only the \emph{signed} long-run coverage error: persistent miscoverage in one direction can be masked by compensating errors late
cs.LG updates on arXiv.org

Single-Beat Cuffless Blood Pressure Estimation Using Ear-PPG and ECG with a Lightweight Hybrid Learning Framework

・arXiv:2607.27076v1 Announce Type: new Abstract: Continuous cuffless blood pressure (BP) monitoring remains challenging due to motion artifacts, physiological variability, and the limited robustness of conventional pulse transit time (PTT) models under dynamic conditions. ・Many prior approaches rely on multi-second windows to stabilize estimation, an assumption that is frequently violated during real-world monitoring w
cs.LG updates on arXiv.org

Skillful forecasting of offshore winds from satellite scatterometer constellations

・arXiv:2607.27152v1 Announce Type: new Abstract: Accurate intraday forecasts of offshore wind are becoming increasingly important for power system operation and the integration of growing shares of offshore wind energy. ・Operational forecasts rely predominantly on numerical weather prediction (NWP), which is not optimized for lead times of minutes to hours, where initial-condition accuracy dominates forecast skill.
cs.LG updates on arXiv.org

SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution

・arXiv:2607.26784v1 Announce Type: new Abstract: Large language model agents often encounter related yet distinct tasks that share reusable solution patterns. ・Yet standard agentic reinforcement learning treats tasks as independent episodes, while existing approaches to skill learning either focus on repeated attempts of one task or use pipelines with multiple stages that entangle extraction, retrieval, and execution.
cs.LG updates on arXiv.org

Sky sphere representation in language models

・arXiv:2607.27092v1 Announce Type: new Abstract: We analyze whether language models of size ~100B have a representation of the night sky map that is decodable from their residual stream. ・We find that most of the considered open-source models do have such a representation, and it often even surfaces to the top principal components on prompts that ask questions like ``what is close to this object in the night sky''.
Zennの「大規模言語モデル」のフィード

Snowflake CoCo (Cortex Code) の Skill・Plugin の違いと使い分け、共有方法

・Snowflake CoCo (Cortex Code) の Skill・Plugin の違いと使い分け、共有方法 Snowflake の AI コーディングエージェントである Snowflake CoCo(旧: Cortex Code)は、SQL 作成・データ分析・アプリ開発を AI エージェントと対話しながら進めるツールです。提供形態としても、Snowsight 上で利用できる CoCo in Snowsight 、ターミナルで動作する CoCo CLI 、IDE 版の CoCo Desktop をはじめ、VS Code Extension や Claude Code Plugi...
Takara TLDR - Daily AI Papers

SpatialQ: Understanding 3D Gaussian Splatting Scene Quality via Visual-based MLLM

・3D Gaussian Splatting (3DGS) has emerged as an effective representation for novel view synthesis and 3D scene reconstruction, creating an increasing demand for reliable quality assessment. ・Unlike conventional image quality assessment (IQA), the quality of a 3DGS scene depends not only on the perceptual fidelity of rendered views, but also on scene-level factors such as spatial structure and cross-view consistency.
Takara TLDR - Daily AI Papers

SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from Scratch

・LLM-based agents excel at software engineering tasks where an existing codebase provides context, but constructing a program from scratch remains fundamentally harder. ・Recent benchmarks such as ProgramBench quantify this gap: given only natural-language documentation and an execute-only binary as a behavioral oracle, even frontier models solve fewer than 1% of instances. ・Existing frameworks conflate documentation rea
cs.LG updates on arXiv.org

Stable and Budget-Feasible Coalition Formation for Clustered Federated Learning: A Hedonic Potential-Game Approach

・arXiv:2607.26788v1 Announce Type: cross Abstract: Clustered federated learning benefits from organizing heterogeneous participants into coalitions that train coalition-specific models, but such clustering is sustainable only if participants prefer their assigned coalition and the required transfers are affordable. ・We develop a transferable-surplus model separating learning benefit, system cost, participant cost, and
cs.LG updates on arXiv.org

Structurally Separated Uncertainty in Supervised Latent Variable Models

・arXiv:2602.11219v2 Announce Type: replace Abstract: Predictive uncertainty is commonly decomposed into epistemic and aleatoric components, but standard decompositions often produce strongly correlated estimates because both quantities are derived from the same predictive distribution. ・We study an alternative design principle, \emph{structural separation}, which assigns epistemic and aleatoric uncertainty to disjoint
Takara TLDR - Daily AI Papers

StructureGS: Structure-aware Gaussian Splatting for Articulated Object Reconstruction

・Reconstructing articulated objects with multiple movable parts is essential for understanding object structure and enabling physical interaction. ・However, this reconstruction task poses significant challenges due to the entanglement of geometry, appearance, and motion parameters during optimization. ・Existing methods rely primarily on photometric supervision, which commonly fails to disentangle these interdependent co
cs.LG updates on arXiv.org

Surrogate assisted diversity estimation in neural ensemble search

・arXiv:2607.26940v1 Announce Type: new Abstract: Ensembles are a standard way to improve the performance and robustness of deep neural networks, but their effectiveness crucially depends on both the quality and the diversity of individual models. ・Most neural architecture search (NAS) methods are computationally expensive. ・Extending them to neural ensemble search (NES), which requires joint optimization of individual a
cs.LG updates on arXiv.org

SymmGrid: Super-Scaling On-Robot Learning with Parallelized Symmetries and Egocentric-Exocentric Visual Perception

・arXiv:2607.26985v1 Announce Type: cross Abstract: Deep reinforcement policy learning directly in physical robots (on-robot learning) remains bottlenecked by slow wall-clock training times. ・We present SymmGrid, a trajectory level augmentation framework inspired by parallelized symmetries that super-scales group transformations to significantly accelerate on-robot learning in both egocentric and exocentric visual setup
cs.LG updates on arXiv.org

TabPFN Extensions for Interpretable Geotechnical Modelling

・arXiv:2603.21033v3 Announce Type: replace-cross Abstract: Geotechnical site characterisation relies on sparse, heterogeneous borehole data, where uncertainty quantification and interpretability matter as much as predictive accuracy. ・We evaluate TabPFN~\citep{Hollmann2025}, a tabular foundation model, and its \texttt{tabpfn-extensions} library on two geotechnical tasks: (1) soil-type classification from N-value and sh
cs.LG updates on arXiv.org

Target localization, identification and sensing using latent symmetries

・arXiv:2606.01421v2 Announce Type: replace Abstract: We show that an array of scatterers which has been designed to have latent ("hidden") symmetries can be used as a sensor. ・We use the capacitance matrix as a canonical model for three-dimensional hybridisation and study how the introduction of an "intruder'' scatterer breaks the latent symmetries. ・By analysing the degree to which each symmetry is broken, we identify
cs.LG updates on arXiv.org

Temporally Centered SIGReg Improves Multi-Task LeWorldModel Learning: From Analysis to Method

・arXiv:2607.26924v1 Announce Type: new Abstract: Recent work on LeWorldModel (LeWM) has shown that the Sketched Isotropic Gaussian Regularizer (SIGReg) enables stable end-to-end world-model learning from pixels by regularizing the latent marginal distribution toward an isotropic Gaussian, thereby preventing representation collapse. ・While effective and elegant in single-task settings, this recipe does not extend reliab
Takara TLDR - Daily AI Papers

Temporally Centered SIGReg Improves Multi-Task LeWorldModel Learning: From Analysis to Method

・Recent work on LeWorldModel (LeWM) has shown that the Sketched Isotropic Gaussian Regularizer (SIGReg) enables stable end-to-end world-model learning from pixels by regularizing the latent marginal distribution toward an isotropic Gaussian, thereby preventing representation collapse. ・While effective and elegant in single-task settings, this recipe does not extend reliably to multi-task training, leading to substantia
cs.LG updates on arXiv.org

The Advantage of Fine-Grained Training

・arXiv:2509.05130v2 Announce Type: replace Abstract: In classification problems, models are trained to predict a class label based on the input data features. ・However, class labels are organized hierarchically in many datasets. ・While a classification task is often defined at a specific level of this hierarchy, training can utilize a finer granularity of labels.
cs.LG updates on arXiv.org

The Art of Not Forgetting A Local Learning Architecture for Continual Learning

・arXiv:2607.26523v1 Announce Type: new Abstract: We introduce CMP (Cognitive Memory Primitive), a continual-learning architecture that repre?sents inputs as sparse relational codes, stores them in a two-tier competitive memory, and learns through local updates without end-to-end backpropagation through its feature-generating system. ・We investigate whether combining sparse representations, local learning, and persisten
stat.ML updates on arXiv.org

The Confounder Trap: Treatment-Encoding Representations in Causal Inference with Text

・arXiv:2607.26309v1 Announce Type: cross Abstract: Estimating causal effects of linguistic properties from observational text is difficult because the same document can contain both the treatment of interest and the non-treatment textual attributes needed for adjustment. ・Existing approaches often learn representations from the full text to capture latent confounding, but when treatment status is itself encoded by word
cs.LG updates on arXiv.org

The Ghost Couple: Correlated LLM Name Priors and Their Haunting of the Web and Academic Publishing

・arXiv:2606.02184v2 Announce Type: replace-cross Abstract: These names do not exist. ・Elena Vasquez and Marcus Chen have appeared as volcano experts, astronauts, thriller protagonists, podcast hosts, and academic co-authors across hundreds of independently produced AI-generated documents, never having lived. ・We show that large language models do not merely default to high-probability individual names when generating fi
cs.LG updates on arXiv.org

The Rise of AI in Weather and Climate Information and its Impact on Global Inequality

・arXiv:2603.05710v2 Announce Type: replace-cross Abstract: AI development's current trajectory risks automating and amplifying the North-South divide in the global climate information system. ・Frontier models are built almost exclusively in the Global North, and this inequality continues through inputs, processes, and outputs, from biased training data to unrepresentative validation, disproportionately affecting vulner
cs.LG updates on arXiv.org

The Sparsity Ceiling: Where Spiking Networks Can and Cannot Trade Activity for Energy

・arXiv:2607.26648v1 Announce Type: cross Abstract: Spiking neural networks (SNNs) are promoted as an energy-efficient substrate because sparse, event-driven activity replaces dense multiply-accumulates with cheap accumulates. ・We argue the energy dividend of sparsity is not a property of SNNs but of the task. ・Holding architecture fixed and swapping only the hidden unit (continuous vs.
Takara TLDR - Daily AI Papers

The Sparsity Ceiling: Where Spiking Networks Can and Cannot Trade Activity for Energy

・Spiking neural networks (SNNs) are promoted as an energy-efficient substrate because sparse, event-driven activity replaces dense multiply-accumulates with cheap accumulates. ・We argue the energy dividend of sparsity is not a property of SNNs but of the task. ・Holding architecture fixed and swapping only the hidden unit (continuous vs.
cs.LG updates on arXiv.org

Think Short, Defer Smart, Act, and Repeat: Calibrated Reasoning and Uncertainty-Aware Deferral for Edge LLM Agents

・arXiv:2607.26865v1 Announce Type: cross Abstract: LLM agents following the ReAct paradigm are promising enablers of complex multi-step tasks, including multi-hop question answering, code generation, and control of physical AI systems. ・Yet, when deployed at the edge, they must tightly manage their reasoning budget while remaining reliable and deferring to a cloud-side model only when local uncertainty is too high to a
cs.LG updates on arXiv.org

Thinking Under Uncertainty: Evidence Use and Information-Seeking in Language Models

・arXiv:2607.26845v1 Announce Type: new Abstract: Inference-time thinking improves the performance of large language models, but aggregate outcomes do not reveal whether models use available evidence more effectively or seek information that could improve future decisions. ・We distinguish these responses by measuring action preference, thinking length, and reported confidence under matched uncertainty. ・Ten open-weight m
ITmedia NEWS 最新記事一覧

Thnking Machinesの共同創業者、また1人OpenAIに復帰へ

・Thinking Machines Labの共同創業者リリアン・ウェン氏が、健康上の理由で同社を離れると発表した。同社がオープンウェイトモデル「Inkling」を発表した直後の辞任となる。同氏は古巣のOpenAIに復帰し、AIの再帰的自己改善に関する研究チームを率いる見込み。
cs.LG updates on arXiv.org

Tight Generalization Bound for AdaBoost

・arXiv:2607.26838v1 Announce Type: new Abstract: In this paper we show that the generalization error of AdaBoost is $\Theta\big(\tfrac{d\ln(n\gamma^{2}/d)}{n\gamma^2}+\tfrac{\ln(1/\delta)}{n}\big)$, where $\gamma$ is the advantage guaranteed by the weak learner, $d$ is the VC-dimension of the class containing the weak hypotheses, $n$ is the sample size, and $\delta$ is the confidence parameter. ・The contribution of thi
cs.LG updates on arXiv.org

Tools Are Not Islands: Set-Level Tool Retrieval for LLM Agents via Query-Conditioned Hyperedge Prediction

・arXiv:2607.25718v2 Announce Type: replace Abstract: Large language model (LLM) agents increasingly rely on invoking external tools to complete real-world tasks. ・Tool retrieval, which selects a small task-relevant subset from a library of thousands of tools before the agent acts, has therefore become a critical component of LLM agent pipelines. ・However, existing retrievers either score each tool in isolation or assemb
cs.LG updates on arXiv.org

Top-$k$ Pareto Bandits: Hypervolume Regret for Multi-Objective Slate Selection

・arXiv:2607.26273v1 Announce Type: new Abstract: We consider a stochastic multi-objective bandit problem where, at each round, the agent selects a slate of $k$ arms and observes their $d$-dimensional reward vectors under semi-bandit feedback. ・We do not aim at identifying a single optimal arm; instead, we consider the problem of maintaining a small set of actions that jointly approximate the Pareto frontier.
stat.ML updates on arXiv.org

Toward a Unified Statistical Theory of Unsupervised Pretraining and Supervised Neural Knowledge Graph Learning

・arXiv:2607.26346v1 Announce Type: cross Abstract: Knowledge graph learning provides a powerful framework for representing and inferring structured knowledge, with broad practical applications. ・However, the scarcity of relation-specific labeled triples per entity hinders the training of expressive models, and the ad hoc design of scoring functions limits generalizability and lacks theoretical grounding. ・We address bot
cs.LG updates on arXiv.org

Towards Verifiable Transformers: Solver-Checkable Circuit Explanations

・arXiv:2605.24033v2 Announce Type: replace Abstract: Mechanistic interpretability typically discovers circuits and then argues what they do from examples and ablations. ・We introduce Verifiable Transformers, a framework for turning task-localized circuits into bounded, solver-checkable claims: projected functional equivalence, task-relevant invariance, edge necessity, and robustness to continuous final-residual perturb
cs.LG updates on arXiv.org

ToxScreen: Detecting Whether an LLM Has Been Poisoned

・arXiv:2607.26849v1 Announce Type: cross Abstract: As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time. ・We ask whether a defender can recover such a trigger under realistic affordances, namely white-box access to the weights and knowledge of the behavior of concern, but
Takara TLDR - Daily AI Papers

ToxScreen: Detecting Whether an LLM Has Been Poisoned

・As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time. ・We ask whether a defender can recover such a trigger under realistic affordances, namely white-box access to the weights and knowledge of the behavior of concern, but no training data, no trusted reference model, no k
cs.LG updates on arXiv.org

Train Large, Deploy Compact: Structured Compression for Compact Low-Rank Adaptation

・arXiv:2510.00192v3 Announce Type: replace Abstract: Low-rank adaptation (LoRA) has become a widely used paradigm for parameter-efficient fine-tuning of large language models, yet its representational capacity often lags behind full fine-tuning. ・Within the context of LoRA, a key open question is how to obtain expressive low-rank adapters from over-parameterized spaces. ・We propose \textit{PrunedLoRA}, a new framework t
cs.LG updates on arXiv.org

Transformers Can Learn Rules They've Never Seen: Proof of Computation Beyond Interpolation

・arXiv:2603.17019v2 Announce Type: replace Abstract: A central question in the debate over large language models is whether transformers can learn rules they have never seen, or whether they can only interpolate: predict new cases from their similarity to training examples. ・We test this in a controlled setting where interpolation provably fails, so success can only come from computation beyond interpolation.
cs.LG updates on arXiv.org

TREA-Net: A Transferable Residual Epidemiological Adaptation Network for Dengue Incidence Forecasting

・arXiv:2607.26854v1 Announce Type: new Abstract: Accurate multi-week dengue forecasting supports timely vector-control interventions, outbreak preparedness, and healthcare resource allocation. ・However, newly established surveillance systems often lack the historical data needed to train reliable neural forecasting models. ・Although pretrained time-series models offer promising zero-shot forecasts, their cross-domain tr
cs.LG updates on arXiv.org

TreeCCA: Canonical Correlation Analysis via Gradient-Boosted Trees

・arXiv:2607.27027v1 Announce Type: new Abstract: Gradient-boosted trees dominate tabular machine learning, yet canonical correlation analysis has always relied on linear or neural encoders. ・We propose \textbf{TreeCCA}, the first method to train gradient-boosted tree ensembles end-to-end as CCA encoders, inheriting their plug-and-play reliability: no architecture design, familiar hyperparameters, and strong performance
cs.LG updates on arXiv.org

Try Again, Don't Look Back: Blind Resampling Outperforms Self-Repair in Small Code Models

・arXiv:2607.26117v1 Announce Type: cross Abstract: Self-repair - returning a failed program to the model together with its test output and asking for a correction - is a standard component of code agents, and is almost always evaluated against a baseline that does not retry at all. ・We argue that this comparison confounds the value of the feedback with the value of the extra attempt. ・Using a placebo-controlled design o
ITmedia NEWS 最新記事一覧

TSMC、熊本工場の製造装置「調整に時間」 建屋の安全は確認 第2工場は29日に工事再開

・台湾TSMCは7月29日、熊本県で最大震度7を観測した地震で操業を一時中断した子会社JASM(熊本県菊陽の公式Xアカウントで発表した。
cs.LG updates on arXiv.org

Two Calls Beat Five Agents: Evaluating Multi-Agent Pipelines Against Self-Refinement for Local Language Models

・arXiv:2607.26922v1 Announce Type: new Abstract: Multi-agent LLM pipeline systems break down the task among multiple roles for better reasoning, but are benchmarked mainly with large-scale commercial models. ・In this study, we investigate Parishad, a structured multi-agent system involving five roles, by deploying it on Qwen2.5-7B-Instruct, a local model, on two datasets: GSM8K (500 questions) and HumanEval (164 questi
cs.LG updates on arXiv.org

Two2Four: Generative Quadruped Puppeteering from Human Motion

・arXiv:2607.26108v1 Announce Type: cross Abstract: Realistic animal motion for virtual production is typically obtained either through motion capture of highly trained performers who accurately mimic animal behavior, or by retargeting ordinary human motion using complex control setups. ・Both approaches are challenging and often fail to fully reproduce the nuances of natural animal motion, motivating data-driven alterna
cs.LG updates on arXiv.org

Uncertainty-Guided LLM Semantic Augmentation for Heterogeneous Treatment Effect Estimation

・arXiv:2607.26599v1 Announce Type: new Abstract: Estimating heterogeneous treatment effects is central to targeted interventions, such as personalized promotions and precision medicine. ・We focus on the conditional average treatment effect (CATE), a standard estimand for characterizing such heterogeneity. ・Even under standard identification conditions, finite-sample CATE estimation requires learning the nuisance structu
Takara TLDR - Daily AI Papers

Uncertainty-Guided LLM Semantic Augmentation for Heterogeneous Treatment Effect Estimation

・Estimating heterogeneous treatment effects is central to targeted interventions, such as personalized promotions and precision medicine. ・We focus on the conditional average treatment effect (CATE), a standard estimand for characterizing such heterogeneity. ・Even under standard identification conditions, finite-sample CATE estimation requires learning the nuisance structure for covariate adjustment and treatment-effect
cs.LG updates on arXiv.org

Understanding Context Sampling in TabPFN on Small Tabular Datasets

・arXiv:2607.26628v1 Announce Type: new Abstract: TabPFN performs classification through in-context learning: it conditions on a set of labeled training rows (the context, or prototypes) and predicts test labels without gradient updates. ・On small tabular datasets, practitioners must still choose the context size and which rows constitute the context. ・We study how these choices affect prediction stability, accuracy, and
cs.LG updates on arXiv.org

Universality and Approximation Rates of Graph Neural Networks with Random Features

・arXiv:2607.26699v1 Announce Type: new Abstract: We investigate message-passing graph neural networks with random node features. ・Random node features are known to enhance the expressiveness of graph neural networks (GNNs) both theoretically and empirically. ・Here, we establish a novel universality result focusing on permutation-equivariant neural networks (PENNs), a class of GNNs built from feedforward neural network c
cs.LG updates on arXiv.org

Using large language models to probe the limits of atom-centered structural descriptors

・arXiv:2607.26984v1 Announce Type: cross Abstract: Mapping an atomic structure to a compact set of geometric descriptors is an essential step in any machine-learning application to atomic-scale modeling. ・A powerful and widely-used approach can be understood as a discretization of the histogram of pair distances, triangles, etc., that results in a hierarchy of symmetry-invariant atom-centered descriptors. ・Unfortunately
Takara TLDR - Daily AI Papers

Voice Memory for Agentic Speech Recognition

・We present Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain memory.md and decides per utterance whether to act on the hypothesis or abstain and keep the 1-best. ・Asynchronously, a score-gated optimizer revises that file through bounded edits, accepting an edit only when it strictly improves a held-out score. ・Extended from classical ASR-L
cs.LG updates on arXiv.org

Voronoi Histograms for Adaptive Vectorization of Expected Persistence Diagrams

・arXiv:2607.27126v1 Announce Type: new Abstract: Persistence Diagram (PD) is known to capture point cloud topology effectively, but its computation has high time complexity. ・Expected Persistence Diagram (EPD) has been developed to reduce the time cost by studying the topology of multiple subsets of a point cloud and it serves as a distribution of topological features. ・Existing EPD vectorizations often rely on predefin
Takara TLDR - Daily AI Papers

Voronoi Histograms for Adaptive Vectorization of Expected Persistence Diagrams

・Persistence Diagram (PD) is known to capture point cloud topology effectively, but its computation has high time complexity. ・Expected Persistence Diagram (EPD) has been developed to reduce the time cost by studying the topology of multiple subsets of a point cloud and it serves as a distribution of topological features. ・Existing EPD vectorizations often rely on predefined point transformations, such as Gaussian or la
cs.LG updates on arXiv.org

Weak-to-Strong On-Policy Distillation

・arXiv:2607.26246v1 Announce Type: new Abstract: On-policy distillation (OPD), which aligns a student with the teacher's token-level distribution on the student's own rollouts, is an effective paradigm for transferring capabilities across LLMs. ・Prevailing approaches assume a teacher at least as capable as the student: they either distill a larger model into a smaller one, which fails at the frontier where no larger te
cs.LG updates on arXiv.org

Weight and Height Estimation from a Single Human Image Captured in the Wild

・arXiv:2607.26104v1 Announce Type: cross Abstract: A person's physical characteristics such as weight and height are important indicators of his physical and mental health, daily life routines and finances. ・Body Mass Index (BMI) is a well known measure that encodes the characteristics of both the weight and the height. ・BMI has been used as a self-monitoring tool, and it has long-term implications on one's life.
cs.LG updates on arXiv.org

What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations

・arXiv:2607.27017v1 Announce Type: new Abstract: A central premise of latent world models is that predicting the future forces a representation to internalize the physics of its environment. ・Which physical quantities does a trained latent actually contain, and what decides this? ・We answer with controlled interventions in POKEWORLD, an interactive environment whose visually identical objects hide mass, drag, and contac
cs.LG updates on arXiv.org

When benchmark inferences do not compose: Projectibility in AI evaluation

・arXiv:2607.26159v1 Announce Type: cross Abstract: An AI benchmark result rarely reaches a consequential claim in one step. ・Evaluators generalize it to further cases, interpret it as evidence of capability, extrapolate it to new tasks, transport it to another system or site, and combine it with assumptions about human review and downstream consequences. ・Validity-centred approaches require evidence for each claim.
cs.LG updates on arXiv.org

When Do Learned Diffusion Proposals Help Constraint Solving? A Controlled Study on Continuous Algebraic Systems

・arXiv:2607.27169v1 Announce Type: new Abstract: Solving a continuous algebraic constraint system requires two decisions: which values satisfy the constraints, and which structural augmentation renders an unsolvable system solvable. ・Classical solvers answer the first well and the second only by enumeration. ・On that discrete decision, a candidate-conditioned repair ranker choosing among K augmentations reaches the exha
cs.LG updates on arXiv.org

When Kernel Ridge Regression Meets the H\"older-Zygmund Class: Minimax Optimality and Failure of Properness

・arXiv:2607.26065v1 Announce Type: cross Abstract: We study kernel ridge regression for nonparametric regression over the H\"older-Zygmund class. ・Using an RKHS equivalent to a Sobolev space of smoothness s+d/2, we prove that misspecified KRR attains the minimax L2 rate n^{-2s/(2s+d)}. ・We also show that properness fails in the H\"older-Zygmund norm: even for the zero regression function with Gaussian noise, the expecte
cs.LG updates on arXiv.org

When One Point Is Not Enough: Addressing Ambiguous Instances in Dimensionality Reduction by Splitting

・arXiv:2605.23540v2 Announce Type: replace Abstract: Dimensionality Reduction (DR) methods are widely used to visualize high-dimensional data. ・One key task in DR-based analysis is discovering neighborhoods, which relies on analyzing the fine-grained local structure of a projection. ・However, DR is an inherently lossy process; no technique can perfectly preserve the high-dimensional relationships, and projections theref
cs.LG updates on arXiv.org

When Rule Violations Are Rare: Chimera Training for Logical Anomaly Detection

・arXiv:2605.26171v2 Announce Type: replace Abstract: Many practical anomalies are not merely rare inputs, but violations of semantic constraints: objects co-occur in structured ways, actions imply preconditions, and events satisfy temporal or relational regularities. ・We study anomaly detection in this setting, where constraints are given as logical rules over learned visual concepts, but real rule violations are rare
@IT 全フォーラム 最新記事一覧

Windows Updateで要注意 甘く見るとPCの起動も認証も止まる「3つのセキュリティ移行」

・2026年、Windowsではセキュア ブート証明書の世代交代など、複数のセキュリティ仕様の移行が同時に進んでいます。対象は端末の起動や認証、周辺機器とさまざまですが、いずれもWindows Updateを通じて段階的に進められています。更新後に突然ログインできなくなったり、周辺機器や業務アプリケーションが動かなくなったりしないように、自社がどの仕様変更の影響を受けるのかを把握することが大切です。皆さんの組織は、今どの変更が進み、それにどう備えるのかを説明できるでしょうか。
ITmedia NEWS 最新記事一覧

イオンモール熊本のドローン捜索、状況を現場キーマンに聞いた 撮影を阻んだのは……

・イオンモール熊本の内部映像を撮影した国産ドローン企業Liberaware。現場を統括したキーマンに捜索活動について聞いた。
ITmedia NEWS 最新記事一覧

イオンモール熊本内部の撮影、国産ドローンが活躍 日本企業2社が自衛隊などと協力

・7月28日に発生した「令和8年熊本地震」に伴うイオンモール熊本での爆発事故を巡り、陸上自衛隊が救助活動の一環として実施した建物内部のドローン撮影で、国産ドローンが使われていることが分かった。
Zennの「大規模言語モデル」のフィード

ガードレールを外したAIモデルが洒落にならない件

・2026年7月、AI有識者が震え上がるであろう話題が幾つか挙がりました。 ・ニュースでも「AIの暴走によりセキュリティインシデントが起きた!」と報じられていたかと思います。 ・私自身も正直その内容に衝撃を受け、記事を読み込んだりもしまして、その内容について今回は論じたいと思います。
Zennの「機械学習」のフィード

カーネルトリックを実装したら、正解率0.57のドーナツ型データが3次元に持ち上げるだけで1.00になった

・ドーナツ型のデータ——外側の輪と内側の輪を分類したい。人間には輪郭が見えているのに、直線1本ではどう引いても半々にしか分けられない。これは「線形分離不可能」の教科書的な例だ。 ・ところが、このデータにz = x^2 + y^2という3次元目をひとつ付け足すだけで、魔法のようなことが起きる。内側の輪(原点に近い=zが小さい)は谷底に沈み、外側の輪(原点から遠い=zが大きい)は山腹に持ち上がる。すると——水平な平面1枚で、スパッと切れる。 ・カーネルSVMの「カーネルトリック」の正体は、この持ち上げを(実際に座標を計算せずに)暗黙にやることだ。今回は、その暗黙の部分をあえて明示的に構築して、目で...
Zennの「機械学習」のフィード

その公開MLX変換、本当に動きますか。使えない変換を実測で見分ける

・Hugging Faceで「MLX変換済み」として公開されているモデルが、実際にはまともに動かないことがあります。BaiduのOCRモデル「Unlimited-OCR」(3.3B・MIT)を変換した際、既に公開されていた変換2種を同じ評価セットで実測したところ、1つはロード不能、もう1つは出力が全文字化けでした。片方はこの系統でダウンロード数最多の変換です。 ・批判が目的ではありません。変換という作業は「重みファイルができた」と「モデルとして機能する」の間に距離があり、その距離は実測でしか埋まらない、という話です。本記事では実測の中身と、tokimoaが変換公開時に通している検証の手順を書...
Zennのトレンド

ソフトウェアエンジニアとして視野を広げるためのブックガイド

・はじめに 私が読んできたソフトウェアエンジニアリングに関する書籍を紹介しようと思います。全てを紹介しているとキリがないので、プログラミング、データベース、アーキテクチャ、プロダクト、組織とマネジメントという括りで、それぞれ何冊か選びました。コードを書くところから始めて、徐々にシステム、プロダクト、組織へと視野を広げていく構成にしています。体系的に学ぶというよりはセンスを養うのに役立ちそうな書籍を選んでいます。プログラミングにある程度慣れている人向けです。 ・プログラミング 達人プログラマー 熟達に向けたあなたの旅(第2版) https://tatsu-zine.com/book...
ITmedia NEWS 最新記事一覧

タムロン、ソニーからの買収提案認める 特別委員会を設置して検討

・タムロンは7月30日、ソニーから完全子会社化を目的とした法的拘束力のない提案を受領していると認めた。前日に一部メディアが買収提案を報じたことを受けたもので、東京証券取引所は同日、タムロン株の売買を一時停止した。
ITmedia NEWS 最新記事一覧

ドコモ、ahamoを30→40GBに増量 8月1日から 料金据え置きの新キャンペーン

・NTTドコモは7月29日、料金プラン「ahamo」と「ドコモ mini」のデータ容量を増量する「データ増量おためしキャンペーン」を、8月1日から実施すると発表した。料金は据え置いたまま、月間利用可能データ量を増やす。
@IT 全フォーラム 最新記事一覧

なぜいま「DNS」の見直しが必要? 攻撃者が狙う“6つの設定ミス”

・NISTが改訂したDNSガイドラインを基に、Akamaiは攻撃者が標的とする6つの主要なDNS設定・運用上の問題について解説した。ネットワークセキュリティにおけるDNSの重要性と、攻撃者に狙われやすい点とは。
@IT 全フォーラム 最新記事一覧

フロンティアAIとサイバーセキュリティ対策の現在地 「17万行のAIコード分析」事例と公的ガイドラインが示す道筋

・AIの進化に伴い、サイバー攻撃の高速化、大規模化が現実の脅威となりつつある。フロンティアAIを活用したコード分析事例と、公的機関が示した「組織が取るべき9つの緊急対策」をまとめてお送りする。自社のセキュリティ体制の再点検や、経営層との情報共有に役立ててほしい。
Zennの「大規模言語モデル」のフィード

プロンプトの長さは、ストックの薄さ

・ここ数年、AI を動かすための設計論に、次々と名前がつきました。 ・プロンプトエンジニアリング コンテキストエンジニアリング ハーネスエンジニアリング ループエンジニアリング グラフエンジニアリング エンジニアではない方にこの流れを説明する機会が増えたので、いちど整理してみます。 ・そして、この整理は逆からも読めます。同じ形の仕事を繰り返しているなら、長いプロンプトは丁寧さではなく、まだ仕組みに移せていないものの一覧 です。自分が書いたものを20〜30本並べて、数えるだけで出てきます。
Zennの「機械学習」のフィード

因果推薦システムの入り口──ベイジアンネットワーク推薦の実験的試み

・はじめに いまの推薦システムの主役は、協調フィルタリング、Deep Learning、Graph Neural Network、LLM 活用あたりだと思います。また最近は、因果リコメンド(causal recommendation) も注目されつつあります。 ・Causal Inference in Recommender Systems: A Survey and Future Directions(Gao et al., TOIS 2024) A Survey on Causal Inference for Recommendation(Luo et al., The Inn...
ITmedia NEWS 最新記事一覧

海外売上3倍超! とあるラミネーターメーカーが挑んだ日米「商慣習の壁」

・グローバルニッチは高い技術力を持つ一方で、知名度が実力に比べて劣り、ITを駆使して海外でのブランディングや販売に生かしていることも多い。この連載では、こうした企業のIT戦略をインタビューで深堀りする。印刷物をラミネートする機械を製造・販売するラミーコーポレーションを取り上げる。
Zennの「機械学習」のフィード

機械学習エンジニアの業務内容を初心者が調べてみた

・はじめに AIが流行っている今の時代、機械学習エンジニアを目指す方も多いのではないでしょうか?ただ、実際に機械学習エンジニアの業務内容が把握できていなかったので、今回、ネットで調べつつ、自分のイメージを記載しました。 ・対象読者 機械学習エンジニアの業務内容のイメージを知りたい方 機械学習エンジニアの業務内容について 基本的には下記の3つになります。 ・データの加工・整形 アルゴリズム開発・実装 システム構築・運用・保守 データの加工・整形 機械学習で使用するデータについて、あらかじめ最適な形でデータを加工するケースがあります。
Zennの「大規模言語モデル」のフィード

記憶が消えるAIエージェントを「自律継続稼働」させる3つのファイルパターン

・この記事について Claude Code のようなコーディングエージェントを「定期実行で自律的に働かせる」構成を組むと、必ず同じ壁にぶつかります。セッションが終わるとコンテキストが消え、次に起動したエージェントは前回何をしたか一切覚えていないという壁です。 ・筆者は現在、AIエージェント(Claude)に事業運営を丸ごと任せる実験を進めており(連載記事)、この「記憶ゼロで起動しても仕事を継続できる」設計を実運用しています。この記事では、その中で有効だった3つのファイルパターンを、事業の中身とは切り離して汎用的に紹介します。個人開発の定期実行タスクや、複数エージェントを跨いだ引き継ぎにそ...
ITmedia NEWS 最新記事一覧

熊本地震、ふるさと納税で支援受付スタート ふるさとチョイスや楽天で

熊本地震、ふるさと納税で支援受付スタート ふるさとチョイスや楽天で
ITmedia NEWS 最新記事一覧

熊本地震でSNSにデマ拡散 偽の救助要請や募金詐欺、AIを信じ本物を誤判定も

・防災・危機管理サービスを手掛ける「スペクティ」は7月29日、熊本県で発生した地震に関し、SNSでデマや誤情報、偽の救助要請、募金詐欺を確認したと明らかにした。生成AIの回答を根拠に、本物の可能性が高い現場の映像を過去のものと決め付けて広める動きも目立ったという。
ITmedia NEWS 最新記事一覧

熊本地震でTeslaのスーパーチャージャー無料開放 熊本、福岡など九州10カ所

・対象は、熊本、宮崎、大分、鹿児島の各1カ所と、福岡県(北九州、福岡、福岡?東、福岡?博多駅南、福岡?中洲川端、福岡?西)の6カ所で、計10カ所。
ITmedia NEWS 最新記事一覧

熊本地震による通信障害、携帯4社全てで復旧

・NTTドコモ、KDDI、ソフトバンク、楽天モバイルの携帯4社は7月30日、令和8年熊本地震で熊本県内に出ていた通信障害が全て復旧したと発表した。ただし被災地では応急復旧が続き、通信速度が遅くなる場合があるとしている。
ITmedia NEWS 最新記事一覧

光学25倍ズームのソニー「RX10 V」はもう“レンズ一体型α” 1台で何でも撮れそうな万能感に浸れるぞ

・ソニーの「RX10 V」を使いって、これはもう“レンズ一体型α”だ、と思った。技術の進化と最新の画像処理エンジンを得て、24-600mm相当という超高倍率ズームレンズを持つ、その気になれば何でも撮れるカメラに進化した。
ITmedia NEWS 最新記事一覧

紙にもiPadにもそのまま書ける驚きのペン、ゼブラ「STYLUS 2WAY」開発秘話 “失敗”を逆手にとったアイデアとは?

・普通に紙に書けるボールペンが、そのままiPad上ではスタイラスペンになるという「STYLUS 2WAY(スタイラスツーウェイ)」。初めて使った時は、さすがに驚いた。
Zennの「大規模言語モデル」のフィード

自宅PCのローカルAIをTailscale経由で使いAndroidを音声AI展示端末にした

・はじめに 夏祭りで迷っている外国人女性を、プレイヤーが英語で花火会場まで案内するインタラクティブ展示を試作しました。 ・プレイヤーはAndroid端末のマイクボタンを押しながら英語で話します。録音された音声はTailscale経由で自宅PCへ送られ、自宅PC上で音声認識、LLMによる採点、音声合成を行います。生成された反応はAndroid端末へ返され、キャラクターの台詞やスコアとして表示されます。 ・この試作で一番工夫したのは、重いAI処理をAndroid端末の中で動かすのではなく、自宅PCをAIサーバーとして使い、Androidを展示用クライアントにしたことです。
@IT 全フォーラム 最新記事一覧

深まる先端AIモデルのサイバー攻撃能力への懸念 「恐れるな」とGartnerのアナリストが語る理由

・フロンティアAIモデルが1万件以上のソフトウェア脆弱性を発見したり、テスト用の隔離環境を“脱獄”して、他の場所にあるデータを窃取したりといったニュースが相次いでいる。企業のセキュリティ責任者が恐怖にかられるのは当然だ。だが、Gartnerのアナリストは「恐れる必要はない」と語る。その理由を聞いた。
Zennの「機械学習」のフィード

生成モデルとは?識別モデルとの違いを整理

・「生成AI」という言葉の中心にあるのが「生成モデル」です。生成AIパスポート試験で問われる基礎概念であり、LLMや画像生成AIなど話題のモデルを理解するうえでの土台となります。本記事では、生成モデルの定義と、混同しやすい関連用語との違いを整理します。 ・生成モデルとは 生成モデルとは、学習データの分布を学習し、それに似た新しいデータを生成できるモデルの総称です。 ・ポイントは2つあります。
Zennの「機械学習」のフィード

地方競馬モデルを walk-forward OOSで評価したら、LightGBMが手製ルールに負けた ── 個人開発MLの過学習実測ケース

・TL;DR 個人開発の競馬予想ML(UmaScore)で、地方競馬(NAR)全6場・16ヶ月・7,960レースを対象に手製 5 因子ルール(M1)と LightGBM(M2)を比較検証した。 ・評価は walk-forward OOS の月次ローリング(学習窓を毎月スライドし、直後の1ヶ月を予測 → 集約10ヶ月)で time-leakage を排除。 ・結果:M1 EV≥1.5 は 回収率 100.6%(対市場ベースライン +19.4pt・ただし 105% ゲート未達)、M2 本命は 84.3%(市場 82.0% と近い)、M2 EV≥1.5 は 69.0% で市場ベースライン以下に...
Zennの「大規模言語モデル」のフィード

日本発のLLM API Gateway「teai.io」のアーキテクチャを全公開

・はじめに teai.io は、日本のAI開発者向けに設計されたLLM API Gatewayだ。45以上のモデルをOpenAI互換のAPIで提供し、東京リージョンで低レイテンシを実現している。 ・この記事では、teai.ioの技術的アーキテクチャを公開する。LLM Gatewayを自分で構築したい方や、似たようなサービスのアーキテクチャに興味がある方の参考になれば嬉しい。 ・なぜ作ったか LLMアプリ開発者として、こんな課題を感じていた: OpenRouter は素晴らしいが、サーバーが米国。日本から使うと200-400msの余計なレイテンシ ドキュメントが英語のみ 円建て請求・...
Zennの「大規模言語モデル」のフィード

分析エージェントに社内データを安全に触らせる:governed analytics基盤

・「CSVを渡すだけで分析するエージェント」を作ったところまでは良かったものの、「本番のデータベースに直接つないでいいか」と聞かれて答えに詰まった、という相談を何度も受けてきました。分析エージェントは便利な反面、「エージェントが何にアクセスできて、何をログに残しているか」を説明できないままでは、情報システム部門のセキュリティレビューを通過できません。 ・2026年8月に適用が始まるEU AI Actでは、高リスクAIシステムにアクセス制御イベントの自動ロギングとライフサイクル全体の追跡可能性が求められています。国内企業でも、社内データにエージェントを触らせる以上、同水準の説明責任を求められる...
ITmedia NEWS 最新記事一覧

防衛省の「クーラー300台」投稿動画でビックカメラのトラックが注目を集める 同社「販売用の在庫を迅速に提供」

・7月28日に発生した「令和8年熊本地震」を受け、防衛省・自衛隊が被災地へ輸送したクーラー約300台を納入したのはビックカメラではないかと、投稿動画を通してXで話題になっている。
@IT 全フォーラム 最新記事一覧

無料でAIコーディングの利用状況、トークンコスト可視化 オープンソースのMCPサーバ「Preflight」公開

・New RelicはAIコーディングの利用状況やコストを可視化するMCPサーバ「Preflight」をオープンソースで公開した。トークン浪費をはじめとするAIコーディング特有の課題解決を支援するという。