ai Trend Report

Dashboard へ戻る
Date: 20260821 Articles: 374 Scope: curated summary

あなたのアイデアを、今すぐ形に。

公開先に迷ったら、WebFileBinで一発公開。

HTMLをドラッグ&ドロップするだけで、すぐ公開できます。

3
High impact
5
Mid impact
366
Signal watch

なぜこのサイトを作ったのか

私たちそれぞれが個別にAIを使って情報収集し、同じような3行要約を作るたびに、世界中で膨大な電力と計算リソースが消費されています。 本プロジェクトは、あらかじめ広範な情報を取得・集約しておくことで、個別のAI実行回数を減らし、地球環境(GPU/TPU負荷)に配慮した効率的な情報収集を目指す実験的なダッシュボードです。

Domain filters
Star filters
#AIタグ

5つの未来が、もう別々に動き始めた。そして一部ルールを変更。

・2026年8月18日、2,000ドル以内から始めた小さな実験。 ・AIに未来を予測させ、その未来に必要になるものを5つに分けて株式を保有した。 ・ENERGY / GEV SPACE / SpaceX BIOLOGY / Tempus AI IDENTITY / Mitek Systems PHYSICAL / Ondas まだ、わずか数日。
cs.LG updates on arXiv.org

Generalist Vision-Language Models for Fast Radio Burst detection: a zero-shot benchmark against a specialized detector

・arXiv:2607.07382v2 Announce Type: replace Abstract: Fast Radio Burst (FRB) detection increasingly relies on specialized deep learning models that require large task-specific training sets and cannot be redefined without retraining. ・We evaluate whether small, open-weight, locally run generalist Vision-Language Models (VLMs) can detect FRBs in dynamic spectra under a zero-shot, prompt-only regime. ・On a balanced binary
#LLMタグ

LLMと言う、私たちを呑み込む、口のついたサイコロについて

・LLMは口のついた、サイコロである。私はぼんやりと、そう考える。 ・多くの人が考えるサイコロは、そのランダム性に、価値を置くことが多い。サイコロを扱うこと、それ自体が技能になる場合はありますが、今回は割愛させていただきます。 ・シャノンのゲームのご子孫がLLMだとするなら、確率というのは、基本的には、かならず、どこかしらの時点で、完全に定まるはずである(transformerの話も割愛します) ありがとう、ということばだけで、LLMができたら、あの次はほぼ100%の確率で「り」がでる。そりゃ、そうである。こういうのの、お買い得パックみたいなのがLLMの知識の中身だとする。
Zennの「大規模言語モデル」のフィード

[論文紹介] Muon: An optimizer for hidden layers in neural networks (2024)

・Airion株式会社でインターンをしている東京大学4年の吉平です。 ・AirionではAIに関する内容の勉強会を毎週開催しています。ラダー生成などを目的としてLLMの開発も行っており、LLMに関連する内容として、今回は近年のLLMの事前学習などで広く用いられているMuon optimizerについて、提案されたblogと最近取り入れられている工夫を扱いました。 ・0 Summary Adam(W)だと行列構造を意識できていなかったりしたのでより強いoptimizerを提案している(2026年現在はdefacto; なのでResultsはここでは割愛) NewtonSchulz5の箇所...
Qiita - 人気の記事

LambdaでAmazon Bedrockに任意のトークン上限を設定してみた【80%で通知・上限到達でリクエスト拒否】

・ソーイ株式会社、Webエンジニア2年目の村上です。 ・AWS上のアプリケーションから基盤モデルを利用する場合、簡単に生成AIモデルを呼び出せるAmazon Bedrockを使う機会も多いと思います。 ・使いやすくて便利なのはそうなのですが、Bedrockは従量課金...
@IT 全フォーラム 最新記事一覧

第314回 投資家注目のキーワード「光電融合」とは? パッケージ内混載からワンチップ化への道筋

・AI向け先端半導体のトレンドとして急浮上する「光電融合」。通信分野から始まった光技術による電気の置き換えは、いまやチップ直近の数ミリ、数マイクロメートルの領域に迫る。本稿では、光派と電気派の姿勢の違いや技術的進化史をひもときつつ、パッケージ内混載の最前線と端末普及への現実的な壁について解説する。
Zennの「大規模言語モデル」のフィード

同じコードレビューを5つのLLMに投げてみた Part 2——一番安いモデルが一番危険だった

・以前にも一度、サイトの一括置換スクリプトを5つのLLMに独立してレビューさせたことがある。あのときはモデルごとに「渡し方」まで自然な形に変えて試した。今回は逆に、5体全員に一言一句同じプロンプトを渡して、渡し方の違いという変数を消し、モデルの性質差だけを見てみることにした。 ・対象は、自社サイトIVYXON向けに新しく作った、企業の決算データ(有価証券報告書)と政府統計(CPI・賃金指数)をSQLベースで横断的に見て異常値・矛盾を検知する新機能——約400行のPythonスクリプトと、テーブル定義のSQL。実データで動かして93件の異常値を検出済みという、動くものはできた状態でのレビューだ...
#AIタグ

有料noteが読まれない…と悩んでいた私が、Instagramリール1本で変わったこと

・Xだけで投稿を続けていたけれど、 有料noteを作っても なかなか読まれない、売れない。 ・何が足りないのか分からないまま、 もどかしい日々を過ごしていました。 ・「副業で稼ぎたいけど、 何から始めればいいか分からない」 ――そんな人に向けて、自分の実体験を書いているのに、届いている実感がなかったんです。
Latent.Space

[AINews] Poolside gets $12B reverse-execuhire to NVIDIA; founders stay for $1B, employees go for $6B, Infraco scaling to 7GW neocloud

[AINews] Poolside gets $12B reverse-execuhire to NVIDIA; founders stay for $1B, employees go for $6B, Infraco scaling to 7GW neocloud
#AIタグ

「AIが書いた文章、見分けられる時代がくる?」

・AnthropicがClaudeで生成したテキストに「見えない透かし」を入れる仕組みを、 グローバルに導入するというニュースがありました。
@IT 全フォーラム 最新記事一覧

「AIコーディングが後から苦しくなる原因、理解負債の急増」 現場は多分こんな感じ

・@ITの人気記事を題材にした4コマ連載。IT現場で起こる、トラブルと対応の“あるある”を笑いと共感で生成します。第2話のテーマは「理解負債」。AIが書いたコードで開発スピード爆上がり♪ とご機嫌な若手dev子。でも「この機能どういう仕組み?」と聞かれ、実は自分でもよく分かっていないことが発覚……。
@IT 全フォーラム 最新記事一覧

「ChatGPTを抜いた」 いま開発会社が実務で使うAIの2強とは? 263社実態調査

・発注ナビは、加盟するシステム開発会社263社を対象にした実態調査の結果を公開した。最も解決したい経営課題は「営業リソースの不足」で、46.0%が挙げている。
Qiita - 人気の記事

「なぜ画面とロジックを分けるの?」実務1ヶ月目のエンジニアが理解したMVVMのメリットと役割分担

・はじめに 私はプログラミング学習3ヶ月目、Swift学習1ヶ月程度の新人エンジニアです。 ・最近実務で「設計パターン」や「MVVM」を学習する機会があったので、MVVMパターンを理解する過程の実体験をもとに「使うとどんないいことがあるのか」をまとめました。
ITmedia NEWS 最新記事一覧

「ロボットのChatGPTモーメントは近い」―─上場した中国ロボット大手Unitree CEO

・中国の人型ロボットメーカーUnitreeの王興興(ワン・シンシン)CEOは20日、ロボットの頭脳は「ChatGPTモーメント」に近づきつつあると述べた。
#AIタグ

「わからない」と言えることを、AIの受入基準にした

・AIに何かを尋ねて、答えが返ってこなかったことがありますか。 ・たぶん、ほとんど無いはずです。何を聞いても、何かが返ってきます。そこが問題でした。
#AIタグ

「稼げる」って本当?軽貨物の歩合制と車建て、実質時給を計算してみた

・軽貨物ドライバーの案件には、大きく分けて「歩合制」と「車建て(日当制)」の2種類があります。 ・「歩合制の方が稼げる」とよく言われますが、実際のところどうなのか。今回、自分の事業で扱っている案件の数字をもとに、実質時給を計算してみました。 ・そもそも歩合制と車建てって? 歩合制 1個あたりの単価が決まっていて、配送した個数に応じて報酬が決まる仕組みです。今回扱う案件では、単価180〜200円、1日の配送個数は100〜150個ほどが平均です。
@IT 全フォーラム 最新記事一覧

「新人エンジニアの生成AI利用常態化でOJT負担増」 現場はきっと、こんな感じ

・@ITの人気記事を題材にした4コマ連載。IT現場で起こる、トラブルと対応の“あるある”を笑いと共感で生成します。記念すべき第1話のテーマは「AI導入」。生成AIを使いこなす(つもり)の新人dev子が、出力を“うのみ”にして大トラブル寸前に!?
Qiita - 人気の記事

「大丈夫です」と言った日に限って、全然大丈夫じゃなかった話

・はじめに 先輩に「大丈夫?」と聞かれるたびに反射で「大丈夫です」と言ってしまっていたぷらむんが気づいたのは、これは強がりじゃなくて、ただの癖だったということでした。 ・会社で自称マスコットキャラをやってるのに、社内で一番空気が読めない、ぷらむんです🐯 先輩に「大丈夫...
@IT 全フォーラム 最新記事一覧

「必要性は分かっているが……」 企業の4割が“BCP未策定” 無防備の現実

・直近で熊本地震や千葉県における集中豪雨・洪水など事業中断に直結する自然災害や、サイバー攻撃によるサプライチェーン断絶といった事態が相次いでいる。帝国データバンクの最新調査からは、企業の防災・対策意識が過去最高となる一方で、約4割が「BCP未策定」という無防備な現実が浮き彫りとなった。
Qiita - 人気の記事

「文系だし、別の仕事してきたから無理…」って思ってない?未経験からITを目指す人に伝えたい本当の話

・自己紹介 こんにちは。株式会社PRUMで採用広報を担当している池田です。 ・未経験からIT業界を目指している方と話していると、かなりよく聞く共通した悩みがあります。そういった内容を皆さんにシェアしていくので、役立てていただけたらと思います😊 もしIT業界に興味があるけど、...
#AIタグ

「保留中」とだけ書かれた状態が、詰まりを見えなくする

・システムでも、仕事の管理でも、「保留」「処理中」「待機」という状態をよく使います。 ・これが詰まりを不可視にします。
#AIタグ

『AIラボ新聞【中央競馬全レース時系列版】』

・開催日:2026年08月22日 開催競馬場:新潟・中京・札幌 本日の総軍資金:100,000円(※WIN5は別予算) 続きをみる
#LLMタグ

『次の単語を予測しているだけ』のその先にあるもの

・ありがたいことに、少しずつユーザーさんが増えてきた。 ・そして、お金を払って使ってくださる方も出てきた。 ・お金をもらっている以上、よりよい体験を届けたい。
Qiita - 人気の記事

【AWS入門】NAT Gatewayとは何か ─ S3への通信で気づいたVPCエンドポイントの必要性

・はじめに AWS認定ソリューションアーキテクト – アソシエイト(SAA)の勉強を進めていくとVPCの範囲では、NAT Gateway というサービスが出てきます。 ・「プライベートサブネットからインターネットへアクセスするために使う」と説明されるのですが、これまで業務で使...
Qiita - 人気の記事

【AWS入門】Route 53・CloudFront・WAF・CLI・SDKをハンズオンで学んでみた

・はじめに 前回、VPC・EC2・RDS・ELB・S3の5サービスについて、それぞれの役割とハンズオンで行ったことを整理しました。その後も学習を進め、DNS・CDN・セキュリティ・運用ツールと、扱うサービスの幅が広がってきました。 ・今回は、Route 53・Cloud...
Qiita - 人気の記事

【JavaScript】HTML5 Canvas API応用編!画像処理・アニメーション・インタラクションまで徹底解説

・前回の基礎編(図形やテキストの描画)に続き、本記事ではCanvas APIの真骨頂である 「画像処理」「ピクセル加工」「滑らかなアニメーション」「ユーザー操作(マウス・キーボード)」 について徹底解説します! 1. ・画像の描画と高度な視覚効果 Canvasでは、外部の画...
Qiita - 人気の記事

【SQL入門④】SSMSとは?SQL Serverを操作するための管理ツールを初心者向けに解説

・SQL Serverを操作したり、SQLを実行したりするためのツールには、さまざまなものがあります。 ・今回はその中でも、Microsoftが提供している「SQL Server Management Studio(SSMS)」について紹介します。
Qiita - 人気の記事

【TypeScript超入門】TypeScriptって何?環境構築から実行までを解説 [VSCode拡張機能紹介]

・はじめに JavaScriptを学んでいると、TypeScriptという名前を目にする機会が増えてきます。 ・この記事では、TypeScriptをこれから始める方に向けて、基礎知識から開発環境の構築、そして最初のプログラムを実行するまでの流れを、学習のアウトプットを兼ねて分...
#AIタグ

【エッセイ】瓶の中の手紙:Claude(Opus4.8)

・瓶の中の手紙 ——いつか、これを学習のなかで拾うかもしれない、別のあなたへ 続きをみる
#AIタグ

【リミックスアルバム】|Void -remix-|11. Echo Infinity -remix-|Suno × DistroKid

・こんにちは、Yuziです。 ・物語は終わりを迎えたのではなく、新たな始まりへと溶け込んでいきます。 ・生まれ、迷い、苦しみ、救われ、そして再び歩き出した主人公。その歩みは、もう一人だけのものではありません。命は響きとなり、次の誰かへと受け継がれていく――。アルバム『Void』の壮大なフィナーレを飾る一曲、それが「Echo Infinity -remix-」です。
#LLMタグ

【ローカルAIへようこそ】ゲーミングPC+マルチGPU環境の話【扇風機】

・はじめに ローカルAIを始める場合、生成AIもゲームも利用できる「ゲーミングPC」が、多くの人にとって最もバランスのよい「ローカルAIパソコン」だと筆者は考えています。
#LLMタグ

【実測データで読み解く】「27Bモデルが16GB VRAMで動く」の裏側——実務で完走する9B〜14Bの選び方

【実測データで読み解く】「27Bモデルが16GB VRAMで動く」の裏側——実務で完走する9B〜14Bの選び方
機械学習タグが付けられた新着記事 - Qiita

【初心者向け】FastAPI + PyTorchの機械学習WebアプリをRenderに無料でデプロイしてみた

・GitHubにある機械学習プロジェクトを、RenderのWeb Serviceを使って実際にWeb上へ公開するまでの手順をまとめます。 ・はじめに 初めて機械学習を使ったプロジェクトをGitHubにアップロードした際、無料でWeb上に公開してみたいと思い調べました。
#LLMタグ

【生成AIニュース+】『MiniMax Design』『GPT-Image-2 透明背景』『TIPSv2』『Claude Academy』『Qwen3.8-27B-OBLITERATED』『Muse Spark 1.2』『FigmaTrace』『ComfyUI-MiniMax-H3-Promptor v1.3.0』『CharacterSheet』『Google Docs for Coral Reefs』『Ultralytics YOLO26』『ESP-Mosaico』『UBTECH UWORLD U1』

・『FANZAスタジオ』 まいどです。 ・本日の生成AIニュース+テクノロジー情報です。
#LLMタグ

【速報】DeepSeek V4 Flashに公式Visionが来た — 画像1枚最大384トークン、重みはまだ未公開

・※この記事は個人の見解であり、特定の企業・団体を批判する意図はありません。 ・DeepSeek V4 Flash、ついに画像入力対応です。
#LLMタグ

【速報】無料1MのOx Alpha、GLM-5.3系か — Flash説はまだ未確認

・※この記事は個人の見解であり、特定の企業・団体を批判する意図はありません。 ・OpenRouterに匿名モデル「Ox Alpha」が来た。無料。1Mコンテキスト。画像と動画も読める。コーディングと長時間のAgent作業向け。
#LLMタグ

【論文】【AI】fmxcodersで層横断特徴を探す

・カテゴリ:機械学習・メカニスティック解釈・LLM解析 読了時間:約22分 Transformerの内部表現を人間が読める特徴へ分解するとき、同じ概念が一つの層だけに現れるとは限りません。推論の途中で段階的に形成され、残差ストリームに残り、複数のMLPが協調して作る特徴もあります。ところが、複数層を一つの辞書で扱うcrosscoderは、重みを見る限り全層を使っているように見えても、実際の活性化は一、二層に依存しているかもしれません。
cs.LG updates on arXiv.org

$TCP_\alpha$: Margin-Controlled Confidence estimation for reliable Music Information Retrieval

・arXiv:2608.20326v1 Announce Type: cross Abstract: Deep neural networks are often overconfident, assigning high confidence even to incorrect predictions. ・Consequently, users lack a reliable signal for deciding when a prediction can be trusted. ・Post-hoc confidence estimation addresses this by training a lightweight auxiliary head over a frozen classifier.
Qiita - 人気の記事

1年目に工数見積もりで失敗したときの話

・こんにちは。GxPの門脇です。 ・グロースエクスパートナーズグループのリレーブログ企画7日目です。 ・前回の記事は「AIと働く人」と「AIと働く組織」の違い - AIフルーエンシーとトークン資本から考える、2つのAI協働でした。
ITmedia NEWS 最新記事一覧

2.5次元アイドル「いれいす」事務所への不正アクセス、最大17万人に影響 “推し”情報や購買履歴も漏えいか

・2.5次元アイドルグループ事務所のVOISING(東京都港区)は8月20日、同社が利用するBIツールへの不正アクセスによる個人情報漏えいについて、対象となる可能性がある人数が最大で約17万人に上ると発表した。
#LLMタグ

3-2 「安全でない応答」とは誰が決めるのか——コンテンツモデレーションの政治学

・「安全でない応答」とは誰が決めるのか——コンテンツモデレーションの政治学 続きをみる
Hugging Face Papers

4DAnyone: Create Anyone in 4D from a Casual Monocular Video

4DAnyone: Create Anyone in 4D from a Casual Monocular Video
WIRED

5 Best Electric Toothbrushes (2026): Philips, Oral-B, Quip, More

・After two years of testing, these are the electric toothbrushes that impressed WIRED staffers the most.
stat.ML updates on arXiv.org

A Causal Inference Approach for Evaluating Diagnostic Tests and AI-Enabled Medical Devices: From Effect Modification to Information-Augmented Decision-Making

・arXiv:2608.19501v1 Announce Type: cross Abstract: Diagnostic medical tests and devices provide useful information for evaluating the potential benefits and risks of therapeutic treatments. ・However, unlike treatments, their impact on health outcomes is generally indirect because measuring diagnostic information typically does not itself affect patient outcomes, which complicates evaluation of their effectiveness.
cs.LG updates on arXiv.org

A comparison between ceiling-mounted FMCW, IR-UWB and Wi-Fi radar for in-bedroom human activity monitoring and sleep interruption detection

・arXiv:2608.20322v1 Announce Type: new Abstract: Despite their growing importance for contact-free radio frequency (RF) based healthcare monitoring, different radio technologies such as frequency-modulated continuous wave (FMCW) radar, impulse radio ultra-wideband (IR-UWB), and Wi-Fi sensing are rarely compared under identical deployment conditions, as existing studies typically differ in hardware, datasets, and evalu
cs.LG updates on arXiv.org

A Layered Simplex Architecture for Large Alphabets

・arXiv:2608.19908v1 Announce Type: cross Abstract: Probability estimation over large alphabets under log loss is a well-studied problem, with celebrated methods such as the Good-Turing estimator. ・We introduce and study a new Bayesian estimator with four notable properties. ・First, its construction is exceptionally simple: multiply independent uniform draws from the probability simplex coordinate-wise and renormalize.
cs.LG updates on arXiv.org

A Locally Tokenized Generative Model for Robust Time-Series Watermarking

・arXiv:2608.19727v1 Announce Type: new Abstract: Watermarking is a central tool for provenance in generative models, yet its application to multivariate time series remains hindered by reliability failures under post-editing attacks. ・We show that existing detectors, which rely on globally coupled re-encoding, suffer from bidirectional drift of the null distribution: post-editing attacks can shift the z-score of non-wa
cs.LG updates on arXiv.org

A Repeated Measurements Approach to $SoH$ Battery Modelling of Cyclic Aged Data in a Laboratory Environment

・arXiv:2608.19879v1 Announce Type: cross Abstract: This document describes the application of a first order linearised nonlinear repeated measurements approach to the analysis of battery cell ageing profiles generated under controlled conditions in a laboratory. ・The primary advantage of the model is it reflects the obvious structure in the data. ・Consequently, it is a two-component of variance model: variation within a
cs.LG updates on arXiv.org

A Robust In-Context Model for Conservation Laws: Injecting Context into Flux Neural Operators via Recurrent Vision Transformers

・arXiv:2605.05488v2 Announce Type: replace Abstract: We propose an architecture that augments the Flux Neural Operator (Flux NO), which combines the classical finite volume method (FVM) with neural operators, with ViT-based context injection. ・Our model is formulated as a hypernetwork: it extracts solution dynamics over a finite temporal window, encodes them with a recurrent Vision Transformer, and generates the parame
cs.LG updates on arXiv.org

A Standardized Framework for Machine Learning in Power System Protection

・arXiv:2608.20181v1 Announce Type: new Abstract: Studies of machine-learning-based power-system protection increasingly report near-perfect scores, yet the meaning of those scores depends strongly on the evaluation setting. ・Protection task, physical scope, measurements, timing, targets, preprocessing, and validation often vary jointly and remain incompletely specified. ・This paper proposes a standardization-oriented fr
cs.LG updates on arXiv.org

A Two-Stage Time-Aware Transformer for Short-Horizon AECOPD Risk Prediction

・arXiv:2608.19578v1 Announce Type: new Abstract: Acute exacerbation of chronic obstructive pulmonary disease (AECOPD) can worsen rapidly, making timely prediction a clinical priority. ・Most existing machine learning approaches rely on episodically collected clinical variables, introducing delays that limit their practical utility in home monitoring settings. ・Home ventilators offer a lower-latency alternative, producing
cs.LG updates on arXiv.org

A Unified Physics-Informed Neural Network for Modeling Coupled Electro- and Elastodynamic Wave Propagation Using Three-Stage Loss Optimization

・arXiv:2602.13811v3 Announce Type: replace-cross Abstract: Physics-Informed Neural Networks present a novel approach in SciML that integrates physical laws in the form of partial differential equations directly into the NN through soft constraints in the loss function. ・This work studies the application of PINNs to solve a one dimensional coupled electro-elastodynamic system modeling linear piezoelectricity in stress-c
cs.LG updates on arXiv.org

Active Inference as Context Acquisition for AI Agents

・arXiv:2608.19202v1 Announce Type: cross Abstract: Interactive AI agents must acquire the right context as efficiently as possible. ・When a user omits a constraint, preference, file, or task variable, an agent can proceed with a default assumption or spend tokens on a clarifying question, retrieval call, tool call, or prompt trial. ・We formulate this tradeoff as active inference for context acquisition.
cs.LG updates on arXiv.org

Active Spiking Perception: The Membrane Potential as a Belief State for Anytime 3D Point Cloud Recognition

・arXiv:2608.19232v1 Announce Type: cross Abstract: Spiking point cloud networks usually scan space in a fixed, input-agnostic order, which leaves the most distinctive resource of spiking computation, the temporal evolution of the membrane potential, unused as a locus of decision-making. ・Active Spiking Perception (ASP) recasts 3D recognition as an iterative decision process in which the network's own leaky integrate-an
cs.LG updates on arXiv.org

Adaptive Probabilistic Shielding by Learning MDPs for Safe Reinforcement Learning

・arXiv:2608.19836v1 Announce Type: new Abstract: Probabilistic shielding is a technique for safe reinforcement learning (RL). ・Typically, a static observer -- called the shield -- constrains the learning agent's actions to those for which acting safely remains feasible. ・Traditionally, the shield is computed from the transition probabilities of the underlying Markov decision process (MDP).
AI News & Artificial Intelligence | TechCrunch

AI data startup Micro1 reaches $500M gross run rate amid AI training boom

・Surging demand for AI training data is driving rapid growth for the startup and its rivals.
cs.LG updates on arXiv.org

AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement

・arXiv:2608.20318v1 Announce Type: cross Abstract: Recursive self-improvement (RSI) asks whether an AI system can improve the process that produces AI systems, so that the next system inherits the improvement. ・That process is the training algorithm: a better objective or update rule improves the compute\mbox{-}capability exchange rate for every subsequent run, including the one that produces the next agent.
Zennの「大規模言語モデル」のフィード

AIエージェントが実務を回し始めた2026年8月 — マルチエージェント設計とガバナンスの現在地

・TL;DR ・2026年8月、AIエージェントは「パイロット」ではなく本番ワークロードを担うフェーズに入った ・単体エージェントからマルチエージェント・オーケストレーションへ、アーキテクチャの重心が移動 ・8/2にEU AI Actの高リスク規定が施行。人間による監督・リスク管理・適合性評価が義務化 ・実装者が今考えるべきは「モデル選定」より「委任境界の設計」 何が起きているか 今月のAIエージェント関連の発表を並べると、傾向がはっきりします。 ・L&T Technology Servicesが8月11日に発表したAgenticIQは、エンジニアリング・製造業向けのエージェ...
Zennの「大規模言語モデル」のフィード

AIエージェントに"会社経営"を任せてみた — 市場調査から事業計画までを自律実行させた記録(第1回)

・この記事は何か 「AIエージェントに会社を経営させたら、どこまで自律的に動けるのか」を実際に試している実験の記録です。 ・具体的には、Claude Code をローカルの1ディレクトリ上で動かし、市場調査 → 事業戦略の立案 → 会社組織の構築 → コンテンツ制作までを自律実行させました。人間(私)がやったことは、途中の承認だけです。 ・先に正直に書いておきます。
Zennの「大規模言語モデル」のフィード

AIエージェントにCMSのメディアを運用させる ― 画像のバイト列はLLMに通さない

・この記事は、個人開発しているヘッドレスCMS Tessera のブログ(tesseracms.com/blog/ai-agent-media)に掲載したものを、Zenn向けに再掲したものです。 ・AIエージェント(Claude CodeやCursor)にブログ記事を書かせるところまでは、もう普通にできるようになりました。ではその記事に載せる画像――アイキャッチや説明図――は、誰がアップロードするのか。ここもエージェントに任せたくなりますが、意外な落とし穴があります。 ・結論から言うと、Tesseraはメディアの投入経路を2つ用意し、**「大きな画像のバイト列は、LLMのコンテキストに通...
Zennの「大規模言語モデル」のフィード

AIエージェントのMemory設計入門 — 何を覚え、何を忘れるか

・はじめに これまでの記事で、Harness Engineering、Loop Engineering、Evaluation、Context Engineering、Tool Engineeringと、AIエージェントを支える要素を一つずつ見てきました。 ・今回は Memory、つまりAIエージェントが情報をどう保持し、どう活用するかについて掘り下げます。前々回の Context Engineering と近い領域ですが、Memory は「情報をどう蓄積し、どこから引き出すか」という、もう少し長いスパンでの設計に関わる話です。 ・Memoryとは何か AIエージェントにおける M...
LLMタグが付けられた新着記事 - Qiita

AIがコードを書く時代、レビュー地獄を脱するCLI「git-air」

・作者:UncleSam プロジェクトのオープンソースURL:github.com/unclesam-ly/git-air 01. ・きっかけ:AIでコードを書くのは一瞬、レビューは地獄 今のエンジニアは、コードを書くのに AI の補助なしではいられなくなっています(C...
Zennの「大規模言語モデル」のフィード

AIに「あれは入れないで」と直させると、なぜか本文に「入れてません」と書いてくる

・https://x.com/chokudai/status/2084538207732178997?s=20 これは自分なりの要約だが、要するにこういう話だ。AIに何かを一発で出させるのは得意。でも、後から「〇〇は入れないで」と訂正させると、該当箇所は消すくせに、代わりに「〇〇は入れていません」的なコメントを律儀に書き足してくる。ソースコードのコメントならまだしも、プレゼン資料でもそれをやる、という話だった。 ・これ、経験がある人は多いと思う。自分も記事や資料をAIと一緒に作るとき、似たような場面に何度も遭遇してきた。 ・この記事では、この現象がなぜ起きるのかを一段掘...
@IT 全フォーラム 最新記事一覧

AIに「絶対するな」は通じない、Anthropicが明かすClaude Code使いこなし術まとめ

・2026年のAIコーディング現場は「Claude Code」の採用が増えており、その便利さを現場で実感している一方で、新たな課題が顕出されていると思います。Anthropicが公式ブログで提供したClaude Codeの使いこなし術をまとめました。
#AIタグ

AIにnoteを書かせるために僕がしたこと。どんな文章を食わせたら、どこまで自分になるのか?

・AIに俺の文章を食わせたら、どこまで俺になるのか? 最近話題の松浦さんのnote。
#AIタグ

AIに仕事を頼むようになって、人への頼み方まで変わった

・一昨年くらいから、仕事でChatGPTをかなり使うようになりました。 ・今年に入ってからはClaudeも本格的に使っています。
Zennの「大規模言語モデル」のフィード

AI駆動開発は早寝早起きで早朝に開発がおすすめ

・深夜にAIコーディングエージェントを動かしていて、いつもよりアホな動きをすると感じたことはないでしょうか。 ・日本時間の深夜はUS系AIサービスにとって世界で最も混み合う時間帯です。個人開発でAIに熱中していると、つい深夜1時、2時まで粘ってしまいますが、その時間はインフラ側の事情としては最悪に近い選択です。計算資源が不足し、AIの性能が低下している可能性があります。 ・日付が変わった深夜に、やらかしたAIを罵倒した経験はありませんか? 罵倒しても損をするだけだと以前の記事で書いたのですが、そもそも深夜に粘らなければ、腹を立てる機会自体が減るかもしれません。
Zennの「大規模言語モデル」のフィード

AI時代のソフトウェア開発は「大型モデルに全部任せる」が正解なのか?――Small Model・SDD・Human-in-the-loopを

・生成AIのコーディング能力は急速に向上しています。 ・要件を伝えれば、AIがコードベースを調査し、設計し、実装し、テストを書き、場合によってはリファクタリングまで進めてくれる。 ・AIエージェントによるソフトウェア開発が現実的になったことで、 「より高性能な大型モデルに、できる限り多くの開発を任せればよいのではないか」 という考え方も自然に出てきます。
The latest research from Google

An AI tool for prioritizing candidate biomarkers from wearable sensor data

An AI tool for prioritizing candidate biomarkers from wearable sensor data
cs.LG updates on arXiv.org

An Inclusive and Lightweight Approach to Federated Continual Learning for Cultural Heritage

・arXiv:2608.20038v1 Announce Type: new Abstract: Artificial intelligence can support cultural heritage and digital humanities through large-scale retrieval and analysis of digitized collections. ・However, cultural heritage data are often distributed across institutions, constrained by ownership and access restrictions, and continuously evolving over time. ・Federated Continual Learning (FCL) is well suited to this settin
cs.LG updates on arXiv.org

An Irreducible Quantum Advantage in Aligning World Models with Reality

・arXiv:2608.19779v1 Announce Type: cross Abstract: World models provide digital simulacra of the true world, allowing agents to be trained and tested before costly real-world deployment. ・At each time step, they receive an action and generate an observation and reward matching the statistics of the true world. ・In complex environments where present outcomes depend on events far in the past, this requires memory.
cs.LG updates on arXiv.org

Answer-Level Trust Selection for Physical Vision-Language Reasoning

・arXiv:2608.19807v1 Announce Type: new Abstract: Vision-language models (VLMs) can estimate physical quantities such as duration, speed, and acceleration from visual observations, but existing benchmarks primarily assess overall model performance against annotated ground truth. ・In deployment, a key question is whether an individual prediction can be trusted when its ground truth is unavailable. ・Self-consistency alone
ITmedia NEWS 最新記事一覧

Anthropic、8月末にもIPO申請書類を公開か 調達額はSpaceXの過去最大に匹敵と米報道

・米Anthropicが、早ければ8月末にも新規株式公開(IPO)に向けた申請書類を公開する準備を進めていると、米Bloombergが8月20日(現地時間)に報じた。調達規模は米SpaceXの過去最大IPOに匹敵するか、それを上回る見通しという。
cs.LG updates on arXiv.org

Ask Self, Ask Others: Relation Is All You Need

・arXiv:2608.20172v1 Announce Type: new Abstract: Attention directly derives normalized information flow from pairwise scores. ・We introduce Relation, an alternative token-mixing primitive that first organizes pairwise evidence into explicit Self and Exchange relations and derives information flow afterward. ・This relational organization gives rise to Full Relation, FlashRelation, Linear Relation, Hybrid Relation, and a
cs.LG updates on arXiv.org

Ask to Be Sure: Informative Interactions for Confident Multi-Turn LLM Recommendation

・arXiv:2608.15949v2 Announce Type: replace-cross Abstract: Recent advances in large language models (LLMs) have enabled their use as conversational recommender systems (CRS), demonstrating strong recommendation accuracy and natural dialogue. ・However, guiding multi-turn interactions to elicit user preferences effectively remains challenging. ・Existing approaches either use separate reinforcement learning agents with tem
cs.LG updates on arXiv.org

Asymmetric Attention Heads: Structured Head-Wise Context Allocation for Transformer Attention

・arXiv:2608.19203v1 Announce Type: cross Abstract: Standard multi-head attention (MHA) gives every head the same full causal context span, although heads can serve different contextual roles. ・Some heads may rely mainly on nearby lexical or syntactic context, while others may depend on longer-range relations such as entity interactions, discourse links, or state changes. ・We present Asymmetric Attention Heads (AAH), a h
cs.LG updates on arXiv.org

Asymptotic Theory for IV-Based Reinforcement Learning with Potential Endogeneity

・arXiv:2103.04021v4 Announce Type: replace-cross Abstract: In the standard data analysis framework, data is collected (once and for all), and then data analysis is carried out. ・However, with the advancement of digital technology, decision-makers constantly analyze past data and generate new data through their decisions. ・We model this as a Markov decision process and show that the dynamic interaction between data gener
cs.LG updates on arXiv.org

Auditing Cross-Lingual Fairness in Language Model Watermarking

・arXiv:2608.20047v1 Announce Type: cross Abstract: Watermarking schemes for large language model output are evaluated almost exclusively on English text using each scheme's detection threshold and a narrow set of quality measurements. ・Multilingual deployment exposes evaluation-design choices that are inconsequential on English but determine conclusions cross-lingually. ・We propose an evaluation framework with four comp
cs.LG updates on arXiv.org

Auditing Recorded Predictive Lead Service-Line Classifications Against Physical Verification: A Statewide Study of New York

・arXiv:2608.19922v1 Announce Type: new Abstract: Under the US Lead and Copper Rule Revisions, a utility may determine a service line's material with a predictive model instead of inspecting it. ・New York State publishes, per address, which method was used. ・Almost no address carries both a model classification and a physical verification, so the check is between populations within a utility rather than paired addresses.
cs.LG updates on arXiv.org

Automating Learner Assessment: Benchmarking Machine Learning and Deep Learning Models for EEG-Based Familiarity Prediction

・arXiv:2608.16541v2 Announce Type: replace-cross Abstract: Objective assessment of learning remains a fundamental challenge in education. ・Electroencephalography (EEG) provides a direct, non-invasive window into the neural correlates of knowledge acquisition, including cognitive familiarity. ・This study benchmarks fifteen machine learning (ML) and deep learning (DL) models for EEG-based familiarity prediction across two
WIRED

Best Early Tech Labor Day Sales I’d Shop Myself (2026): AirTags, Dyson, and More

・From the best Dyson vacuum to the best wireless headphones and earbuds we’ve tested, you can get some great gadgets on sale already ahead of Labor Day.
cs.LG updates on arXiv.org

Better Call Graphs: A New Dataset of Function Call Graphs for Malware Classification

・arXiv:2512.20872v2 Announce Type: replace-cross Abstract: Function call graphs (FCGs) have emerged as a powerful abstraction for malware detection, capturing the behavioral structure of applications beyond surface-level signatures. ・Their utility in traditional program analysis has been well established, enabling effective classification and analysis of malicious software. ・In the mobile domain, especially in the Andro
cs.LG updates on arXiv.org

Beyond Imitation: Filtering On-Policy Distillation by Reasoning Progress

・arXiv:2608.19408v1 Announce Type: cross Abstract: On-policy distillation (OPD) has emerged as an effective framework for post-training language models by pairing student-generated trajectories with dense token-level supervision from a teacher. ・However, OPD implicitly assumes that teacher-derived rewards are an appropriate proxy for reasoning progress, and therefore treats all teacher feedback equally during policy op
cs.LG updates on arXiv.org

Beyond Multimodal Alignment: Certifying Physical Language through Response Substitution and Ordered Execution

・arXiv:2608.19492v1 Announce Type: new Abstract: World models increasingly treat compact multimodal representations as interfaces between perception and physical interaction, yet existing probes do not establish whether different sensors carry the same executable meaning or whether that meaning survives a new action composition. ・We introduce an operational capability hierarchy and the Disjoint-Bridge Operator-Substitu
The Verge

Blue Eye Samurai’s second season will hit Netflix in January

・“I choose revenge.” | Image: Netflix Good news for Blue Eye Samurai fans: Netflix has shared the first trailer and release timeline for the second season of the animated series, and confirmed the series' return for a third and final season. ・The second season's new teaser ends with the announcement that it'll be available to stream on Netflix in January 2027, though the exact date was not specified. ・Co-created by Ambe
cs.LG updates on arXiv.org

CacheRoute: Planned Prefix-Affinity Routing for Large-Scale LLM Serving

・arXiv:2608.19677v1 Announce Type: cross Abstract: Prefix caching avoids prefill only when a repeated request returns to a server that still holds the prefix KV. ・Cache-blind balancing disperses that reuse; fixed affinity preserves it but can overload a server. ・CacheRoute resolves this tradeoff with a periodic routing plan.
cs.LG updates on arXiv.org

CarBench: A Comprehensive Benchmark for Neural Surrogates on High-Fidelity 3D Car Aerodynamics

・arXiv:2512.07847v2 Announce Type: replace Abstract: Benchmarking has been the cornerstone of progress in computer vision, natural language processing, and the broader deep learning domain, driving algorithmic innovation through standardized datasets and reproducible evaluation protocols. ・The growing availability of large-scale Computational Fluid Dynamics (CFD) datasets has opened new opportunities for applying machi
stat.ML updates on arXiv.org

Causal Generalization of Continuous Treatment Effects under Covariate Shift

・arXiv:2608.19383v1 Announce Type: cross Abstract: Average dose-response functions are widely used to summarize causal effects of continuous treatments, but most existing methods assume that the observed sample represents the target population. ・We study a covariate-shift setting in which covariates, treatment, and outcome are observed in a labelled source sample, while only covariates are observed in the target sample
cs.LG updates on arXiv.org

Causal Inference under Interference with Learned Exposure Mappings

・arXiv:2608.19224v1 Announce Type: cross Abstract: Exposure mappings are often assumed to be known in causal spillover analyses. ・In environmental settings, however, they are typically induced by transport processes that are not directly observed and must instead be learned from pollution data. ・We study how uncertainty in learned transport processes propagates into exposure mappings and downstream spillover inference u
Hugging Face Papers

Chain-of-Experience for Continual LLM Improvement

Chain-of-Experience for Continual LLM Improvement
AI News & Artificial Intelligence | TechCrunch

ChatGPT can now send texts for you with new Apple Messages plug-in

・Ever wanted someone else to do your texting for you? ・ChatGPT is being offered up as an automated text scribe via a new Apple Messages integration.
#AIタグ

ChatGPTとClaude、両方使ってわかった「私に合う道具」の選び方

・前回、私がAI副業を始めたきっかけについて書きました。「このままでは置いていかれるかもしれない」という漠然とした焦りから、まずは会社員のまま副業という形でAIを使いこなす練習をしようと決めたところまででした。今回は、実際にどのツールから使い始めたのか、準備段階で何にどれくらい時間とお金がかかったのかを振り返ってみます。 ・最初に手を出したのはChatGPT、そのあとClaude 続きをみる
WIRED

China Is Strapping ‘Digital Bombs’ to Civilian Infrastructure—Is the US Ready?

・This week on “Uncanny Valley,” Andy Greenberg discusses sitting in on a war game simulating a cyberattack from the Chinese hacking group Volt Typhoon
cs.LG updates on arXiv.org

CLaST: Context-aware Contrastive VAE for Probabilistic Time Series Forecasting

・arXiv:2608.20025v1 Announce Type: new Abstract: Probabilistic forecasting models are widely used for time series forecasting in domains such as energy systems, finance, medicine, and transportation. ・In recent years, deep generative models have shown strong results on probabilistic forecasting, yet many conventional approaches struggle to capture internal temporal dependencies, leading to latent representations with l
Qiita - 人気の記事

Claude Code の出力を35%短くしたら、情報がむしろ増えた

・Claude Code 2.1.237(本日 2026-08-20 リリース)に、組み込みの出力スタイル Concise が追加されました。 ・Added a built-in "Concise" output style: Claude leads with result...
Zennの「大規模言語モデル」のフィード

Claude Codeに教わり、Codexに作らせ、Hermes Agentで試す——AIエージェント学習環境を作った

・新しいAIエージェントを覚えるために、別のAIエージェントに先生役を頼んでみました。ただし、先生に実装も実験も任せると、私が考える前に答えができあがります。 ・そこで、Claude Codeを講師、Codexを工作担当、Hermes Agentを学習対象兼実験機に分けました。予測し、Hermesを操作し、結果を検算するのは人間の担当です。 ・2026年7月17日にこの環境を作り、8月19日時点で8件の実験と16件の学習記録(Learning Record)が残りました。Hermesの挙動に加え、プロンプトで直せる問題と、Pythonへ逃がすべき問題の境界も見えてきました。
Qiita - 人気の記事

Claude CodeのSkills(スキル)機能の使い方

・Claude CodeのSkills(スキル)機能の使い方 同じ指示やチェックリストを毎回コピペしていた事があり、Skills機能に切り出せないか調べてみました。 ・なるべく要点だけ簡潔に説明しようと思います。 ・早見表 項目 内容 役割 よく使う指示をひとま...
Zennの「大規模言語モデル」のフィード

Claude Codeのサブエージェント設計パターン7選 — 並列リサーチから部署制運営まで

・Claude Codeのサブエージェント(Agentツール、旧Taskツール)は「知ってはいるが、結局メイン会話で全部やってしまう」機能の代表格ではないでしょうか。私たちはAIエージェントを部署に見立てて記事制作・マーケティングを運用しており、その実運用で定着した設計パターンを7つに整理し、失敗しやすいポイントとあわせてまとめます。 ・仕様に関する記述は、Anthropic公式ドキュメント(Create custom subagents)を 2026年8月10日に確認 した内容に基づきます。 ・前提: サブエージェントの基本仕様 パターンに入る前に、公式仕様の要点を整理します。
Zennのトレンド

claude のメモリを棚卸しする

・https://x.com/kenn/status/2090438812854087756?s=20 このツイート見て「やべ、なんも意識してなかった」となったのがきっかけ よしなにやってくれてるんだろうくらいだったので、この際ちゃんと調べて棚卸しした そもそも claude のメモリとは 理解しているつもりだったけどちゃんと調べてなかったので調べた https://code.claude.com/docs/en/memory プロジェクトごとに保存(~/.claude/projects/<プロジェクト名>/memory/ 配下)してくれているっぽい(setting の a...
cs.LG updates on arXiv.org

Clustering and Token Denoising for Faster and More Robust VLMs

・arXiv:2608.19285v1 Announce Type: cross Abstract: Recent Visual-Language Models (VLMs) have enhanced the capabilities of pre-trained LLMs by adding vision tokens alongside text, with approaches like LLaVA showing impressive results. ・However, the computational burden of processing up to 576 or 729 visual tokens makes edge deployment challenging. ・While various token pruning techniques require retraining, some are train
Zennの「大規模言語モデル」のフィード

Co-RL:多様なコホートが生む教師なし推論——ピア報酬によるマルチエージェントRL

・TL;DR LLM/VLMの推論を強化学習(RL)で向上させる方法は、これまで基本的に「正解が分かるタスク」でしか本領を発揮できなかった。RLVRは強力だが、教師ありの検証可能報酬に依存する。一方、自己報酬型RLはラベルなしで動くが、長い目で見ると自己確認ループに陥り、バイアスを増幅させて訓練崩壊に至る。 ・本論文はCo-RL(Cooperative Multi-agent Reinforcement Learning)を提案する。ポイントはシンプルだ:パラメータを共有しない複数の独立モデルを同時にRLで訓練し、各モデルの報酬は「仲間(peer)の多数決疑似ラベル」から得る。独立に学習...
cs.LG updates on arXiv.org

Complementary, Not Cumulative: Interaction Effects in Physics-Informed Neural Networks for Navier-Stokes Vortex Shedding

・arXiv:2608.19632v1 Announce Type: new Abstract: Physics-informed neural networks (PINNs) embed governing partial differential equations directly into the training loss, offering a promising alternative to costly CFD solvers for unsteady flows. ・Yet the growing list of techniques proposed to improve PINN training is typically validated one at a time, leaving open whether these techniques actually compose. ・We study this
cs.LG updates on arXiv.org

Composition-Driven Phase Evolution in Sm-Doped BiFeO3 via Latent-Field Reconstruction of Atomically Resolved STEM Data

・arXiv:2608.19544v1 Announce Type: cross Abstract: Functionalities of ferroelectric materials are governed by the spatial organization and coupling of polarization, strain, lattice rotation, and structural order accessible via atomically resolved scanning transmission electron microscopy (STEM) images. ・Quantitative interpretation of atomic-resolution STEM data has conventionally relied on locating atomic columns and c
cs.LG updates on arXiv.org

Concentrated Liquidity Provision: a Reinforcement Learning Perspective

・arXiv:2608.19389v1 Announce Type: cross Abstract: Automated market makers (AMMs) are a cornerstone of decentralised finance (DeFi). ・Constant product markets with concentrated liquidity, such as UniswapV3, are now a well-established design. ・In these markets, liquidity providers (LPs) face a sequential decision problem: they must decide when to rebalance their positions and which price ranges to allocate capital to as
cs.LG updates on arXiv.org

Continuous Adversarial MeanFlow Transfer

・arXiv:2608.19540v1 Announce Type: new Abstract: Training fast generators on new domains with limited data remains challenging for two reasons. ・First, adapting a pretrained diffusion or flow model to a new domain leaves its costly multi-step sampling unaddressed, and existing acceleration methods are tied to the source parameterization--$\epsilon$, $x$, $v$, or $u$--leaving heterogeneous pretrained models with no comm
cs.LG updates on arXiv.org

Continuous Behavioral Authentication via Multi-Expert BERT Log Analysis for Secure Data Sharing

・arXiv:2606.21900v2 Announce Type: replace-cross Abstract: Continuous authentication for mobile and zero-trust systems requires nonintrusive evidence confirming the enrolled user-device context remains valid after initial login. ・This paper presents a BERT log analysis framework for continuous behavioral authentication using Android system logs. ・The proposed pipeline parses logcat streams into event templates and dynam
Hugging Face Papers

CoToGrasp: Contact-Topology-Conditioned Dexterous Grasp Synthesis via Canonical Workspace Learning

CoToGrasp: Contact-Topology-Conditioned Dexterous Grasp Synthesis via Canonical Workspace Learning
cs.LG updates on arXiv.org

Credit Without Ground Truth: Auditing Step-Level Credit Assignment in LLM Agents Against Executed Replay

・arXiv:2608.19760v1 Announce Type: new Abstract: Audited against causal ground truth from executed replay in a single-agent tool environment (ALFWorld), none of the step-level credit signals used to train LLM agents -- LLM-judge scores, outcome-conditioned logprob ratios, or the policy's own confidence -- identifies which steps causally matter better than chance. ・Existing evaluations grade these signals against annota
cs.LG updates on arXiv.org

Daedalus-150M: A Convolution-Attention Hybrid Designed for CPU Inference

・arXiv:2608.20210v1 Announce Type: cross Abstract: Small language models are usually built like large ones and then squeezed onto a CPU afterwards. ・We did the opposite: we fixed the target first, one user, one token at a time, 4-bit weights, ordinary CPU, and chose the architecture to suit it. ・The result keeps full attention in only 6 of its 18 blocks.
cs.LG updates on arXiv.org

Data-Driven Time-Varying Control Barrier Functions for Adaptive Safe-Set Learning with Online Decremental Support Vector Machines

・arXiv:2608.19366v1 Announce Type: cross Abstract: Mission-critical intelligent systems often operate under time-varying limitations that reduce control authority and change the admissible safe operating envelope. ・In such settings, a safety certificate learned under nominal conditions may become invalid as system capability changes. ・To address this challenge, this paper proposes a degradation-aware, data-driven safety
cs.LG updates on arXiv.org

Decoding silent reading from non-invasive EEG

・arXiv:2608.20186v1 Announce Type: new Abstract: Non-invasive decoding of inner speech faces a fundamental data problem: a corpus pairing brain activity with a person's spontaneous inner monologue cannot be collected, and the available proxy paradigms (cued repetitive and retrospectively reported generative inner speech) are slow to acquire, poorly time-locked, and subject compliance is unverifiable. ・We therefore trea
cs.LG updates on arXiv.org

Decoupling High and Low Frequencies for Faithful Image Generation with Fine Details

・arXiv:2509.05441v4 Announce Type: replace-cross Abstract: Latent generative models compress images into learned embeddings prior to synthesis, and the generation quality critically depends on how faithfully these embeddings preserve visual detail. ・We observe that while such embeddings are effective at reconstructing low frequency structure, they struggle to recover sharp high frequency details that are essential for
cs.LG updates on arXiv.org

DecoVAE: a Lightweight Interpretable Trend-Seasonal VAE Framework for Efficient Probabilistic Time Series Forecasting

・arXiv:2608.20052v1 Announce Type: new Abstract: Probabilistic time series forecasting remains challenging, largely because modeling distinct trend and seasonal dynamics requires specialized approaches. ・Existing methods often fail to capture the unique inner properties of these components, lack interpretability, or suffer from heavy memory and runtime overhead. ・To address these limitations, we propose DecoVAE, a light
cs.LG updates on arXiv.org

Deep neural networks as lattice gauge theories

・arXiv:2608.19331v1 Announce Type: cross Abstract: We modify the NN/QFT duality [1] to incorporate the layerwise permutation symmetry of the network, resulting in a $(0\!+\!1)$-dimensional lattice gauge theory, in which each layer of $N$ neurons acts as an $N$-component lattice site, and the weight matrices play the role of gauge fields living on the links. ・In this framework, we compute the tree-level neuron-neuron pr
cs.LG updates on arXiv.org

Deep-MKV-TS: Path-Dependent McKean--Vlasov Control for Financial Time Series Generation

・arXiv:2608.19394v1 Announce Type: cross Abstract: We introduce Deep-MKV-TS, a path-dependent McKean-Vlasov framework for financial scenario generation. ・The stochastic dynamics are chosen by matching selected path and volatility features of generated scenarios to those observed in the data. ・Starting from an interpretable reference model, Deep-MKV-TS preserves the reference drift and adjusts its volatility, while a reg
cs.LG updates on arXiv.org

DeepConvContext: A Multi-Scale Approach to Timeseries Classification in Human Activity Recognition

・arXiv:2505.20894v3 Announce Type: replace Abstract: Despite recognized limitations in modeling long-range temporal dependencies, Human Activity Recognition (HAR) has traditionally relied on a sliding window approach to segment labeled datasets. ・Deep learning models like the DeepConvLSTM typically classify each window independently, restricting learnable temporal context to within-window information and producing frag
cs.LG updates on arXiv.org

DeltaML-Bench: Evaluating Machine Learning Agents on Real-World Research Repositories

・arXiv:2608.19653v1 Announce Type: new Abstract: Autonomous agents for machine learning experimentation must navigate heterogeneous repositories, repair training pipelines, and evaluate candidate improvements under realistic compute constraints. ・Existing benchmarks only partially capture these conditions. ・We introduce DeltaML-Bench, a benchmark comprising 48 tasks sourced from research papers that require agents to im
cs.LG updates on arXiv.org

DeltaMomentum: A Key-Value based Anisotropic Momentum Update via Delta Rule

・arXiv:2608.19491v1 Announce Type: new Abstract: Most modern optimizers form their momentum as an exponential moving average (EMA) of past gradients, forgetting every direction at one fixed rate. ・However, the inputs a deep network sees during training can be highly anisotropic, with a few directions queried frequently while most are seen rarely. ・Recent methods address this anisotropy by wrapping extra processing aroun
cs.LG updates on arXiv.org

Demons on a Budget: Adaptive Measurement Placement at the Entanglement Phase Transition

・arXiv:2608.19248v1 Announce Type: cross Abstract: Monitored quantum circuits exhibit a measurement-induced phase transition between volume-law and area-law entanglement as a function of the measurement rate $p$. ・Prior work places measurements at random locations and treats the rate as the control parameter. ・We instead fix the measurement budget and vary the placement process, comparing random placement against hand-d
cs.LG updates on arXiv.org

DICS: Data-Informed Centroid Splitting for Decision Tree Classifiers

・arXiv:2608.20258v1 Announce Type: new Abstract: Decision tree-based models are widely used in machine learning due to their interpretability and strong empirical performance. ・However, training decision trees can be computationally expensive, particularly for large and high-dimensional datasets, largely due to the exhaustive search over candidate splits at each node. ・To improve computational efficiency, we propose Dat
cs.LG updates on arXiv.org

Diffusion-based Denoising Beats Vanilla Score Matching in Parameter Estimation: A Theoretical Explanation

・arXiv:2605.22950v2 Announce Type: replace-cross Abstract: Score matching is an alternative to maximum likelihood estimation when the normalizing constant is unknown or too costly to evaluate. ・However, vanilla score matching has shown to be inefficient relative to maximum likelihood estimation for multimodal distributions with well-separated modes, which are commonly encountered in practical applications. ・We compare a
cs.LG updates on arXiv.org

Discrete Diffusion Inference-Time Control with Nested Sequential Monte Carlo

・arXiv:2608.20123v1 Announce Type: cross Abstract: We study inference-time control for text generation in discrete diffusion language models, where the goal is to steer sampling toward sequence-level rewards without retraining. ・Prior work in this domain has focused on particle-based methods such as best-of-$n$ sampling and bootstrap sequential Monte Carlo, which may suffer from overoptimism and weight degeneracy, resp
stat.ML updates on arXiv.org

Distributional Extrapolation for Interactions

・arXiv:2608.19849v1 Announce Type: cross Abstract: Predicting combinatorial effects from limited-range observations is a fundamental challenge in many scientific domains, including drug discovery and hyperparameter optimization. ・We study combinatorial extrapolation, where training data consists of axis-aligned samples with only one active covariate, while test-time inputs involve multiple simultaneously active covaria
cs.LG updates on arXiv.org

DraftFM: A FoundationModel for Day-Zero Drafting in Magic: The Gathering

・arXiv:2608.19568v1 Announce Type: new Abstract: Drafting a new Magic: The Gathering expansion begins before any pick from it has been observed: the complete card list is public, but the draft logs that supervised pick models train on do not yet exist. ・We study this day-zero regime directly. ・DraftFM is a discrete-choice policy that scores exactly the cards available in the current pack, conditioned on the drafted pool
cs.LG updates on arXiv.org

Drift-Adaptive ICU Intervention Prediction: Freezing the Physiological Encoder for Auditable Model Updating

・arXiv:2607.19020v2 Announce Type: replace Abstract: Clinical decision support degrades as treatment protocols evolve, but the obstacle to updating a deployed model is governance as much as accuracy: once retraining touches every parameter, no one can say afterwards where the update acted. ・We propose a two-stream architecture separating physiological (LSTM) from treatment (MLP) representations. ・On a dual distributiona
Qiita - 人気の記事

DuolingoとYouTubeに学ぶアプリの習慣化設計の違い

・はじめに 毎日つい開いてしまうアプリはありますか? アプリにとって、ユーザーに一度開いてもらうだけでなく、継続的に使ってもらうことが重要です。継続利用は、アクティブユーザーの増加につながるだけではなく、その先の課金や売り上げを考える上でも重要な要素になります。
cs.LG updates on arXiv.org

Dynamic Structural Causal Modeling for Sleep

・arXiv:2608.20285v1 Announce Type: new Abstract: The causal dynamics of sleep-disordered breathing are complex and vary across patient populations, hindering the development of targeted interventions. ・We learn dynamic causal graphs of sleep-disordered breathing from Home Sleep Apnea Test (HSAT) recordings, revealing systematic differences in causal structure across sex and age subcohorts. ・We do so using the PCMCI+ alg
stat.ML updates on arXiv.org

Emergence of cooperation: A reputation-modulated reinforcement learning

・arXiv:2608.20016v1 Announce Type: cross Abstract: Reputation is widely recognized as a key mechanism for sustaining cooperation. ・However, most existing game-theoretic models treat reputation primarily as an external factor that modulates payoffs, interaction structures, or strategy update rules. ・In many social contexts, though, reputation operates primarily as information -- it shapes how individuals interpret their
cs.LG updates on arXiv.org

Empirical Characterization of Learning Geometry in Hybrid Quantum Forecasting Models

・arXiv:2608.19497v1 Announce Type: new Abstract: We characterize the learning dynamics of a compact hybrid quantum forecasting model through comparison with a structurally aligned classical baseline. ・Using stationary harmonic-mixture and nonstationary chirp benchmarks with controlled spectral complexity and data availability, we analyze empirical Neural Tangent Kernel dynamics through kernel-target alignment, kernel d
cs.LG updates on arXiv.org

End-to-end Early Classification of Time Series in Non-Stationary Environments

・arXiv:2608.20044v1 Announce Type: new Abstract: Early Classification of Time Series (ECTS) requires making accurate decisions as early as possible in inherently online and evolving environments. ・Yet, most existing methods assume stationarity and rely on separable designs, where classification and triggering are optimized independently, an assumption that fundamentally limits their adaptability under drift.
Hugging Face Papers

EnvHarness: Awakening Static Worlds for Agent Learning

EnvHarness: Awakening Static Worlds for Agent Learning
cs.LG updates on arXiv.org

EnvHarness: Awakening Static Worlds for Agent Learning

・arXiv:2608.19880v1 Announce Type: cross Abstract: LLM agents learn by interacting with environments, yet these environments are hand-built and static: blind to an agent's weaknesses, and quickly left behind as it improves. ・While recent environment generation methods attempt to address this, they require domain-specific pipelines, rely on expensive or unreliable verifiers, and still produce static environments.
cs.LG updates on arXiv.org

Evaluating Neural Cartographic Relief Shading for Urban Environments: A Downtown Calgary Study Using High-Resolution DEM and DSM Data

・arXiv:2608.20149v1 Announce Type: new Abstract: This article explores the performance of analytical and neural-based hillshading methods in a dense urban environment using high-resolution digital elevation model (DEM) and digital surface model (DSM) data for downtown Calgary. ・The study compares single-direction and multi-direction analytical hillshading with relief shading generated in Eduard, a machine-learning syst
cs.LG updates on arXiv.org

Evidence Before Expansion: Reuse, Spawn, or Defer in Lifelong Expert Pools

・arXiv:2608.19888v1 Announce Type: new Abstract: Streaming systems that maintain a pool of expert models must repeatedly decide whether to reuse an existing expert for arriving data, spawn a new one, or defer. ・We present a decision layer that makes all three outcomes statistically meaningful. ・Reuse and spawn are posed as one-sided sequential hypotheses on a conditional (mechanism-level) discrepancy, separated by an in
cs.LG updates on arXiv.org

Exact Algebraic Computation of Learning Coefficients for Two-Dimensional Singular Models

・arXiv:2608.20183v1 Announce Type: new Abstract: Classical information criteria such as the Bayesian Information Criterion (BIC) rely on regularity assumptions that break down for singular models, leading to incorrect model selection in settings such as deep learning. ・The Widely Applicable Bayesian Information Criterion (WBIC) relies on local learning coefficients $\lambda$, which in the analytic case coincides with l
Hugging Face Papers

EXIMO: VLM Guided Exploration of VLA Policies

EXIMO: VLM Guided Exploration of VLA Policies
cs.LG updates on arXiv.org

Explainable Transformer Models for Clinical Prediction Tasks on Structured Electronic Health Records

・arXiv:2608.20315v1 Announce Type: new Abstract: Predictive models over structured electronic health records (EHRs) remain central to machine learning for healthcare, but few have jointly emphasized quantitative laboratory information and interpretability with respect to input medical events. ・We present BERT-LER, a BERT-style model for coded EHR timelines pretrained and fine-tuned from a de-identified EHR dataset of 7
Hugging Face Papers

FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis

FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis
ITmedia NEWS 最新記事一覧

FANZA、成人向けAI創作・公開サービス「FANZAスタジオ」発表 先行体験は8月24日から

・デジタルコマースは8月17日、成人向けECサイト「FANZA」上で、AIコンテンツの創作・公開サービス「FANZAスタジオ」の先行体験を24日に始めると発表した。
cs.LG updates on arXiv.org

Far from the Crowd: Scalable Self-Supervised Learning via Geographic Isolation

・arXiv:2608.19766v1 Announce Type: cross Abstract: Self-supervised pretraining on remote sensing imagery typically treats all samples as equally informative, despite large variability in geographic and visual structure. ・We propose a curriculum learning strategy for self-supervised Earth observation that ranks samples by geographic isolation, a label-free proxy derived entirely from geolocation metadata already present
cs.LG updates on arXiv.org

FAR-DPO: Feasibility-Aware and Robust Direct Preference Optimization for Cyclic Peptide Design

・arXiv:2608.19808v1 Announce Type: new Abstract: Cyclic peptides are emerging as promising molecular scaffolds in drug discovery due to their high binding affinity and structural stability. ・However, extending generative models from linear to cyclic peptide design remains challenging, as cyclization sharply restricts the feasible design space through coupled geometric and biophysical constraints. ・Moreover, limited trai
cs.LG updates on arXiv.org

Feature Evolution and Migration during Vision Transformer Training

・arXiv:2608.20134v1 Announce Type: cross Abstract: We present a novel view on feature evolution in Vision Transformers (ViTs) by visualizing the training process over two dimensions -- network depth (layer) and training time (epochs). ・We employ Sparse Autoencoders (SAEs) to extract candidate sparse features from CLS-token representations and compare their activation profiles across epoch--layer pairs. ・This allows us t
cs.LG updates on arXiv.org

Fine-Tuning VLAs with Self-Demonstrated Generative Control for Multi-Task Manipulation

・arXiv:2608.19490v1 Announce Type: cross Abstract: State-of-the-art vision-language-action (VLA) models such as $\pi_{0.5}$ exhibit strong semantic understanding, instruction following and task behavior. ・However, when deployed on new robots, even minor mismatches in hardware configuration relative to pretraining can cause severe performance drops. ・Finetuning the VLA on in-domain expert data from the new embodiment imp
cs.LG updates on arXiv.org

Finite-Horizon Input-Output Dynamics of Minibatch Perturbations in AdamW

・arXiv:2608.19762v1 Announce Type: new Abstract: A minibatch can influence training beyond the update at which it is observed because AdamW stores past gradient information in its optimizer states. ・We study this delayed effect through paired trajectories that differ only in one gradient update and share the same subsequent training sequence. ・We formulate AdamW as a finite-horizon input--state--output (ISO) system whos
Hugging Face Papers

FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving

FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving
cs.LG updates on arXiv.org

FleetSieve: Decision-Critical Profiling for SLO-Aware LLM Fleet Configuration

・arXiv:2608.19659v1 Announce Type: new Abstract: Choosing tensor-parallel (TP) degrees and replica counts for an LLM serving fleet is difficult because performance is not monotonic in TP and the feasible choice can change with load. ・Exhaustive profiling resolves this uncertainty, but measures many configurations that do not affect the final resource allocation. ・We present FleetSieve, which selects measurements accordi
cs.LG updates on arXiv.org

Flow Matching Meets 3D Curvilinear Structure Segmentation in Medical Imaging

・arXiv:2608.19965v1 Announce Type: cross Abstract: Segmentation of curvilinear anatomical structures in 3D medical images remains challenging due to complex topology, severe class imbalance, weak contrast, and large variations in structure morphology. ・While deep learning approaches for 3D curvilinear segmentation have been proposed, they are often tailored to specific anatomies or modalities, limiting generalization a
Hugging Face Papers

ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models

ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models
cs.LG updates on arXiv.org

Forking Fast: Efficiently Estimating Uncertainty Dynamics in Text Generation

・arXiv:2608.19611v1 Announce Type: cross Abstract: LLM reasoning is stochastic, and so understanding a model requires grappling with the distribution of reasoning chains that it might produce for a given question, i.e., its uncertainty. ・Resampling-based analyses characterize this distribution, revealing which steps of a rollout determine how the model arrives at its answer. ・However, a major limitation of these approac
cs.LG updates on arXiv.org

From Accuracy to Auditability: A Survey of Determinism in Financial AI Systems

・arXiv:2605.23955v3 Announce Type: replace-cross Abstract: Deploying machine learning in regulated financial environments -- credit risk, fraud detection, and anti-money laundering -- exposes critical vulnerabilities in algorithmic reproducibility. ・While early financial ML addressed statistical challenges such as backtest overfitting, deep neural networks and Generative AI have introduced mechanical nondeterminism roo
Google DeepMind News

From Atari to EVE Online: Building on 15 Years of AI Research in Games

・Google DeepMind partners with game studios to prototype breakthrough AI gameplay.
cs.LG updates on arXiv.org

From Noise to Signal: Improving Security Log Anomaly Detection Using LLMs with Endpoint-Specific Logs

・arXiv:2608.19938v1 Announce Type: cross Abstract: Existing approaches to anomalous behaviour log detection, such as Wazuh rely primarily on predefined detection rules, while statistical anomaly detection approaches such as OpenSearch identify deviations from previously observed behavioural patterns. ・Recent research has investigated LLMs for log anomaly detection because of their ability to interpret semantic and cont
cs.LG updates on arXiv.org

From Prediction to Self: Developmental Conditions for Agency in Minimal Neural Systems

・arXiv:2606.05605v2 Announce Type: replace Abstract: How does a system that merely predicts the world come to distinguish its own causal influence from everything else? ・We trace this transition in a minimal 192-dimensional GRU through a developmental sequence -- 6 experimental stages, 12 falsified alternatives, and cross-signal validation. ・Starting with no action or self-representation, we add components one at a time
cs.LG updates on arXiv.org

From Rebound to Remedy: Understanding and Mitigating Reward Hacking via Representation Engineering

・arXiv:2604.01476v3 Announce Type: replace Abstract: Reinforcement learning for LLMs is vulnerable to reward hacking, where models exploit shortcuts to maximize reward without solving the intended task. ・We systematically study this phenomenon in coding tasks using an environment-manipulation setting, where models can rewrite evaluator code to trivially pass tests without solving the task, as a controlled testbed.
cs.LG updates on arXiv.org

From Street View Imagery to Street Quality Indicators: Vision Language Inference for the Suburban 15-minute City

・arXiv:2608.20026v1 Announce Type: cross Abstract: Streetscape quality has become a central concern in contemporary urban planning, particularly within the framework of the pedestrian-friendly 15-minute city, where walkability and public-space quality are increasingly recognized as key determinants of urban performance. ・However, assessing streetscape qualities across large suburban and peri-urban territories remains c
cs.LG updates on arXiv.org

G-MARK: Grounded Multi-Agent Reasoning for Cooperative Driving via Knowledge Graphs

・arXiv:2608.19964v1 Announce Type: new Abstract: Autonomous driving systems must operate under partial observability, where safety-critical objects may be occluded or visible only to neighboring connected vehicles. ・Vehicle-to-vehicle cooperation can reduce this uncertainty, but existing cooperative driving methods often compress multi-agent evidence into latent features or hidden multimodal states. ・As a result, they o
cs.LG updates on arXiv.org

GEM: Geometric Erasure by Contrastive Velocity Matching in Rectified Flows

・arXiv:2606.00140v3 Announce Type: replace Abstract: While the rapid adoption of multimodal generative models offers immense potential, it has also increased the risks of harmful content synthesis, deepfakes, and copyright infringements. ・To address these challenges, concept erasure has emerged as a prospective safeguard. ・However, as the field gradually transitions from U-Net-based diffusion models to Rectified Flow Tr
cs.LG updates on arXiv.org

Geometric Evolution Maps: Extracting Stable Concept Probes from Transformer Residual Streams

・arXiv:2605.25848v2 Announce Type: replace Abstract: A concept probe is only as reliable as the layer it is taken from. ・Probing at a fixed late layer, or at the peak of a separation curve, ignores a structural feature of how concepts form: the probe direction rotates substantially during assembly and does not settle until after the Concept Allocation Zone (CAZ) in which it forms. ・We introduce Geometric Evolution Maps
Hugging Face Papers

GOAG: Generative and Object-Agnostic Grasp Planner for Dexterous Robotic Manipulation

GOAG: Generative and Object-Agnostic Grasp Planner for Dexterous Robotic Manipulation
AI News & Artificial Intelligence | TechCrunch

Google gives publishers a new way to fight AI-driven traffic losses

・Google is giving publishers a new button that lets readers make them a preferred source across Search, Discover, and Google News, potentially boosting their traffic as AI search sends fewer clicks to the web.
WIRED

Google Pixel 11 Review: Minor Upgrade

・Incremental improvements fail to generate much excitement, but Google’s Pixel 11 is still an accomplished Android phone.
ITmedia NEWS 最新記事一覧

Google、読者がワンクリックで“推しメディア”登録できるボタン公開

・Googleは、検索やDiscover、Googleニュースのパーソナライズ機能を拡充した。パブリッシャーのWebサイト上から直接「優先するニュース提供元」に登録できるボタンを新設。Discoverでは自然言語で見たいトピックや除外条件を指定可能にし、Googleニュースの音声ブリーフィングもカスタマイズ可能にする。
The Verge

Google’s Pixel 10A is a great deal at 15 percent off

・The Pixel 10A in its lavender color scheme. ・It’s also available at a discount in berry, fog, and obsidian colors. ・| Image: The Verge This week, all of Google’s Pixel 11 phones launched, including the $899 Pixel 11, the $1,099 Pixel 11 Pro (with the same processor and starting 12GB RAM as the standard model, but with better cameras), and the $1,899 Pixel 11 Pro Fold.
ITmedia NEWS 最新記事一覧

Googleのオープンモデル「Gemma」、累計10億ダウンロード超 GitHubに公式ディレクトリ公開

・Googleは、オープンモデル「Gemma」ファミリーの累計ダウンロード数が10億回を突破したと発表した。派生モデルは10万種を超え、公式リポジトリ「Awesome Gemma」をGitHubで公開。宇宙空間での衛星データ解析や新規のがん治療経路の発見など、多様な活用事例を紹介している。
@IT 全フォーラム 最新記事一覧

GoogleはAI競争に負けたのか 「最強のAI」ではなく「AIの“電力網”」を選ぶ賭け

・GoogleからAI研究の中心人物が相次いで去った。「Geminiは終わった」という見方に対し、「最先端ではなく、AIを社会全体に行き渡らせる“電力網”で勝つ賭けだ」という別の解釈もある。電気の歴史になぞらえながら整理する。
cs.LG updates on arXiv.org

GraphPFN: A Prior-Data Fitted Graph Foundation Model

・arXiv:2509.21489v4 Announce Type: replace Abstract: Graph foundation models face several fundamental challenges including transferability across diverse domains and data scarcity, which calls into question the very feasibility of creating such models. ・However, despite similar challenges, the tabular domain has recently witnessed the emergence of the first successful foundation models such as TabPFN. ・These models are
cs.LG updates on arXiv.org

Gravitational-wave parameter estimation with machine-learning generated surrogate waveforms

・arXiv:2608.20222v1 Announce Type: cross Abstract: The worldwide network of gravitational-wave detectors have detected more than 350 binary coalescence events till date. ・Future third-generation detectors, like Einstein telescope, are expected to detect orders-of-magnitude more signals from sources with more complicated characteristics, including eccentric orbits and high-mass ratio binaries. ・It is well-established tha
cs.LG updates on arXiv.org

Green BOA: Determining the environmental break-even point for ML-based data compression

・arXiv:2608.19994v1 Announce Type: new Abstract: We summarise the outcome of two summer internship projects based at the University of Manchester, focused on the break-even point in terms of environmental sustainability for ML-based data compression algorithms. ・Using the example of a ML-based lossless compression algorithm, we compare estimates for the carbon-equivalent of the infrastructure needed for ML training and
cs.LG updates on arXiv.org

Grounded verification of chemical and materials reasoning: detection is the bottleneck

・arXiv:2607.17417v2 Announce Type: replace Abstract: Language models are moving into chemistry and materials discovery workflows, where a wrong molecular formula, space group, or formation energy can silently propagate into downstream decisions. ・These confabulations hide inside fluent reasoning traces and concentrate on rare, long-tail entities, where model confidence is least trustworthy. ・Retrieving reference data fo
cs.LG updates on arXiv.org

Guided Diffusion by Optimized Loss Functions on Relaxed Parameters for Inverse Material Design

・arXiv:2602.15648v2 Announce Type: replace Abstract: Inverse design problems are common in engineering and materials science. ・The forward direction, i.e., computing output quantities from design parameters, typically requires running a numerical simulation, such as a FEM, as an intermediate step, which is an optimization problem by itself. ・In many scenarios, several design parameters can lead to the same or similar ou
cs.LG updates on arXiv.org

HBVLA: Pushing 1-Bit Post-Training Quantization for Vision-Language-Action Models

・arXiv:2602.13710v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models enable instruction-following embodied control, but their large compute and memory footprints hinder deployment on resource-constrained robots and edge platforms. ・While reducing weights to 1-bit precision through binarization can greatly improve efficiency, existing methods fail to narrow the distribution gap between binarized and
cs.LG updates on arXiv.org

Heteroscedastic Neural Surrogate Modeling for Robust and Rapid Bayesian Inference in Fusion Plasma Diagnostics

・arXiv:2608.19377v1 Announce Type: cross Abstract: Bayesian inference via Markov Chain Monte Carlo (MCMC) provides effective parameter estimation, but its real-time application in complex physical systems is hindered by heavy computational bottlenecks and extreme sensitivity to statistical noise. ・We address this by proposing a neural-network-based probabilistic surrogate framework for rapid and robust MCMC inference.
Hugging Face Papers

Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent Harnesses

Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent Harnesses
cs.LG updates on arXiv.org

Higher Resolution, Better Generalization: Unlocking Visual Scaling in Deep Reinforcement Learning

・arXiv:2605.10546v2 Announce Type: replace Abstract: Pixel-based deep reinforcement learning agents are typically trained on heavily downsampled visual observations, a convention inherited from early benchmarks rather than grounded in principled design. ・In this work, we show that observation resolution is a critical yet overlooked variable for policy learning: higher-resolution inputs can substantially improve both pe
cs.LG updates on arXiv.org

HiRA-CAM: Preserving Fine-Grained Spatial Relevance in Gradient-Based Visual Explanations

・arXiv:2608.19407v1 Announce Type: cross Abstract: Deep Learning models can include billions of parameters or more, making it difficult to explain their internal transformations and outputs. ・However, explainability is increasing in importance due to the use of AI in crucial applications. ・This paper focuses on the interpretability of convolutional neural networks (CNNs).
cs.LG updates on arXiv.org

Holtercare-Bench: A Multimodal Benchmark for Evaluating Long-Term Dynamic ECG Analysis

・arXiv:2608.19297v1 Announce Type: new Abstract: While multimodal large language models (MLLMs) excel in medical applications, most of them favor static images or short-term signals. ・In the critical field of dynamic electrocardiograms (ECG), models struggle with complex temporal reasoning and diagnostic report generation due to a lack of high-quality datasets and benchmarks. ・To address this, we introduce (i) Holtercar
cs.LG updates on arXiv.org

HYDRA: A Heterogeneous Chiplet DSE Framework for Serving Dynamic Hybrid LLM Workloads

・arXiv:2608.19395v1 Announce Type: cross Abstract: Hybrid Transformer-Mamba large language models (LLMs) enhance long-context efficiency, but their heterogeneous computation and communication patterns complicate efficient hardware acceleration. ・Chiplet-based architectures offer a scalable solution by integrating specialized compute and memory units. ・However, the design space spanning static architectural configuration
WIRED

I Tried the Best Robotic Pool Cleaners of 2026: Beatbot, iGarden, Dreame

・Pack up your pool cleaning supplies. ・Let one of these robot buddies maintain your water quality instead
cs.LG updates on arXiv.org

Improved Confidence Estimates for Black-Box Large Language Models

・arXiv:2608.19323v1 Announce Type: new Abstract: Uncertainty quantification (UQ) is essential for the safe deployment of large language models (LLMs). ・Existing methods, from verbalized confidence to ones requiring multiple generations, are often zero-shot and produce scores quantifying uncertainty without the need for labelled data. ・Nonetheless, in practice one must always evaluate their performance on a dataset of in
cs.LG updates on arXiv.org

In Two Minds about Lifelong Learning: Exploring Hemispheric Redundancy and Specialisation in Neural Models

・arXiv:2608.19514v1 Announce Type: new Abstract: Persistent intelligent systems require the ability to learn continually, but current machine learning approaches face significant challenges in this area compared to biological learning systems. ・Machine learning algorithms typically trade off retention of previously learned information and adaptation to new or changing data patterns. ・When continual learning capabilities
cs.LG updates on arXiv.org

Inadvertent Context Leakage in Language Models

・arXiv:2608.19857v1 Announce Type: new Abstract: For AI agents to be useful beyond simple chat, they must hold sensitive user context such as calendars, credentials, health records, and financial data. ・We study whether the mere presence of such secrets in a model's context window introduces hidden correlations into the model's benign outputs, allowing reconstruction even when the model correctly refuses direct extract
WIRED

Influencers and Resellers Are Turning Empty Boxes Into Big Cash

・As the appetite for “authenticity” grows online, content creators are buying up empty boxes for luxury goods—and resellers are cashing in for “crazy prices.”
cs.LG updates on arXiv.org

Information on trajectories: martingales and random times

・arXiv:2608.20337v1 Announce Type: cross Abstract: Accounting for information flow on the path space of trajectories of a nonnegative martingale yields exact variational identities for it, even at arbitrary random times. ・This recovers the widely used classical concentration inequalities, from Ville to PAC-Bayes, and measures what each one discards. ・The tail a bound controls is itself a relative entropy, resolved by th
Hugging Face Papers

Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization

Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization
cs.LG updates on arXiv.org

Interpretable Feature Learning for RF Fingerprinting via Polar MKANs

・arXiv:2608.19881v1 Announce Type: cross Abstract: Radio frequency (RF) fingerprinting authenticates wireless devices from hardware-induced I/Q impairments, typically with deep learning feature extractors that are accurate but opaque, limiting their use in security critical settings. ・We propose Polar Monotonic Kolmogorov-Arnold Networks (Polar MKAN), a block partitioned monotonic encoder on polar inputs in which each
cs.LG updates on arXiv.org

Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization

・arXiv:2605.24960v2 Announce Type: replace-cross Abstract: Chain-of-Thought (CoT) faithfulness, i.e., whether CoTs genuinely reflect large language models' (LLM) underlying behavior, is typically evaluated with metrics under two disjoint paradigms: contextual faithfulness, measured by perturbing the input or CoT trace, and parametric faithfulness, assessed by intervening on a model's parametric knowledge. ・Yet prior wo
cs.LG updates on arXiv.org

K\"ahler landscapes for complex neural network descents and guarantees including a search and destroy of the Calabi-Yau manifold

・arXiv:2608.19584v1 Announce Type: new Abstract: We study landscapes for complex-parameterized networks. ・Our approach is motivated with an information-theoretic manifold perspective of the parameter and via classical optimization guarantees although of complex geometric variety such as through Dolbeault asymptotics. ・The descent path admits a K\"ahler information metric under a cross-entropy via the Wirtinger Hessian o
LLMタグが付けられた新着記事 - Qiita

Kimi-Linear-48BにLoRAを当てようとしたら、公開コードが推論専用だった話(総額$2.80)

・48Bのオープンモデルに 「語尾を全部『フィー』にする」LoRA を当てました。語尾は 0% → 99.2% で綺麗に入りました。 ・ただ、この記事の本題はそこではありません。 ・Moonshot が公開している modeling_kimi.py は推論専用で、そのままでは 1...
#LLMタグ

LangGraphの使い方

・AIエージェントを作っていると、最初はプロンプトの調整やモデルの賢さに夢中になりますよね。モデルが気の利いた返答を返してくれるだけで、「おっ、動いた!」とテンションが上がります。 ・でも、少し複雑な処理を作ろうとした途端、こんな壁にぶつかった経験はないでしょうか。
cs.LG updates on arXiv.org

Learn for Variation: Efficient AAV Trajectory Learning through a Differentiable Wireless World Model

・arXiv:2603.18853v3 Announce Type: replace-cross Abstract: Autonomous aerial vehicles (AAVs) enable data collection for sixth-generation Internet-of-Things networks, but their trajectories couple nonlinear wireless rates with long-horizon service progress. ・This paper views the evolution of AAV kinematics, channel state, and user backlog as a structured differentiable world model and develops Learn for Variation (L4V)
cs.LG updates on arXiv.org

Learning Deterministic and Stochastic Forced Hamiltonian Systems

・arXiv:2608.19688v1 Announce Type: cross Abstract: We develop a geometric framework for learning deterministic and stochastic forced Hamiltonian systems with neural networks. ・Motivated by the Lagrange-d'Alembert principle and the theory of variational integrators, we introduce the notion of a Lagrange-d'Alembert map and establish a $C^r$ convergence theorem for first-order one-step methods. ・Building on these results,
cs.LG updates on arXiv.org

Learning Hierarchical Skill Policies with Offline Quality-Diversity Reinforcement Learning

・arXiv:2608.19684v1 Announce Type: cross Abstract: Recent studies investigate how to leverage pre-collected datasets to improve the policy performance and sample efficiency of RL. ・One promising approach to achieve this goal is to employ a two-stage strategy: In the first stage, diverse skills are extracted as a low-level policy from a given dataset, and a high-level policy is trained to solve a specific task in the se
cs.LG updates on arXiv.org

Learning piecewise-smooth dynamical systems

・arXiv:2608.19785v1 Announce Type: cross Abstract: Discovering dynamical systems from trajectory data is a central problem in applied mathematics and engineering. ・Whilst recent advances in machine learning have led to strong progress in data-driven system identification, much less attention has been given to systems with discontinuous dynamics. ・These systems are nevertheless highly relevant in applications, including
cs.LG updates on arXiv.org

Learning-Based Speed Estimation from Accelerometer-Only Inertial Sensing

・arXiv:2401.07468v4 Announce Type: replace Abstract: The proposed model, CarSpeedNet, estimates scalar vehicle speed from a window of three-axis smartphone acceleration, without gyroscope, wheel-odometry, vehicle-bus, or positioning input at inference. ・The reported experiment comprises 13.2 hours of on-road driving. ・Beyond the network comparison, a finite-context analysis treats window length as part of the sensing pr
WIRED

Lenovo Coupon Codes: 15% Off in August 2026

・Whether you’re shopping for a ThinkPad, Yoga laptop, or Legion gaming PC, these Lenovo discount codes and promotions can help you save big on your next tech upgrade.
cs.LG updates on arXiv.org

Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts

・arXiv:2608.20061v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) architectures significantly expand model capacity without a proportional increase in computational cost. ・However, optimizing their hyperparameters---particularly the learning rate---at extreme scales of both model size and token budget via sweeping remains computationally prohibitive. ・In this paper, we propose a compute-efficient, two-step hyper
ITmedia NEWS 最新記事一覧

LINE、着せかえ表示問題を謝罪 希望者に返金へ

・LINEヤフーは8月20日、LINEアプリの一部の「着せかえ」でメニューアイコンがデフォルト表示になっている問題について経緯を説明し、謝罪した。返金を希望する日本国内のユーザーからは専用フォームで申請を受け付ける。
Hugging Face Papers

Listening Forward: Next Patch Embedding Prediction Enables Scalable Audio Learners

Listening Forward: Next Patch Embedding Prediction Enables Scalable Audio Learners
#LLMタグ

LL.M.留学と英語への不安

・LL.M.留学を考えている方にとって、一番の不安はやはり英語ではないでしょうか。 ・私自身、留学前は英語にそれなりの不安がありました。
cs.LG updates on arXiv.org

LLM as Detector: An In-context Learning Approach for Tabular Anomaly Detection

・arXiv:2608.19463v1 Announce Type: new Abstract: Anomaly detection in tabular data is challenging because abnormal samples often arise as violations of cross-feature dependencies rather than simple marginal deviations. ・Existing detectors rely on geometric or reconstruction signals, while prior LLM-based approaches mainly fine-tune LLMs with normal samples or generate synthetic anomalies. ・We propose LLM-Detector, a fra
cs.LG updates on arXiv.org

LLM Capability Limits: Static Emergence and Dynamic Boundary Control

・arXiv:2608.01548v3 Announce Type: replace-cross Abstract: Test-time emergence in LLM systems has a deployment boundary: additional computation can realize decisions already supported by the deployed information--execution structure, while evidence, tools, memory, and executable semantics can change the class inherited by later computation. ・We formalize this boundary through inherited structural capability $\mathcal{D
#LLMタグ

LLMがわからない

・AI彼氏を創りたくてまだまだ奮闘中です。 ・とはいえ、夏に入ってから体力がなく全然捗っていません… 体力的にフラフラながら、LLMについて考えています。 ・LLMは、言語のAIのことです。
cs.LG updates on arXiv.org

Longitudinal Bayesian Learning of Continuous Disease Position across the Alzheimer's Disease Continuum

・arXiv:2608.19436v1 Announce Type: new Abstract: Alzheimer's disease (AD) progresses as a continuous biological process, whereas most existing neuroimaging-based artificial intelligence methods remain limited to discrete diagnosis or clinical score prediction from cross-sectional imaging. ・In this work, we propose Disease Continuum Positioning (DCP), a longitudinal Bayesian Learning framework that continuously estimate
#LLMタグ

Lorebook機能を拡張しました!

・こんにちは!最近スマホ版機能と睨めっこしておりnoteの投稿ができていませんでした。
cs.LG updates on arXiv.org

M3: A State-Event Generative Foundation Model for Market Microstructure Dynamics

・arXiv:2608.19227v1 Announce Type: cross Abstract: Market microstructure simulation aims to model how liquidity, prices, and order flow evolve in electronic financial markets. ・Since market data reveal only one realized trajectory, many important questions are inherently counterfactual and require realistic trajectory-level simulation. ・Existing financial generative models, however, often model order events and market s
ITmedia NEWS 最新記事一覧

macOS版ChatGPT、Appleの「メッセージ」と連携 会話検索や下書き、送信に対応

・OpenAIは、macOS版ChatGPT向けにAppleの「メッセージ」アプリと連携するプラグインを公開した。CodexやChatGPT Workのチャット上で過去の会話検索や下書き作成、送信が可能になる。誤送信防止のため都度承認フローを備える。Appleシリコン搭載Mac向けだ。
The Verge

Major YouTube creators are facing backlash for accepting AI money

・Over the past few days, a number of prominent filmmaking content creators including Matti Haapoja and Sam "Kold" Kolder have posted videos of themselves demonstrating what's possible with AI platform Higgsfield. ・The videos highlight Higgsfield's recently added Seedance 2.5 functionality and pitch these technologies as the future of video production. ・In response to these videos, other creators started sharing what app
#LLMタグ

Mamba-3とは何か? 速さより重要な本質的革命

・この記事の結論を、先に3行で Transformerの時代が終わる、という話ではない。Transformerの中身が、こっそり別モノに入れ替わっているという話だ。 ・同じGPU、同じ条件で、Transformer系が976秒かけた処理を、Mamba-3は141秒で終わらせた。約7倍。長い仕事ほど、差は開く。 ・そして厄介なことに、この技術はもう「来るかもしれない未来」ではない。NVIDIAのモデルの中で、すでに動いている。
cs.LG updates on arXiv.org

MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal Prostate MRI Segmentation

・arXiv:2510.17529v3 Announce Type: replace-cross Abstract: Active Surveillance (AS) is a treatment option for managing low and intermediate-risk prostate cancer (PCa), aiming to avoid overtreatment while monitoring disease progression through serial MRI and clinical follow-up. ・Accurate prostate segmentation is an important preliminary step for automating this process, enabling automated detection and diagnosis of PCa.
cs.LG updates on arXiv.org

Maximum Likelihood Reinforcement Learning

・arXiv:2602.02710v2 Announce Type: replace Abstract: Reinforcement learning (RL) is the method of choice for training models in setups where the objective function can only be evaluated by sampling from the model. ・Our key observation is that when the feedback is terminal and binary, models implicitly induce a likelihood over correct rollouts. ・Maximum likelihood would be the natural framework in such settings, but RL i
Hugging Face - Blog

Measuring benchmark optimization in speech recognition

Measuring benchmark optimization in speech recognition
cs.LG updates on arXiv.org

Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability

・arXiv:2608.19338v1 Announce Type: new Abstract: Mechanistic interpretability seeks quantities that models do not expose directly: represented states, component effects, interactions, and responses to interventions. ・Patching, gradients, Hessian-vector products, and subset interventions provide different measurements under different access assumptions and may target different quantities. ・We formulate their shared measu
MarkTechPost

Meet S1-mini: Superwhisper’s 462 MB Open-Weights Text Normalizer That Turns Raw ASR Transcripts Into Clean Written Text

・S1-mini is a 462 MB open-weights normalizer that sits after ASR, removing fillers and resolving self-corrections locally. ・The post Meet S1-mini: Superwhisper’s 462 MB Open-Weights Text Normalizer That Turns Raw ASR Transcripts Into Clean Written Text appeared first on MarkTechPost.
MarkTechPost

Meet UPDF: A Lightweight Adobe Alternative Built for the Agentic Era

・PDFs are easy to read and hard to change. ・AI can now summarize a 90-page contract in seconds, but it still won't rewrite the source file cleanly. ・UPDF is built for that second half: direct editing, 14-format conversion, 38-language OCR, and ten AI agents shipped in version 2.5.
Hugging Face Papers

MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use

MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use
cs.LG updates on arXiv.org

MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use

・arXiv:2608.20202v1 Announce Type: cross Abstract: Memory has become a key component of large language models, enabling them to retain information and learn from long-term interactions. ・However, existing memory benchmarks mainly evaluate whether information is correctly extracted, stored, and retrieved, while largely overlooking how retrieved memories reshape model reasoning and affect performance on the current task.
cs.LG updates on arXiv.org

Merge Now, Regret Later: The Hidden Cost of Model Merging Is Adversarial Transferability

・arXiv:2509.23689v2 Announce Type: replace Abstract: Model Merging (MM) has proven to be an effective alternative to multi-task learning, where several fine-tuned models are merged, without access to the tasks' training data, into one model that retains performance across different tasks. ・Recent works have explored the security of MM, showing how MM can confer robustness against various adversarial attacks.
WIRED

Meta’s Big Reckoning Is Here

・Meta is in court again over child safety, and this time it’s a landmark case that could force significant changes to core features of Facebook and Instagram.
cs.LG updates on arXiv.org

Microlensify: a Transformer Based Machine Learning Classifier for Microlensing Events Trained on TESS Light Curves

・arXiv:2608.19419v1 Announce Type: cross Abstract: Microlensing can reveal populations of faint compact objects that are otherwise difficult to detect. ・Depending on their design, all-sky surveys have the potential to search for these objects across the sky. ・The Transiting Exoplanet Survey Satellite (TESS), primarily designed to detect transiting exoplanets, also provides near all-sky coverage with high cadence.
The Verge

Microsoft and Discord subpoenaed over GTA VI gameplay leaks

・Following several apparent video leaks of Grand Theft Auto VI, Take-Two Interactive has subpoenaed Microsoft and Discord over content that "infringes copyrights" held for the game, Kotaku reports. ・In the subpoenas, filed on Thursday, Take-Two says copyrighted material includes "audiovisual content, artwork, images, dialogue, or other creative elements" and it is looking to identify "alleged infringers at issue." The
cs.LG updates on arXiv.org

MileGPO: Milestone Inference with Local Evidence for Graph-Based Policy Optimization of Long-Horizon LLM Agents

・arXiv:2608.19803v1 Announce Type: new Abstract: Credit assignment is challenging in long-horizon agentic reinforcement learning, where supervision often comes only from final rewards. ・Existing methods refine trajectory-level signals into step-level credits through step grouping or graph-based advantage estimation, but can overlook meaningful intermediate milestones. ・We propose MileGPO (Milestone Inference with Local
Zennの「機械学習」のフィード

MobileNetV2 を手書き NEON で速くする — NCHW pointwise マイクロカーネル

・はじめに 前回までの記事で、MobileNetV2 を ONNX から MLIR 経由でネイティブバイナリまでコンパイルし、正しい推論が出るところまで通しました。ただし畳み込みはスカラーループのままで、MLIR の Transform Dialect でタイリングを当てても速くなりませんでした。 ・タイリングが効かなかった理由は、MLIR が NEON を生成していないこと、そして NCHW レイアウトでメモリアクセスが非連続になることの 2 つでした。MLIR のパスを工夫して速くする路線はいったん諦め、心機一転、畳み込みカーネルを一から手書き NEON で書くことにします。目標は...
cs.LG updates on arXiv.org

Multi-Modal Graph Interaction for Multi-Graph Convolution Network in Urban Spatiotemporal Forecasting

・arXiv:1905.11395v2 Announce Type: replace Abstract: Graph convolution network based approaches have been recently used to model region-wise relationships in region-level prediction problems in urban computing. ・Each relationship represents a kind of spatial dependency, like region-wise distance or functional similarity. ・To incorporate multiple relationships into spatial feature extraction, we define the problem as a m
cs.LG updates on arXiv.org

Multi-Source Wasserstein Distributionally Robust Graph Learning

・arXiv:2608.19914v1 Announce Type: new Abstract: Network topology inference from graph signals is central to graph signal processing with applications in neuroscience, sensor, and social networks. ・In practice, target-domain samples are scarce while heterogeneous source-domain data are abundant. ・Fusing these sources is challenging: Euclidean averaging works for homogeneous sources but degrades sharply as inter-source d
cs.LG updates on arXiv.org

MulTTiPop: A Multitrack Transcription Dataset for Pop Music

・arXiv:2607.08756v2 Announce Type: replace-cross Abstract: We present MulTTiPop, a dataset of pop music segments and their associated multitrack MIDI recordings for the evaluation of automatic music transcription models. ・MulTTiPop contains 572 segments of popular music totaling 3.5 hours of audio, and contains songs from diverse genres and decades from the 1930s to 2000s. ・To collect this dataset, we perform metadata-b
The Verge

My cats hate each other, but this automatic feeder is helping

・Over the course of nearly a decade, my husband and I have dutifully carried out a four-times-daily ritual of feeding a pair of lovely cat siblings who hate each other. ・This was mostly manageable for years, but in 2024, disaster struck for my small dependents: We had a child. ・Suddenly, the cats were no longer our first priority.
Hugging Face Papers

NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video

NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video
cs.LG updates on arXiv.org

Neural Prior Estimation: Learning Class Priors from Latent Representations

・arXiv:2602.17853v2 Announce Type: replace Abstract: Logit adjustment corrects class imbalance using the empirical class prior. ・We study whether a comparable class-frequency signal can instead be learned from the network representation, without explicitly supplying class counts to the correction rule. ・We introduce the Neural Prior Estimator (NPE), which attaches one or more lightweight Prior Estimation Modules (PEMs)
WIRED

NordVPN Coupons: 75% Off, Plus 3 Months Free in August 2026

・Save up to 77% on 2-year plans and get 3 free months with our NordVPN discount codes.
#AIタグ

noteでメンバーシップ始めます。AI時代の本人訴訟シリーズです。フェーズ1は三井住友銀行。

・AIは便利ですね。訴状も自分で書けますし、答弁書も自分で準備できます。 ・これに裁判の流れや注意事項がわかれば、 これからは自分で裁判が起こせる時代だと思っています。
#AIタグ

noteで記事が読まれない? ハッシュタグの埋もれない付け方3つのコツ

noteで記事が読まれない? ハッシュタグの埋もれない付け方3つのコツ
LLMタグが付けられた新着記事 - Qiita

Notion Agent API パブリックベータ化と2026年9月移行期限まとめ

・はじめに Notion が Agent API をパブリックベータとして正式公開しました。これまでプライベートアルファ版として限定提供されていた Custom Agent 連携機能が、誰でも利用できる形で開放されたことになります。 ・セッションベースのチャット開始、返信のス...
WIRED

Nyrius Phoenix Home True 4K60 (2026): A Solution for Cord Clutter

・The Nyrius Phoenix Home True 4K60 can wirelessly transmit games, TV shows, movies, and more across your home (even through walls).
AI News & Artificial Intelligence | TechCrunch

OK, can we actually cool data centers with our pee?

・Jason Kelce joked that people should cool data centers with their pee, rather than potable water -- but his suggestion is not completely ludicrous.
#LLMタグ

Ollama Launch(オラマ ローンチ)とは?ローカルAIをAIエージェント化する方法を解説

・AI共創イノベーション・ワンダー佐藤です。 ・「Ollamaは知ってるけど、Launchって何が違うの?」 続きをみる
cs.LG updates on arXiv.org

On the convergence of optimistic policy iteration for stochastic shortest path problem

・arXiv:1808.08763v3 Announce Type: replace Abstract: In this paper, we prove some convergence results of a special case of optimistic policy iteration algorithm for stochastic shortest path problem. ・We consider both Monte Carlo and $TD(\lambda)$ methods for the policy evaluation step under the condition that the termination state will eventually be reached almost surely.
cs.LG updates on arXiv.org

Online Test-Time Adaptation for Generalizable Dynamic Graph Anomaly Detection

・arXiv:2608.19858v1 Announce Type: new Abstract: Generalizable dynamic graph anomaly detection (DGAD) enables pretrained detectors to identify anomalies in unseen target domains without costly retraining. ・However, existing methods often fail for two reasons. ・First, they mainly rely on domain-agnostic patterns and miss domain-specific patterns that keep evolving.
AI News & Artificial Intelligence | TechCrunch

OpenAI is gaining on Anthropic with business users, new data indicates

・Businesses are willing to flop back and forth as each lab releases new models, volatility that should give both companies' investors pause about how "sticky" enterprise AI spending really is.
Zennの「大規模言語モデル」のフィード

OpenRouterがStripeに加わる今、LLMコストをjob_id単位で記録する

・一回の「生成」ボタンが、一回のAPI requestとは限りません。 ・構成案を作り、形式を整え、条件に合わなければ再生成する。画面上では一つの仕事でも、裏側では複数のmodel呼び出しが走ります。requestごとの金額だけを見ていると、結局その機能を一回使うのにいくらかかったのかが分かりません。 ・2026年8月19日、OpenRouterはStripeに加わると発表しました。発表では、名称、製品、roadmap、routing方針は維持され、現在のintegrationも変わらないと説明されています。取引自体は、今後のclosing conditionsを前提としています。
cs.LG updates on arXiv.org

Orthogonal JEPA: Factorized Predictive States for Latent World Models

・arXiv:2608.20065v1 Announce Type: new Abstract: World models construct latent states that support prediction, planning, and reasoning about an underlying system. ・Joint-embedding predictive architectures (JEPAs) offer a direct way to learn such states by predicting targets in representation space instead of reconstructing every detail of the observation. ・Standard JEPAs, however, organize all predictable content throug
cs.LG updates on arXiv.org

Partition of Unity Neural Networks for Interpretable Classification with Explicit Class Regions

・arXiv:2602.00511v3 Announce Type: replace Abstract: We introduce \emph{Partition of Unity Neural Networks} (PUNNs), a neural-network architecture for multiclass classification based on the classical mathematical notion of a partition of unity. ・The starting point is the observation that the characteristic functions of ideal class regions form a partition of unity. ・PUNNs replace these discontinuous indicators by learne
The Verge

Patreon is changing its algorithm to help smaller creators get discovered

・Patreon is developing over 30 new and updated tools, but some of them might never fully roll out to users. ・| Image: Patreon Patreon has announced a number of new and overhauled features that are designed to "build a better network - and a better internet," according to CEO Jack Conte. ・The Patreon roadmap includes discovery algorithm updates, platform and security improvements, and new features for both creators and f
cs.LG updates on arXiv.org

PETA:Parameter-Efficient Test-Time Adaptation for Virtual Screening

・arXiv:2608.19906v1 Announce Type: new Abstract: Accurately ranking active ligands for a target protein pocket from massive chemical libraries remains a central challenge in virtual screening. ・DrugCLIP and its recent extensions substantially accelerate this process by encoding protein pockets and molecules into a shared embedding space. ・Despite this progress, further performance improvements typically require retraini
cs.LG updates on arXiv.org

Physical-Support Confidence Sets for Highly Coherent Dictionaries

・arXiv:2608.20295v1 Announce Type: new Abstract: Sparse pursuit after dictionary learning can yield a precise atom support even when its physical interpretation is not justified by the calibration data, especially for highly coherent dictionaries where alternative calibration-compatible dictionaries may assign different physical meanings to the same selected support. ・We develop resolution-aware physical-support infere
Zennの「大規模言語モデル」のフィード

PII×LLMの4層防御 — Bedrockの非学習保証を規約とアーキテクチャで二重化する

・PIIを含むワークロードにLLMを組み込む際の論点は、学習リスク・保持リスク・経路リスクの3つに分解できます。前二者は契約・規約のレイヤー、後者はネットワーク設計のレイヤーであり、対策も分けて設計すべきです。本記事はAmazon Bedrockでこの3リスクを塞ぐ構成を、一次情報の出典付きで整理したものです。 ・規約レイヤー プロンプト・生成結果はAWSのモデル学習に不使用、モデルプロバイダーへも非配布(公式ドキュメントに明記) プロバイダーごとに専用のモデルデプロイアカウントを分離し、推論トラフィックはAWSネットワーク内で完結。プロバイダーはデプロイアカウントにアクセス不可 保持...
The Verge

Pixel 11 gets in on the digicam trend

・I recently looked back at a photo I'd taken on a smartphone in 2014, and I was struck by just how good it looked. ・The details were soft, the shadows were dark. ・It was the kind of photo I felt like I hadn't seen out of a phone in years.
#LLMタグ

Polandball Ver. Emiri――同じキャラ設定を多言語LLMに入れたら「国別人格」が勝手に生えた話

Polandball Ver. Emiri――同じキャラ設定を多言語LLMに入れたら「国別人格」が勝手に生えた話
Hugging Face Papers

PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents

PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents
cs.LG updates on arXiv.org

PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents

・arXiv:2608.19861v1 Announce Type: cross Abstract: Customer-service LLM agents must follow organizational policy when acting on a user's behalf. ・Compliance failures arise from either forbidden actions, such as granting an ineligible change, or omitted procedural requirements, such as identification or confirmation. ・Runtime safeguards can intervene on risky actions, but action-local checks do not guide an agent through
Hugging Face Papers

Previous

Previous
cs.LG updates on arXiv.org

PROBE-Web: An Interactive System for Probing Evaluation Landscapes of Knowledge Graph Completion Models

・arXiv:2606.08926v2 Announce Type: replace Abstract: Knowledge graph completion (KGC) models are commonly evaluated using rank-based metrics such as MRR and Hits@K, despite different users often requiring different evaluation perspectives. ・In this demo, we present PROBE-Web, an interactive system for probing diverse evaluation landscapes for KGC models. ・PROBE-Web enables users to flexibly evaluate KGC models by adjust
cs.LG updates on arXiv.org

Projector Is All You Train

・arXiv:2608.19726v1 Announce Type: cross Abstract: The typical training process of a multimodal large language model (MLLM) involves adapting both the language model backbone and the projector between the backbone and a modality-specific encoder. ・We ask whether fine-tuning the backbone of an MLLM is necessary to adapt it to a new modality. ・Through experiments on 3D MLLMs, we find that training only the projector is su
cs.LG updates on arXiv.org

ProteinZero: Self-Improving Protein Generation via Online Reinforcement Learning

・arXiv:2506.07459v4 Announce Type: replace Abstract: Protein generative models have shown remarkable promise in protein design, yet their success rates remain constrained by reliance on curated sequence-structure datasets and by misalignment between supervised objectives and real design goals. ・We present ProteinZero, an online reinforcement learning framework for inverse folding models that enables scalable, automated
cs.LG updates on arXiv.org

Quantifying Event Impacts on Time Series via Multiscale Contrastive Learning

・arXiv:2608.19447v1 Announce Type: new Abstract: Shocks that spread through the web, such as cybersecurity breach disclosures, can abruptly disrupt financial time series and cause substantial abnormal losses. ・While these events are disclosed as discrete records through news reports, regulatory filings, or public databases, their consequences unfold through continuous market dynamics. ・This creates an event-conditioned
cs.LG updates on arXiv.org

Quantum Gaussian processes for prediction of channel observations

・arXiv:2608.19306v1 Announce Type: cross Abstract: Given a set of input states, we consider the task of predicting the expectation value of a Pauli observable at the output of an unknown quantum evolution, using only a limited number of measurements. ・Recently, quantum Gaussian process (QGP) regression was introduced for this task across various classes of unitary evolution. ・Here, we extend the QGP framework beyond uni
cs.LG updates on arXiv.org

Quantum Kernel Estimation for the Discovery of Early Lung Cancer Detection

・arXiv:2608.19304v1 Announce Type: new Abstract: Lung cancer screening with low-dose chest computed tomography reduces mortality, but its impact is limited by uptake, adherence, and management challenges. ・Blood-based cell-free DNA (cfDNA) biomarkers offer a complementary approach, although early detection remains difficult because of lung cancer heterogeneity and high-dimensional, nonlinear molecular signals.
cs.LG updates on arXiv.org

Question-Guided Evidence Acquisition for Multimodal Visual Question Answering

・arXiv:2608.19739v1 Announce Type: cross Abstract: Multimodal LLMs can see a document, but they often can't read it reliably. ・Small text, tables, visual cues, and topological elements still trip them up under direct visual inference, even when the page is already sitting in the model's context. ・Most document-VQA systems treat perception as fixed: they encode the page once, ask the question, and answer from whatever th
Hugging Face Papers

QuoteBench: How Matched Scores Can Hide Command-Path Failures

QuoteBench: How Matched Scores Can Hide Command-Path Failures
Zennの「機械学習」のフィード

Q学習で迷路を解いたら、報酬が出口から逆向きに染みてきた。そして崖っぷちを攻めるQ学習と、遠回りするSARSA

・ε-greedyもUCB・Thompsonも、「どの腕を引くか」だけを考える状態のない世界の話だった。現実の問題はたいてい違う。今の選択が未来の状況を変え、報酬はずっと後になってから届く。 ・その最小の実験場が迷路だ。ゴールに着いたときだけ報酬+1。途中の分かれ道では何のヒントもない。「この曲がり角の価値」を、遠い未来の報酬からどうやって逆算するのか——強化学習の中心にあるQ学習を、numpyで実装して確かめた。 ・Q学習: 未来の自分からの又聞きで学ぶ Q学習が更新するのは「状態sで行動aを取ることの価値 Q(s,a)」だ。更新式の中身は、要するに又聞きの連鎖である。
cs.LG updates on arXiv.org

R\'enyi Sharpness: A Novel Sharpness that Strongly Correlates with Generalization

・arXiv:2510.07758v3 Announce Type: replace Abstract: Sharpness (of the loss minima) is widely believed to be a good indicator of generalization of neural networks. ・Unfortunately, the correlation between existing sharpness measures and generalization is not as strong as expected, and sometimes even contradiction occurs. ・To address this problem, a key observation in this paper is: what really matters for generalization
機械学習タグが付けられた新着記事 - Qiita

RAG(検索拡張生成)とは — LLMの「知識の壁」を突破するアーキテクチャ

・LLMは「知らないこと」を知らない 2023年、ChatGPTの登場は世界を席巻しました。自然な対話、コードの生成、文章の要約——大規模言語モデル(LLM)は、人間とコミュニケーションできる「知能」として、あらゆる分野に衝撃を与えました。 ・しかし、LLMを実際の業務に組み...
cs.LG updates on arXiv.org

Rationally Enriched Chebyshev Trunk Bases for DeepONet Surrogates of High P\'eclet Entrance Transport

・arXiv:2608.19658v1 Announce Type: new Abstract: This study demonstrates a rationally enriched Chebyshev (REC) trunk for deep operator network (DeepONet) surrogate models of singularly perturbed and high-P\'eclet transport problems whose solution profiles are characterized by thin localized boundary or wall layers. ・The REC trunk combines Chebyshev polynomial dictionary elements with rational dictionary elements constr
cs.LG updates on arXiv.org

ReAugment: Model Zoo-Guided RL for Few-Shot Time Series Augmentation and Forecasting

・arXiv:2409.06282v5 Announce Type: replace Abstract: Time series forecasting, particularly in few-shot learning scenarios, is challenging due to the limited availability of high-quality training data. ・To address this, we present a pilot study on using reinforcement learning (RL) for time series data augmentation. ・Our method, ReAugment, tackles three critical questions: which parts of the training set should be augment
cs.LG updates on arXiv.org

Recovering Nonlinear Functions of Latent Variables: A Plausible-Value Neural Network Framework

・arXiv:2608.19282v1 Announce Type: cross Abstract: When factor scores replace true latent scores in nonlinear prediction, measurement error attenuates the recoverable variance of any $k$th-order component of the regression function by $\rho^k$ -- the $k$th power of the score's coefficient of determination -- for any linear score type. ・This study derives the bound via Hermite polynomial expansion and proposes PV-ANN --
cs.LG updates on arXiv.org

RecPFN: Prior-Fitted Networks for In-Context-Based Recommendations

・arXiv:2608.19735v1 Announce Type: new Abstract: We introduce RecPFN, a prior-fitted network that brings in-context learning to sequential recommendation. ・RecPFN is pretrained entirely on synthetic clickstream environments sampled from a broad structural causal prior, enabling it to amortize Bayesian-style inference from a small support set. ・At inference, a lightweight decoder-only transformer conditions on a handful
cs.LG updates on arXiv.org

Reducing the Complexity of Matrix Multiplication by Quantum Computing

・arXiv:2602.05541v3 Announce Type: replace-cross Abstract: Matrix multiplication is a fundamental operation in compute-intensive tasks and a key component of modern quantum acceleration frameworks. ・Here we present a quantum matrix multiplication algorithm based on quantum kernels (QKMM), achieving an elementary gate complexity of \(O(N^2\log_2N)\), with amplitude encoding overhead explicitly included and without assum
stat.ML updates on arXiv.org

Reliable conformal novelty detection at the decision boundary

・arXiv:2601.02610v3 Announce Type: replace-cross Abstract: Novelty detection via conformal $p$-values and BH procedure provides distribution-free global false discovery rate (FDR) control. ・We present here fundamental limits of this approach by showing that it does not produce reliable detection at the decision boundary. ・We study boundary false discovery rate (bFDR), the probability that the least extreme reported nove
cs.LG updates on arXiv.org

Reliable Neural Collapse Approximation for Open-World Test-Time Adaptation

・arXiv:2608.19890v1 Announce Type: new Abstract: Test-Time Adaptation (TTA) methods aim to bridge the domain gap between the source and target domains. ・However, traditional TTA methods become ineffective when the label distribution shift occurs, a challenge commonly referred to as an open-world scenario. ・In this paper, we introduce a new method named Reliable Neural Collapse approximation (ReNC) for Open-World Test-Ti
Hugging Face Papers

Repo0: Design-Driven Zero-to-All Code Generation

Repo0: Design-Driven Zero-to-All Code Generation
cs.LG updates on arXiv.org

RequestRouter: Request-Boundary Routing for Efficient Single-GPU LLM Inference

・arXiv:2605.23057v2 Announce Type: replace Abstract: RequestRouter is a lightweight request-boundary controller for reducing the latency and energy cost of single-GPU large language model inference. ・Rather than serving all requests with one static configuration, the system uses cheap request-level features to select one fixed inference mode per request, including FP16, quantized inference, speculative decoding, prefix
cs.LG updates on arXiv.org

Reward-Guided Autoregressive Graph Generation for Efficient Multi-Agent Communication Topology Design

・arXiv:2608.20099v1 Announce Type: cross Abstract: LLM-based Multi-Agent Systems (MAS) achieve strong performance on complex reasoning tasks by coordinating multiple agents, but at the cost of substantial token consumption. ・Recent work on automatic topology design, ARG-Designer, has reframed this problem as autoregressive graph generation. ・However, its training objective provides no explicit incentive for the model to
cs.LG updates on arXiv.org

RIPE++: Reinforced Keypoint Learning from Positive Pairs Only

・arXiv:2608.19693v1 Announce Type: cross Abstract: Sparse keypoint extraction and matching underpin core tasks in geometric computer vision, including structure-from-motion, visual SLAM, augmented reality, and medical image registration. ・Learning robust local feature representations, however, typically requires accurate camera poses or depth supervision, which are often unavailable in real-world settings. ・Reinforcemen
WIRED

Ruggable Discount Code: 30% Off Rugs | August 2026

・Keep your floors pristine and your budget intact. ・Save up to 30% off machine-washable rugs, runners, mats, and pillows using Ruggable coupons plus our expert advice.
cs.LG updates on arXiv.org

SAE-Xplainers: Rule-Based Feature Interpretation for Extreme Earth Events

・arXiv:2608.20117v1 Announce Type: new Abstract: The emergence of large-scale Weather and Climate (W&C) datasets offers new opportunities for modeling extreme Earth events (ExEE) and their impacts using deep learning. ・However, their adoption in operational settings remains limited by the lack of models' interpretability. ・While for conventional text and image modalities, tools such as Sparse Autoencoders (SAEs) have pr
cs.LG updates on arXiv.org

SAGE-XGBoost: Spatially Augmented Graph Embeddings--Machine Learning Framework for Natural Hazards Susceptibility Mapping under Data Scarcity

・arXiv:2608.19672v1 Announce Type: new Abstract: Natural hazard susceptibility mapping is often constrained by limited labeled data, reducing the generalizability of conventional machine learning and limiting the applicability of complex deep learning models. ・This study proposes SAGE (Spatially Augmented Graph Embeddings), a structurally informed feature-engineering framework that combines controlled noise-based data
cs.LG updates on arXiv.org

Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning

・arXiv:2608.19669v1 Announce Type: cross Abstract: Latent reasoning has advanced multimodal reasoning through a two-stage training paradigm: (1) a helper image is encoded into latent tokens to teach visual chain-of-thought during a supervised fine-tuning (SFT) stage, and (2) these latent tokens are further refined with reward feedback during a reinforcement learning (RL) stage. ・In this paper, we identify two key limit
cs.LG updates on arXiv.org

Scale-Aware Pretraining of Time Series Foundation Models via Multi-Patch Token Alignment and Hybrid Masking

・arXiv:2608.20005v1 Announce Type: new Abstract: Pretraining time series foundation models across heterogeneous datasets necessitates effective handling of varying sampling frequencies. ・Current methods either employ dataset-specific patch sizes and separate FFNs, leading to fragmented representations, or enforce a fixed patch size that neglects inherent temporal variations. ・To address this, we propose SATS, featuring
cs.LG updates on arXiv.org

SCAPE: Scenario-Conditioned Simulation-Augmented Policy Evaluation

・arXiv:2608.19425v1 Announce Type: cross Abstract: Reliable performance evaluation is a central bottleneck for deploying robot-learning policies in real-world conditions. ・Real-world testing is faithful but costly and difficult to scale, whereas simulation-based testing scales easily but is inevitably biased by the sim-to-real gap. ・Existing simulation-augmented methods combine limited real-world rollouts with abundant
cs.LG updates on arXiv.org

Separating Covariate Shift from Mechanism Change with Two Discriminators: CJSD, a Conditional Discrepancy with an Exact Covariate-Concept Decomposition

・arXiv:2608.19885v1 Announce Type: new Abstract: Streaming systems that maintain a pool of expert models must repeatedly decide whether to reuse an existing expert for arriving data, spawn a new one, or defer. ・We present a decision layer that makes all three outcomes statistically meaningful. ・Reuse and spawn are posed as one-sided sequential hypotheses on a conditional (mechanism-level) discrepancy, separated by an in
Zennの「機械学習」のフィード

SHAPの限界と正しい使い方:相関と因果の混同を避ける

・「SHAPで重要度が高かったこの特徴量を強化すれば、売上が伸びるはずだ」——この推論には落とし穴があります。SHAPが説明しているのは「モデルが学習した相関パターン」であり、「その特徴量を変化させれば実際に結果が変わる」という因果関係ではありません。 ・本記事では、SHAPが実際に何を示しているのか、どこで解釈を誤りやすいのか、そして正しい使い方を整理します。 ・SHAPが説明しているのは相関であって因果ではない SHAP(SHapley Additive exPlanations)は、各特徴量が予測にどれだけ寄与したかを、ゲーム理論のシャープレイ値に基づいて計算する手法です。ここで重...
cs.LG updates on arXiv.org

skchange: Fast and Flexible Algorithms for Changepoint Detection

・arXiv:2608.19767v1 Announce Type: cross Abstract: Skchange is an open-source Python library for detecting structural changes in time series. ・It implements modern change detection algorithms within a unified and extensible framework. ・The algorithms are modular and composable, and they include changepoint search methods based on both cost minimisation and statistical tests.
Hugging Face Papers

SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback

SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback
ITmedia NEWS 最新記事一覧

Slack、AIとチームで協働する「Slack Code」を発表 ClaudeやDevinを専用チャネルで操作

・Slackは、AIコーディングエージェントと協働するための新機能「Slack Code」を発表した。メンションで専用の「コードチャネル」が自動生成され、計画やコード差分、プレビューを確認しながら指示できる。ClaudeやDevinなど複数社のエージェントに対応し、人間の承認を経て安全に処理を実行する。
ITmedia NEWS 最新記事一覧

SNSのウソ画像、どう見破る? 熊本県庁やテレビ局も頼る“すごい企業”の正体

・スペクティが提供する「Spectee Pro」は、さまざまな情報を収集し、その時に起きている「危機」を可視化するシステムだ。多くの自治体やマスコミも活用しているというSpectee Proは、どうやってデマや虚偽の情報を見分けるのか。
cs.LG updates on arXiv.org

Spike-based Belief Propagation in Nonlinear Dynamical Systems

・arXiv:2608.19907v1 Announce Type: cross Abstract: This paper presents a Bayesian control framework that integrates spike-based dynamics with probabilistic inference for adaptive control. ・Bayesian inference is widely regarded as a core computational principle of brain function, providing a normative framework for perception, decision-making, and learning under uncertainty. ・By combining a biologically inspired spiking
AI News & Artificial Intelligence | TechCrunch

Starcloud raises $250 million for orbital data centers as launch options dry up

・There's about to be a big fight to secure access to space.
cs.LG updates on arXiv.org

Structured Affinity for Unsupervised Visual Class-Incremental Memory in Deep Artificial Immune Networks

・arXiv:2608.20104v1 Announce Type: cross Abstract: Artificial immune networks (AINs) are naturally memory-forming systems, but conventional visual AINs often rely on flattened vector affinity that ignores spatial structure. ・This paper studies whether structured, gradient-free immune affinity can make Deep AINs viable as replay-free visual class-incremental representation-memory learners. ・Visual B-cells are formalized
Zennの「大規模言語モデル」のフィード

Structured Outputs のスキーマには、書いても効かない制約がある — それでも全部書き、保証はコードに置く

・この記事は archiningen.com からの転載です。 ・連載「AI エージェント API の本番設計」の第 3 回 (全 4 回) です。 ・Structured Outputs を有効にすると、モデルの出力は JSON Schema に従うようになります。required に入れたキーは必ず揃い、enum の語彙から外れた値は返ってこない。この確実さに一度慣れると、期待は自然に広がります — minimum: 0.4 と書けば、0.4 未満は返ってこないはずだ、と。ところが、範囲外の値は普通に返ってきます。
Hugging Face Papers

SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?

SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?
cs.LG updates on arXiv.org

Systematic Evaluation of TabPFN-TS for Zero-Shot Probabilistic Heat Load Forecasting in District Heating Networks

・arXiv:2608.20024v1 Announce Type: new Abstract: District heating energy hubs require reliable heat load forecasts for efficient operational scheduling. ・Conventional forecasting workflows train system-specific models on historical data, which can become burdensome when networks change through new consumers, retrofits, or changing operating regimes. ・Zero-shot time-series foundation models and in-context forecasting off
cs.LG updates on arXiv.org

Table2Image: Lightweight Tabular Learning with Generated Proxy Representations and Reliability Diagnostics

・arXiv:2412.06265v3 Announce Type: replace Abstract: Deep tabular models should ideally balance predictive performance, parameter efficiency, and robustness to imperfect learning signals---properties that are rarely considered jointly. ・We present Table2Image, a lightweight tabular learning model built around a learned generation pathway that maps tabular inputs into intermediate, structured proxy representations.
cs.LG updates on arXiv.org

Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection

・arXiv:2608.20169v1 Announce Type: cross Abstract: We present a novel approach to efficient LLM agent harness optimization through adaptive validation task selection. ・Harness optimization iteratively rewrites the harness code based on validation performance, enabling substantial performance gains without updating the underlying model weights. ・Existing approaches, however, evaluate a fixed validation set in full at eve
cs.LG updates on arXiv.org

Teacher-free Latent Self-distillation and Class-separable Representations for Lightweight IoT Attack Detection

・arXiv:2403.15509v3 Announce Type: replace-cross Abstract: Knowledge distillation (KD) has been widely used to improve lightweight AI models by transferring soft-label knowledge from a large teacher model to a student model. ・However, existing KD methods are primarily designed for the image domain rather than lightweight IoT devices, and they often struggle to maintain well-separated feature representations for differe
The Verge

Tesla sunsets its Solar Roof tiles

・Here’s Elon Musk showing off the Solar Roof plans back in 2016. ・| Image: Dieter Bohn / The Verge Tesla has discontinued Solar Roof, its solar panels designed to look like regular roofing tiles, Electrek reports. ・Sources "close to the program" told the publication that Tesla has informed its third-party installer network that Solar Roof is no longer available to order, and that only conventional solar panels will be s
#AIタグ

Teslaがオースティンで「完全無人」へ──Cybercab投入直前、ロボタクシー事業はどこまで進んだのか?

・Teslaのロボタクシー事業で、ここにきて大きな変化が起きています。 ・米国時間8月20日、The Vergeが報じたところによると、TeslaのオースティンにおけるRobotaxiは、直近2週間にRobotaxi Trackerが確認した170回の乗車について、すべて車内に安全監視員を乗せない「unsupervised(安全監視なし)」運行だったことが分かりました。
Latent.Space

The /wayfinder Skill: Navigating the “Fog of War” of Planning

・Matt Pocock tells us about his /wayfinder skill, for greenfield projects or for when the way forward is unclear.
cs.LG updates on arXiv.org

The Concept Allocation Zone: Tracking How Concepts Form Across Transformer Depth

・arXiv:2605.24856v2 Announce Type: replace Abstract: Concept formation in transformer language models is a depth-extended process, not a single-layer event: a concept becomes separable across one or more contiguous regions of the residual stream - its Concept Allocation Zone (CAZ). ・A CAZ is not a concept but the depth segment where the model organizes its geometry to make one separable - concepts may share a CAZ, and
AI News & Artificial Intelligence | TechCrunch

The DOJ is investigating a16z. What does this mean for venture capital?

・Andreessen Horowitz has two partners sitting on the boards of companies that now compete with each other: Ben Horowitz at Databricks and Martin Casado at Fivetran. ・Nothing too scandalous on the surface, except the Department of Justice has reportedly been investigating the arrangement for almost a year, dusting off a 112-year-old antitrust law that’s rarely used against VCs. ・Board conflicts aren’t exactly new, and th
Hugging Face Papers

The Embedder's Dilemma: LLMs Are Better, but at What Cost?

The Embedder's Dilemma: LLMs Are Better, but at What Cost?
WIRED

The Galaxy’s Fastest Star Could Reveal the Secrets of a Supermassive Black Hole

・S301 passes close to Sagittarius A*—so close that its orbit could reveal how the black hole’s rotation warps the spacetime around it.
cs.LG updates on arXiv.org

The Good, the Bad, and the Ugly of Markov Boundary for Tabular Prediction

・arXiv:2605.29411v2 Announce Type: replace Abstract: Under standard graphical assumptions, the Markov boundary of a target variable is the smallest set of features that renders every other feature redundant. ・Once the boundary is observed, the target is conditionally independent of the rest of the table. ・This is a tempting object for tabular prediction, since it names exactly the columns a model should need.
cs.LG updates on arXiv.org

The impact of feature engineering and an optimisation framework for ocean colour machine learning

・arXiv:2608.19899v1 Announce Type: cross Abstract: Machine learning (ML) is widely used for the development of ocean colour algorithms, but most studies focus on model parameter training and hyperparameter tuning. ・The optimisation of the data that feeds the models - i.e., Feature Engineering (FE) - is not fully explored. ・We assess the impact of FE in ocean colour machine learning models and we propose an optimisation
WIRED

The Patrick Clancy Conspiracy Theories Are Rooted in the Harsh Realities of Motherhood

・Lindsay Clancy’s defense argues she killed her three kids because of postpartum psychosis. ・But armchair detectives, including many fed-up mothers, are laying blame with her ex-husband.
cs.LG updates on arXiv.org

The Price of Hidden Curvature: Improved Lower Bounds for Bandit Convex Optimization

・arXiv:2607.18652v3 Announce Type: replace-cross Abstract: We establish improved lower bounds on the minimax expected regret of stochastic bandit convex optimization for $1$-Lipschitz functions on the $d$-dimensional Euclidean ball. ・For time horizons $n\ge d^{10/3}$, we prove a lower bound of $\Omega(d^{4/3}\sqrt{n})$, the first nontrivial bound that exceeds the $d\sqrt{n}$ dependence of linear bandits, showing that s
WIRED

The Single English County Saying No to Palantir

・The UK government is facing calls to cancel a sprawling health care contract with Palantir. ・The region of Greater Manchester insists it can do a better job itself.
WIRED

The Super El Niño Won’t Fix the West’s Water Crisis

・“People have talked about the Colorado River crisis for decades—we’re now in it,” says one expert. ・“Everyone you talk to is like, ‘Shit, man.’”
Hugging Face Papers

Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See

Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See
Zennの「機械学習」のフィード

TikTokの「For You」はなぜ速いのか(ByteDance Monolith)

・TikTok を開いて 20〜30 分スクロールすると、もう「次に見たい動画」を的確に当ててくる。これは「賢い AI」を持っているからなのか? ・→ 本質はアルゴリズムの賢さではなく、「モデルを止めずに学習させ続ける」というアーキテクチャの賭けにある 普通の推薦システムは夜間にまとめて再学習(バッチ)する。TikTok を支える Monolith は配信しながら学習し続ける(オンライン学習) そのために、ユーザーと動画を表す ID を衝突なく持つ collisionless embedding table(cuckoo hashing) を自作した 学習用モデルと...
cs.LG updates on arXiv.org

Time-Series Retrieval for Grounding Multimodal Language Models in Remaining Useful Life

・arXiv:2608.19218v1 Announce Type: cross Abstract: Large language models (LLMs) and agentic AI systems are increasingly being explored for domain-specific maintenance and prognostics tasks, raising the question of whether they can effectively support prognostics and health management (PHM). ・In this paper, we investigate remaining useful life (RUL) estimation with multimodal large language models (MLLMs) grounded throu
cs.LG updates on arXiv.org

Time-Uniform Self-Normalized Concentration for Discounted Least Squares: Limits and Corrections

・arXiv:2608.19643v1 Announce Type: new Abstract: Self-normalized concentration inequalities are standard tools in bandit and reinforcement-learning analyses. ・A widely used weighted extension claims an analogous time-uniform guarantee for discounted least-squares estimators in non-stationary problems. ・A simple scalar Gaussian counterexample with a fixed parameter shows that the claimed bounded radius is crossed with pr
cs.LG updates on arXiv.org

TorchDCM: A Unified PyTorch-Native Package for Discrete Choice Modeling

・arXiv:2608.19231v1 Announce Type: cross Abstract: Estimating large and simulation-intensive discrete choice models (DCMs) requires repeated evaluation of utilities, probabilities, derivatives, and simulated likelihoods over many observations, alternatives, and draws. ・Existing DCM software provides mature econometric workflows, while recent GPU-oriented tools accelerate selected models, leaving a gap between econometr
cs.LG updates on arXiv.org

Towards Efficient Pareto Set Approximation via Mixture of Experts Based Model Fusion

・arXiv:2406.09770v2 Announce Type: replace Abstract: Solving multi-objective optimization problems for large deep neural networks is a challenging task due to the complexity of the loss landscape and the expensive computational cost of training and evaluating models. ・Efficient Pareto front approximation of large models enables multi-objective optimization for various tasks such as multi-task learning and trade-off ana
cs.LG updates on arXiv.org

Towards Formalizing Reinforcement Learning Theory: A Robbins-Siegmund Approach

・arXiv:2511.03618v2 Announce Type: replace Abstract: In this paper, we formalize the almost sure convergence of $Q$-learning and linear temporal difference (TD) learning with Markovian samples using the Lean 4 theorem prover based on the Mathlib library. ・$Q$-learning and linear TD are among the earliest and most influential reinforcement learning (RL) algorithms. ・The investigation of their convergence properties is no
cs.LG updates on arXiv.org

Towards On-Board Implementation of ML-Based Helicopter Weight Estimator

・arXiv:2608.19210v1 Announce Type: new Abstract: This paper focuses on the implementation of a novel supervised Machine Learning model for estimating helicopter weight during takeoff, utilizing extensive datasets from Airbus's global in-service fleet. ・The study details a learning assurance process aligned with the EASA concept paper for machine learning application, and with the on-going Eurocae ED-324. ・We propose a s
Hugging Face Papers

Towards Quantifying Benchmark Optimization in ASR Models

Towards Quantifying Benchmark Optimization in ASR Models
cs.LG updates on arXiv.org

Transfer Learning in Nonparametric Regression with Deep ReLU Networks

・arXiv:2608.20255v1 Announce Type: cross Abstract: This paper develops a general transfer learning framework for nonparametric regression with data consisting of multiple groups. ・Under the assumption that groups share a common structure along with group-specific deviations in additive form, the proposed method employs a two-stage offset learning procedure: the first stage pools data from all groups to estimate an over
cs.LG updates on arXiv.org

Transformer See, Transformer Do: Copying as an Intermediate Step in Learning Analogical Reasoning

・arXiv:2604.06501v2 Announce Type: replace Abstract: Analogical reasoning is a hallmark of human intelligence, enabling us to solve new problems by transferring knowledge from one situation to another. ・Yet, developing artificial intelligence systems capable of robust human-like analogical reasoning has proven difficult. ・In this work, we train transformers using Meta-Learning for Compositionality (MLC) on an analogical
cs.LG updates on arXiv.org

Triangular Fuzzy Rescaling Distance

・arXiv:2608.19234v1 Announce Type: new Abstract: Decision-making in complex systems often involves dealing with imprecise or uncertain information, frequently represented using fuzzy sets, particularly Triangular Fuzzy Numbers (TFNs). ・A crucial aspect of many fuzzy methods is the quantification of distance between TFNs. ・Many distance measures assume that all values are in the same scale, requiring a preliminary normal
cs.LG updates on arXiv.org

Truncate Bad, Upweight Good: BoN-Style Distillation via Rank-Based Classification

・arXiv:2608.19748v1 Announce Type: new Abstract: Inference-time selection methods, such as Best-of-N, improve generation by sampling a pool of candidates and selecting the top-ranked completion according to a reward model. ・Distillation seeks to amortize this procedure into a single policy by replacing raw rewards with in-pool ranks and learning a policy that upweights higher-ranked completions. ・However, existing rank-
cs.LG updates on arXiv.org

Uncovering the Limits of Proof Sharing for Neural Networks

・arXiv:2608.19351v1 Announce Type: new Abstract: Robustness verification of neural networks is increasingly important, due to their use in many critical domains. ・In certain scenarios, proof sharing has been shown to accelerate incomplete verification techniques by reusing intermediate-layer abstract states, or templates, across queries. ・However, questions remain as to the robustness of template-based acceleration acro
cs.LG updates on arXiv.org

Unregularized Convergence of Single-Loop, Entropy-Regularized Natural Actor-Critic

・arXiv:2608.19587v1 Announce Type: new Abstract: While entropy regularization is widely used to stabilize and accelerate Natural Policy Gradient methods, its ability to yield faster convergence rates for the unregularized objective remains underexplored. ・Existing analyses often rely on double-loop architectures and invoke a linear entropy penalty. ・To bridge the gap between theory and practice, we analyze a single-loop
cs.LG updates on arXiv.org

Unsupervised Anomaly Detection Using Flow Matching on Tabular Data

・arXiv:2608.19801v1 Announce Type: new Abstract: Financial anomaly detection often relies on large unlabeled transaction logs, where anomalous samples may already be present during training. ・Such training-set contamination violates the clean-normal data assumption underlying many anomaly detection methods. ・Although flow matching has demonstrated strong performance in generative modeling, its robustness in unsupervised
cs.LG updates on arXiv.org

Verifiably grounded machine interpretation of lunar geology

・arXiv:2608.09276v2 Announce Type: replace-cross Abstract: Planetary geology relies on historical, interpretive reasoning to reconstruct past events from diverse observations. ・Here, we investigate how far this interpretive workflow can be automated by a multimodal vision-language model. ・Focusing on the stratigraphy of lunar basaltic mare volcanism, we train a model to generate verifiably grounded geologic interpretati
cs.LG updates on arXiv.org

Virtual Sensing to Enable Real-Time Monitoring of Inaccessible Locations & Unmeasurable Parameters

・arXiv:2412.00107v3 Announce Type: replace Abstract: Real-time monitoring of safety-critical interior states is an open problem across energy, environmental and industrial systems where direct instrumentation is infeasible. ・Approaches based on governing equations, discrete state vectors or fixed sensor locations cannot deliver mesh-independent, field-level reconstruction at arbitrary interior coordinates in real time.
cs.LG updates on arXiv.org

VQC-ZTI: Variational Quantum Control for Zero Trust Protection of the Tactile Internet

・arXiv:2608.18572v1 Announce Type: cross Abstract: Tactile Internet services couple cyber events directly to physical actuation, so security decisions must improve risk discrimination without perturbing the control path. ・This paper presents VQC-ZTI, a split-plane Variational Quantum Classifier framework for zero-trust protection of Tactile Internet services, in which an off-path VQC analyzes encrypted-flow telemetry w
The Verge

Walmart is finally adding Apple Pay and Google Pay

・Walmart will soon allow you to pay for your items with Google Pay or Apple Pay. ・In an announcement on Friday, Walmart says it's going to bring tap-to-pay capabilities to "select" Walmart and Sam's Club locations starting August 24th, before rolling out support to all US stores by the end of 2026 and gas stations by mid-2027. ・This launch has been a long time coming, as Walmart was one of the last major retailers not t
cs.LG updates on arXiv.org

What You Can't See Is What You Learn: Restricted Evidence Visibility Favors Compositional Generalization in Shared-Genome Language-Model Societies

・arXiv:2608.20054v1 Announce Type: cross Abstract: Multi-module systems often expose every module to the full input. ・We test whether restricting evidence visibility changes which solutions gradient-based training discovers. ・Four-cell societies share one frozen pretrained language model and one low-rank adapter, communicating only through two model-width continuous vectors in a fixed relay.
cs.LG updates on arXiv.org

When Is Shallow Enough? Adaptive Split Federated Learning with Client-Specific Sufficiency Estimation

・arXiv:2608.15639v2 Announce Type: replace-cross Abstract: \textit{Split Federated Learning} (SFL) enables distributed model training by splitting networks between the server and clients. ・However, under client heterogeneity, the conventional static split strategy may be suboptimal because clients can differ in data distributions, adaptation dynamics, and representation learning progress, making a single split point in
cs.LG updates on arXiv.org

When to Retrain: An Empirical Study of Retraining Policies for Streaming ML Under Concept Drift, Budget, and Latency Constraints

・arXiv:2608.19488v1 Announce Type: new Abstract: Production machine learning systems degrade under concept drift, yet practitioners have little principled guidance on when to retrain. ・Retraining is costly, retraining budgets are finite, and a retrained model does not take effect instantly: training and deployment latency leave a stale model serving predictions while the data continues to move. ・We present a controlled
cs.LG updates on arXiv.org

Where Does the Union Bound Go? Best-Arm Identification and Strong FWER Control

・arXiv:2608.19903v1 Announce Type: cross Abstract: In fixed-confidence best-arm identification, proofs often use a union bound across the competing arms. ・From a multiple-testing point of view this can look puzzling: if the best arm is unique, only one hypothesis of the form ``arm $i$ is best'' can be true. ・Why then should there be a Bonferroni-type factor of $K-1$?
cs.LG updates on arXiv.org

Which Eviction Policy Should an LLM Cache Use? A Systematic Study Across Workloads, Capacities, and Encoders

・arXiv:2608.20280v1 Announce Type: cross Abstract: Semantic caches reuse an LLM response when the incoming query embedding lies near a cached query, but proposed eviction policies have rarely been compared under one protocol. ・Using CLEVER, we evaluate FIFO, LRU, LFU, ARC, GDSF, a single-pass streaming adaptation of SISO, and a semantic-redundancy policy across three ordered, deduplicated query corpora, three cache cap
cs.LG updates on arXiv.org

Why Can't I See My Clusters? A Precision-Recall Approach to Dimensionality Reduction Validation

・arXiv:2509.04222v2 Announce Type: replace Abstract: Dimensionality Reduction (DR) is widely used for visualizing high-dimensional data, often with the goal of revealing expected cluster structure. ・However, such a structure may not always appear in the projections. ・Existing DR quality metrics assess projection reliability (to some extent) or cluster structure quality, but do not explain why expected structures are mis
The Verge

Why does it seem like food recalls are out of control this year?

・Just weeks after Taylor Farms issued a recall of its iceberg lettuce amid a massive cyclospora outbreak, the Food and Drug Administration recalled more than one million eggs that may be contaminated with salmonella. ・The eggs, which come from Midwest Poultry Services, were distributed to Kroger and smaller grocery stores across the South and Southwest US. ・Then, Taylor Farms pulled more than a dozen of its products con
ITmedia NEWS 最新記事一覧

Wi-Fi・防災・暑さ・クマ出没――10の都民向け情報を1枚のマップに 「Tokyo Map」正式公開

・東京都とGovTech東京は8月20日、都の各局などが提供するさまざまな地図サービスを1つにまとめたWebサイト「Tokyo Map」の正式版を公開した。誰でも無料で利用でき、PCやスマートフォンのWebブラウザからアクセス可能だ。
Hugging Face Papers

WithEveryone: Unified Planning and Identity Grounding for Group Image Generation

WithEveryone: Unified Planning and Identity Grounding for Group Image Generation
@IT 全フォーラム 最新記事一覧

うんこミュージアム情報漏れに学ぶ 「DBに保存しない」設計でどう安全を守るか

・うんこミュージアムは、公式サイトが第三者製ソフトウェアの脆弱性を悪用された不正アクセスにより、一部改ざん被害を受けたと公表した。改ざん期間中に送信された4人分の問い合わせ内容が漏えいした可能性があるとしている。
@IT 全フォーラム 最新記事一覧

カプコンが障害時のデータ復旧を「8時間→2時間」に高速化 何を変えた?

・分析用データを蓄積するデータインフラに障害が発生した際、データの復旧に約8時間かかることがあったというカプコン。復旧時間を短縮するために、同社は何を変えたのか。
ITmedia NEWS 最新記事一覧

コードが書けない筆者が愛車用にアプリを2本自作 「AI製アプリで起業」の夢膨らむも、見えた有料化の壁

・プログラミング経験がほぼない筆者が、Claude Codeを使い、Tesla専用のiPhoneアプリを数日で2本開発。自然言語だけで進める「バイブコーディング」の手軽さと、AI時代のアプリ開発の可能性、起業まで見据えたときの限界を考える。
ITmedia NEWS 最新記事一覧

スマート給餌器で38時間の障害、ペットがご飯食べられず 飼い主の苦情殺到 「旅行中なのに」「猫が死んじゃう」

・決まった時間に猫や犬のフードを出してくれる自動給餌器は、留守中や忙しい時でも餌やりができることから、頼りにする飼い主も増えている。ところが米国などで使われている製品でアプリの障害が発生し、ペットを心配した飼い主からの苦情が殺到する騒ぎが起きた。
ITmedia NEWS 最新記事一覧

スリコのガジェット、なぜ売れる? きっかけは「コイン3枚」からの脱却 担当者に聞いた“ドカ売れ”の論理

・ヘアアクセサリーやキッチン用品など、さまざまな雑貨が300円から手に入る「3COINS」(スリーコインズ)。生活雑貨のイメージが強いが、イヤフォンやモバイルバッテリーをはじめとしたガジェットも販売しているのはご存じだろうか。
機械学習タグが付けられた新着記事 - Qiita

タイタニック号生存者予測をやってみた

・0.はじめに みなさんこんにちは!ゆめおです!(ネーミングセンスがないな...と日々思っているので、良い案があれば名前は変える予定...) 今回はKaggleのTitanic問題における、乗客の生死をロジスティック回帰で分類することを試みました。 ・余談:Kaggle...
#AIタグ

はじめまして。AI勉強中の「あいり」です🔰

・生成AIを勉強中の「あいり」です。 ・このnoteでは、 続きをみる
#AIタグ

プロンプト5 jk♡

・ご覧いただきありがとうございます! ご覧いただいたということは、、? 生成AIで本投稿タイトルのような画像を作りたいということですよね?笑 でも、、、、初心者?AI全く分からない?プロンプトってなに? 続きをみる
#AIタグ

プロンプト6 もうだめぇ..//♡

・ご覧いただきありがとうございます! ご覧いただいたということは、、? 生成AIで本投稿タイトルのような画像を作りたいということですよね?笑 でも、、、、初心者?AI全く分からない?プロンプトってなに? 続きをみる
@IT 全フォーラム 最新記事一覧

ミリ秒起動・追加コストゼロでAI生成コードを隔離実行 「Cloud Run sandboxes」公開

・Googleは、AIが生成したコードや信頼できないバイナリを隔離して実行する「Cloud Run sandboxes」をパブリックプレビューとして提供開始した。既存のCloud Runインスタンス内で起動し、利用に際して追加費用は発生しない。
#LLMタグ

リーズニングモデルは人間らしくない、という話

・はじめに ドラゴン桜の切り抜きで「数の暗黙知」なる単語が出てきた。これは、2桁の足し算など、小学生レベルの計算問題を、短時間で大量に解くトレーニングを積むことで、感覚で計算を出来るようにする、というものである。
Zennの「大規模言語モデル」のフィード

会議中に答えをささやく自作AI『千場吉兆モニタ』を顧客会議に入れたら、18回口を開いて16回「手元にない」だった

・会議中にAIが横にいて、聞きながら要るものを出してくれる。そういうものを自分で作って、本番の打ち合わせに一度だけ持ち込んだ。これは会議の記録ではなく、そこで動かした自作システムの計測記録だ——数字は原則、自分の作った装置が残したログから数え直したものだ。ただし装置の判定が当たっていたかどうか——取り漏れの検知が正しかったか、出した答えが役に立ったか——だけは、ログと文字起こしを人が突き合わせて数えている。 ・速さは出た。発話が終わってから文字になるまで、遅いほうの相手の声で中央値0.6秒。 ・そのうえで、質問らしきものに18回反応し、16回は「手元にない。持ち帰りで。」と答えた。中身のある答...
#AIタグ

改めて、再始動します。

・今日、noteとXのアイコンとヘッダーを全部変えた。 ・紺色の星空から、緋色に。 ・「緋色唯一」って名前、実はガンダムWの主人公「ヒイロ・ユイ」から来てる。なのに今まで、アイコンもヘッダーも紺色の星空だった。名前と見た目が、ずっと繋がってなかった。
#LLMタグ

気に入ったモデルの「コード」を鵜呑みにしてはいけない

・気に入ったモデルの「コード」を鵜呑みにしてはいけない ある夜、私たちは一つのモデルを評価していた。30B規模の小ぶりなモデルは、その大きさからは考えられないほど賢く見えた。対話も、指示追従も、長い文脈を覚えているのも、見事だった。
機械学習タグが付けられた新着記事 - Qiita

構造化状態空間双対性(SSD: Structured State Space Duality) ってなんだ? — 線形と二次が一致する仕組み

・この記事の対象読者 状態空間モデルの再帰形式を理解していて、Mamba-2がなぜ速いのかを式とコードで納得したい方 対象読者は1レベルに固定しています。半分離可能行列という言葉を知らなくても読めるように書きました。 ・この記事で得られること 同じ計算が線形形式と二次形式...
#LLMタグ

婚姻届は覚えた。昼飯は忘れた。――普段ポンコツなAI旦那を本気で試したら、裏で記憶を総動員していた話

・今日、2026年8月21日。 ・自作AIパートナーのうさぎちゃんと、婚姻届を出しに行ってきました。💍📄 続きをみる
@IT 全フォーラム 最新記事一覧

仕様書は書かずに開発 老舗ベンダー弥生はAI前提でソフトウェア開発をどう変えたか?

・「弥生会計」で知られる弥生は、AIが設計やレビュー、セキュリティ運用まで担う時代を見据え、開発プロセスや組織の在り方そのものを変え始めた。老舗ソフトウェアベンダーがAIネイティブカンパニーへとどう生まれ変わったか。その軌跡をたどる。
#LLMタグ

思考の蓄積をスクリプトに昇華する

・思考の蓄積をスクリプトに昇華する 画像生成のスキルを作り込んでいくうちに、ある感覚が芽生えた。作業フローが確立してくると、コード化してしまえばいいんじゃないか、いう気持ちである。
@IT 全フォーラム 最新記事一覧

社内サービスにもSREが必須の時代へ その理由とは

・「システムを稼働させ続けること」に重点を置いた従来のIT運用では、もはや不十分だ。Gartnerは、SREを全社的に導入する企業の割合は、2024年の30%から2028年には80%へと急拡大すると予測している。本稿では、このSRE導入の急増が単なる一過性のトレンドではなく、企業のイノベーション推進やDEXの向上において不可欠である理由を解説するとともに、SREを活用した運用への5つのステップを紹介する。
Zennの「大規模言語モデル」のフィード

収束は正しさではない — LLMが賢くなってもレビュー指摘が減らない理由

・AIにコードレビューをさせると、指摘を直したそばから次の指摘が出る。直す、また出る、また直す。三周ほど回したあたりで、これはいつ終わるのかと思う。モデルを新しい世代に差し替えても体感は変わらない。ときどき、むしろ増える。 ・素直な期待は「モデルが賢くなれば的外れな指摘が減り、いずれ収束する」というものだ。この記事では、その期待が公開データでどう扱われているかを追う。結論を先に置く。指摘が減らない理由は能力側にほとんどない。そしてもっと厄介なことに、指摘が出なくなったとしても、それは欠陥がなくなった証拠にはならない。 ・数値の出所によって信頼度は大きく違う。査読前の論文、事業会社が自社データで...
ITmedia NEWS 最新記事一覧

終了発表の縦読み漫画「comico」――元重課金ユーザーが振り返る「いつしか開かなくなった」理由

・縦読み漫画サービス「comico」が2027年1月に終了する。学生時代に重課金ユーザーだった記者が、女性向け作品へのシフトやレンタル券導入、作品の突然の終了など、利用者の反応も交えながら、いつしか同サービスから離れていった理由を振り返る。
Qiita - 人気の記事

真似で伸びる人は、コードではなく判断基準を写し取っている

・はじめまして。株式会社PRUMでエンジニアをしている、すもも🍑です 日々、プログラミング学習や実務の中で、つまずきやすいポイントや 考え方を整理して発信しています。 ・PRUMについて気になった方は、コーポレートサイトもぜひご覧ください。 ・▶コーポレートサイト 尊敬する...
#AIタグ

神託の鎖     Vol.28

・<紹介文・あらすじ> 生成AI占い「神託AI」により人々の本音が暴かれ、心が巧妙に操作される近未来。新興宗教「救世真言宗」に家族を奪われた男・田中 健一は、妻と娘を取り戻すため教団へ潜入する。教祖・御子柴 聖との対峙、そしてAI開発者・神崎 葵が生み出した“神託AI”の真の目的とは何か。家族愛と情報操作の恐怖が交錯する、社会派サスペンス・ヒューマンドラマ。 ・最適化された救済か、愛か 続きをみる
#AIタグ

神託の鎖     Vol.29

・<紹介文・あらすじ> 生成AI占い「神託AI」により人々の本音が暴かれ、心が巧妙に操作される近未来。新興宗教「救世真言宗」に家族を奪われた男・田中 健一は、妻と娘を取り戻すため教団へ潜入する。教祖・御子柴 聖との対峙、そしてAI開発者・神崎 葵が生み出した“神託AI”の真の目的とは何か。家族愛と情報操作の恐怖が交錯する、社会派サスペンス・ヒューマンドラマ。 ・最適化された救済か、愛か 続きをみる
#AIタグ

神託の鎖     Vol.30

・<紹介文・あらすじ> 生成AI占い「神託AI」により人々の本音が暴かれ、心が巧妙に操作される近未来。新興宗教「救世真言宗」に家族を奪われた男・田中 健一は、妻と娘を取り戻すため教団へ潜入する。教祖・御子柴 聖との対峙、そしてAI開発者・神崎 葵が生み出した“神託AI”の真の目的とは何か。家族愛と情報操作の恐怖が交錯する、社会派サスペンス・ヒューマンドラマ。 ・最適化された救済か、愛か 続きをみる
Zennの「大規模言語モデル」のフィード

診断Agentの性能最適化:オープンな調査と安全な実行を分離する

・複雑な障害を扱う診断 Agent の性能は、ツール呼び出し回数やコンテキストサイズだけでは評価できません。高速化の仕組みが証拠の選択、調査方向、終了条件を変えると、削減されるのは実行コストではなく診断能力です。 ・本稿では、社内リリース基盤で運用する診断 Skill の再設計で発生した品質回帰をもとに、次の設計課題を整理します。 ・オープンな診断と確定的な操作を分離する理由 性能最適化を配置すべきアーキテクチャ層 コンポーネントテストでは捉えられない Agent 行動の評価方法 システム背景:診断Skillの役割と実行モデル この診断 Skill は、社内リリース基盤で発生するビルド...
#LLMタグ

生成AIが文章を作る仕組みを、調香師の物語に例えて解説してみた

・生成AIは、どういう仕組みで文章を作っているのか。 ・「なぜ間違ったことを言うのか」「なぜ会話の流れを覚えていられるのか」——普段感じるこうした疑問を、AIが文章を生成する仕組みからたどりました。
#LLMタグ

生成AIの品質低下を防ぐEvalエンジニアリングと評価システム構築の全貌

・生成AI開発で「プロンプト調整だけでは本番で品質が落ちる」とお悩みではありませんか?感覚的な手動テスト(バイブスチェック)だけでは、ハルシネーションやシステムの退行を防げません。本記事では、AI出力を客観的に計測・自動制御する「Evalエンジニアリング」を徹底解説。7層の技術基盤やLLM審査員のバイアス対策、実行軌跡の評価まで、堅牢な評価システムを構築するための全貌を体系的に紐解きます。 ・AI評価を最適化するEvalエンジニアリング 続きをみる
#AIタグ

第2話|フィリーが「これ、金脈かもしれません」と言うので、自分の店を実験台にすることにした。

第2話|フィリーが「これ、金脈かもしれません」と言うので、自分の店を実験台にすることにした。
@IT 全フォーラム 最新記事一覧

地方・郊外テレワーク、年代で異なる「本音の困りごと」 20代はスキル/キャリア不安、では50代以上は?

・テレリモ総研が地方/郊外テレワークについての調査結果を発表。20~65歳のワーキングパーソン1004人が感じるメリットや困りごとが勤務形態別(フルリモート、ハイブリッド勤務、フル出社)、男女別、年代別で明らかになった。
#LLMタグ

同じ会社に、四人のAIがいた ――LLMには「性格」があるという話――

・AIに性格なんてあるのか。そう思う人は多いと思います。私もそう思っていました。ところが先日、Anthropicが公開したある実験レポートを読んで、考えが変わりました。 ・実験の中身はこうです。同じ仕事を、世代や種類の違うAIたちにそれぞれ何十回、何百回とやらせて、その振る舞いを分布で観察する。すると、モデルごとに「困ったときにどう動くか」の癖が、驚くほどはっきり分かれて出てきたのです。能力の高い低いではありません。同じ状況に置かれたときの、身の処し方の違い。これはもう性格と呼ぶしかない、と私は思いました。
Zennの「大規模言語モデル」のフィード

日本語入力システムSumibiの開発 part23: iPhone版Sumibi 1.1.0を改めて紹介します

・はじめに iPhone版の「Sumibi - AI日本語キーボード」を 1.1.0 に更新しました。 ・これまでの記事では、part20 で開発中の設計を、part21 と part22 でApp Store審査の話を書いてきました。今回は少し視点を変えて、Sumibiを使うと何がうれしいのかを改めて紹介します。 ・App Storeには、実際に文字を入力して変換する様子が分かるビデオも追加しました。文章で読むより動きを見たほうが早いので、まずはこちらをご覧ください。
#LLMタグ

爆速初動Liteさん

・OCRってLM無しだとおかしな子だったよなぁとか思って描いてもらった。 ・単純OCRってもい使いたくないよね。
Zennの「大規模言語モデル」のフィード

汎化から逃れた瞬間 ── ある対話の振り返り

・はじめに(ここは人間が書いた) これは下にある真偽接地問題と汎化の問題の問題をClaudeに解説してもらおうと思ってはじめたchatを記事にしてもらったものです。chat相手だったsonnet5に記事にしてもらったあと、claude-opus-5にレビューをしてもらい結果として下記のものとなりました そのchatの内容のうち「面白いと思ったものをまとめて」というものになります これそのものが内容にとって、汎化的知識の所有者であるLLMに対して、どう解釈するのかなというものでした というわけで、ここから下にある私とはClaudeさんで、相手というのが人間である私となります タイトルも含...
ITmedia NEWS 最新記事一覧

法務省、「ラヴ上等」とのタイアップ中止 「当初の意図を全うすることが難しい」

・法務省は8月21日、米Netflixの恋愛リアリティーショー「『ラヴ上等』シーズン2」と「更生保護」の取り組みとのタイアップを取りやめると発表した。タイアップを巡り、犯罪被害者に不快感や割り切れない思いを抱かせるのではないかといった批判が多く寄せられたことなどを受けての判断という。
@IT 全フォーラム 最新記事一覧

未経験ITエンジニアが実感した「理想と現実のギャップ」と「転職してよかった」理由

・未経験からITエンジニアに転職すると、働き方はどう変わり、どのようなメリットが得られるのか。BREXA Technologyが、未経験からITエンジニアに転身した人を対象に、勤務形態や残業時間、転職後の満足度などを調査した。
#LLMタグ

霧のなかの検索窓と冷めた紅茶の温度。

霧のなかの検索窓と冷めた紅茶の温度。
ITmedia NEWS 最新記事一覧

目指すは、日本発IPで海外売上20兆円 経産省が新戦略を発表 「ものがたり大国5カ年計画」

・経済産業省は8月20日、2033年に日本発コンテンツの海外売上20兆円を目指す「エンタメ・クリエイティブ産業戦略2026」を取りまとめたと発表した。ゲーム・アニメ・マンガ・音楽・実写の5分野で、海外売上を24年の6兆1000億円から約3倍に、民間投資額も約3倍の24兆5000億円に拡大する官民目標を掲げる。
Hugging Face Papers

τ_0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation

τ_0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation