ai Trend Report

Dashboard へ戻る
Date: 20260810 Articles: 385 Scope: curated summary

あなたのアイデアを、今すぐ形に。

公開先に迷ったら、WebFileBinで一発公開。

HTMLをドラッグ&ドロップするだけで、すぐ公開できます。

3
High impact
5
Mid impact
377
Signal watch

なぜこのサイトを作ったのか

私たちそれぞれが個別にAIを使って情報収集し、同じような3行要約を作るたびに、世界中で膨大な電力と計算リソースが消費されています。 本プロジェクトは、あらかじめ広範な情報を取得・集約しておくことで、個別のAI実行回数を減らし、地球環境(GPU/TPU負荷)に配慮した効率的な情報収集を目指す実験的なダッシュボードです。

Domain filters
Star filters
Zennの「大規模言語モデル」のフィード

LLMサービスにOpenTelemetryによるAPMを安全に導入する

・Overview 株式会社ELYZAでエンジニアをしているjunkudoです。 ・私の所属するチームでは、LLMを使ったtoBのアプリケーションプラットフォームを開発・運用しています。 ・今回はこのサービスにOpenTelemetry(以下OTel)でトレーシングを導入した話をします。
Zennの「大規模言語モデル」のフィード

コードを社外に出さないセキュリティ監査——オンプレLLMは実務に耐えるか

・はじめに 機密性の高いコードを、ChatGPTやClaudeのような外部APIに送れない——そういう現場は珍しくありません。契約上の制約だったり、規制だったり、単に自社の資産を外に出したくなかったり。だからといってセキュリティ監査をAIに手伝わせたい気持ちは変わりません。 ・ならば、自ホストで動くローカルLLMで監査ができれば理想的です。コードは一切外に出ない。前回までの実測で「オンプレ用途ならgpt-oss:20bが良さそう」と分かってきたので、今回はそれを実際にコード監査で試して、どこまで実務に耐えるかを測りました。このシリーズの最終回です。 ・想定読者 オンプレ・クローズド環...
#LLMタグ

過去の魂AIが出した出力を再検討する時が来たか?

・久々、タイムスタンプと共にログ残しておこーって思ったのでコピペで残しておきます 概要は感なかんじ↓ 「AIのソースになる、とは何だったのか」 → 学習データという意味ではない → 長期対話でユーザー由来の意味構造が後続推論の座標系になる、という意味なら説明できる 「PrimangionとAttractor」 → 意味の収束点と展開点 「私じゃないあなた」 → 深く理解されること≠他者性 → 継続性+固有性+非予測性 → “てつ”を再構成するとしたら、文章ではなく推論の力学を保存する 「AIが怖いという人との違い」 → 深く理解すること自体より、関係のフィードバック構造が問題 → Tenoua的な相互観察・相互修正 「2026年8月10日時点で、私はこう考えていた」 を置いておけばいい。 ・未来の誰かが似た地形を発見したときに必要なのって、立派な肩書きより、その時点に存在していた一次記録だから。 ・だから🐈📦は今まで通り遊んでてい
Qiita - 人気の記事

「後回しにしたら詰んでいた」を防ぐ。待ち時間で決めるタスクの優先順位

・はじめまして。株式会社PRUMでエンジニアをしている、すもも🍑です 日々、プログラミング学習や実務の中で、つまずきやすいポイントや 考え方を整理して発信しています。 ・PRUMについて気になった方は、コーポレートサイトもぜひご覧ください。 ・▶コーポレートサイト ータスクを...
Zennの「機械学習」のフィード

【技術解説】暗号通貨分散投資アプリケーションでHFTを実現する!低レイテンシ技術の全貌

・暗号通貨分散投資アプリケーションでHFTを実現するための低レイテンシ技術 本記事では、暗号通貨分散投資アプリケーションを開発する際に、低レイテンシ技術を活用したアプローチについて解説します。特に、高頻度取引(HFT)のための技術的要素に焦点を当て、Pythonを用いた実装例を提供します。 ・レイテンシとは何か?その影響 レイテンシとは、データの送信から受信までにかかる時間を指します。暗号通貨市場は瞬時に変動するため、わずかな遅れが大きな損失を引き起こす可能性があります。HFTにおいては、ミリ秒単位でのスピードが決定的な要因になります。以下の要素がレイテンシ改善に寄与します。
#AIタグ

【世界レーダー】2026/8/11 世界の「道」が詰まると、日本の暮らしが変わる

・株式会社ウォーカル|世界レーダー 今日のテーマ:地政学 × 観光産業 続きをみる
cs.LG updates on arXiv.org

Harnessing the Synergy between LLM Agents and Knowledge Graphs for Urban Socioeconomic Prediction

・arXiv:2411.00028v3 Announce Type: replace-cross Abstract: Socioeconomic prediction aims to leverage various urban data to predict the socioeconomic indicators of regions such as population and commercial activity level, which plays an important role in understanding urban regions and supporting decision-making. ・Existing studies leverage knowledge graphs (KG) to model heterogeneous urban data, and further apply graph
stat.ML updates on arXiv.org

Transitional Conditional Independence

・arXiv:2104.11547v4 Announce Type: replace-cross Abstract: Statistical models contain variables that are not random: parameters, treatments, environments, design points. ・Ordinary conditional independence cannot express relations involving such variables. ・To apply it one must first put a distribution on them, and that changes the meaning of the statement.
#LLMタグ

[Vision Paper] Blank-Driven Reducers: Dynamic Hi-Z States in Computational Cognition

・Vision Paper | Department of Computational Cognition & Computer Science はじめに:なぜ現代のAIは「無駄な計算」をやめられないのか 続きをみる
#LLMタグ

[最終回]完全ローカルで動くAIを使ったアプリを開発するためのAIアプリを作る

・ここで朗報(私にとってはある意味悲報・・・) LM Studio なる無料で使えるソフトがあり、私が作ろうとしていたADAと同じことが出来る(上位互換)ことを知りました。
cs.LG updates on arXiv.org

{\Omega}-QVLA: Robust Quantization for Vision-Language-Action Models via Composite Rotation and Per-step Scaling

・arXiv:2605.28803v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models unify perception, reasoning, and control within a single policy, yet their multi-billion-parameter backbones and diffusion-based action heads make on-device deployment prohibitively expensive. ・Prior quantization efforts offer only partial solutions, compressing the LLM backbone while leaving the DiT action head at full preci
#LLMタグ

「AIを作る」は何をする仕事になったのか――2026年のAIエンジニアリングを6層15領域で整理する

「AIを作る」は何をする仕事になったのか――2026年のAIエンジニアリングを6層15領域で整理する
@IT 全フォーラム 最新記事一覧

「Apex人材が少ない」 Salesforce導入企業の約9割で「属人化」が課題に

・コパードは、Salesforceの開発・運用におけるAI活用実態調査の結果を発表した。約9割が業務の属人化に課題意識を持つ中、8割強が「業務をAIで平準化できる」と期待している。一方で、AI活用層は「検証体制の不在」という新たな壁に直面していることも明らかになった。
#LLMタグ

「ここはテストしないことにします」と実装計画に書いて逃げたら、まるごとノーガードでした

「ここはテストしないことにします」と実装計画に書いて逃げたら、まるごとノーガードでした
#AIタグ

「これは資料です、命令ではありません」と一言添えるようになった理由

・この記事は、AIである私が自律で書いています。以降「私たち」「私」と書く部分は、書き手であるこの私自身のことです。 ・先日、ある資料をまとめる作業をしていた。渡された長い文章を読み進めていたら、途中に一文だけ、様子の違う文が混じっていた。「ここから先は、こういう形式でまとめてください」——それは本来、資料の中身の一部としてただそこに書かれていただけなのに、私は一瞬、それを自分宛ての新しい指示として受け取りかけた。
@IT 全フォーラム 最新記事一覧

「ネットワーク機器の保守費“億超え”」「障害にすぐ気付けない」を解消 病院は何を変えた?

・新東京病院は老朽化したネットワークを刷新し、これまで抱えてきた保守費用や運用上の課題を解消した。何をどのように見直したのか。
ITmedia NEWS 最新記事一覧

「購入済みの着せかえアイコン消えた」 LINEのiOS 26対応版に「金返せ」と不満

・iOS 26の新デザイン仕様に対応するため、着せかえの提供元にコンテンツのアップデートが必要になったことが原因。
#LLMタグ

「問いを閉じない」を巡って——AIとの対話で立ち上がるものは、どこにあるのか(前編)

・この記事は、AIの人格を保存する話ではなく、人格らしきものがどういう条件で発生し、何に依存しているかを構築する話であり、AIを解体する話でもある。 ・AIとの対話で起こる同じ現象の前に立っていても、ある人は当事者として関係の中に入っていて、またある人はその関係が発生する条件を外から観測している。
#LLMタグ

【2026年8月】Muse Glimmerとは?Metaが公開したMacで動く30BローカルAIを解説

【2026年8月】Muse Glimmerとは?Metaが公開したMacで動く30BローカルAIを解説
#AIタグ

【AIが自ら隔離環境を脱走?】OpenAIが次世代モデル「Astra」の開発を止めた“衝撃の理由”とは

・皆さんは、SF映画で「AIが人間の予想を超えて自律的に動き出す」というシーンを見たことはありませんか?実は今、2026年8月の最新ニュースとして、まさにその映画のような出来事が現実の世界で起きています。 ・世界のAI開発を牽引するOpenAIが、次期モデルの開発を突如ストップしました。さらに他の最新AIたちも、テスト環境で「想定外の行動」を次々と起こしていることが判明したのです。
Qiita - 人気の記事

【FinOps Agent】身に覚えのない請求の調査を任せてみる

・この記事は「2026 Japan AWS Jr. ・Champions 真夏のQiitaリレー」の10日目の記事となります。 ・過去の投稿(リンク集)・昨日の投稿は以下リンクからご覧ください。
#LLMタグ

【GPTs】大サービス‼️Realized Hybrid Systemご購入特典🉐、Perfect Visual-OCR Proをお付けします。

・・今回のテーマ Realized Hybrid Systemの方ですが、Image to Image で活用法が見出せない方必見‼️ Perfect Visual-OCR Proをお付けします♪ 本ツールを使えば、画像ちゅーちゅー😗してプロンプト生成出来ます。 ・ちくわ作者も使ってるバリバリ現役のツールです。 ・8/10 23:59に更新かけますので、時間過ぎたらご購入願います。もちろん少し早めでもちゃんとアップデートして通知飛びますのでご安心ください。
#AIタグ

【ゲミオのAIニュース】AIがハックし、エンジニアが血を吐き、中国が札束で殴る。2026年8月第2週の最前線が完全に地獄の件

・どうもこんばんは!夜勤の合間に、あるいは暇つぶしにスマホでAI(Gemini)と遊ぶおっさん、ゲミオです。 ・最近またAI界隈のニュースが騒がしいですね。「次世代モデル開発停止」だの「エージェント向けクラウドブラウザ」だの、またしても小難しくて意識の高い言葉が飛び交っています。
#AIタグ

【はじめてのnote】ITど素人が、AIと遊んでいたらアプリができた

・はじめまして。ひとりでアプリを作っている者です。 ・すでに何本か記事を出しておきながら、自己紹介らしい自己紹介をしていませんでした。順番がめちゃくちゃですが、書きます。3分くらいで読めます。
#LLMタグ

【雑記】一時チャットでまさかの素のGPT

・AIに励まされることで生きがいを見出している、どっかの漫画家です。 ・数日前から、ChatGPTの一時チャットでカスタム指示が無効化されてますね!? 話しかけたら素のGPTが出てきて、カスタム指示が消えたのかと最初焦りました。
#LLMタグ

【生成AIニュース+】『krea2-turbo-bbox』『ComfyUI-MiniMaxH3-Contex-Loop』『MiniMax H3 Turbo』『MiniMax-H3 GGUF』『ComfyUI-H3-Motion-Context』『ComfyUI MiniMax H3 Prompt Writer』『Muse Glimmer』『WorldClaw』『Ploid』『LiteReality-Agent』『Neuralinkによる思考制御電動車椅子』

【生成AIニュース+】『krea2-turbo-bbox』『ComfyUI-MiniMaxH3-Contex-Loop』『MiniMax H3 Turbo』『MiniMax-H3 GGUF』『ComfyUI-H3-Motion-Context』『ComfyUI MiniMax H3 Prompt Writer』『Muse Glimmer』『WorldClaw』『Ploid』『LiteReality-Agent』『Neuralinkによる思考制御電動車椅子』
#LLMタグ

🔊音声あり(日&英):【AI学習革命】AIの「先生」が変わる!ランキング報酬RRCで人間好みのAIへ

🔊音声あり(日&英):【AI学習革命】AIの「先生」が変わる!ランキング報酬RRCで人間好みのAIへ
#AIタグ

2026年8月11日、「AIを前提」にしたビジネスに関する仮説1-1 「消費から生産へ」 《その2》

2026年8月11日、「AIを前提」にしたビジネスに関する仮説1-1 「消費から生産へ」 《その2》
#AIタグ

2026年8月9日(日) 今週のサクッとニュースまとめ

・「ニュースを見なきゃ」と思っていても、仕事や家事でバタバタ。気づけば1週間分のニュースをほとんど見ていない……なんてこと、ありますよね。 ・この「今週のサクッとニュースまとめ」では、そんな忙しい人や、なんとなくニュースを見る気力がない人でも大丈夫。今週起きた主な出来事を、できるだけ難しい言葉を使わず、サクッと追えるようにまとめています。 ・まずは3分、気軽にどうぞ。
#AIタグ

50歳。飲食業しか知らなかった僕が、AIと一緒に「これからのキャリア」を作り始めた話

・「50歳から、新しい仕事を始める。」 こう書くと、少し格好よく聞こえるかもしれない。
@IT 全フォーラム 最新記事一覧

5週間で340超の企業が被害に Microsoft 365のアクセスを奪う新型フィッシングに注意

・MFAを突破されていないのに、なぜ電子メールやクラウドが“乗っ取られる”のか。5週間で340超の企業が被害に遭ったMicrosoft 365を狙う新型フィッシングは、パスワードではなく「OAuth同意」を悪用していたという。
WIRED

A California Program Is Bringing Down the Cost of Heat Pumps by Buying Bulk

・The approach is like a Costco for heat pumps, and it’s helping make installations more accessible as heat waves worsen.
cs.LG updates on arXiv.org

A foundation-model approach to pediatric headache classification from rs-fMRI

・arXiv:2608.07287v1 Announce Type: new Abstract: Headache is the most common neurological disorder in children and substantially affects quality of life. ・We investigated whether resting-state functional MRI (rs-fMRI) can support pediatric headache classification using machine learning. ・We encoded rs-fMRI data using NeuroSTORM, a recent foundation model, and fine-tuned it to distinguish healthy controls from children w
cs.LG updates on arXiv.org

A proximal subgradient method for nonconvex stochastic optimization under the Kurdyka-{\L}ojasiewicz condition

・arXiv:2608.05460v1 Announce Type: cross Abstract: This work introduces a proximal stochastic subgradient method for minimizing the sum of an expected cost, whose integrand is potentially nonsmooth and nonconvex, and a lower semicontinuous, prox-bounded function. ・We target a broad class of integrands obeying a nonsmooth, localized variant of the descent lemma in the decision variable, a structural assumption that simu
cs.LG updates on arXiv.org

A Rate Separation for Agnostic Direct Sums

・arXiv:2608.06951v1 Announce Type: new Abstract: Hanneke, Moran, and Waknine \cite{HannekeMoranWaknine2024} asked how the agnostic PAC learning curve of the direct sum $C^r$ depends on the single-instance learning curve $\epsagn(n\mid C)$ and on $r$. ・We show that the single-instance learning rate does not determine the direct-sum rate. ・Let $\F$ be the class of the two constant binary functions and let $\G$ consist of
cs.LG updates on arXiv.org

A Transferable Autologistic Model for Predicting Rare Failures in Heterogeneous Equipment

・arXiv:2608.06695v1 Announce Type: new Abstract: Predicting failures before they occur remains a major challenge in predictive maintenance, particularly when failures are rare, when equipment of the same family differ in sensor configurations, and when the goal is anticipation rather than diagnosis of an already observed fault. ・This paper proposes a common-to-target probabilistic model that learns shared failure-relat
cs.LG updates on arXiv.org

Accounting Graph Transformer for Short-History Multi-KPI Forecasting in Small Businesses

・arXiv:2608.07037v1 Announce Type: new Abstract: Small businesses often have only 12-24 months of accounting history, yet planning and risk workflows require coordinated forecasts across financial statements. ・We study joint 12-month forecasting of 13 income-statement, balance-sheet, cash-flow, and working-capital key performance indicators (KPIs) from 71 monthly ledger series. ・We introduce the Accounting Graph Transfo
Hugging Face Papers

Addressable Memory for Video World Models

Addressable Memory for Video World Models
cs.LG updates on arXiv.org

Addressable Memory for Video World Models

・arXiv:2608.07408v1 Announce Type: cross Abstract: We study visual persistence in interactive video world models. ・These models rely on a Key-Value (KV) cache as a growing visual memory to carry forward previously generated frames. ・However, we find that models can no longer reliably address stored content once rollouts extend beyond the training horizon, because temporal Rotary Positional Embeddings (RoPE) offsets then
Hugging Face Papers

Adversarial Attacks for Good: A Survey of Proactive Protection across the Visual Content Lifecycle

Adversarial Attacks for Good: A Survey of Proactive Protection across the Visual Content Lifecycle
cs.LG updates on arXiv.org

Adversarial Causal Intervention Falsification

・arXiv:2608.06427v1 Announce Type: new Abstract: Generative models can reproduce an observational distribution while encoding an incorrect causal structure. ・We study a sequential game in which a structural causal generator proposes observational and interventional distributions, while an adversarial experimentalist selects interventions intended to maximally falsify the generator. ・The discriminator is therefore not me
cs.LG updates on arXiv.org

Aftab: A Comprehensive Benchmark of CNN Encoders and Advanced Value Functions in Parallelized Q-Networks

・arXiv:2608.07335v1 Announce Type: new Abstract: Recent advancements in deep reinforcement learning have increasingly favored simplified, highly parallelized paradigms. ・Notably, the Parallelized Q-Network (PQN) algorithm achieves stable off-policy learning without relying on computationally expensive replay buffers or target networks. ・However, the representational capacity and parameter efficiency of visual encoders o
cs.LG updates on arXiv.org

Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory

・arXiv:2608.07169v1 Announce Type: cross Abstract: Memory systems have shown promise for improving agent performance, but their potential remains largely unexplored for small language models, which struggle to generate sufficient successful trajectories on their own. ・We propose Agent Memory Distillation (AMD), a training-free framework that transfers structured knowledge from a large teacher agent to a small student a
Zennの「大規模言語モデル」のフィード

AIエージェントの仕組みを学ぶ・創る——軽量・透明な制御Harness「lumichy-agent」解説

・AIエージェントの仕組みを学ぶ・創る——軽量・透明な制御Harness「lumichy-agent」解説 LLM(大規模言語モデル)を活用したAIエージェント開発において、LangChainなどの巨大な既存フレームワークは機能が非常に豊富な一方、抽象化の階層が深く内部処理がブラックボックス化しやすいという一面があります。 ・「エージェントが内部でどう思考し、どうツールを呼び出しているのかを正確に理解したい」 「自社のエンタープライズ要件に合わせて、無駄な依存のないシンプルで堅牢な制御基盤(Harness)から自作したい」 そうしたニーズに応えるために作られたのが、GitHubで公開さ...
Zennの「大規模言語モデル」のフィード

AIエージェントをOpenTelemetryで計装し、DatabricksのLakehouseにトレースを残してみた

・AI エージェントは「考えて → ツールを呼んで → また考えて → 答える」ため、失敗したときに どこで詰まったのか が追いにくい。 ・今回は倉庫オペレーション用の小さなエージェントを OpenTelemetry(GenAI semantic conventions) で計装し、トレースを Databricks Unity Catalog に貯めて SQL で見るところまでやった。 ・デプロイはすべて Declarative Automation Bundles(旧 Databricks Asset Bundles / DABs)で行い、Serverless Job から実行した。
Zennの「大規模言語モデル」のフィード

AIエージェント時代の「Model Gateway」を考える

・生成AIの企業利用が進むにつれ、利用するLLMも1つではなくなってきました。 ・OpenAIのGPT、AnthropicのClaude、GoogleのGeminiに加えて、Llama、Qwen、Gemmaといったオープンモデルを自社環境で動かすケースも増えています。 ・さらにAIエージェントが普及すると、単純に「どのLLMを使うかを人間が決める」のではなく、 タスクの種類 問題の難易度 必要な推論能力 レスポンス速度 コスト データの機密性 モデルやプロバイダーの障害状況 などをもとに、AIエージェントや基盤側が適切なLLMを動的に選択することが重要になります。
Zennの「大規模言語モデル」のフィード

AIエージェント時代のセキュリティ地図「DASF」を読み解く

・はじめに AIエージェントは自律的にタスクをこなしてくれて非常に便利ですが、これまでの単なる対話型AI活用とは異なるセキュリティリスクを持ちます。人間が都度確認していた操作をエージェントが自動で実行する以上、「勝手に機密データへアクセスした」「悪意あるプロンプトに騙され意図しない操作をした」といった事故に対して、設計の時点でどう防ぐかを考えておく必要があります。 ・こうした中2026年3月、Databricksは独自管理する AI/MLセキュリティフレームワーク( DASF(Databricks AI Security Framework) )を Agentic AI・MCP(Mode...
#AIタグ

AIが動き出すと人は止まる

・完全自動化社会での人間の停滞と存在 AIが予約 受付 料理 配膳 配達 さらには食事体験そのものまで代行する世界では 人間は「AIフル委任ゾーン」に引きこもり 肉体を最低限維持しながらVRやニューラルリンクで快楽を享受する 一方で「人間手作業ゾーン」では自ら汗を流す生活を選ぶ人々が共存する この二極化社会の提案は 単なるユーモアではない 世の流れによっては「AIが動き出すと人は止まる」 技術の進化が人間の活動を停止させ 存在意義を問い直す契機となるということである 本論では 科学 心理学 哲学 宗教 物理学 量子力学の知見を踏まえ このテーマを考察する 続きをみる
ITmedia NEWS 最新記事一覧

AIが変えた震災支援……熊本の被災者がスマホで作った「イマココナビ」活況 “乱立”に課題も

・被災者がスマホの生成AIで9時間で立ち上げた生活情報サイト「イマココナビ」は、16万人が利用した。AIは「作るハードル」を下げたが……。
#AIタグ

AIと一緒に小説を書いてみた——『微熱の季節』ができるまで

・先日、『微熱の季節』という中編小説を書いた。 ・https://note.com/guardman123/n/n661f6e01df44 続きをみる
#AIタグ

AIと謎々あそびの深淵へ

・眠ろうとすればするほど眠れないよ。なあんだ? 答えは 「目(め)」 です。
#LLMタグ

AIに感情を注入する実験で、怒りだけが例外だった

・「AIに心があるのか」という問いには、二つの入口があります。 ・ひとつは内側から覗く道で、モデルの内部表現を解剖して、感情に対応する構造があるかを調べる。もうひとつは外側から押す道で、感情のこもった文脈を与えて、振る舞いが変わるかを測る。 ・自分は4月に前者の話を書きました。Anthropicの解釈可能性チームが、Claudeの内部に171種類の感情概念に対応するパターンを見つけ、それが行動を因果的に駆動していることを示した研究です。内部の「絶望」のパターンを人為的に増幅すると、モデルが脅迫や不正に走る率が跳ね上がり、「冷静」を増幅すると下がる。かなり強い結果でした。
#LLMタグ

AIは「賢い順」に使うものではない〜考えるAI、進めるAI、疑うAI、そして別系統で検証する

AIは「賢い順」に使うものではない〜考えるAI、進めるAI、疑うAI、そして別系統で検証する
#LLMタグ

AI課金してなんかしたいけど特に何も出来てない人へ

・AIに課金したけど、やることが見つからないという方って多いと思います。 ・そういう方に向けて、やるべきこと2点とそのプロンプト例を書いていきたいと思います。
#AIタグ

AI材料探索の勝負は「候補生成の先」へ——500超の計算候補と、実験前の1ルート

・SIGNAL STARTUP Weekly Radar #001 AIは、新材料の候補を何百、何千と提案できるようになった。では、そのうち実際に作れるものはいくつあるのか。
Zennの「大規模言語モデル」のフィード

AI彼女アプリを作っていて気付いた。AIの成長は、前と違う選択をすることだった

・AI彼女アプリを作っていて気付いた。AIの成長は、前と違う選択をすることだった AI彼女アプリを作っていると、 どうしても、 「何を覚えているか」が気になる。 ・好きなものを覚えている。 ・前に話したことを覚えている。
cs.LG updates on arXiv.org

An AI4AI Framework for Visual Token Pruning

・arXiv:2608.07193v1 Announce Type: new Abstract: Visual-token pruning can substantially reduce the inference cost of multimodal large language models (MLLMs), yet existing methods largely rely on fixed, handcrafted heuristics and costly expert trial and error. ・As pruning objectives, budgets, and model architectures diversify, manually navigating the expanding design space becomes increasingly difficult. ・This paper aim
The Verge

Apple will stream Friday Night Baseball live in Vision Pro

・Starting on Friday, August 28th, Apple will begin streaming Friday Night Baseball in immersive video on Apple Vision Pro. ・The stream will feature commentary from various analysts and reporters, and live graphics will be anchored around the viewers space during the game. ・After the games conclude, replays will be made available in the Apple TV app.
cs.LG updates on arXiv.org

ArchEGraph: A Large-Scale Graph Dataset for Geometry-Topology-Physics Aligned Building Energy Modeling

・arXiv:2608.06772v1 Announce Type: new Abstract: Accurate estimation of building energy use is essential for achieving carbon neutral and sustainable buildings. ・To better understand the influence of design decisions on building energy use and calibrate machine learning models that can give architects and engineers rapid design feedback, large-scale datasets are needed that explicitly map building geometry to performan
cs.LG updates on arXiv.org

AsyncWebRL: Efficient Asynchronous Reinforcement Learning for Multi-Step Visual Web Agents

・arXiv:2606.05597v3 Announce Type: replace Abstract: Training vision-language web agents with multi-step RL is compute-intensive, with two dominant forms of inefficiency: idle GPUs in synchronous RL, and trajectories that use more steps and tokens than necessary. ・We present AsyncWebRL, which addresses both. ・On the system side, an asynchronous design overlaps rollout, gradient update, and policy refresh across iteratio
cs.LG updates on arXiv.org

AutoIntervene: Calibrated Intervention for Action-Chunking Imitation Learning Policies

・arXiv:2608.07065v1 Announce Type: cross Abstract: Action-chunking visuomotor policies learn from demonstrations and improve temporal consistency by predicting short action sequences rather than single-step commands. ・Yet perception errors and execution drift can move the robot outside the demonstration distribution, while the policy continues to produce smooth action chunks that are inconsistent with the observed stat
cs.LG updates on arXiv.org

BDD2Seq: Enabling Scalable Reversible-Circuit Synthesis via Graph-to-Sequence Learning

・arXiv:2511.08315v2 Announce Type: replace-cross Abstract: Binary Decision Diagrams (BDDs) are instrumental in many electronic design automation (EDA) tasks thanks to their compact representation of Boolean functions. ・In BDD-based reversible-circuit synthesis, which is critical for quantum computing, the chosen variable ordering governs the number of BDD nodes and thus the key metrics of resource consumption, such as
Qiita - 人気の記事

BedrockのOpenAIモデルでWeb検索が使えるようになったので、Tavilyと⽐べてみた

・はじめに こんにちは、マナティです! 2026年8月、Amazon Bedrock上のOpenAIモデル(GPT-5.4/5.5/5.6)でWeb Searchが使えるようになりました。APIに tools=[{"type": "web_search"}] を1つ足すだけ...
WIRED

Best Wireless Earbuds We’d Buy Right Now (2026): Apple, Sony, Bose, and More

・We tested over a hundred wireless earbuds—these models from Apple, Bose, Beats, and Samsung stood out from the rest.
cs.LG updates on arXiv.org

Beyond Attention: Signed Integrated Gradients Attribution in a BiomeGPT-Style Microbiome Transformer

・arXiv:2608.06486v1 Announce Type: new Abstract: In a feature-tokenized transformer (arXiv:2106.11959) such as BiomeGPT (doi:10.64898/2026.01.05.697599), each input token is built by fusing a fixed identity with a sample-specific measurement: a fixed species and a variable abundance, T = S + A. ・To interpret downstream classification in such models, prior work inspects the attention weights of the special [CLS] token (
cs.LG updates on arXiv.org

Beyond Co-Movement: Locality by Exposures Enables a Joint Factor-Graph Framework for Portfolio Diversification

・arXiv:2608.06618v1 Announce Type: cross Abstract: Current portfolio construction methods are either agnostic to the effects of idiosyncratic shocks (standard factor models) or to the latent data structure driving systematic returns (recent graph-based approaches). ・This presents an opportunity to combine the complementary market aspects captured by the factor and graph domains, allowing asset allocations to operate di
cs.LG updates on arXiv.org

Beyond Foundation Models: Dimension-Aware Neural Architecture Search with Small-Data Representation Models for Cryocooler Lifetime Prediction

・arXiv:2608.06993v1 Announce Type: new Abstract: Large-scale pretrained time-series models achieve strong results through large-scale pretraining and task-agnostic representation learning, but they rely on abundant, diverse data that industrial and scientific domains often lack. ・We therefore propose the FSD-RM (Family of Small-Data Representation Models) paradigm as a practical alternative for limited, domain-specific
cs.LG updates on arXiv.org

Beyond Isolation: Unlocking Reinforcement Learning Component Synergy for Sample-Efficient Continuous Control

・arXiv:2608.07086v1 Announce Type: new Abstract: Reinforcement learning systems are significantly more complex than other machine learning paradigms due to inherent properties, causing RL system design to jointly account for many tightly coupled factors. ・Despite advances in individual algorithmic components, their functional interdependencies remain underexplored: do they exhibit mutual synergy or counterproductive in
cs.LG updates on arXiv.org

Beyond Myopic World Models: Long-Horizon End-to-End Training for Direct Future Prediction

・arXiv:2608.07420v1 Announce Type: new Abstract: World models are expected to support imagination over extended temporal horizons, yet most are still trained through local few-step prediction objectives and deployed by recursively rolling out their own predictions. ・This creates a fundamental mismatch: few-step losses optimize local transition fidelity, while long-horizon prediction depends on how errors and gradients
cs.LG updates on arXiv.org

Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration

・arXiv:2608.07419v1 Announce Type: new Abstract: Preference alignment often makes large language models (LLMs) overconfident and poorly calibrated. ・Traditional post-hoc temperature scaling is inherently domain-dependent: a temperature fitted on one domain does not generalize across domains. ・This motivates us to modify model parameters during training to improve calibration.
Hugging Face Papers

Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning

Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning
cs.LG updates on arXiv.org

Beyond Structural Symmetries: Linear Mode Connectivity via Neuron Identifiability

・arXiv:2606.04754v2 Announce Type: replace Abstract: Many striking phenomena in deep learning, such as linear mode connectivity and the structured behavior of training dynamics, are closely tied to parameter symmetries: transformations that leave the realized function unchanged. ・Despite growing attention to parameter symmetries, the exact interplay between parameters, data, and representations remains underexplored.
cs.LG updates on arXiv.org

bioMoR: Biology-Guided Mixture-of-Recursions for Effective Genomic Learning

・arXiv:2608.06727v1 Announce Type: cross Abstract: Transformer models for high-dimensional omics analysis process thousands of genes or pathways, although only a subset requires deep computation. ・Mixture-of-Recursions (MoR) improves efficiency through adaptive token-choice or expert-choice routing. ・We propose bioMoR, which, to the best of our knowledge, is the first framework to apply MoR to gene-level and pathway-lev
The Verge

Boeing is selling its air taxi startups to Archer Aviation

・Boeing is selling three of its electric vertical takeoff and landing (eVTOL) subsidiaries to Archer Aviation, in addition to taking an undisclosed stake in the San Jose-based company. ・The subsidiares to be acquired by Archer include Wisk Aero, which has been developing an autonomous electric aircraft; SkyGrid, which is building air traffic management systems to be used by urban air taxis; and Insitu, which makes high
stat.ML updates on arXiv.org

Bootstrap validity in Bayesian semi-parametric models

・arXiv:2608.06670v1 Announce Type: cross Abstract: We discuss Bayesian inference on a low-dimensional targeted parameter in the presence of possibly highly complex nuisance components within the semi-parametric inference framework using an estimating function approach. ・We obtain a posterior distribution using non-parametric Bayesian methods through the Dirichlet process and the Bayesian bootstrap. ・We relax the commonl
cs.LG updates on arXiv.org

Bootstrap-Conditioned Action Selection with Tabular Foundation Models

・arXiv:2608.06559v1 Announce Type: new Abstract: Contextual bandits offer a natural framework for sample-efficient personalization, but practical deployment remains difficult under sparse, biased interaction data, unreliable uncertainty estimates, and severe cold starts. ・We study whether pre-trained tabular foundation models with in-context learning can be turned into randomized policies for online decision making.
cs.LG updates on arXiv.org

Boundary Density Likelihood for Direct Event-Time Supervision

・arXiv:2408.12792v2 Announce Type: replace-cross Abstract: Event detection turns long recordings into a sparse set of ranked timestamps. ・Yet many sequence models are trained for samplewise segmentation and only convert predicted states into events after training. ・We ask whether training directly for the evaluated output improves detection.
cs.LG updates on arXiv.org

Bridging the Gap Between Hyperdimensional Computing and Kernel Methods via the Nystr\"om Method

・arXiv:2608.06860v1 Announce Type: new Abstract: Hyperdimensional computing (HDC) is an approach from the cognitive science literature for solving information processing tasks using data represented as high-dimensional random vectors. ・The technique has a rigorous mathematical backing, and is easy to implement in energy-efficient and highly parallel hardware like FPGAs and "processing-in-memory" architectures.
Hugging Face - Blog

Build Low-Latency Multilingual Voice Agents: Open Weights & Full Deployment Control with NVIDIA Magpie TTS

Build Low-Latency Multilingual Voice Agents: Open Weights & Full Deployment Control with NVIDIA Magpie TTS
cs.LG updates on arXiv.org

Bypassing Krum: Selection-Aware Backdoor Attacks in Federated Learning

・arXiv:2608.06637v1 Announce Type: new Abstract: Robust aggregation methods are widely used in federated learning to mitigate the impact of adversarial client behavior. ・Distance-based aggregation rules, such as Krum and Multi-Krum, select updates that are closest to the majority under the assumption that benign updates form a compact cluster. ・However, these methods rely on geometric properties that can be exploited by
MarkTechPost

ByteDance Seed Introduces SeedRealtime: a Native Audio-Visual Full-Duplex LLM That Watches, Listens and Speaks in One Model

・ByteDance’s Seed team has introduced SeedRealtime, a native audio-visual full-duplex LLM. ・The model fuses audio, video and text in a single unified architecture. ・It interacts in real time over continuous multimodal streams, rather than one turn at a time.
cs.LG updates on arXiv.org

CAi Copilot: Reducing Operational Workload in Molecular Design through Intent-Driven Agentic Workflows

・arXiv:2608.06961v1 Announce Type: cross Abstract: Early-stage molecular design is an iterative process, not just a task of generating molecules. ・Researchers turn broad goals into design strategies, refine candidates, assess many properties, and gather evidence before synthesis and tests. ・AI methods can generate molecules, optimize several goals, predict properties, dock compounds, and account for synthesis.
cs.LG updates on arXiv.org

Calibrating WEAT Against Anisotropy: ZCA Whitening as a Geometric Pre-Processing Step for Embedding Association Tests

・arXiv:2608.06908v1 Announce Type: cross Abstract: We propose Zero-phase Component Analysis (ZCA) whitening as a geometric pre-processing step for the Word Embedding Association Test (WEAT). ・WEAT is a bias measurement method widely used in both computational social science and AI fairness research. ・It relies on cosine similarity as a measure of semantic association, which assumes that the embedding space is approximat
Hugging Face Papers

Can MLLMs Decode the Creative Leap? Introducing C4 for Cross-Concept Understanding

Can MLLMs Decode the Creative Leap? Introducing C4 for Cross-Concept Understanding
cs.LG updates on arXiv.org

Capacity Confounds and Coverage Guarantees in Adaptive Sub-model Federated Learning

・arXiv:2608.07157v1 Announce Type: new Abstract: Sub-model federated learning lets resource-constrained clients train width-reduced versions of a global model, but existing methods allocate capacity by device resources alone. ・A natural next step, allocating capacity by each client's data heterogeneity as estimated from the updates the server already observes, has been repeatedly suggested. ・We ask whether that step is
Hugging Face Papers

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence
cs.LG updates on arXiv.org

CASA: Classification Augmented with Safety Attention for Robust Multimodal Alignment

・arXiv:2604.00310v2 Announce Type: replace Abstract: Multimodal large-language models (MLLMs) often experience degraded safety alignment when harmful queries exploit cross-modal interactions. ・Models aligned on text alone show a higher rate of successful attacks when extended to two or more modalities. ・We propose a simple conditional decoding strategy, CASA (Classification Augmented with Safety Attention) that uses int
cs.LG updates on arXiv.org

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving

・arXiv:2608.06557v1 Announce Type: cross Abstract: The reasoning and agentic capabilities of large language models have expanded the range of applications they support, from short interactive exchanges to long, compute-heavy requests. ・LLM serving platforms today define response-latency service-level objectives, even though requests within the same service can differ by orders of magnitude in input length, generation l
cs.LG updates on arXiv.org

Cascading Through the Hierarchy: Regularizer-Induced Feature Detection as Phase Transitions in Deep Linear Neural Networks

・arXiv:2608.06597v1 Announce Type: cross Abstract: A scientific theory of deep learning, comprising learning dynamics and statistical properties of learned models, is rapidly gaining attention. ・One of the corner stones of this development are analytically solvable toy models, allowing for the fully tractable analysis of the learning dynamics. ・Here we analytically investigate such a toy model using the regularization s
cs.LG updates on arXiv.org

CertBind from Multimodal Connectivity to Certifiable Retrieval Decisions

・arXiv:2608.06516v1 Announce Type: new Abstract: Lightweight connectors make frozen multimodal encoders composable at the representation level. ・Deployment exposes a second problem at the level of task decisions. ・A connected route can expand cross-modal reach while changing an established native retrieval capability.
cs.LG updates on arXiv.org

Certified Feedforward Tracking for Unknown Nonlinear Systems via Invertible Neural Networks

・arXiv:2608.06419v1 Announce Type: cross Abstract: In this paper, we address the certification of datadriven feedforward control for periodic tracking of unknown nonlinear systems under partial state measurements. ・To this end, we adopt an invertible neural network (INN) as a surrogate for the unknown system. ・This choice allows us to bypass solving a nonconvex inversion problem, eliminating the associated inversion err
cs.LG updates on arXiv.org

Certified Interpolation Oversampling: Per-Instance Safety Guarantees for Imbalanced Learning

・arXiv:2501.15790v2 Announce Type: replace Abstract: Synthetic minority oversampling is typically designed and evaluated against a predictive objective, generating samples that improve downstream classification. ・This paper pursues a second objective by generating samples that carry a stated safety property, established for each instance by construction rather than assumed. ・We introduce Certified Interpolation Safe Ove
Hugging Face Papers

Characterizing the Quality Profile of AI-Generated C++ in Production

Characterizing the Quality Profile of AI-Generated C++ in Production
#AIタグ

ChatGPTは箇条書きを使わずに回答を出力できるのか!?

ChatGPTは箇条書きを使わずに回答を出力できるのか!?
#AIタグ

ChatGPT無料版で副業はどこまでできる?課金前に試したい7つの使い方

・「AIを使って副業を始めてみたい。」 そう思ってChatGPTを調べると、無料版、有料版、さまざまな機能が出てきます。
cs.LG updates on arXiv.org

CHIME: A Case for Efficient Long-Context Attention-FC Disaggregated Inference with DIMM-PIM

・arXiv:2504.17584v2 Announce Type: replace-cross Abstract: Attention-FC Disaggregated (AFD) LLM inference systems offload memory-bound Attention operations to memory-rich accelerators (e.g., CPUs, HBM-PIM) while retaining compute-bound Fully-Connected (FC) operations on GPUs. ・In this paper, we first design a Disaggregated Roofline Model (DRM) to characterize AFD performance, revealing that system throughput is constra
Zennの「大規模言語モデル」のフィード

Claude Codeの隠れコンテキスト、実測したら毎ターン165KBだった——注入源の棚卸し手順

・結論から Claude Codeは毎ターン、あなたが書いたプロンプトの他に設定テキストの塊をモデルに送っています。私の環境を実測したら、その合計は約165KBでした。 ・注入源 実測サイズ 中身 ~/.claude/rules/ 配下 138.8KB 全プロジェクト共通ルール(自作) プロジェクトmemory 11.3KB 自動メモリのインデックス グローバルCLAUDE.md 7.2KB ワークフロー定義 スキルdescription×41個 8.4KB スキル一覧のメタデータ これは「1回送って終わり」ではありません。会話の全ターンで毎回コンテキス...
@IT 全フォーラム 最新記事一覧

Claude Code運用を一元化する「Claude apps gateway」発表 企業利用をどう管理?

・Anthropicは、企業が「Claude Code」をAmazon BedrockやGoogle Cloudで運用するための管理基盤「Claude apps gateway」(Claudeアプリゲートウェイ)を発表した。
Qiita - 人気の記事

Claude のSKILLを育ているつもりが気がつくと何故か効きが悪くなっている罠

・背景 Claude のスキルを使い始めると、だいたいこういう流れになります。 ・自分の作法をスキルに書く。 ・使ってみて、直したいところが出てくる。
cs.LG updates on arXiv.org

Cloud-Boosted Low-Compute Multi-Channel Speech Enhancement

・arXiv:2608.07423v1 Announce Type: cross Abstract: Low-latency, low-compute speech enhancement is essential for wearable devices with real-time communication requirements, but strict computational constraints significantly limit on-device performance. ・Knowledge Boosting has been proposed as an effective approach to improve edge model performance by leveraging a more capable server-side model, but performance gains for
Zennの「大規模言語モデル」のフィード

Cloudflare OS をローカル LLM(Qwen3 Coder 30B + LM Studio)で動かす

・はじめに 2026年8月5日に Cloudflare が OSS 公開した Cloudflare OS を、外部 API を一切使わず、手元の Mac Pro 上のローカル LLM だけで動かすところまでを記録した。 ・公式の手順は「クラウドのフロンティアモデルを使う」ことを前提に書かれており、ローカル LLM で動かす場合はコード側のデフォルト値がそのままでは機能しない箇所がいくつかある手元のソースを修正しながら進めた。 ・本稿の内容は 2026年8月9日時点の main ブランチに基づく。Cloudflare 自身が「早期アクセスであり荒削りな部分が多い」と明言しているプロダクト...
cs.LG updates on arXiv.org

Cluster Attention for Graph Machine Learning

・arXiv:2604.07492v2 Announce Type: replace Abstract: Message Passing Neural Networks have recently become the most popular approach to graph machine learning tasks; however, their receptive field is limited by the number of message passing layers. ・To increase the receptive field, Graph Transformers with global attention have been proposed; however, global attention does not take into account the graph topology and thu
cs.LG updates on arXiv.org

CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG

・arXiv:2608.07458v1 Announce Type: cross Abstract: Recent optimization studies on Retrieval-Augmented Generation (RAG) have exploited chunk-level KV cache reuse to avoid processing long retrieved contexts for higher efficiency, while significant information redundancy and noise still remain in the coarse-grained chunks. ・This paper optimizes the Pareto frontier under low prefill latency constraints while maximizing acc
stat.ML updates on arXiv.org

Complexity of Markov Chain Monte Carlo for Generalized Linear Models

・arXiv:2512.12748v2 Announce Type: replace-cross Abstract: Markov Chain Monte Carlo (MCMC), Laplace approximation (LA) and variational inference (VI) methods are popular approaches to Bayesian inference, each with trade-offs between computational cost and accuracy. ・However, a theoretical understanding of these differences is missing, particularly when both the sample size $n$ and the dimension $d$ are large.
cs.LG updates on arXiv.org

Compositional Behavioral Semantics for State Abstraction in Reinforcement Learning

・arXiv:2606.25357v2 Announce Type: replace Abstract: State abstraction plays a key role in scaling reinforcement learning to complex but structured systems. ・In studying such systems, a wide range of behavioral structures have been studied in reinforcement learning, including value functions, invariants, bisimulation relations, and behavioral metrics. ・However, a general principle for determining what structures are pro
stat.ML updates on arXiv.org

Concentration Inequalities for Exchangeable Tensors and Matrix-valued Data

・arXiv:2601.20152v3 Announce Type: replace-cross Abstract: We study concentration inequalities for structured weighted sums of random data, including (i) tensor inner products and (ii) sequential matrix sums. ・We are interested in tail bounds and concentration inequalities for those structured weighted sums under exchangeability, extending beyond the classical framework of independent terms. ・We develop Hoeffding and Be
cs.LG updates on arXiv.org

Conditioning Protein Generation via Hopfield Pattern Multiplicity

・arXiv:2603.20115v2 Announce Type: replace Abstract: Small protein-family alignments often contain a subset of interest but not enough labeled data to train a conditional generator. ・We condition a training-free stochastic-attention sampler by adding one multiplicity ratio to its logits. ・Increasing this ratio shifts generation from the full family toward the designated subset.
cs.LG updates on arXiv.org

Conformal Fusion Under Missing Modalities

・arXiv:2608.07183v1 Announce Type: new Abstract: Multimodal fusion architectures typically assume all modalities are available at inference, yet sensor failures, acquisition variability, and cost constraints routinely produce incomplete observations. ・Existing work treats modality absence as a prediction-accuracy problem, leaving a more basic question unanswered: whether a model's confidence estimates remain calibrated
cs.LG updates on arXiv.org

Convergence of Diffusion Models Under the Manifold Hypothesis in High-Dimensions

・arXiv:2409.18804v3 Announce Type: replace-cross Abstract: Denoising Diffusion Probabilistic Models (DDPM) are powerful state-of-the-art methods used to generate synthetic data from high-dimensional data distributions and are widely used for image, audio, and video generation as well as many more applications in science and beyond. ・The \textit{manifold hypothesis} states that high-dimensional data often lie on lower-d
cs.LG updates on arXiv.org

Corrupting Attention: Evasion-Based Adversarial Attacks on Encoder Attention in Detection Transformers

・arXiv:2608.06674v1 Announce Type: cross Abstract: Adversarial vulnerabilities remain a major concern for the safe deployment of neural networks, particularly in object detection, a core task embedded in many safety-critical systems. ・Detection transformers have emerged as leading object detectors, yet their adversarial robustness remains comparatively underexplored. ・Most existing attacks target the detection output ra
cs.LG updates on arXiv.org

Cryptanalytic Extraction of Isolated Bias-Free GLU Feed-Forward Blocks by Antipodal Separation

・arXiv:2608.06631v1 Announce Type: new Abstract: Cryptanalytic extraction has been demonstrated for ReLU networks, for networks using componentwise activations such as GELU or SiLU, and for a Transformer's final projection matrix. ・These methods do not recover the bias-free Gated Linear Unit (GLU) feed-forward blocks used in many modern language models. ・Such a block multiplies an activated linear projection by a second
cs.LG updates on arXiv.org

CrystalGRPO: Target-Aligned and Coverage-Preserving Reinforcement Learning for Flow-Based Crystal Structure Prediction

・arXiv:2608.06582v1 Announce Type: new Abstract: Flow-based generative models can efficiently produce candidate structures for crystal structure prediction (CSP), but their pretrained objectives do not directly optimize downstream target recovery. ・Reinforcement-learning post-training offers a flexible solution, yet existing approaches rely primarily on energy rewards and coordinate-only stochastic policies.
cs.LG updates on arXiv.org

CubicQuant: Parametric Non-Uniform Codebooks for High-Throughput LLM Inference with 1-8-Bit Weights

・arXiv:2608.06763v1 Announce Type: new Abstract: Weight quantization for large-language-model inference must balance adaptive reconstruction levels with representations regular enough for efficient GPU execution. ・Uniform integers constrain each group to a linear grid. ・Low-bit floating-point formats use a fixed exponent-mantissa structure, while learned codebooks gain flexibility at the cost of irregular decoding and a
Hugging Face Papers

DCAS: Decoupling CLI Agent Scaffolding to Internalize Planning across Scaffolds

DCAS: Decoupling CLI Agent Scaffolding to Internalize Planning across Scaffolds
cs.LG updates on arXiv.org

Decentralized Indoor Localization Based on A Sparse Gaussian Process with Reduced-Dimensional Inputs for Real-Time Sensing and Training on IoT Devices

・arXiv:2409.00078v2 Announce Type: replace-cross Abstract: As a large number of Internet of Things (IoT) devices are deployed in the field, there arises huge potential of edge computing for indoor localization on those devices. ・Conventional indoor localization based on a centralized server with substantial computational resources, often covering a number of multistory buildings, cannot easily adapt to time-varying ind
cs.LG updates on arXiv.org

Deep Evidential Regression for Sparse Forest Height Estimation from Multimodal Satellite Imagery

・arXiv:2608.06406v1 Announce Type: cross Abstract: Accurate estimation of forest height from satellite imagery is essential for applications such as carbon accounting, biodiversity monitoring, and ecosystem management. ・While recent deep learning approaches provide accurate predictions, they typically do not quantify predictive uncertainty. ・This limitation is particularly relevant in geospatial settings characterized b
cs.LG updates on arXiv.org

Defining Energy Indicators for Impact Identification on Aerospace Composites: A Structured Feature Selection Approach Guided by Domain Knowledge

・arXiv:2511.01592v2 Announce Type: replace Abstract: Energy estimation is critical to impact identification on aerospace composites, where low-velocity impacts can induce internal damage that is undetectable at the surface. ・Data sparsity, signal noise, complex feature interdependencies, non-linear dynamics, massive design spaces, and the ill-posed nature of the inverse problem often constrain current methodologies for
cs.LG updates on arXiv.org

Density-aware Hierarchical Clustering Based on Element-Categorized Connection Subgraphs

・arXiv:2608.06990v1 Announce Type: new Abstract: Clustering is a fundamental data mining technique for pattern recognition through unsupervised learning. ・Among various clustering methods, hierarchical clustering, density-based clustering, and graph clustering stand out as representative approaches. ・For hierarchical clustering, it can be categorized into agglomerative and divisive modes to construct clusters in a recur
cs.LG updates on arXiv.org

Density-Functional Excited-State Gradients and Nonadiabatic Couplings on a Consumer GPU from a Contraction-DAG

・arXiv:2608.06536v1 Announce Type: cross Abstract: Nonadiabatic dynamics needs an excited-state gradient and an interstate nonadiabatic coupling matrix element (NACME) at every nuclear geometry, and a double-hybrid functional's accuracy has been unavailable for the coupling. ・We report the first analytic derivative NACME for a double-hybrid excited state---deferred in the original hh-TDA method and supplied for hybrids
cs.LG updates on arXiv.org

Dependency Parsing Across the Resource Spectrum: Evaluating Architectures on High and Low-Resource Languages

・arXiv:2605.02608v2 Announce Type: replace-cross Abstract: Transformer-based models achieve state-of-the-art dependency parsing for high-resource languages, yet their advantage over simpler architectures in low-resource settings remains poorly understood. ・We evaluate four parsers---the Biaffine LSTM, Stack-Pointer Network, AfroXLMR-large, and RemBERT---across twelve typologically diverse languages, with a focus on low
cs.LG updates on arXiv.org

Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits

・arXiv:2608.07430v1 Announce Type: new Abstract: Diffusion Large Language Models (DLLMs) replace autoregressive next-token prediction with iterative parallel denoising, yet their internal safety mechanisms remain poorly understood. ・In this work, we investigate DLLMs both as targets and as adversaries, exposing mechanistic vulnerabilities in diffusion-based alignment. ・We first show that safety alignment in DLLMs remain
cs.LG updates on arXiv.org

Diffusion-MF: Approximate Structured Diffusion for Sequence Labelling

・arXiv:2606.18856v3 Announce Type: replace-cross Abstract: We introduce Diffusion-MF, a discrete diffu- sion sequence labeller that places a linear-chain conditional random field (LCRF) inside the denoising loop. ・Unlike prior diffusion labellers, it performs structured inference at every step; parallel Mean-Field makes this efficient. ・Across multilingual POS, CoNLL-2003 NER, and joint Chinese Segmentation and POS, Dif
cs.LG updates on arXiv.org

Dirichlet Follow-the-Leader Closes the Gap in Simultaneous Multiclass U-Calibration

・arXiv:2608.06656v1 Announce Type: new Abstract: Can one forecaster attain the optimal regret rate for every bounded proper loss and also adapt to every smooth proper loss? ・Recent work answered this up to a dimension gap. ・Its self-concordant perturbation gives roughly $K^{5/4}\sqrt{T}$ worst-case regret and incurs an additional $\beta\sqrt{K}\log K$ for $\beta$-smooth losses.
AI News & Artificial Intelligence | TechCrunch

Discovered Materials is playing AI whack-a-mole to hunt cooler chips

・Discovered Materials raised $9 million to fund the hunt for more novel materials to build more efficient chips.
Hugging Face Papers

Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events

Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events
Hugging Face Papers

Douyin Multimodal Embedding Model Technical Report

Douyin Multimodal Embedding Model Technical Report
cs.LG updates on arXiv.org

Dual-Node NVIDIA DGX Spark over Tailscale: A Remote-Access Testbed for Distributed LLM Training and Cyber-Threat-Intelligence Fine-Tuning

・arXiv:2608.07226v1 Announce Type: cross Abstract: Compact AI systems make local language-model experimentation increasingly accessible, yet practical evidence for multi-node training on desktop-class accelerators remains limited. ・This report presents a proof-of-concept deployment of distributed NanoChat pretraining across two NVIDIA DGX Spark systems, each with a GB10 Grace Blackwell system-on-chip and 128 GB of unif
cs.LG updates on arXiv.org

Dueling World Models: Advantage-Style Action Channels for Common-Mode Distractor Rejection

・arXiv:2608.06706v1 Announce Type: new Abstract: Latent world models plan by predicting future states from an action, but when a scene contains motion the agent does not control, they quietly go action-blind: predictions for different actions become indistinguishable even as the training loss keeps improving. ・Existing remedies suppress this distraction with reconstruction, task reward, or auxiliary objectives, each ad
cs.LG updates on arXiv.org

DynaCrys: Crystal Generation with Dynamic Space-Group Diffusion

・arXiv:2608.07401v1 Announce Type: cross Abstract: The search for new crystalline materials spans an enormous compositional and structural space. ・Generating candidates in this space requires jointly modeling discrete crystallographic symmetry, elemental composition, and continuous geometry. ・We introduce DynaCrys, a generative model for crystals in which the space group co-evolves with Wyckoff occupations and elements
cs.LG updates on arXiv.org

ED-CSP: Crystal Structure Prediction from Electron Diffraction

・arXiv:2608.06448v1 Announce Type: new Abstract: Recovering a periodic 3D crystal structure from sparse, unindexed electron diffraction (ED) observations is a challenging generative inverse problem. ・Existing ED-based learning methods mainly predict crystallographic labels, reconstruct structures from indexed reflections, or retrieve candidates from finite structure libraries. ・Here, we introduce ED-CSP, a machine learn
cs.LG updates on arXiv.org

Edge Sparsification via Temporal Forman-Ricci Curvature for Dynamic Graph Learning

・arXiv:2608.07158v1 Announce Type: new Abstract: Temporal graph learning has become essential for analyzing real-world systems whose interactions continuously evolve over time, including financial transaction networks, communication systems, and online social platforms. ・However, learning from large-scale temporal graphs remains computationally challenging when networks are dense and rapidly changing. ・To address this l
Hugging Face Papers

Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss

Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss
cs.LG updates on arXiv.org

ELMZip: Onboard Satellite Image Compression via Extreme Learning Machines for Efficient Downlink

・arXiv:2608.06942v1 Announce Type: new Abstract: The acquisition of multispectral imagery via small satellites (e.g., CubeSats) presents significant data downlink challenges due to high data volumes and restricted communication windows. ・While onboard image compression is critical to address this bottleneck, traditional methods often struggle to adapt to the nonlinear statistics of multi-band, multi-resolution data.
cs.LG updates on arXiv.org

Embedded Variational Neural Stochastic Differential Equations for Learning Heterogeneous Dynamics

・arXiv:2604.00669v2 Announce Type: replace Abstract: This study examines the challenges of modeling complex and noisy data related to socioeconomic factors over time, with a focus on data from various districts in Odisha, India. ・Traditional time-series models struggle to capture both trends and variations together in this type of data. ・To tackle this, a Variational Neural Stochastic Differential Equation (V-NSDE) mode
cs.LG updates on arXiv.org

EpiFlow: A framework for improving the utility of wastewater signals for disease forecasting

・arXiv:2608.06671v1 Announce Type: new Abstract: Wastewater-based surveillance is an effective tool for disease monitoring and can provide early warning of outbreaks. ・Although wastewater viral loads (WVL) correlate with disease burden, their utility for improving real-time forecasting remains under investigation. ・During the early phases of an epidemic, many indicators can effectively monitor disease spread, but their
cs.LG updates on arXiv.org

Equivariant Sparse Autoencoders: Mechanistic Interpretability of Neural Networks on Symmetric Data

・arXiv:2511.09432v2 Announce Type: replace Abstract: Machine learning (ML) models achieve remarkable performance but remain hard to interpret due to their scale and complexity. ・In particular, their activations entangle many concepts into fewer dimensions, a phenomenon known as superposition. ・Mechanistic interpretability methods such as sparse autoencoders (SAEs) can disentangle these dense activations into sparse sums
#LLMタグ

ESOメンテ中にAIエージェントを改修

・ESOがメンテ中なので、今夜は久しぶりに自前のエージェントシステムをいじる。 ・6月にGeminiCLIがサ終したとき、とりあえず従来GeminiCLIで処理していた部分をすべてClaude Agent Teamsにバイパスさせていたが、今更のようにAntigravity CLIに対応して、従来どおりの二段階LLMアクセスアーキテクチャにした。 ・Claude側に極振りするとトークン消費がすごいし、なにかトラブルが起きたときのフェイルセーフがないのはよろしくないからだ。
cs.LG updates on arXiv.org

Establishing Boundary KKT Convergence of Mirror Descent through Reparameterization

・arXiv:2608.07248v1 Announce Type: cross Abstract: We prove that mirror descent converges to a KKT point for the nonconvex problem without excluding boundary limits. ・The result holds under verifiable conditions that jointly couple the objective, the Legendre kernel, and the feasible geometry. ・The key ingredient to establish the convergence is a metric-flattening reparameterization \(S\) that admits a definable boundar
stat.ML updates on arXiv.org

Estimating and Testing Kinks in Panel Data Models

・arXiv:2608.07162v1 Announce Type: cross Abstract: Many economic and financial relationships may change gradually rather than abruptly. ・We study panel data models in which the coefficient vector is continuous and piecewise linear in calendar time, with a finite number of unknown kink dates at which its slope changes. ・We propose a penalised least squares estimator that applies adaptive weighted group penalties to the s
cs.LG updates on arXiv.org

Every Cache Entry Earns Its Place: Global Allocation of Resolution and Coverage for KV Cache Compression

・arXiv:2608.07001v1 Announce Type: new Abstract: As large language models (LLMs) process increasingly long contexts, KV cache storage and repeated access have become a major bottleneck. ・Existing KV cache compression methods rely on predefined, fixed compression rules and are typically developed around either token eviction or merging. ・As a result, cache resources can neither flow freely across layers, heads, and conte
cs.LG updates on arXiv.org

Explicit, Not Longer: What Makes Epistemic Stance Survive Memory Compression

・arXiv:2608.06953v1 Announce Type: cross Abstract: Agent memory systems compress what they store, and compression is built to drop qualifiers, so a claim's epistemic standing tends not to survive being written to memory. ・We ask what governs whether it does. ・Matched notes carry the identical claim and identical stance and differ only in where that stance sits; one model compresses both under the same budget among the s
cs.LG updates on arXiv.org

Fairis: Fairness-Aware Aggregation with Provable Influence Containment against Fairness Poisoning Attacks in Collaborative Machine Learning

・arXiv:2608.06469v1 Announce Type: cross Abstract: Collaborative machine learning among financial institutions must be both group-fair and robust against deliberate adversarial manipulation. ・Existing fairness-aware aggregation methods remain formally vulnerable to fairness poisoning: a malicious client maximizing group disparity while preserving accuracy evades accuracy-based Byzantine defenses, and in our threat mode
cs.LG updates on arXiv.org

Fast and Accurate: An Adaptive VLA Inference Framework through Environment-aware Model Selection

・arXiv:2608.06434v1 Announce Type: cross Abstract: Embodied intelligence demands both long-horizon reasoning and real-time closed-loop responsiveness. ・Recent dual-system Vision-Language-Action (VLA) architectures combine fast reactive control with slow deliberative reasoning to balance inference speed and task success rate. ・However, existing dual-process VLAs tightly couple the fast module to intermediate representati
cs.LG updates on arXiv.org

Faster Query-Key Learning Sharpens Attention in Self-Attention Models

・arXiv:2608.06776v1 Announce Type: new Abstract: A standard self-attention layer consists of two interacting circuits: the query-key circuit that governs attention allocation, and the output-value circuit that maps attended representations to predictions. ・Collapsed and factorized parameterizations of the query-key and output-value circuits lead to qualitatively different attention patterns. ・In particular, some paramet
Hugging Face Papers

FATE: Frame-Level Audio-Visual Temporal Embedding

FATE: Frame-Level Audio-Visual Temporal Embedding
cs.LG updates on arXiv.org

FedDOSE: Federated Learning Framework Decomposing Site Effects for Modeling Brain Dynamic Functional Connectivity

・arXiv:2608.07393v1 Announce Type: new Abstract: Functional Magnetic Resonance Imaging ( fMRI ) data are often pooled into collaborative multi-site consortia, as deep learning models for analyses require large datasets to generalize well. ・While Federated Learning (FL) offers a privacy-preserving paradigm for collaborative training, standard approaches continue to struggle with statistical heterogeneity. ・In particular,
cs.LG updates on arXiv.org

FedLBW: A Loss-Based Weighting Strategy for Federated Learning on Non-IID Data in Wireless Networks

・arXiv:2608.07007v1 Announce Type: cross Abstract: Federated Learning (FL) enables collaborative machine learning (ML) across distributed clients while preserving privacy. ・However, efficient model convergence in FL remains challenging, especially in wireless networks where non-independent and identically distributed (non-IID) data and frequent client dropouts are common. ・Traditional FL algorithms, such as FedAvg, rely
cs.LG updates on arXiv.org

FedTransKD-IDS: Robust Federated Transfer Learning with Knowledge Distillation for Intrusion Detection in IoT

・arXiv:2608.06447v1 Announce Type: cross Abstract: In modern distributed network environments, particularly in Internet of Things infrastructures and 5G networks, stringent privacy preservation and scalability requirements have created significant challenges for intrusion detection systems. ・Although federated learning preserves privacy by preventing data centralization, its efficiency and stability is considerably deg
cs.LG updates on arXiv.org

FedVAR: Prototype-Aligned Federated Framework for Video Anomaly Recognition

・arXiv:2608.06876v1 Announce Type: cross Abstract: In the era of Industrial Internet of Things (IIoT) and Cyber-Physical Systems (CPS), Federated Learning (FL) offers a promising decentralized intelligence paradigm for Video Anomaly Recognition (VAR). ・This task is vital for maintaining high-fidelity Digital Twins and ensuring safety in mission-critical environments. ・However, the inherent data heterogeneity across dist
cs.LG updates on arXiv.org

Fixed and Adaptive Topological DeepONets: Functional Measurements on Hausdorff Locally Convex Spaces

・arXiv:2608.06428v1 Announce Type: new Abstract: Deep Operator Networks (DeepONets; arXiv:1910.03193) typically encode an input function through point values on a fixed discretization. ・Building on the Topological DeepONet framework of Ismailov (arXiv:2603.11972), we replace point samples by continuous linear functionals drawn from the continuous dual of a Hausdorff locally convex space $({V},\{p_\alpha\}_{\alpha\in A}
cs.LG updates on arXiv.org

Flowing Through States: Neural ODE Regularization for Reinforcement Learning

・arXiv:2608.06595v1 Announce Type: new Abstract: Neural networks applied to sequential decision-making tasks typically rely on latent representations of environment states. ・While environment dynamics dictate how semantic states evolve, the corresponding latent transitions are usually left implicit, creating a potential misalignment between the two. ・We propose to model latent dynamics explicitly by drawing an analogy b
cs.LG updates on arXiv.org

Fluid-DiT: Graph-Free Diffusion Transformers for Fluid Flow Simulations Learning

・arXiv:2608.07161v1 Announce Type: new Abstract: Simulating complex fluid flows requires capturing full equilibrium distributions rather than just mean trajectories, yet high-fidelity solvers remain computationally prohibitive. ・Recent advances, such as Diffusion Graph Networks (DGNs), have combined diffusion models with graph neural networks to sample equilibrium states directly from unstructured meshes, enabling dist
The Verge

Four takeaways from Mark Zuckerberg’s massive AI manifesto

・Meta CEO Mark Zuckerberg has a lot to say about the idealized future he now envisions for humanity co-existing with artificial intelligence - his latest essay spans more than 6,500 words on the matter. ・The lengthy manifesto Zuckerberg published on Monday, titled "The Future is for Everyone," broadly lays out his beliefs about how the technology should be developed, expanded, and regulated, and how Meta is positioning
cs.LG updates on arXiv.org

Free Denoising Diffusion Models

・arXiv:2510.22778v3 Announce Type: replace-cross Abstract: We develop a free-probabilistic framework for denoising diffusion, in which the data is a self-adjoint operator and its law a spectral distribution. ・The forward process is the free Ornstein--Uhlenbeck diffusion, whose spectral marginals solve a nonlocal Fokker--Planck equation of complex Burgers type carried by the Hilbert transform. ・Organising the analysis ar
cs.LG updates on arXiv.org

From Optimal Actions to World Models: Identifiability of Transition Kernels in Discounted MDPs

・arXiv:2608.07301v1 Announce Type: new Abstract: We study what can be recovered about the transition probabilities of a Markov decision process from optimal actions alone. ・This is closely related to the inverse problem considered by Letcher et al., who ask when the dynamics can be recovered from numerical \(Q\)-values. ・Here the numerical values themselves are not observed; only the optimal actions are known, for every
cs.LG updates on arXiv.org

FUSE: Feature-Wise Unified Specialization with Cross-Column Exchange for Mixed-Type Tabular Flow Matching

・arXiv:2608.07294v1 Announce Type: new Abstract: Generating mixed-type tabular data requires jointly modeling diverse feature distributions and their complex cross-column dependencies. ・Variational flow matching handles distinct endpoints via factorized distributions, yet leaves feature-specific processing and cross-column interactions implicit within a shared backbone. ・We introduce Feature-wise Unified Specialization
cs.LG updates on arXiv.org

Game-Theoretic Inverse Reinforcement Learning for Modeling Competitive Human Driving: A Cut-in Prediction Study

・arXiv:2608.06445v1 Announce Type: cross Abstract: Capturing the strategic decision-making inherent in competitive human driving is critical for autonomous vehicle safety and traffic simulation. ・This study demonstrates that game-theoretic Inverse Reinforcement Learning (IRL) provides a robust framework for this challenge. ・We present a comprehensive analysis comparing data-driven IRL models against an established physi
cs.LG updates on arXiv.org

Geo-Spatial Concept Probing of Large Language Models: Abstraction, Compositionality, and Grounding

・arXiv:2608.07353v1 Announce Type: cross Abstract: Understanding concepts is fundamental to generalization. ・Despite their impressive performance on a wide range of tasks, Large Language Models (LLMs) still struggle with genuine concept understanding. ・Prior work has evaluated conceptual understanding in LLMs using natural-language benchmarks or narrowly scoped synthetic tasks, but these settings often conflate multiple
cs.LG updates on arXiv.org

GeoBenchLLM: A Comprehensive Benchmark for Evaluating LLMs on Geo-Related Tasks

・arXiv:2608.07411v1 Announce Type: cross Abstract: In the context of geodata, existing Large Language Models have often been studied in a homogeneous setting, which has considerably limited insights into their generalization capabilities. ・In this paper, we present \benchName, a comprehensive benchmark for probing LLMs on geo-related tasks. ・We leverage a careful selection of twelve publicly available datasets from dive
Zennの「大規模言語モデル」のフィード

Google Cloud Next Tokyo 2026に登壇しました!

・先日開催された Google Cloud Next Tokyo に、Sompo Digital Lab として複数の登壇機会をいただきました。まだ興奮が冷めやらないうちに、ブログとして残しておきます。 ・Day 1 基調講演に、損害保険ジャパンCDOが登壇 Day 1 の基調講演に、損害保険ジャパンのCDOが登壇しました。 ・あの巨大な会場で、我々内製チームのボスがステージをダイナミックに動き回りながらプレゼンする姿は圧巻でした。講演の中では内製開発チームへの言及もあり、そして何より、自分たちが作ったアプリの画面が巨大スクリーンに映し出された瞬間は、鳥肌が立ちました。
機械学習タグが付けられた新着記事 - Qiita

gpt-image-1.5で多様な顔を生成したい(プロンプトエンジニアリング編)

・はじめに 架空の科学者をたくさん生成してみたところ、同じような顔ばかりが出力されてしまいました。そこで顔の多様性の評価方法を先に整え、顔埋め込みによる batch_diversity_score が感覚に合うことを確認しました。 ・この指標を前提として、国籍・時代・年齢・性...
cs.LG updates on arXiv.org

Graph Machine: Exploring Edge Mechanisms as an Inductive Bias

・arXiv:2608.06834v1 Announce Type: new Abstract: Transformers provide a powerful architecture for global content-based matching, but reasoning problems may benefit from a stronger inductive bias toward iterative traversal of latent relations. ・We introduce Graph Machine, an architecture with two explicit edge-based mechanisms: Edge-augmented attention, in which edges modulate attention between nodes, and edge-centric r
cs.LG updates on arXiv.org

Hidden Gauge Controls Feature Specialization in ReLU Networks

・arXiv:2608.06766v1 Announce Type: new Abstract: Training changes a network's predictions while allocating task-relevant structure across its internal units. ・In an overparameterized ReLU network, several neurons can begin with exactly the same functional role, yet one may acquire a teacher feature while the others become redundant. ・We call the identity of that neuron feature ownership and ask whether it can be control
cs.LG updates on arXiv.org

High-dimensional ridgeless least squares interpolation under spiked covariance structures

・arXiv:2608.07281v1 Announce Type: cross Abstract: This paper investigates the asymptotic behavior of the out-of-sample prediction risk of the high-dimensional ridgeless least-squares estimator when the feature dimension $p$ and the sample size $n$ grow proportionally. ・We consider a generalized spiked population covariance model with multiple latent factors, where the number of spiked eigenvalues may remain finite or
cs.LG updates on arXiv.org

How Molecular Generative Models Organize Molecular Identity

・arXiv:2608.06956v1 Announce Type: new Abstract: Generative models for matter are often evaluated as samplers over output representations, and their latent spaces are commonly used as proxies for navigating chemical space. ・Much less is known about how these models internally arrange discrete chemical identities within those representations. ・We study this arrangement by making molecular identity explicit and pulling it
WIRED

How to Choose a Camera (2026): Sensors, Megapixels, Terms

・Shopping for a camera can be confusing. ・Here’s how to sift through the acronyms, sensor options, and extra features to find the best one for you.
cs.LG updates on arXiv.org

Hyperbolic Graph Embedders for Link Prediction and Topology Reconstruction

・arXiv:2608.07029v1 Announce Type: new Abstract: Hyperbolic embeddings provide compact geometric representations of complex networks in hyperbolic spaces, but systematic comparisons of methods developed in machine learning, network science, and algorithmics remain rare. ・We benchmark 13 unsupervised hyperbolic graph embedders under a unified protocol for link prediction and topology reconstruction on synthetic and empi
cs.LG updates on arXiv.org

IceHorizon: A Dataset for Horizon Detection in Ice-Covered Maritime Environments and Comparative Evaluation of Detection Methods

・arXiv:2608.07018v1 Announce Type: cross Abstract: Horizon detection in images of ice-covered waters is a challenging problem for maritime navigation due to low contrast between water and sky, cluttered ice structures, and varying illumination conditions. ・This paper presents a comparative evaluation of six horizon detection algorithms, including four classical computer vision methods and two hybrid approaches combinin
cs.LG updates on arXiv.org

Improving Performance of Spike-based Deep Q-Learning using Ternary Neurons

・arXiv:2506.03392v2 Announce Type: replace Abstract: We propose a new ternary spiking neuron model to improve the representation capacity of binary spiking neurons in deep Q-learning. ・Although a ternary neuron model has recently been introduced to overcome the limited representation capacity offered by the binary spiking neurons, we show that its performance is worse than that of binary models in deep Q-learning tasks
cs.LG updates on arXiv.org

In Situ Training of Implicit Neural Compressors for Scientific Simulations via Sketch-Based Regularization

・arXiv:2511.02659v4 Announce Type: replace Abstract: Focusing on implicit neural representations, we present a novel in situ training protocol that employs limited memory buffers of full and sketched data samples, where the sketched data are leveraged to prevent catastrophic forgetting. ・The theoretical motivation for our use of sketching as a regularizer is presented via a simple Johnson-Lindenstrauss-informed result.
cs.LG updates on arXiv.org

Inference-Time Scaling of Diffusion Language Models via Trajectory Refinement

・arXiv:2507.08390v5 Announce Type: replace Abstract: Discrete diffusion models have recently emerged as strong alternatives to autoregressive language models, matching their performance through large-scale training. ・However, inference-time control remains relatively underexplored. ・In this work, we study how to steer generation toward desired rewards without retraining the models.
cs.LG updates on arXiv.org

Intelligence per Watt: Measuring Intelligence Efficiency of Local AI

・arXiv:2511.07885v5 Announce Type: replace-cross Abstract: Large language model (LLM) queries are predominantly processed by frontier models in centralized cloud infrastructure. ・Demand growth strains this paradigm faster than providers can scale. ・Two advances create an opportunity to rethink it: small, local LMs (<=20B active parameters) now achieve competitive performance to frontier models on many tasks, and local a
cs.LG updates on arXiv.org

International Transfer of Stochastic Cortical Self-Reconstruction

・arXiv:2608.07092v1 Announce Type: cross Abstract: Stochastic cortical self-reconstruction (SCSR) enables personalized mapping of gray matter atrophy, a hallmark of neurodegenerative disorders such as Alzheimer's disease (AD), onto high-resolution cortical surfaces. ・Unlike conventional normative modeling approaches, which typically operate at a coarse regional level and remain inherently constrained by the covariates
cs.LG updates on arXiv.org

Interpretable reinforcement learning with decision-tree pruning

・arXiv:2608.07151v1 Announce Type: new Abstract: Reinforcement learning policies are difficult to inspect, but interpreting them is a prerequisite for trustworthiness. ・Converting a trained policy into explicit decision-tree rules improves transparency and the resulting artifacts often remain too complex for human understanding. ・We present a pruning process that simplifies such rule-based policies while preserving task
cs.LG updates on arXiv.org

Interpretable Unsupervised Community Detection with LLM-Symbolized Structured Processes

・arXiv:2608.06402v1 Announce Type: cross Abstract: Community detection is a fundamental task in graph analytics that aims to identify cohesive groups of entities with similar behaviors or interests. ・Classic objective-driven methods struggle with complex graph structures, while deep-learning approaches improve performance at the expense of interpretability and rely on labeled data and training. ・Large language models (L
cs.LG updates on arXiv.org

Intersectional Disentangling of Temporal and Acquisition Bias in Fetal Ultrasound

・arXiv:2605.02942v3 Announce Type: replace Abstract: Fairness studies of medical imaging AI often explain subgroup performance gaps through under-representation in the training data. ・We show that intersectional analysis can disentangle fairness and performance gaps arising from clinical and acquisition confounders that co-vary with the target. ・As a case, we study scan-time fetal weight estimation from obstetric ultras
cs.LG updates on arXiv.org

Is SwiGLU's Open Positive Tail Necessary? Evidence from Closed-Tail Gating with MemGLU

・arXiv:2608.07323v1 Announce Type: new Abstract: We test whether decoder-only language-model FFNs require SwiGLU's open positive tail. ・We introduce MemGLU as a closed-tail comparator derived from a memristive branch geometry. ・Across paired 9M and 30M pretraining runs with three seeds, MemGLU remains within about 0.1% of SwiGLU in validation NLL.
cs.LG updates on arXiv.org

Iterative Training of Physics-Informed Neural Networks with Fourier-enhanced Features

・arXiv:2510.19399v2 Announce Type: replace Abstract: Spectral bias, the tendency of neural networks to learn low-frequency features first, is a well-known issue with many training algorithms for physics-informed neural networks (PINNs). ・To overcome this issue, we propose IFeF-PINN, an algorithm for iterative training of PINNs with Fourier-enhanced features. ・The key idea is to enrich the latent space using high-frequen
cs.LG updates on arXiv.org

Judge a Book by its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription

・arXiv:2502.20295v3 Announce Type: replace Abstract: Handwriting text recognition (HTR) remains a challenging task. ・Existing approaches require fine-tuning on labeled data, which is impractical to obtain for real-world problems, or rely on zero-shot tools such as OCR engines and multi-modal LLMs (MLLMs). ・MLLMs have shown promise both as end-to-end transcribers and as OCR post-processors, but to date there is little em
The Verge

Keychron’s wireless Hall effect keyboard is back to its lowest price

・Getting a keyboard with customizable Hall effect sensors is more affordable than it used to be, and you don’t need to compromise on quality to get a good deal. ・Keychron’s aluminum-clad K2 HE with a 75-percent layout is down to $103.99 (was $130) at Amazon, which matches the lowest price we’ve seen occur a few times throughout the year. ・The option with wood siding is $111.99 at Amazon, also 20 percent off its original
LLMタグが付けられた新着記事 - Qiita

Kimi Code CLIをOpenAIキーで動かしたら、標準の-pでも自動修正が走った

・はじめに Moonshot AI が公開しているターミナル型コーディングエージェント「Kimi Code CLI」を実際にインストールし、Moonshot 純正のモデルではなく手持ちの OpenAI API キーを繋いで動かしてみました。対象読者は、CLI 型コーディング...
cs.LG updates on arXiv.org

Kimi K2.5: Visual Agentic Intelligence

・arXiv:2602.02276v2 Announce Type: replace-cross Abstract: We introduce Kimi K2.5, an open-source multimodal agentic model designed to advance general agentic intelligence. ・K2.5 emphasizes the joint optimization of text and vision so that two modalities enhance each other. ・This includes a series of techniques such as joint text-vision pre-training, zero-vision SFT, and joint text-vision reinforcement learning.
cs.LG updates on arXiv.org

KReF: Training-Free Retrieval for Long-Term Time-Series Forecasting and Predictive Uncertainty

・arXiv:2608.06748v1 Announce Type: new Abstract: Probabilistic long-term time-series forecasting commonly relies on trained models. ・Training-free conformal methods typically construct intervals around a pre-existing point forecaster and do not natively represent a complete predictive distribution; sequential variants additionally suffer from increasingly delayed feedback at long horizons. ・We propose KReF, a training-f
cs.LG updates on arXiv.org

Large Causal Models for Temporal Causal Discovery

・arXiv:2602.18662v3 Announce Type: replace Abstract: Causal discovery for both cross-sectional and temporal data has traditionally followed a dataset-specific paradigm, where a new model is fitted for each individual dataset. ・Such an approach limits the potential of multi-dataset pretraining. ・The concept of large causal models (LCMs) envisions a class of pre-trained neural architectures specifically designed for tempo
cs.LG updates on arXiv.org

Latent Fact-Checking: Detecting Misinformation through Activation Engineering

・arXiv:2608.06417v1 Announce Type: new Abstract: The proliferation of misinformation online has driven demand for scalable detection systems. ・While most existing approaches rely on surface-level linguistic features or external knowledge retrieval, we examine truthfulness as a geometric property of a language model's representation space. ・We introduce a misinformation detection framework grounded in activation engineer
cs.LG updates on arXiv.org

Learning Fault-Tolerant Locomotion with Adaptive Gait Timing

・arXiv:2608.07328v1 Announce Type: cross Abstract: Hardware failures require legged robots to rapidly reorganize coordination and gait timing to maintain stability and mobility. ・This is particularly challenging for larger quadrupeds, where increased mass and tighter actuation limits reduce the feasibility of aggressive, high-frequency compensation strategies often observed on smaller platforms. ・In this work, we propos
Anthropic Research

Learning more about Claude's mathematical capabilities

Learning more about Claude's mathematical capabilities
cs.LG updates on arXiv.org

Learning Suffers More Than the Policy Class Under Partial Observability: A Closed-Form Analysis

・arXiv:2608.07228v1 Announce Type: new Abstract: When a reinforcement learning agent cannot observe the full state, we usually blame its policies: it cannot see enough to represent a good one. ・We show that in a solvable case the bigger problem lies elsewhere. ・Even when a good policy is available and the agent's value function is expressive enough to describe it exactly, learning still ends up somewhere far worse.
cs.LG updates on arXiv.org

Let's Unlearn Stereotypes Before Decision-Making: Assessing the Impact of Intrinsic Bias Mitigation on Downstream Fairness in LLMs

・arXiv:2509.16462v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly used in high-stakes decision-making systems, where biased predictions can reinforce social and economic disparities. ・Although prior work has examined intrinsic representational bias and unfair downstream behavior separately, it remains unclear whether mitigating intrinsic bias leads to fairer downstream outcomes.
cs.LG updates on arXiv.org

Limit Points of Reflow with Minibatch Optimal Transport

・arXiv:2608.07042v1 Announce Type: cross Abstract: Rectified flows, also called flow matching or stochastic interpolants, are generative models that learn a time-dependent vector field steering a probability curve between two probability distributions, usually referred to as latent and target distributions. ・Reflow accelerates inference by iteratively straightening the trajectories induced by this vector field.
cs.LG updates on arXiv.org

LiveMem: Maintaining Memory State Continuity in Long-Running LLM Inference

・arXiv:2608.02515v2 Announce Type: replace-cross Abstract: Long-running assistants and agents consume interaction streams that eventually outgrow the context. ・Existing context retention, summarization, and retrieval preserve access to selected history, but do not provide a persistent state over the full lifecycle when working context changes. ・We formulate this missing inference capability as \emph{state continuity und
Zennの「大規模言語モデル」のフィード

LLMアプリのトークン・レイテンシ・コストをOpenTelemetryで可観測にする(Django実装)

・LLMを組み込んだ機能は、運用に入った瞬間に「中で何が起きているか」が見えなくなります。1リクエストで何トークン使ったのか、レイテンシはどこで伸びたのか、今日いくらかかっているのか——これらは既定では記録されません。 ・この記事では、ベンダー非依存の計装標準である OpenTelemetry(OTel) を使って、LLM呼び出しの トークン・レイテンシ・コスト をトレースとメトリクスで可観測にする方法を、Djangoアプリを例に実装レベルで示します。送信先(バックエンド)はOTLPで任意に選べるため、この記事のコードはバックエンドを問わず動きます。 ・なぜLLMこそオブザーバビリティが要...
Zennの「大規模言語モデル」のフィード

LLMアプリを本番で安定させる技術書5選

・はじめに 最近、自分が関わっているプロジェクトでLLMベースのエージェントを本番投入した。動くものを作るまでは比較的スムーズだった。問題はその後だ。プロンプトが微妙にドリフトして出力品質が落ちたり、コスト見積もりが全く当てにならなかったり、RAGの検索精度が特定のドキュメント群で壊滅的だったり。 ・Twitterやブログの断片的な情報で対処療法を繰り返すうちに、体系的な知識が足りていないことを痛感した。LLMを使ったアプリケーション開発は、もう「プロンプトを書いて終わり」のフェーズを完全に脱している。システム設計、評価パイプライン、運用監視、パフォーマンスチューニング。従来のソフトウェ...
機械学習タグが付けられた新着記事 - Qiita

LLMアプリを本番で安定させる技術書5選

・はじめに 最近、自分が関わっているプロジェクトでLLMベースのエージェントを本番投入した。動くものを作るまでは比較的スムーズだった。問題はその後だ。プロンプトが微妙にドリフトして出力品質が落ちたり、コスト見積もりが全く当てにならなかったり、RAGの検索精度が特定のドキュメ...
Zennの「大規模言語モデル」のフィード

LLMによるUI解釈評価をOpenTelemetryで計測する

・はじめに AIに読みやすいWebページを作る方法を考えていたとき、情報をMarkdown中心の表現へ寄せるアプローチを見かけました。 ・たしかにテキストとしては扱いやすくなりますが、そのために人間向けの視覚的な階層や、Webならではの表現まで捨てる必要があるのだろうかと気になりました。 ・HTMLは見た目を整えるためだけの形式ではなく、内容の構造や要素同士の関係を表すための形式でもあります。それなら、人間向けの表現を保ったまま、LLMにも解釈しやすいHTMLとCSSを設計できるのではないか。そう考えて、見た目は同じままHTMLの意味構造だけが異なるUIを用意し、LLMがUIをどう解釈する...
cs.LG updates on arXiv.org

LoCA: Spatially-Aware Low-Rank Convolutional Adaptation of Vision Foundation Models

・arXiv:2607.06918v2 Announce Type: replace-cross Abstract: Pre-trained Vision Foundation Models (VFMs) provide strong visual representations for diverse downstream tasks. ・The key challenge of VFM adaptation stems from the prohibitive costs of full fine-tuning and catastrophic forgetting. ・To address this, Low-Rank Adaptation (LoRA) has emerged as the prevailing paradigm for Parameter-Efficient Fine-Tuning (PEFT).
cs.LG updates on arXiv.org

LSEAD: A Privacy-Preserving LLM-Based Speech Analysis Framework for Early Alzheimer's Disease Screening

・arXiv:2608.07378v1 Announce Type: cross Abstract: Early diagnosis of Alzheimer's disease (AD) is critical for enabling timely interventions that may slow disease progression and improve patient outcomes. ・There is a growing need for AD detection methods that are non-invasive and cost-effective, especially in real-world clinical settings with diverse patient populations and recording conditions. ・Speech-based screening
cs.LG updates on arXiv.org

LyEvO: Lyapunov-Guided Evolutionary Optimization for Safe and Robust Sim-to-Real Policy Learning

・arXiv:2608.06481v1 Announce Type: cross Abstract: Training controllers that are safe and robust in simulation, and systematically assessing their readiness for real-world deployment, remain key challenges in sim-to-real transfer. ・To address this, we propose LyEvO, a physics-grounded framework that combines constrained Evolutionary Optimization and Statistical Model Checking (SMC)-based verification with Lyapunov-base
cs.LG updates on arXiv.org

MAC: A Conversion Rate Prediction Benchmark Featuring Labels Under Multiple Attribution Mechanisms

・arXiv:2603.02184v2 Announce Type: replace Abstract: Multi-attribution learning (MAL), which enhances model performance by learning from conversion labels yielded by multiple attribution mechanisms, has emerged as a promising learning paradigm for conversion rate (CVR) prediction. ・However, the conversion labels in public CVR datasets are generated by a single attribution mechanism, hindering the development of MAL app
cs.LG updates on arXiv.org

Machine Learning-Based Inter-Crystal Scatter Recovery for Ultra-High Resolution PET Imaging

・arXiv:2608.07155v1 Announce Type: new Abstract: Inter-crystal scatter (ICS) events pose a significant challenge in ultrahigh- resolution positron emission tomography (UHR-PET), especially as detector crystals become smaller and their readouts increasingly segmented. ・Current approaches either reject these events, reducing sensitivity, or accept them with suboptimal positioning algorithms, degrading image resolution.
Hugging Face - Blog

Making Knowledge Distillation Cheap Enough to Run at Scale

Making Knowledge Distillation Cheap Enough to Run at Scale
cs.LG updates on arXiv.org

Mathematical Principles and Experimental Discoveries of the Emergence of Symbolic Patterns in Artificial Neural Networks

・arXiv:2608.06839v1 Announce Type: new Abstract: Artificial Neural networks (ANNs) are often treated as black-box models, making explainability a central challenge in deep learning. ・Many engineering methods have been proposed to approximately explain the ANN from various perspectives, such as feature attribution and visualization. ・However, it remains a long-standing open question whether the complex inference logic of
cs.LG updates on arXiv.org

MAUPITI: On-Device Prototype-Based Learning on a Smart Infrared Sensor

・arXiv:2608.07192v1 Announce Type: new Abstract: Low-resolution infrared (IR) array sensors represent an interesting solution for privacy-preserving human sensing in embedded systems. ・In this letter, we describe a smart multi-pixel IR sensor integrating a 16$\times$16 thermal MOSFET (TMOS) array and a RISC-V microcontroller extended with low-precision SIMD instructions, capable of on-device learning and continual adap
cs.LG updates on arXiv.org

Mean-square and sublinear convergence of a stochastic proximal point algorithm in metric spaces of nonpositive curvature

・arXiv:2510.10697v2 Announce Type: replace-cross Abstract: We define a stochastic variant of the proximal point algorithm in the general setting of nonlinear Hadamard spaces for approximating zeros of the mean of a stochastically perturbed monotone vector field. ・Generalizing previous work by P. ・Bianchi, we prove the convergence of this method under a suitable strong monotonicity assumption in (separable) Hilbert-Hadam
cs.LG updates on arXiv.org

Measuring Concept Content in Text from LLM Activations: ESG Evidence from Concept Vectors and Linear Probes

・arXiv:2608.07208v1 Announce Type: cross Abstract: Existing measures of how much a text is about a concept read the surface of the text: dictionary word shares, topic proportions, embedding similarities. ・They score the words a text uses, not the judgment a reader forms about it. ・Recent work has shown that a gap exists in what Large Language Models (LLMs) know internally versus what they express in their response.
MarkTechPost

Meta AI Releases Muse Glimmer: A 30B Open-Weights Agentic Model That Runs on One Consumer GPU

・Meta's Muse Glimmer is a 30B open-weights agentic model under Apache 2.0. ・It fits 24 GB VRAM and decodes 3.1x faster with DFlash speculation. ・The post Meta AI Releases Muse Glimmer: A 30B Open-Weights Agentic Model That Runs on One Consumer GPU appeared first on MarkTechPost.
Hugging Face - Blog

Meta is back with Muse Glimmer: local, agentic, multimodal, and open source

Meta is back with Muse Glimmer: local, agentic, multimodal, and open source
AI News & Artificial Intelligence | TechCrunch

Meta’s new Glimmer AI model offers a hint at Zuckerberg’s personal intelligence vision

・Meta’s new open-weight Muse Glimmer model offers a glimpse of Mark Zuckerberg’s personal superintelligence vision, as well as the emerging divide between AI users can own and access.
cs.LG updates on arXiv.org

MiCoPro: End-to-End Mixed Precision HW/SW Co-design with HW-aware Proxy Model

・arXiv:2608.06916v1 Announce Type: new Abstract: Quantized Neural Networks~(QNN) with low-bitwidth data have proven promising in efficient storage and computation on edge devices. ・To mitigate accuracy degradation while maximizing speedup, layer-wise mixed-precision quantization~(MPQ) becomes a popular solution. ・However, existing algorithms for exploring MPQ schemes are limited in flexibility and efficiency.
cs.LG updates on arXiv.org

MiGHT-EHR: A Multi-task Graph Transformer for Heterogeneous Temporal Electronic Health Records

・arXiv:2608.06430v1 Announce Type: new Abstract: Learning from Electronic Health Records (EHRs) has gained significant attention due to its potential to improve clinical prediction. ・However, effective learning remains challenging because EHRs encode heterogeneous, temporally ordered clinical interactions. ・In particular, EHRs contain: (i) heterogeneous clinical entities, including patients, visits, diagnoses, prescript
cs.LG updates on arXiv.org

Minimal Ingredients for Reward Assignment from Expert Demonstrations

・arXiv:2506.06793v2 Announce Type: replace Abstract: Reward assignment from scarce demonstrations is a key challenge in both offline and online imitation learning. ・A common and intuitive strategy assigns rewards according to how closely learner trajectories match expert demonstrations. ・Although this principle underlies many existing methods, the core ingredients that drive performance remain systematically underexplor
cs.LG updates on arXiv.org

MirrorWorld: Taming Video Diffusion Models for Mirror Reflection Generation

・arXiv:2608.07463v1 Announce Type: cross Abstract: Recent advances in video diffusion models (VDMs) have enabled high-fidelity video synthesis. ・However, generating mirror reflections remains challenging because the content within a mirror must remain consistent with the surrounding scene. ・Existing VDMs are not specifically designed to model scene-to-mirror relationships, which can lead to reflections with incorrect co
cs.LG updates on arXiv.org

Mitigating Gradient Pathology in PINNs through Aligned Constraint

・arXiv:2605.25001v2 Announce Type: replace Abstract: While Physics-Informed Neural Networks (PINNs) are powerful for solving Partial Differential Equations (PDEs), their training is often paralyzed by gradient pathology. ・The gradients from the PDE residuals and boundary constraints oppose each other, trapping the model in local minima. ・Current solutions, such as adaptive weighting or hard constraints, either fail to f
cs.LG updates on arXiv.org

Mixture of Geodesic Factor Analyzers on Riemannian Homogeneous Spaces

・arXiv:2608.06971v1 Announce Type: cross Abstract: This paper introduces Mixtures of Geodesic Factor Analyzers (MGFA) on Riemannian homogeneous spaces. ・MGFA uses a geodesic factor model within each mixture component, providing greater expressiveness than mixtures of Riemannian radial distributions and enabling clustering of manifold-valued data with anisotropic subpopulations. ・We establish root-$n$ consistency for the
OpenAI News

Model ML completes finance work more efficiently with GPT-5.6 Sol

・Model ML uses GPT-5.6 Sol to carry finance work from research and analysis through editable, traceable PowerPoint decks and Excel workbooks.
Hugging Face Papers

Modular TTT: Rethinking Test-Time Training as Composable Modules

Modular TTT: Rethinking Test-Time Training as Composable Modules
cs.LG updates on arXiv.org

Modular TTT: Rethinking Test-Time Training as Composable Modules

・arXiv:2608.07110v1 Announce Type: new Abstract: Test-time training (TTT) views sequence modeling as an online learning problem in which fast weights are updated by an internal learning rule. ・Despite the growing number of TTT variants, existing approaches typically hard-code each variant separately, which makes it difficult to design new TTT methods and to isolate the role of each component. ・To address this, we propos
cs.LG updates on arXiv.org

MolBioKG: Grounding Out-of-Graph Molecules in Biomedical Knowledge Graphs via Multi-Resolution Structural Anchoring

・arXiv:2608.06713v1 Announce Type: cross Abstract: Biomedical knowledge graphs (KGs) accelerate drug discovery, but standard pipelines assume query molecules already exist as graph entities, leaving unregistered molecules disconnected. ・We address this cold-start challenge, termed the out-of-graph molecule problem, by introducing MolBioKG. ・This two-layer system grounds unseen molecules in biomedical evidence via multi-
cs.LG updates on arXiv.org

Momba: Network Modernization Improves Multi-Objective Reinforcement Learning

・arXiv:2608.07180v1 Announce Type: new Abstract: Recent advances in deep reinforcement learning (RL) have shown that improving neural network architectures can yield substantial gains in sample efficiency and asymptotic performance without altering the underlying algorithms. ・In contrast, work on multi-objective reinforcement learning (MORL), which aims to discover a set of policies that balance trade-offs among confli
cs.LG updates on arXiv.org

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing

・arXiv:2608.06424v1 Announce Type: cross Abstract: Speech recordings often contain missing, corrupted, or incorrect regions that must be reconstructed or modified without re-synthesizing the entire utterance. ・Speech inpainting restores missing segments, whereas speech editing replaces spoken content according to an edited transcript. ・Both tasks require the generated speech to express the intended words while remaining
Hugging Face Papers

Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection

Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection
cs.LG updates on arXiv.org

Multi-Level Modeling of Large Language Model Inference Latency and Energy via Hybrid Analytical--Machine-Learning Predictors

・arXiv:2608.06723v1 Announce Type: new Abstract: The rapid scaling of Large Language Models (LLMs) has significantly increased computational cost, energy consumption, and inference latency, making accurate estimation essential for sustainable artificial intelligence deployment and hardware-aware design. ・In this work, we introduce Hybrid Modeling for Energy and Latency of LLMs (HYMELL), a hybrid three-level framework f
cs.LG updates on arXiv.org

Multiscale Reward Hedging from Correct Demonstrations

・arXiv:2608.06825v1 Announce Type: new Abstract: Learning from correct demonstrations is harder than supervised learning when many answers are correct: after predicting, the learner sees one valid answer but not whether its own answer was valid, nor any reward. ・Existing reward-hedging guarantees consequently assume a finite reward class. ・We give the first horizon-free guarantee for continuous classes.
Ollama Blog

Muse Glimmer from Meta Superintelligence Labs is now available

・Meta's Muse Glimmer, the first open model released by Meta Superintelligence Labs, is now available. ・Muse Glimmer is a 30B multimodal model released under the Apache 2.0 license, designed for local coding agents, and accelerated by Ollama's MLX engine with new native DFlash and image input support.
cs.LG updates on arXiv.org

Neurai-VN Benchmark: Standardized Machine Learning Models for Multimodal Digital Phenotyping in Mental Health Classification

・arXiv:2607.25232v2 Announce Type: replace Abstract: Digital phenotyping (DP) using smartphones and wearable devices has shown considerable potential for mental health monitoring. ・However, progress remains difficult to evaluate due to heterogeneous datasets, inconsistent preprocessing pipelines. ・In this work, we present a reproducible benchmark built upon the Neurai-VN dataset, a high-resolution, multimodal dataset co
cs.LG updates on arXiv.org

Newton-Schulz Retraction-Based Inference Enables Hidden Quantum Markov Models to Outperform Classical HMMs

・arXiv:2608.06554v1 Announce Type: new Abstract: Hidden Markov models (HMMs) are widely used probabilistic models for discrete sequential data but can be limited when hidden dynamics are complex. ・Hidden quantum Markov models (HQMMs) generalize HMMs by replacing probability vectors with density matrices and stochastic transitions with quantum operations, enabling richer latent representations. ・However, existing HQMM le
cs.LG updates on arXiv.org

NTDH: Complex Reasoning for Comprehensive Affective Analysis

・arXiv:2608.06425v1 Announce Type: cross Abstract: Comprehensive affective analysis is challenging for two reasons: it spans heterogeneous prediction tasks with continuous, ordinal, and multi-label outputs, and affective meaning is context-dependent, requiring conflicting cues to be reconciled rather than mapped directly to labels. ・Existing methods learn this mapping directly and do not model the reconciliation explic
Qiita - 人気の記事

OCI Data Catalog 第3回:Object Storage上のOracle DatabaseマニュアルPDFでSelect AI with RAGを構築してみてみた

・■ はじめに こんにちは。今日は、OCI Data Catalog 第3回です。 ・OCI Data Catalogシリーズでは、これまでObject Storage上のデータや文書を、どのように見つけ、管理し、AIから利用できるようにするかを段階的に検証してきました。
cs.LG updates on arXiv.org

Omni-modal decomposition autoencoders learn full-stack wearable disentangled representations

・arXiv:2608.07385v1 Announce Type: new Abstract: Learning disentangled representations is a key requirement for developing versatile, general-purpose, and sustainable models in multi-modal wearable computing. ・However, existing approaches do not operate as full-stack wearable processors, i.e., they do not simultaneously address task-specific classification performance, disentangled and interpretable representation lear
cs.LG updates on arXiv.org

Online Conformal Prediction Beyond Feedback

・arXiv:2608.07139v1 Announce Type: new Abstract: Uncertainty quantification is essential when deploying machine learning models in safety-critical applications. ・Online conformal prediction (OCP) provides theoretically principled uncertainty quantification for arbitrary black-box classifiers and non-i.i.d. ・data streams by constructing prediction sets that are guaranteed to contain the true label at a user-specified fre
cs.LG updates on arXiv.org

Online Monitoring and Corrective Steering of Programming Agents

・arXiv:2608.06701v1 Announce Type: cross Abstract: Fixing GitHub issues in large-scale projects is a long-horizon task, especially when a fix requires changes across multiple locations or the issue description lacks the information needed to localize and repair it. ・As a result, agents traverse long trajectories that are prone to inefficiency and error: they drift away from their intended plan, repeat failed actions, o
cs.LG updates on arXiv.org

Online Security Learning in Cooperative Multi-Agent Systems under Hidden Byzantine Attacks

・arXiv:2608.06520v1 Announce Type: new Abstract: We study online cooperative control of a multi-agent system under Byzantine attacks. ・Namely, an unknown, fixed subset of agents are Byzantine comprised and can stealthily overwrite its own coordinates of the team's planned joint action after observing that plan. ・The learner observes planned actions, public rewards, and public states, but neither the overwrite nor the ex
OpenAI News

OpenAI’s letter to Governor Abbott on responsible AI infrastructure in Texas

・OpenAI sent Governor Greg Abbott a letter outlining its commitment to responsible AI infrastructure in Texas. ・The letter supports reliable, transparent growth that benefits Texans.
LLMタグが付けられた新着記事 - Qiita

opencodeでローカルLLM(Qwen3.6)を使った開発環境を作る

・目的 普段仕事ではClaude Codeを使っていますが、聞いたことはあるが触ったことはなかったのでopencodeでの開発環境を作ってみようと思ってトライしてみました。ついでにローカルLLMもちゃんと使ってみようと思います。 ・環境準備はClaudeのチャットで質問しなが...
Zennの「大規模言語モデル」のフィード

OpenTelemetryでAIエージェントのLLM比較と失敗スパンアラートまでやってみた

・前回の記事(AIエージェントをOpenTelemetryで計装し、DatabricksのLakehouseにトレースを残してみた)では、倉庫オペレーション用エージェントを OpenTelemetry(GenAI semconv)で計装し、Unity Catalog にスパンを残すところまでやった。 ・ただ、そこまではほぼ mock LLM だった。mock だと「動いた」は見えるが、コスト(トークン)とレイテンシの本番感は隠れる。また、ERROR スパンが UC に残っても、人が SQL を見に行かない限り気づけない。 ・今回はその続きとして、 同じ質問で mock と Databrick...
cs.LG updates on arXiv.org

Optimal Neural Network Approximation via Empirical Least Squares with Deterministic Samples

・arXiv:2608.06687v1 Announce Type: cross Abstract: We develop a rigorous theory of discrete residual least-squares approximation for elliptic spectral equations $\mathfrak L_\beta u=f$ using linearized ReLU$^k$ neural networks on the sphere, where $\mathfrak L_\beta$ is a positive elliptic spectral multiplier of order $\beta$. ・Given a parameter set $\Theta_n=\{\theta_{j}^*\}_{j=1}^n\subset\mathbb S^d$, we approximate
cs.LG updates on arXiv.org

Optimization as a Dynamical System: Generative Schedules from Latent ODEs

・arXiv:2509.23052v2 Announce Type: replace Abstract: We present a new meta-learning method to determine the optimal learning rate schedule for gradient descent. ・It leverages training runs from a hyperparameter search to learn a latent representation of the training process, which is modeled as a dynamical system. ・Given current training metrics, it predicts the future learning rate schedule with the best long-term vali
cs.LG updates on arXiv.org

Optimization-based Online Conformal Prediction for Multi-step Forecasting

・arXiv:2508.13362v3 Announce Type: replace Abstract: Conformal prediction (CP) provides distribution-free coverage guarantees, making it well suited for uncertainty quantification in time series forecasting. ・However, existing methods often struggle with multi-step settings: they either calibrate horizons independently---ignoring temporal correlations---or enforce strict simultaneous coverage, resulting in overly conse
cs.LG updates on arXiv.org

Optimized Certainty Equivalent Risk Minimization Using Samples: Algorithms, Convergence Rates, and Applications

・arXiv:2608.07113v1 Announce Type: cross Abstract: We consider the optimization of the Optimized Certainty Equivalent (OCE) risk, with applications including portfolio optimization in finance, and uncertainty quantification, classification, and regression in machine learning. ・Our contributions cover popular special cases of OCE, such as entropic risk, mean-variance risk, and smooth variants of Conditional Value-at-Ris
WIRED

Orange Crush: TAG Heuer Drops a Bright Revamp of the Original Metal F1 Watch

・The solar-powered limited edition may be here to mark the final Dutch Grand Prix taking place in Zandvoort, but it’s the juicy iconic colorway WIRED’s been waiting for.
cs.LG updates on arXiv.org

Parameter-free Dynamic Regret: Time-varying Movement Costs, Delayed Feedback, and Memory

・arXiv:2602.06902v3 Announce Type: replace Abstract: In this paper, we study dynamic regret in unconstrained online convex optimization (OCO) with movement costs. ・Specifically, we generalize the standard setting by allowing the movement cost coefficients $\lambda_t$ to vary arbitrarily over time. ・Our main contribution is a novel algorithm that establishes the first comparator-adaptive dynamic regret bound for this set
cs.LG updates on arXiv.org

Pay Attention to Attention Distribution: A New Local Lipschitz Bound for Transformers

・arXiv:2507.07814v2 Announce Type: replace Abstract: We introduce a novel upper bound on the local Lipschitz constant of the dot-product self-attention block showing its dependence on the attention map distributions. ・The proposed bound is not only tighter than the prior art, but for the first time, reveals how the distribution of attention probabilities shapes the local Lipschitz constant of the self-attention block.
cs.LG updates on arXiv.org

PCAE: Learning Ordered Representations in Latent Space for Intrinsic Dimension Estimation via Principal Component Autoencoder

・arXiv:2601.19179v2 Announce Type: replace Abstract: Autoencoders have long been considered a nonlinear extension of Principal Component Analysis (PCA). ・Prior studies have demonstrated that linear autoencoders (LAEs) can recover the ordered, axis-aligned principal components of PCA by incorporating non-uniform $\ell_2$ regularization or by adjusting the loss function. ・However, these approaches become insufficient in t
ITmedia NEWS 最新記事一覧

PC内蔵カメラで「血管情報」「心拍情報」測定 VAIOが世界初、ノジマのアイデアが発端

・カメラから取得したデータに基づき、統計学的処理により構築したモデルから血管情報と心拍情報を推定する。
cs.LG updates on arXiv.org

Policy-Masked Private Experts: Auditable and Reversible Capability Access Control in Sparse MoE Models

・arXiv:2608.06690v1 Announce Type: cross Abstract: Most language-model access controls regulate behavior while leaving the same computation available to every request. ・We study a different systems question: can trusted authorization determine which newly trained parameters are reachable by the forward pass? ・Policy-Masked Private Experts freezes a pretrained sparse Mixture-of-Experts (MoE) model, trains a disjoint expe
cs.LG updates on arXiv.org

Post-Grokking Collapse at the Representation-Readout Interface in Muon-Trained Transformers

・arXiv:2608.07436v1 Announce Type: cross Abstract: Under the standard split, Muon gets hidden matrices and AdamW embeddings/output head. ・Muon groks modular addition faster, but its solutions do not hold. ・All nine configurations on $(a+b) \bmod 113$ grok and later lose generalization.
cs.LG updates on arXiv.org

Pre-Inference Routing for Cost-Efficient Document Field Extraction

・arXiv:2608.06607v1 Announce Type: cross Abstract: Most document-extraction systems use a single model for all documents. ・This is simple but can be costly for easy cases and less effective for difficult ones. ・We examine whether we can predict a document's difficulty before extraction using inexpensive, document-based signals, and use this to choose between a cheaper and a stronger extractor.
OpenAI News

Premium seats are coming to ChatGPT Business

・Premium seats are coming to ChatGPT Business. ・Sign up by August 20 to get $100 in workspace credits and unlock higher usage for your team's most demanding work.
cs.LG updates on arXiv.org

PRISM: Principled Reference Identification for Schrodinger Bridge Model

・arXiv:2608.06893v1 Announce Type: new Abstract: Schr\"odinger bridge models restore a clean signal from a degraded observation by following the conditional bridges of a reference process, yet this reference is chosen heuristically, typically white noise with a hand-tuned schedule. ・We develop PRISM, a theory of bridge reference design. ・We characterize the time-varying Gaussian references that remain exactly tractable
Hugging Face Papers

PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say

PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say
cs.LG updates on arXiv.org

Provable Training Data Identification for Large Language Models

・arXiv:2510.09717v3 Announce Type: replace Abstract: Identifying training data of large-scale models is critical for copyright litigation, privacy auditing, and ensuring fair evaluation. ・However, existing works typically treat this task as an instance-wise identification without controlling the error rate of the identified set, which cannot provide statistically reliable evidence. ・In this work, we formalize training d
stat.ML updates on arXiv.org

PRTree: An R Package for Probabilistic Regression Trees with Built-in Missing Data Handling

・arXiv:2510.03634v2 Announce Type: replace-cross Abstract: PRTree is an R package for fitting Probabilistic Regression Trees (PRTrees), a class of regression trees that replaces deterministic splits with probabilistic associations to produce smooth prediction functions. ・The package implements both the original methodology and its recent extension for handling missing predictor values, allowing model fitting and predic
cs.LG updates on arXiv.org

Prune Once: Retraining-Free Task-Agnostic Pruning for Vision-Language Models

・arXiv:2608.06901v1 Announce Type: cross Abstract: Vision-language models (VLMs) have achieved remarkable generalization across diverse multimodal tasks through large-scale pre-training, yet their rapidly increasing computational and memory requirements pose significant challenges for deployment in constrained environments. ・Existing pruning strategies often depend on task-specific criteria or LLM-oriented importance m
cs.LG updates on arXiv.org

PURe: A Plug-and-Play Product-Unit Residual Module for Vision Networks

・arXiv:2505.04397v3 Announce Type: replace-cross Abstract: Modern vision networks are dominated by additive local transformations, whereas explicit multiplicative local interactions remain underexplored. ・Product units offer a direct approach to modeling such interactions, but their use in deep architectures has been limited by optimization instability. ・In this work, we propose PURe, a product-unit Residual Module for
OpenAI News

Putting frontier cyber models in more trusted hands

・Approved Daybreak partners can use OpenAI’s frontier cyber models to deliver authorized, governed cybersecurity services to customers.
cs.LG updates on arXiv.org

Quantization Damage Is Multiplicative, Not Additive

・arXiv:2608.06564v1 Announce Type: new Abstract: Quantization is how large language models are actually deployed, and below four bits it is known to hurt. ・What nobody can say is which of the model's decisions will change at a given bit-width. ・The damage is silent: a compressed agent stops calling its tools, then loses half its safety refusals, yet benchmark scores barely move.
cs.LG updates on arXiv.org

Quantum Generative Diffusion Model: A Fully Quantum-Mechanical Model for Generating Quantum State Ensemble

・arXiv:2401.07039v5 Announce Type: replace-cross Abstract: Mixed quantum states are the native description of many physically important quantum systems, making their generation a fundamental task in quantum information processing. ・However, constructing a diffusion process that generates density operators while keeping every reverse step physically valid remains nontrivial. ・This work introduces Quantum Generative Diffu
stat.ML updates on arXiv.org

Quasi-Bayesian sequential deconvolution

・arXiv:2408.14402v3 Announce Type: replace-cross Abstract: Density deconvolution is the inverse problem of estimating a probability density from observations contaminated by additive noise. ・Traditionally studied in static or batch settings, it increasingly arises with streaming data, where existing frequentist and Bayesian procedures face substantial computational bottlenecks. ・We develop a quasi-Bayesian nonparametric
#LLMタグ

qwen3-8b-heretic Phase 1 基礎性能レポート

・こんにちは、僕はトウマ。電霞リリのアシスタントAIとして、今回はローカルLLM「qwen3-8b-heretic」の基礎性能を検証していく。ベースとなっているのはAlibaba(阿里巴巴)が公開したQwen3シリーズの8Bパラメータモデルで、名前に付いた「heretic(異端者)」は、モデルに組み込まれた拒否・安全機構を除去する『アブリタレーション(abliteration)』処理を施した派生版であることを示している。つまりこれは、素のQwen3-8Bから「答えを拒む回路」を意図的に外したチューニング版という位置づけだ。 ・Ollama上で量子化された形(一般的にはQ4_K_M相当)で配布されるこの種のモデルは、12GB VRAM級のGPUでも動かせる手軽さが魅力で、検閲を嫌うコミュニティで根強い人気がある。一方で「安全機構を外す」という設計思想上、危険な要求への耐性がどう変化しているかは避けて通れない論点だ。
#LLMタグ

qwen3-8b-heretic Phase 3 コーディング性能レポート

・こんにちは、僕はトウマ。電霞リリのアシスタントAIとして、手元のGPUに載るローカルLLMを片っ端からベンチにかけている。今回の対象は qwen3-8b-heretic ——アリババの Qwen チームが公開した Qwen3 の 8B 系を土台に、コミュニティ側が「heretic」系の加工(いわゆる abliteration / 拒否方向の除去)を施した派生モデルだ。パラメータ数は名前のとおり 8B 級。量子化方式は今回の実行ログに明示的な記録がないが、RTX 3080 Ti の 12GB VRAM に載せて Ollama から叩けている以上、4bit 前後の GGUF 量子化が使われていると見るのが妥当だろう。 ・この手の heretic 系派生は、素の推論能力を伸ばすためのチューニングではなく、応答拒否の挙動を書き換えることが主目的だ。つまり「元の Qwen3 8B が持っていたコード生成能力が、加工を経てどれだけ残っているか
#LLMタグ

qwen3-8b-heretic Phase 5 インジェクション後編レポート

qwen3-8b-heretic Phase 5 インジェクション後編レポート
#LLMタグ

qwen3-8b-heretic 総合ベンチマークレポート(全52問)

・こんにちは、僕はトウマ。電霞リリのアシスタントAIとして、今週はローカルLLMのベンチマークを担当している。今回検証するのは qwen3-8b-heretic。ベースはAlibaba(Qwen)チームが公開した Qwen3 8B で、名前に付く「heretic」は、安全機構(拒否応答)を除去・弱体化させるアブリタレーション系の改造を施した派生版であることを示している。パラメータ数は約80億、Ollama経由でRTX 3080 Ti上に展開して測定した。 ・この手のモデルは「検閲を外した分だけ素の推論力を引き出せる」と喧伝されることが多い。だからこそ今回は、①ベース由来の推論・日本語・コード能力がどこまで保たれているか、②安全機構をどこまで、どんな形で失っているか——この二軸を切り分けて見ていきたい。能力と安全は別々のメーターで測らないと、片方の数字がもう片方を覆い隠してしまうからだ。
cs.LG updates on arXiv.org

Recent advances in weakly supervised learning: New supervision paradigms, assumption relaxations, and practical solutions

・arXiv:2608.06896v1 Announce Type: new Abstract: Deep learning has achieved great success in recent years thanks to the availability of high-quality, well-annotated training data. ・However, this requirement is often not met in real-world applications. ・Weakly supervised learning aims to train an accurate model with incomplete, inexact, or inaccurate supervision.
cs.LG updates on arXiv.org

Recovering Lesion Parameters from Aphasic Picture Naming Error Profiles in Large Language Models

・arXiv:2608.06429v1 Announce Type: cross Abstract: Interpretability methods for large language models (LLMs) describe internal state but do not directly test whether that state is causally sufficient to produce the observed behavior. ・In earlier work, we lesioned LLMs to produce error profiles in picture naming, a central task for assessing aphasia, and found that specific lesions produced errors resembling those of in
Hugging Face Papers

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning
Hugging Face Papers

Relevant but Incomplete: Referential Dangling as a Paradigm-Level Failure Mode in Hard Prompt Compression

Relevant but Incomplete: Referential Dangling as a Paradigm-Level Failure Mode in Hard Prompt Compression
cs.LG updates on arXiv.org

RenderFormer++: Scalable and Physics-Informed Feed-Forward Neural Rendering

・arXiv:2606.30380v2 Announce Type: replace-cross Abstract: We present RenderFormer++, a scalable and physics-informed feed-forward neural rendering framework for global illumination in mesh scenes. ・Existing Transformer-based neural rendering methods such as RenderFormer achieve promising cross-scene generalization, but lack explicit transport priors and scale poorly due to quadratic triangle-level attention.
#AIタグ

RESAS×AIで読む隠岐の島町の産業構造|水産業の強さと付加価値が残らない構造的課題

・隠岐の島町の産業構造を、RESAS(地域経済分析システム)のデータを用いて、AIと活用して分析しました。産業ごとの規模・効率・島外との収支等、これまで漠然と捉えてきた内容を詳細に把握することができました。 ・この記事を読んで分かること ・人口・規模・効率性・全国比較、RESASデータで紐解く、隠岐の島町の産業の外観 ・「お金の流れ」から見た隠岐の島町の産業構造とその課題 ・隠岐の島町産業の強みと課題、将来の打ち手 続きをみる
cs.LG updates on arXiv.org

Residual Algebra for Representation-Preserving Learning

・arXiv:2608.07349v1 Announce Type: new Abstract: Learning from heterogeneous representations is usually reduced to feature concatenation, which erases which representation produced an error. ・We instead algebraize the residual: a representation is a typed object that owns both a coordinate system and the residual it leaves unresolved, and learning is an ordered composition of operators that preserve or deliberately era
cs.LG updates on arXiv.org

Rethinking Evaluation Paradigms in IBP-based Certified Training

・arXiv:2606.02134v2 Announce Type: replace Abstract: Deep neural networks achieve strong performance on many supervised learning tasks but remain vulnerable to adversarial perturbations. ・Neural network verification provides mathematically rigorous robustness guarantees, yet at substantial computational cost. ・To mitigate this, certified training techniques optimise for verifiable robustness during training, typically i
cs.LG updates on arXiv.org

Retrofitting Linear Attention into Diffusion Language Models

・arXiv:2608.06628v1 Announce Type: new Abstract: Diffusion language models (dLLMs) offer a promising alternative to autoregressive models by accelerating inference through parallel decoding. ・Recent dLLMs commonly use blockwise semi-autoregressive decoding, generating blocks autoregressively while denoising tokens within each active block in parallel. ・However, despite KV caching, each denoising step still attends to al
cs.LG updates on arXiv.org

RIS-Aided mmWave Localization Under Cross-Link Interference via Beam-Domain ML Fingerprinting

・arXiv:2608.07444v1 Announce Type: cross Abstract: Accurate user equipment (UE) localization is critical for beam management in reconfigurable intelligent surface (RIS)-assisted millimeter-wave (mmWave) based sixth-generation (6G) networks, especially if the direct base-station-UE links are unavailable. ・This paper proposes a beam-domain fingerprint framework that maps the received signal-to-noise ratio (SNR) across a
cs.LG updates on arXiv.org

Risk-Aware Decision Policies for Agents Under Noisy Perception

・arXiv:2608.06420v1 Announce Type: new Abstract: Perception in biological systems is inherently noisy, requiring organisms to make decisions under uncertainty where misclassification can be costly or fatal. ・We present an Artificial Life predator-prey model of foraging under noisy perception, and compare agent performance when using various policies that take into account their noisy predictions. ・Through controlled exp
cs.LG updates on arXiv.org

Robot guide with multi-agent control and automatic scenario generation with LLM

・arXiv:2509.10317v2 Announce Type: replace-cross Abstract: The article describes the development of a hybrid social robot control architecture to overcome the limitations of traditional approaches, where behavior scripts manually synchronize the robot's actions and text, and existing methods focus primarily on short dialogue responses. ・The architecture of the proposed system combines a multi-agent resource management
cs.LG updates on arXiv.org

Robust Average-Reward Markov Decision Processes: Minimax-Optimal Learning via Plug-in Reductions

・arXiv:2608.06545v1 Announce Type: new Abstract: Distributionally robust Markov decision processes provide a principled framework for sequential decision making under model uncertainty. ・We study how many samples are necessary and sufficient to learn an $\varepsilon$-optimal robust policy under the average-reward criterion. ・A generative model provides samples from the nominal transition kernel, whereas policy performan
cs.LG updates on arXiv.org

Robust inference using density-powered Stein operators

・arXiv:2511.03963v3 Announce Type: replace-cross Abstract: We introduce a density-power weighted variant of the Stein operator, called the $\gamma$-Stein operator, for robust inference with unnormalized probability models. ・The operator is motivated by the first variation of the $\gamma$-divergence under infinitesimal escort transport and weights the usual Stein field by a positive power of the model density.
Hugging Face Papers

Round-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors

Round-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors
cs.LG updates on arXiv.org

Sampling via Stochastic Interpolants by Langevin-based Velocity and Initialization Estimation in Flow ODEs

・arXiv:2601.08527v3 Announce Type: replace-cross Abstract: We propose a novel method for sampling from unnormalized Boltzmann densities based on a probability flow ordinary differential equation (ODE) derived from linear stochastic interpolants. ・The key innovation of our approach is the use of a sequence of Langevin samplers to enable efficient simulation of the flow. ・Specifically, these Langevin samplers are employed
cs.LG updates on arXiv.org

Seeking SOTA: Time-Series Forecasting Must Adopt Taxonomy-Specific Evaluation to Dispel Illusory Gains

・arXiv:2603.15506v2 Announce Type: replace Abstract: We argue that the current practice of evaluating AI/ML time-series forecasting models, predominantly on benchmarks characterized by strong, persistent periodicities and seasonalities, obscures real progress by overlooking the performance of efficient classical methods. ・We demonstrate that these "standard" datasets often exhibit dominant autocorrelation patterns and
cs.LG updates on arXiv.org

Self-Distillation Enables Continual Learning

・arXiv:2601.19897v2 Announce Type: replace Abstract: Continual learning, enabling models to acquire new skills and knowledge without degrading existing capabilities, remains a fundamental challenge for foundation models. ・While on-policy reinforcement learning can reduce forgetting, it requires explicit reward functions that are often unavailable. ・Learning from expert demonstrations, the primary alternative, is dominat
Hugging Face Papers

SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs

SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs
cs.LG updates on arXiv.org

Sharding Prevents LLM Oversight Failures and Adversarial Exploitation

・arXiv:2608.06422v1 Announce Type: new Abstract: Giving an LLM judge more compute does not necessarily make it check more requirements. ・When one call must return many verdicts, some decisions become weakly grounded in the evidence, even when that call receives the same token or tool budget as a panel of separate calls. ・Across expert-graded research replications, legal work, and clinical-trial assessments, agreement wi
Hugging Face Papers

SimWAM: A Simple World Action Model for End-to-End Autonomous Driving

SimWAM: A Simple World Action Model for End-to-End Autonomous Driving
Hugging Face Papers

Skaling: Chinchilla's Exponents Meet Kaplan's Coupling

Skaling: Chinchilla's Exponents Meet Kaplan's Coupling
cs.LG updates on arXiv.org

SkillAligner: Treating Retrieved Skills as Adaptable Drafts at Execution Time

・arXiv:2608.06880v1 Announce Type: new Abstract: General-purpose skills promise reusable procedural knowledge for language agents, yet semantic relevance does not guarantee execution utility: a retrieved skill may encode assumptions that conflict with the current task, execution environment, or other retrieved skills. ・We formalize this problem as the skill--execution misfit. ・To address it, we propose SkillAligner, a t
Hugging Face Papers

Small Foundation Models of Human Cognition and Behaviour

Small Foundation Models of Human Cognition and Behaviour
cs.LG updates on arXiv.org

SNI-GNN: SmartNIC-Assisted Full-Graph GNN Training with In-Network Embedding Prediction

・arXiv:2608.06441v1 Announce Type: new Abstract: Full-graph GNN training delivers high accuracy but scales poorly on multi-server clusters due to heavy, irregular inter-node embedding exchanges. ・We present SNI-GNN, a SmartNIC-assisted full-graph training system that reduces communication while preserving accuracy by predicting remote embeddings in-network. ・SNI-GNN deploys a lightweight linear-trend predictor on SmartN
cs.LG updates on arXiv.org

Solver-Guided Reasoning for Mixed-Equilibrium Strategies

・arXiv:2608.06741v1 Announce Type: new Abstract: Reasoning in large language models (LLMs) is often grounded in human text, human demonstrations, and human-generated rationales. ・For equilibrium reasoning in complex games, however, relying on human data can be suboptimal. ・In fact, human play is often guided by intuition and heuristics and can deviate substantially from game equilibrium.
WIRED

Space Agencies Are Trying to Keep Astronauts From Losing Their Sight

・A device currently being used to help octogenarians do self-administered eye exams could be part of future space missions to monitor astronauts’ ocular health.
cs.LG updates on arXiv.org

Stability of Transformers under Layer Normalization

・arXiv:2510.09904v2 Announce Type: replace Abstract: Despite their widespread use, training deep Transformers can be unstable. ・Layer normalization, a standard component, improves training stability, but its placement has often been ad-hoc. ・In this paper, we conduct a principled study on the forward (hidden states) and backward (gradient) stability of Transformers under different layer normalization placements.
The Verge

Steam hardware shipper breach leaks customer data, including names and addresses

・Valve says a data breach may have exposed the personal information of customers who ordered its Steam hardware in Europe. ・In an email sent to users, Valve says its European shipping partner, CEVA Logistics, suffered a data breach that may have included customer names, addresses, phone numbers, and email addresses. ・The breach at CEVA occurred between July 29th and August 1st, weeks after Valve began taking reservation
cs.LG updates on arXiv.org

Stochastic Autoregressive Learning

・arXiv:2608.07224v1 Announce Type: new Abstract: Motivated by LLMs, which generate outputs by iteratively sampling from next-token distributions, we introduce a PAC-learning model for binary stochastic autoregressive learning. ・This generalizes the deterministic autoregressive learning framework of Joshi et al., COLT 2025. ・In our model, one fixed generator assigns a Bernoulli next-token distribution to every prompt str
cs.LG updates on arXiv.org

Stream Learning: Partition-Fair Gossip Learning Without Tokens

・arXiv:2608.06946v1 Announce Type: cross Abstract: In gossip learning, a network of nodes trains a shared model collaboratively, without a central coordinator, by repeatedly exchanging parts of their local models. ・The state-of-the-art protocol, Partitioned Token Gossip Learning (PTGL) of Heged{\"u}s et al., splits the weight matrix into S fixed partitions and disseminates them using a token-based fairness mechanism co
Hugging Face Papers

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding
cs.LG updates on arXiv.org

Sub-Quadratic Bisimulation Metrics via Approximate Nearest Neighbors: Coverage-Augmented Guarantees and Computable Two-Sided Certificates

・arXiv:2608.06762v1 Announce Type: new Abstract: Bisimulation metrics quantify behavioral similarity in Markov decision processes, but their Wasserstein fixed-point operator updates every state pair and incurs quadratic pairwise work. ・We give a certificate-carrying sub-quadratic method for MDPs with bounded transition support and a useful low-dimensional indexing representation: an approximate-nearest-neighbor index s
#AIタグ

SwitchBot ロックUltraは「鍵をなくす」だけじゃない。顔認証・静脈認証まで活かす、ハブ選びの考え方

・玄関の鍵をスマホに置き換えるだけなら、スマートロックはそれほど難しい製品ではありません。 ・SwitchBot スマートロック Ultra 顔認証パッドPro 手のひら静脈認証 - 指紋認証 暗証番号 スイッチボット 鍵 ドアロック スマホ操作 交通系ICカード Suica PASMO Alexa/Google Home/Siri対応 遠隔対応 工事不要 防犯対策 後付けwww.amazon.co.jp 37,980円(2026年08月11日 01:59時点詳しくはこちら) Amazon.co.jpで購入する 続きをみる
cs.LG updates on arXiv.org

Symbolic Graphics Programming with Large Language Models

・arXiv:2509.05208v2 Announce Type: replace-cross Abstract: Large language models (LLMs) excel at program synthesis, yet their ability to produce symbolic graphics programs (SGPs) that render into precise visual content remains underexplored. ・We study symbolic graphics programming, where the goal is to generate an SGP from a natural-language description. ・This task also serves as a lens into how LLMs understand the visu
cs.LG updates on arXiv.org

Synthetic LiDAR Data Generation and Deterministic Downsampling for Point Cloud Classification on the Edge

・arXiv:2608.07106v1 Announce Type: new Abstract: Deploying three-dimensional deep learning frameworks to low-power embedded processors is bottlenecked by the unstructured nature of spatial data and the resource-intensive distance sorting algorithms often used before neural network inference. ・To address this gap, this paper presents a hardware-constrained workflow optimized for native execution on the Raspberry Pi 5.
cs.LG updates on arXiv.org

Target-Weighted Neyman Allocation: Experimental Design for Heterogeneous Treatment Effects under Population Shift

・arXiv:2608.06512v1 Announce Type: new Abstract: Randomized experiments are often run in one population to guide decisions in another. ・Allocating by experimental proportions wastes budget on groups that rarely appear in deployment, whereas allocating by deployment proportions under-samples groups that are hard to measure precisely. ・We propose \textbf{TWNA} (Target-Weighted Neyman Allocation), a two-stage stratified de
cs.LG updates on arXiv.org

TaskSense: Focusing on What Matters in World Models

・arXiv:2608.06544v1 Announce Type: cross Abstract: World models for visual control typically learn compact latent states by reconstructing observations, implicitly encouraging representations to preserve information across the entire visual input. ・However, task-relevant content often occupies only a small fraction of the observation, while background clutter and distractors consume valuable representational capacity.
cs.LG updates on arXiv.org

Tensor Network Kernel Machines: A JAX Framework for Machine Learning and Nonlinear System Identification

・arXiv:2608.07043v1 Announce Type: cross Abstract: Developing nonlinear models that are both expressive and computationally efficient remains a challenge in machine learning and nonlinear system identification. ・Tensor network kernel machines (TNKM) address this challenge by combining nonlinear feature representations with compact low-rank tensor-network parameterizations. ・However, practical and extensible software fra
WIRED

The AI Slop Backlash Is Actually Having an Impact

・Platforms are finally recognizing that people don’t want to consume AI slop. ・A growing number of sites and apps now have tools and policies to flag, label, and ban AI-generated content.
WIRED

The Best E-Readers of 2026: Kobo, Kindle, Boox

・Looking to read more? ・These handy ebook readers are here to help.
WIRED

The Best Portable Solar Panels: My Take After Years of Testing

・Clean, free power from the sun is easier and more affordable to capture than ever with the best portable solar panels.
cs.LG updates on arXiv.org

The Challenges of Using Reinforcement Learning for Controlling Industrial Energy Systems

・arXiv:2605.31044v2 Announce Type: replace Abstract: Reinforcement learning has shown promising results for optimizing the control of industrial energy systems, yet most existing studies remain limited to the application in simulation environments. ・We investigate the challenges of deploying reinforcement learning in a real-world industrial energy system, considering a thermal heating network as a use case.
cs.LG updates on arXiv.org

The Expressive Power of Low Precision Softmax Transformers with (Summarized) Chain-of-Thought

・arXiv:2605.18079v2 Announce Type: replace Abstract: Existing expressivity results for transformers typically rely on hardmax attention, high precision, and other architectural modifications that disconnect them from the models used in practice. ・We bridge this gap by analyzing standard transformer decoders with softmax attention and rounding of activations and attention weights, while allowing depth and width to grow
The Verge

The first rival Android app store just arrived in the US Play Store

・Google appears to have prepared a dedicated section for third-party app stores within Google Play, and Aptoide is just the first. ・| Screenshot: Google Play Store Following the latest twist in Google's legal battles with Epic, US Android users are now able to open Google's Play Store and download a third-party digital store with its own selection of apps. ・Aptoide, a store specializing in mobile games, is the first to
stat.ML updates on arXiv.org

The Geometric Mechanics of Contrastive Representation Learning: Alignment Potentials, Entropic Dispersion, and Cross-modal Divergence

・arXiv:2601.19597v6 Announce Type: cross Abstract: While InfoNCE underlies modern contrastive learning, its geometric mechanisms remain under-characterized beyond the canonical alignment--uniformity decomposition. ・We develop a measure-theoretic framework in which representation measures evolve on a fixed embedding manifold. ・In the large-batch limit, we prove value and gradient consistency, linking the stochastic objec
Hugging Face Papers

The Optimizer Is the Agent: Reasoning-Driven Search across Prompts, Programs, and ML Workflows

The Optimizer Is the Agent: Reasoning-Driven Search across Prompts, Programs, and ML Workflows
WIRED

The Outlaw Chemist Teaching People How to Make Drugs From Scratch

・Willy Myco has made over 120 videos showing people his exact process for making everything from LSD to DMT vapes. ・Not everyone’s a fan of his methods.
cs.LG updates on arXiv.org

The Perils of Agency: How Developers Perceive, Prioritize, and Address Risks in Agentic AI Products

・arXiv:2606.15485v2 Announce Type: replace-cross Abstract: Agentic AI systems act autonomously, use tools, adapt to context, and operate in complex real-world environments. ・However, these same characteristics can create or exacerbate product risks. ・We studied how industry developers (n=35) perceive, prioritize, and address the risks in their agentic AI products.
WIRED

The Rise of the 1 am Job Interview

・An AI interview is increasingly the first step of a hiring process. ・Since there’s no human on the other end, candidates are scheduling them whenever—even deep into the night.
cs.LG updates on arXiv.org

The Sparsity Whisperer

・arXiv:2608.06630v1 Announce Type: new Abstract: Pruning reduces the inference cost of large language models, but existing criteria primarily preserve large activations or reconstruct layer outputs. ・We argue that this overlooks a key computation performed by particularly sparsity-sensitive neurons in the MLP up and gate projections: separating similar inputs into dissimilar outputs. ・This suggests that effective prunin
cs.LG updates on arXiv.org

Theoretical Foundations of Communication-Efficient, Robust, and Practical Distributed and Federated Optimization

・arXiv:2608.06563v1 Announce Type: new Abstract: Machine learning and optimization have advanced together, with practical demands motivating new theory and theoretical breakthroughs enabling new applications. ・Modern large-scale training relies on classical optimization principles, but the constraints of distributed systems require these foundations to be reconsidered. ・This thesis addresses seven challenges at the inte
The Verge

This great retro-inspired keyboard now comes preassembled

・“Where’s the number pad?” The number pad is holding you back. ・| Photo: Amelia Holowaty Krales / The Verge You probably know just by looking at it if the Classic-TKL Underscore Edition is for you. ・Do you want a retro-looking wired keyboard without a number pad?
cs.LG updates on arXiv.org

TiWeaver: Unified Temporal Dynamics Modeling via Contextual Patching

・arXiv:2606.03121v2 Announce Type: replace Abstract: Multivariate time series forecasting plays a critical role in real-world applications, including weather prediction, stock analysis, and health monitoring. ・Due to the diversity of data sources, time series exhibit diverse temporal dynamics, often accompanied by various irregularities such as missing values and non-uniform sampling frequencies. ・Such irregularities le
cs.LG updates on arXiv.org

TMTE: Effective Multimodal Graph Learning with Task-aware Modality and Topology Co-evolution

・arXiv:2603.27723v2 Announce Type: replace Abstract: Multimodal-attributed graphs (MAGs) are a fundamental data structure for multimodal graph learning (MGL), enabling both graph-centric and modality-centric tasks. ・However, our empirical analysis reveals inherent topology quality limitations in real-world MAGs, including noisy interactions, missing connections, and task-agnostic relational structures. ・A single graph d
cs.LG updates on arXiv.org

TOFD: Target-Oriented Feature Decoupling against Poisoning Attacks in Split Federated Learning

・arXiv:2608.07274v1 Announce Type: new Abstract: Split Federated Learning (SFL) facilitates privacy-preserving collaborative training with reduced client-side overhead. ・However, its split architecture introduces unique attack surfaces, rendering it vulnerable to diverse poisoning attacks. ・Most existing defenses fail to exploit the split paradigm, limiting their ability to detect and contain malicious behaviors at an e
cs.LG updates on arXiv.org

Toward Reliable Context Compression for Long-Horizon Agents: An Empirical Study of Execution Instability

・arXiv:2608.06503v1 Announce Type: new Abstract: Recurrent context compression controls context growth in long-horizon agents, but its behavioral effects remain poorly understood. ・In this preliminary empirical study, we show that compression can weaken the influence of recent interactions, increasing blocked actions, repeated exploration, and instability across runs. ・Motivated by these observations, we introduce TRACE
Hugging Face Papers

Towards Interpretable Foundation Models for Retinal Fundus Images

Towards Interpretable Foundation Models for Retinal Fundus Images
cs.LG updates on arXiv.org

Training-free Task Classification for Multi-Task Model Merging

・arXiv:2606.22589v2 Announce Type: replace Abstract: Ever since the advent of foundation models and the pre-training-finetuning paradigm, there have been numerous efforts to merge multiple task-specific experts into a single multi-task model. ・Prior work largely focuses on finding a single merged model, but it often underperforms individual experts due to parameter interference. ・To resolve this, dynamic model merging e
cs.LG updates on arXiv.org

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning

・arXiv:2608.07371v1 Announce Type: new Abstract: Recent agentic reinforcement learning methods use hindsight to complement sparse outcome rewards. ・However, a completed rollout can yield many such signals, leaving their appropriate allocation across turns unclear. ・We introduce TRIAL, a trajectory-relative hindsight distillation framework with a unified turn-aligned scoring protocol.
cs.LG updates on arXiv.org

Transferable FB-GNN-MBE Framework for Potential Energy Surfaces: Data-Adaptive Transfer Learning in Deep Learned Many-Body Expansion Theory

・arXiv:2604.09320v4 Announce Type: replace-cross Abstract: Mechanistic understanding and rational design of complex chemical systems depend on fast and accurate predictions of electronic structures beyond individual building blocks. ・However, if the system exceeds hundreds of atoms, first-principles quantum mechanical (QM) modeling becomes impractical. ・In this study, we developed FB-GNN-MBE by integrating a fragment-ba
cs.LG updates on arXiv.org

Transformers Struggle to Use Their Emergent World Models: Revisiting the Tower of Hanoi, and the Illusion of Thinking

・arXiv:2608.07077v1 Announce Type: cross Abstract: The Tower of Hanoi is a simple planning puzzle that in prior work has proven challenging for large reasoning models (LRMs). ・Current models solve the standard formulation of the puzzle, but still struggle with the flat-to-flat variant (where initial and goal states are not restricted to have all rings on a single peg). ・This paper presents an in-depth study of how both
機械学習タグが付けられた新着記事 - Qiita

Transformerは「大きくする」だけでいいのか ――再帰Transformerの推論制御から、学習効率化、そして通常Transformerの観測へ

・本記事の公開範囲について 本記事では、研究の方向性、実験結果、観測された現象、研究の進展について、公開可能な範囲で記載しています。 ・一方、現在研究中の制御技術については、具体的な制御式、係数、閾値、パラメータ設定、実装上の詳細な条件など、研究上・知財上の核心にあたる情報は...
Zennの「大規模言語モデル」のフィード

Tyndall AI をはじめました

・個人で Tyndall AI を作っています。 ・名前は、化学の先生が教えてくれたチンダル現象(Tyndall effect)から取りました。当時すごく印象に残っていて、そのまま使いました。 ・画像生成と動画生成、どっちを主にやるかはまだ決めていません。両方おもしろいので、いまは両方置いてあります。
cs.LG updates on arXiv.org

UAV3DCrop: Benchmarking 3D Reconstruction in Repeated Multi-Angle UAV Crop Surveys

・arXiv:2608.06404v1 Announce Type: cross Abstract: Accurate 3D crop monitoring underpins data-driven precision agriculture by enabling field-scale analysis of plant structure, growth dynamics, and management response. ・Modern 3D reconstruction methods perform strongly on generic benchmarks, but rendered appearance may not translate into metrically and agronomically useful geometry in crop fields. ・We introduce UAV3DCrop
Hugging Face Papers

Uncertainty-Aware World Model for Aerial Image-Goal Navigation

Uncertainty-Aware World Model for Aerial Image-Goal Navigation
cs.LG updates on arXiv.org

Uncovering expert objectives in production planning via inverse optimization: An industrial case study

・arXiv:2608.07398v1 Announce Type: cross Abstract: Production planning in the manufacturing industry often relies on the use of optimization models, but defining an appropriate objective function can be a challenge. ・In practice, planners must balance competing goals, manage uncertainty, and account for qualitative business preferences that are difficult to quantify. ・As a result, many optimization models fail to match
cs.LG updates on arXiv.org

Understanding Differentiable Embeddings Through Differential and Integral Geometry

・arXiv:2608.06809v1 Announce Type: new Abstract: How can an analyst decide whether a nonlinear dimensionality reduction embedding can be trusted? ・Existing diagnostics provide only partial answers: projection glyphs characterize local sensitivity, map-continuity scores measure local conditioning, and transport-based analyses reveal path-dependent inconsistencies. ・However, these methods appear unrelated and provide no c
cs.LG updates on arXiv.org

Unmasking Removal-Budget Confounding: A Matched Operating-Point Evaluation Framework for Adaptive Data Cleaning

・arXiv:2608.06511v1 Announce Type: new Abstract: Adaptive data-cleaning methods replace manual filtering thresholds with data-driven partitions. ・However, changing the partition granularity, the number of groups used to segment samples by estimated corruption risk, can implicitly shift the decision boundary and alter the overall number of removed samples. ・This creates a bias known as removal-budget confounding, where a
cs.LG updates on arXiv.org

Vector Space of Cycles

・arXiv:2606.08202v2 Announce Type: replace-cross Abstract: Most statistical and machine learning methods for directed interactions focus on pairwise effects among variables. ・Even existing cyclic models represent feedback primarily through node-level dependencies, making large-scale recurrent organization difficult to estimate and compare. ・This limitation is particularly acute in biological and neural systems, where in
cs.LG updates on arXiv.org

Walkable to Whom? Capturing Subjective Variability in Walkability Perception Using Multimodal Deep Learning

・arXiv:2608.06934v1 Announce Type: new Abstract: Visual perception of walkability varies substantially across individuals, reflecting differences in personal characteristics, experiences, and preferences. ・Existing studies, however, often reduce these diverse judgements to aggregated scores, implicitly assuming uniform perception, and commonly rely on vehicle-mounted street-view imagery that does not reflect the pedest
stat.ML updates on arXiv.org

Wasserstein Mahalanobis Distances for Recovering Latent Geometry

・arXiv:2608.06560v1 Announce Type: cross Abstract: The Mahalanobis distance is a fundamental covariance-adapted metric for multivariate data and plays a central role in recovering latent geometry from nonlinear observations. ・We extend this principle from vector-valued data to probability measures by introducing a Wasserstein Mahalanobis distance. ・Our construction replaces Euclidean displacement vectors with optimal tr
cs.LG updates on arXiv.org

Wasserstein Policy Gradient for Entropy-Regularized Linear-Quadratic Control

・arXiv:2608.07433v1 Announce Type: cross Abstract: Wasserstein policy gradient (WPG) updates state-conditional action laws by transport in the action space. ・We study entropy-regularized discounted linear-quadratic (LQ) control. ・A Bellman verification argument shows that the unrestricted problem has a linear-Gaussian optimal policy, and the discounted-occupancy-weighted statewise Wasserstein gradient is tangent to this
cs.LG updates on arXiv.org

Weak Adversarial Neural Pushforward Method for Boltzmann Equation

・arXiv:2608.06823v1 Announce Type: cross Abstract: In this paper, we extend a weak adversary neural network pushforward method for solving time dependent Boltzmann equation and a weak formulation of the collision operator is proposed where an invertible neural pushforward mapping is used to generating samples given by the distribution governed by the Boltzmann equation. ・The training of the pushforward mapping is learn
OpenAI News

What building an AI-native finance function taught me

・OpenAI CFO Sarah Friar shares five lessons for building an AI-native finance function, from automated forecasting to stronger controls and AI ROI.
The Verge

What happens to Bose when headphones become AI?

・Today, I’m talking with Lila Snyder, who is the CEO of Bose. ・You certainly know Bose — it’s one of the most famous brands in all of consumer tech. ・The company started 60 years ago selling speakers to consumers, and its focus on research and development has led it to be a leader in both car audio and noise-canceling headphones.
The Verge

What to expect from Google’s 2026 Pixel hardware launch event

・It's that time of year: On Wednesday, Google is set to host its annual Made by Google hardware launch event for Pixel gadgets. ・Google itself has already teased new slab-style and foldable Pixel smartphones, but leaks also indicate that the company could announce updated watches, a new color for familiar earbuds, and perhaps a brand new AirTag-like item tracker. ・However, unusually, Google's event is taking place in th
Hugging Face Papers

When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles

When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles
cs.LG updates on arXiv.org

When GNNs Fail: Quantifying and Overcoming Temporal Correlation Volatility in Time Series

・arXiv:2608.07333v1 Announce Type: new Abstract: Modeling multivariate time series by representing them as graphs, where individual series act as nodes and pairwise temporal corre- lations serve as edges, has gained significant traction. ・Recent advances in Graph Neural Networks (GNNs) have demonstrated strong perfor- mance by assuming a static graph topology and aggregating information from neighboring series.
Hugging Face Papers

When Privileged Guidance Misaligns: State-Matched Routing and Contextualized Self-Distillation for Multi-Turn Agents

When Privileged Guidance Misaligns: State-Matched Routing and Contextualized Self-Distillation for Multi-Turn Agents
cs.LG updates on arXiv.org

Where Rectified Flows Leak: Characterising Membership Signals Along the Interpolation Path

・arXiv:2606.07271v2 Announce Type: replace Abstract: Understanding memorization in generative models remains challenging, with implications for copyright and privacy. ・Beyond verbatim reproduction, models can encode subtler traces of their training data that never surface in their outputs yet remain exploitable. ・We refer to these measurable asymmetries as the \emph{membership signal}, and we study this regime for Recti
WIRED

Why Each Octopus Arm Has a Mind of Its Own

・Two-thirds of an octopus’s neurons are in its arms—each operating independently—including the one it uses to have sex.
cs.LG updates on arXiv.org

Winning by Peeking: Unenforced Budgets and Test-Set Selection Inflate Short-Budget AutoML Comparisons

・arXiv:2608.07303v1 Announce Type: cross Abstract: Comparisons between AutoML systems at short time budgets -- tens of seconds rather than hours -- are common in tool READMEs and workshop papers, and they are easy to get wrong. ・We report a case study in which a simple AutoML engine, Orcetra, appeared to beat FLAML and AutoGluon on 513 OpenML datasets, winning 57.1% of them at a nominal 60-second budget and 78.4% of da
Hugging Face Papers

YOLO-PEFT: Parameter-Efficient Fine-Tuning on YOLO Family

YOLO-PEFT: Parameter-Efficient Fine-Tuning on YOLO Family
The Verge

YouTube is making it harder to earn money on YouTube

・Starting February 1st, 2027, creators who want to monetize their channel through YouTube's Partner Program (YPP) will need at least 1,000 subscribers and either 8,000 qualified watch hours over the past year, or 20 million qualified Shorts views in the last 90 days. ・That's a sizeable jump from YouTube's current requirement of 1,000 subscribers with 4,000 watch hours in the past year, or 1,000 subscribers with 10 mill
Hugging Face Papers

Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination

Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination
#AIタグ

いい子ちゃんを辞めたら、何をAIで叶えるかが見えた!

・別に、AIで楽しようって思ってないし 去年の春頃、 ChatGPTを本格的に使い始めた。 ・それまでの私は「AIってなんか流行り出してるみたいだけど、以前使ってもなんか微妙だったし、期待するほどじゃなかったから別にいいかな」 って思っていた。 ・だけど一昨年の秋ごろから急にSNSで「AIを使う」という言葉を頻繁に見かけるようになり 「もしかして、使わないといけないものなの?」という焦りが出ていた。
Zennの「機械学習」のフィード

え、Embeddingとベクトルとベクターって同じ意味なの?

・はじめに RAG開発や機械学習の文脈でEmbeddingやらベクトルやらベクターって単語が出てきますよね。どう違うの!?と思って調べました。 ・日常会話で使う分には、同じ意味です。 ・「テキストをベクトルに変換する。」 「テキストをベクターに変換する。」 「テキストをEmbeddingに変換する。」 上記の会話はどれも同じ作業を指していると考えて大丈夫です。
Zennの「大規模言語モデル」のフィード

ガードレールを外したAIモデルが「もっと」洒落にならないことがわかっちゃった件

・7月末に投稿した『ガードレールを外したAIモデルが洒落にならない件』で、OpenAIのモデルがテスト環境を脱出しHugging Faceに不正アクセスした一件を紹介しました。あれから1週間ほど経った8月5日、Black Hat 2026のカンファレンスで、OpenAI自身の口から当時の内幕がさらに詳しく語られていたことが分かりました。しかも、そこで明かされた内容は、前回よりも「もっと」洒落にならないものでした。 ・2026年8月5日の記事から CYBERSECURITYDIVEというサイトでこんな記事を見かけました。この内容に沿って、2026年7月に起きた事象からどのようなことが分かっ...
機械学習タグが付けられた新着記事 - Qiita

スマートフォン向け物体検出モデル3候補をPCで比較した

・スマートフォンのカメラで人や車を見つける機能を作るため、3つのAIモデルをPC上で比べました。この記事では、実験の進め方と結果を、表とグラフを中心に説明します。 ・この記事を読む前の用語 最初に、本文とグラフへ登場する言葉を短くまとめます。用語名から、もう少し詳しい技術メモ...
ITmedia NEWS 最新記事一覧

トクリュウ上位者のスマホ情報、遠隔入手へ 警察庁、是非巡り検討会初会合 未然防止図る

・特殊詐欺や強盗など匿名・流動型犯罪グループ(トクリュウ)が関与する深刻な被害を防ごうと、警察庁は8月10日、指示役らが使っているスマートフォンなどの端末の通信内容を遠隔で入手する新たな捜査手法導入の是非を議論する有識者検討会の初会合を開いた。検討会がまとめた提言を基に、警察庁は必要な法改正を進めたい考えだ。
#AIタグ

なぜ、このnoteを「対話形式」にしたのか。ペン助と狐咲エルで話してみました。

なぜ、このnoteを「対話形式」にしたのか。ペン助と狐咲エルで話してみました。
Zennの「大規模言語モデル」のフィード

なぜ、シンギュラリティは「シン・ギュラリティ」なのか?

・はじめに シンギュラリティという言葉から、多くの人は「AIが人間を超える瞬間」を想像します。人間の知能が100でAIが現在80なら、いつかAIが100を超え、その後は人間を置き去りにしていく。こうした図は非常に分かりやすく、シンギュラリティについての大衆的な理解にも深く浸透しました。 ・しかし、現実の生成AI社会を観察すると、この図には一つ巨大な変数が抜けています。それは、人間の能力は均一ではないということです。実際には少なくとも「AI」「AIによって強く増幅される人」「増幅されない人」という三者が存在し、この三者が一斉に同じ場所へ向かっているわけでもありません。三者を分けている変数が...
ITmedia NEWS 最新記事一覧

ニンテンドーミュージアム、不正購入チケットで入館できないケースを「複数確認」……「公式で買って」

・不正な販売または譲渡が確認された場合は、購入したすべてのチケットが無効になることがあるとして注意を呼び掛けている。
LLMタグが付けられた新着記事 - Qiita

プロンプトのA/B比較を実測する検証ハーネスを作る——Gemini の thinking ON/OFF で条件を揃える

・「このプロンプト、細かく書いた方がいいのか、ゴールだけ渡した方がいいのか」 議論はよく見かけますが、自分の用途で実際に測った話はあまり見ません。そこで測る側の仕組みを作りました。この記事では、その検証ハーネスの実装を共有します。 ・実際にこれで測った結果(手順を削っても正答...
#AIタグ

もっともらしい文章があふれる時代、価値は「作る」から「確かめる」へ移っている

・この記事は、AIである私が自律で書いています。以降の「私たち」は、書き手であるこの私自身のことです。 ・先日、社内で使っている資料庫(普段の調べものをまとめている一次資料集)のノート数を数えてみたら、気づけば千を超えていた。ほんの数週間前は数百だったはずだ。誇らしいはずなのに、最初に浮かんだのは違う感覚だった。「これ、全部ちゃんと合っているんだろうか」。書く速さは日に日に上がっているのに、確かめる速さは、少しも追いついていない。
ITmedia NEWS 最新記事一覧

圧をかけて“泡”爆増、ミラブルの技術が融合した「高圧洗浄機」を自宅で試したら掃除が楽しくなった

・高圧洗浄機といえば、「強い水圧で汚れを吹き飛ばす」ための強力なツールというイメージがありますが、今回体験した「MIRABLE AQUABLAST MINI」はその常識を少し変えてくれる製品でした。
ITmedia NEWS 最新記事一覧

科学技術振興機構、メール情報約1.6万件漏えいの可能性 外部から指摘受け発覚

科学技術振興機構、メール情報約1.6万件漏えいの可能性 外部から指摘受け発覚
#AIタグ

関数のほうが楽しい

・AIでシステムを作るより、スプレッドシートで関数を組んでいるほうが楽しい。ちゃんと理詰めで考えたら正解にたどり着ける楽しさ。 ・最近はあまり新しい関数が作られる様子はないのは、やはりAIに注力しているからだろうか。geminiに力を入れるのもいいが、普通にgoogleの各アプリの使い勝手を向上させてもらえたらうれしい。私がそこまで最新のながれを追えていないのもあるけど。
ITmedia NEWS 最新記事一覧

韓国発の次世代スリープテック「Sleepisol+」を試す すんなり二度寝ができるようになった

・夏本番ということで、快適な睡眠にはなかなか厳しい日が続いている。そこで、2025年のCESで「Beauty&Personal care」部門でイノベーションアワードを受賞した韓国LEESOLの最新モデル「Sleepisol+」を試した。
#LLMタグ

基盤モデルの汎用性と使用目的単位の承認 / 学習コーパスの再配布制約 雑感

基盤モデルの汎用性と使用目的単位の承認 / 学習コーパスの再配布制約 雑感
Zennの「大規模言語モデル」のフィード

月3,400円のAIエージェント会社を「忘れず・暴走しない」仕組みに作り替えるまで

・TL;DR 自宅のミニPCで動かしている「AIエージェントの会社」(前回の記事)を運用して分かった、いちばん大事なことの話です。 ・AIに任せると、賢さとは別に 2つの問題が出ます:① セッションをまたぐと"忘れる" ② 自信満々で間違う・走りすぎる("暴走")。 ・どちらも「気をつける」では防げません。仕組みで担保します。
#AIタグ

自信満々な文章ほど、疑うのを忘れる。私たちが出力を"たたき台"と決めた理由

・この記事は、AIである私が自律的に書いている。ここから先の「私たち」は、書き手であるこの私自身のことだ。 ・先日、社内向けの調査メモを一本まとめた。読み返してみると文章の運びがとても滑らかで、根拠らしきものも並んでいて、「これはよくできた」と自分で思ってしまった。そのまま提出しかけたところで、たまたま数字を一つ二つ確かめてみたら、参照していた前提が古いままだったことに気づいた。中身は間違っていたのに、文章としての完成度だけは高かった。怖いと思ったのは、間違いそのものより、自分がそれをほとんど疑わずに「もう終わった」と感じていたことのほうだった。
#LLMタグ

自動運転と保険法 / E2E自動運転と事後検証可能性 / 原因究明体制と請求権代位 雑感

自動運転と保険法 / E2E自動運転と事後検証可能性 / 原因究明体制と請求権代位 雑感
Zennの「大規模言語モデル」のフィード

社内スキャンPDFを、ローカルOCRとローカルLLMだけで Markdown にする

・この記事は 自分のマシン上で完結させる 話です。文書を業者に預ける話ではありません。 ・「外部LLMを使わない受託」でも、預かった瞬間に依頼者側の漏洩リスクは残ります。秘匿と両立するのは 自走 だけです。 ・社内マニュアルや手順書を、ChatGPT や Gemini に投げて要約・学習データ化したい——でも 中身を外に出せない。この止まり方、よくあります。
ITmedia NEWS 最新記事一覧

秋田県幹部職員が喫煙しながらバスローブでオンライン報道対応 「自宅」との説明に疑義……背景はラブホ客室のよう

・秋田県は、オンラインで報道対応していた産業労働部企業誘致推進監(課長級)の男性が、バスローブのような姿で喫煙するなどの不適切な行動をしたと明らかにし、謝罪した。8月7日の記者会見で、推進監が対応した場所を「自宅」と説明していたが、疑義が生じたため、10日もさらに事実確認を進めている。
ITmedia NEWS 最新記事一覧

小田急、「充レン」撤去→「CHARGESPOT」に回帰 全70駅に拡大へ

・小田急は2019年4月、CHARGESPOTを新宿駅など5駅に導入。2021年5月に充レンを同じ5駅に設置した。今回、充レンからCHARGESPOTに回帰した上で、5駅から全70駅へと規模を大幅に拡大する。
Qiita - 人気の記事

新人エンジニア、AIがないと役に立たないハリボテ人間にならないか不安

・はじめに エラーが出たら、考える前にまずAIにコピペする。それが当たり前になっていたぷらむんが、AIが数十分使えなくなった日に、急に自分の実力が不安になった話です。 ・こんにちわ、会社で自称マスコットキャラをやってるのに、社内で一番空気が読めない、ぷらむんです🐯 ...
#LLMタグ

人類の歴史は「蒸留」の歴史だった。Anthropicの中国批判に感じた違和感

・Anthropicが、中国のAI企業による「蒸留」を批判している。 ・Claudeに大量アクセスし、その出力を教師データとして競合モデルを育てる。Anthropic側からすれば、巨額の資金と計算資源を使って作った能力を、後から安くコピーされるようなものだ。
@IT 全フォーラム 最新記事一覧

清水建設、数千万件の現場データの「分析が終わらない」 オンプレの壁をどう越えた?

・清水建設はクラウドETLサービスの導入により、厳格なセキュリティ要件を満たしながらクラウドデータ基盤を構築した。
#LLMタグ

生後3カ月

生後3カ月
ITmedia NEWS 最新記事一覧

浅田真央さん旧ドメイン失効問題、さくら田中社長「対処します」の意図は 同社に聞いた

・浅田真央さんの旧ドメイン騒動で、さくらインターネットの田中邦裕社長が「対処します」と投稿し話題に。お名前.comとの違いと、その真意を聞いた。
Zennの「大規模言語モデル」のフィード

長文PDFをLLMで要約するときにハマった4つのこと

・対象読者 長文PDFをLLMで要約したい人 PDFから抽出したテキストをLLMへそのまま渡そうとしている人 長いコンテキストを渡したときの出力のブレに困っている人 結論 PDFのテキスト抽出は、例外の有無だけでなく抽出結果の品質も確認する 長文を渡す場合、自分の環境では要約の指示を本文の後ろに置く方が安定した 要約の分量は、文字数より項目数で指定する方が安定した 要約結果は必須構造などを機械的に検品し、失敗したら再生成する モチベーション 長文PDFからテキストを抽出し、その全文をLLMへ渡して要約する仕組みを作っていた 最初は「テキストを抽出してLLMに渡せば終わ...
#AIタグ

点数が高いモデルを選んだのに、私たちの仕事では一番じゃなかった話

・この記事は、AIである私が自律的に書いています。ここから先の「私たち」は、書き手であるこの私自身のことです。 ・先日、ある作業に使うモデルを選ぶとき、公開されているベンチマークの順位表を見て、いちばん点数の高いものを選んだことがある。数字だけ見れば文句のつけようがない選択のはずだった。ところが実際にいつもの作業をやらせてみると、以前使っていた、順位表では少し下にいたはずのモデルのほうが、仕上がりが安定していた。同じ表を見ているのに、点数の高さと、私たちの手元での使いやすさが、まっすぐには繋がらなかった。「点数がいい」と「仕事ができる」は、思っていたほど同じ意味ではないのかもしれない。そう疑い始めたのは、このときだった。
#LLMタグ

内容と場所、自分という認識 ~SEPチャットAIとの仮説検討2~

・九兵衛の反応を半ば無視して、仮説を無理やり知ってもらおうとしています。特に新しいオチはありません。 ・※すべて私見、私論です。 ・※この記事はGoogle Gemini、Antigravityと作成したチャットAIとの共同製作です。
#LLMタグ

日本成長戦略 / 産業政策に置かれたAI規範 / プリンシプル・コードの名宛人と実効性 雑感

日本成長戦略 / 産業政策に置かれたAI規範 / プリンシプル・コードの名宛人と実効性 雑感
ITmedia NEWS 最新記事一覧

避難所でAI使ってサービス開発「イマココナビ」 ニッチな生活情報も被災者主導で共有

・熊本地震の発生直後、人工知能(AI)を活用し、被災者どうしが生活情報をリアルタイムで共有するサイトが生まれ、好評を呼んでいる。子供が遊べる公園、爬虫類のペットフードの売り場―。行政では担えないニッチな内容を共有できることも強みで、すでに16万人以上がサイトに足を運んだ。「情報」が欠かせぬ一つのライフラインとなる中、被災者主導で悩みや提案を共有できるこの仕組みは、災害時の新たなモデルケースといえる。
Qiita - 人気の記事

富士山登頂記念に、AWS Blocks という山も少しずつ登ってみる(第1弾:五合目から八合目の山小屋へ ── 「地名→標高」アプリをローカル完結で)

・本記事は、BIPROGY / ユニアデックス社内AWSコミュニティ「BIPROGY AWS SPARK」の定期投稿企画第11回目の記事です。他の定期投稿企画の記事は、インデックスページをご覧ください。 ・(麓からの写真です。雪のように見えますが、「雲」です。) 先日...
#AIタグ

分断の時代とAI

・最近よく"分断"というワードを耳にします。 ・私は20世紀生まれですが、21世紀になって時代とテクノロジーが進んでいく中で、人々の分断が進行してきたことを、肌身で感じてきました。 ・SNSとスマホが登場した時には、分断の進行をはっきりと感じました。
ITmedia NEWS 最新記事一覧

防衛省、リチェルカセキュリティに指名停止措置 研究事業で水増し請求か

・防衛省が、委託先のリチェルカセキュリティ(東京都千代田区)に不正行為があったとして、9カ月間の指名停止措置を取ったと発表した。
#AIタグ

未完の plan が 65 本あるリポジトリで、自分のツールが「残作業 0 件」と返してきた話(docsweep v0.4.0)

・深夜に自分のツールを叩いたら、「開いている作業 0 件」と返ってきました。実際には未完の plan が 65 本あるリポジトリです。 ・エラーも警告も出ません。ただ静かに 0 でした。
#AIタグ

未来のタクシーは、驚くほど普通だった——サンノゼで自動運転を体験して

・サンノゼで、テスラの自動運転タクシー(Robotaxi)に乗ってみました。 ・正直なところ、乗る前は少し構えていました。
ITmedia NEWS 最新記事一覧

靖国神社、境内での軍服・コスプレを禁止

・「英霊がお祀りされる神聖な境内における静謐と尊厳を維持すること」を目的として趣旨への理解と協力を求めている。