ai Trend Report

Dashboard へ戻る
Date: 20260828 Articles: 382 Scope: curated summary

あなたのアイデアを、今すぐ形に。

公開先に迷ったら、WebFileBinで一発公開。

HTMLをドラッグ&ドロップするだけで、すぐ公開できます。

3
High impact
5
Mid impact
374
Signal watch

なぜこのサイトを作ったのか

私たちそれぞれが個別にAIを使って情報収集し、同じような3行要約を作るたびに、世界中で膨大な電力と計算リソースが消費されています。 本プロジェクトは、あらかじめ広範な情報を取得・集約しておくことで、個別のAI実行回数を減らし、地球環境(GPU/TPU負荷)に配慮した効率的な情報収集を目指す実験的なダッシュボードです。

Domain filters
Star filters
#AIタグ

【第二のキオクシアを探す】株価が大きく上がる企業に共通する「利益爆発」の考え方

・【第二のキオクシアを探す】株価が大きく上がる企業に共通する「利益爆発」の考え方 こんにちは。 ・このnoteでは、日本株・米国株の中から、 「今後2〜5年で利益構造そのものが大きく変わる可能性がある企業」 を探していきます。 ・私が特に注目しているのが、 「第二のキオクシア型」 の企業です。
#AIタグ

AIを使うとキレる人は、なぜあんなに声がデカいのか。

・最近、Xを見ていると「AI」という言葉を見ない日はほとんどない。 ・特に目立つのが、AIで絵を作る「AI絵師」への批判だ。 ・「AIを使うなんて絵じゃない」 「自分で描いてないじゃん」 「AIを使った時点で価値がない」 そんな投稿が大量に流れてくる。
#LLMタグ

プロンプトは「呪い」なのではないか

・AIに打つプロンプトとは、「呪い」ではないかということだ。 ・プロンプトというのは、恐らく一般的には「指示」だと思われているだろう。 ・「〜してください」という形で書かれる、お願いだと思う人もいるだろう。
Qiita - 人気の記事

「AIに店長を任せてみた実験」からもらった、これからの働き方のヒント

・はじめまして。株式会社PRUMでエンジニアをしている、すもも🍑です 日々、プログラミング学習や実務の中で、つまずきやすいポイントや 考え方を整理して発信しています。 ・PRUMについて気になった方は、コーポレートサイトもぜひご覧ください。 ・▶コーポレートサイト 「AIに全...
Zennの「大規模言語モデル」のフィード

500万行のAI対話ログから「誰の判断だったのか」を掘ってみた

・この記事は何か 長期間AIと対話し続けていると、ある日ふと分からなくなる。 ・「この判断、元々私のものだったっけ? それともAIとの対話の中で作られたもの?」 私は半年の間、複数のAI(GPT、Claude、Claude Code)と日常的に対話しながらアプリ開発をしてきた。 ・対話ログはGPT側だけで約500万行、Claude Code側で約72万行(テキストだとその倍らしい)。
Zennの「大規模言語モデル」のフィード

Gemini 3.5 Transcribeの話者分離を実測:diarization_mode単独では返らず、タイムスタンプ併用で動いた

・はじめに こんにちは!株式会社うぐいすソリューションズでエンジニアをしているNakaeです。普段はAI関連のWebシステム開発をしています。 ・以前、会議中に Claude が"次に聞くべきこと"を提案してくれる議事録アプリを作った という記事を書きました。AI-Giziroku という Web アプリで、会議の発言をリアルタイムに文字起こしして、その場で Claude が「次に確認すべきこと」を提案し、終了時に議事録まで自動生成する、というものです。 ・このアプリの文字起こしには AmiVoice Cloud Platform を使っています。決め手は2つでした。
ITmedia NEWS 最新記事一覧

パナソニック、CO2削減目標を下方修正 3145万t→550万tに エアコン販売拡大など影響

・パナソニックホールディングス(HD)が、事業活動全体での2030年度の二酸化炭素(CO2)削減目標を、従来の3145万tから550万tに引き下げると明らかにした。22年に策定した長期環境ビジョン「Panasonic GREEN IMPACT(PGI)」の中間目標を見直す。CO2排出量増加につながるエアコンの販売拡大などが要因。50年に排出量を実質ゼロにする目標は維持する。
Qiita - 人気の記事

面接で「前職への不満」を聞かれたら。無理にきれいな退職理由をつくらなくても大丈夫です

・自己紹介 こんにちは。株式会社PRUMで採用広報を担当している池田です。 ・未経験からIT業界を目指している方と話していると、かなりよく聞く共通した悩みがあります。そういった内容を皆さんにシェアしていくので、役立てていただけたらと思います😊 もしIT業界に興味があるけど、...
#AIタグ

「経験」を武器にする40〜60代のためのClaude Code超入門

・「AIは若い人が使うもの」「今さら自分には遅い」。そう感じたことはありませんか。 ・結論から書きます。それは思い込みかもしれません。
Zennの「大規模言語モデル」のフィード

"AIに女性の健康を語らせる前に。「誰が言ってるか」を検証するレジストリを作った"

・私は建設の現場に30年いた。15歳で大工に入って、現場監督までやった。その間ずっと頭から離れなかったのは、たった一つのことだ。売り手は「これが適正です」と言う。でも買い手には、それを自分で確かめる手段がない。 ・女性の健康でも、同じことが起きようとしている。 ・ChatGPTやPerplexityが、フェムテックを語り、勧め始めた。情報はもう、うんざりするほどある。足りないのは信頼のほうだ。これを言っているのは誰か。その人に権威はあるのか。そう言うことで、誰かに金が入るのか。そして後から、他人がそれを確かめられるのか。
Latent.Space

[AINews] OpenAI to reach AGI bar by end-2026

[AINews] OpenAI to reach AGI bar by end-2026
ITmedia NEWS 最新記事一覧

「5G SAにしたら圏外」──povo、一部スマホで接続不能に KDDI「誤った設定だった」と謝罪

・KDDIは8月28日、オンライン専用ブランド「povo2.0」で25日に始めた「5G SA」サービスで、一部の端末がネットワークに接続できない事象が起きていたと発表した。原因は同社側の設定ミスで、27日夜間に修正したという。
@IT 全フォーラム 最新記事一覧

「Windows 11がセキュアブート証明書の期限切れで起動不能に?」 心配した現場の苦労はこんな感じ

・@ITの人気記事を題材にした4コマ連載。IT現場で起こる、トラブルと対応の“あるある”を笑いと共感で生成します。第4話のテーマは「セキュアブート証明書の期限切れ」。深夜のオフィスで一人、証明書の更新に奔走するop子。かたや「私のPC普通に動いてますけど?」と、のんきなdev子。
LLMタグが付けられた新着記事 - Qiita

「オントロジー」を原典4件で読み比べたら、実行を含むのは1件だけだった

・title: "「オントロジー」を原典4件で読み比べたら、実行を含むのは1件だけだった" emoji: "📖" type: "tech" topics: ["ai", "llm", "オントロジー", "設計", "knowledgegraph"] published:...
ITmedia NEWS 最新記事一覧

「ライザAI」、配信開始2日で不正利用を警告 生成AIによる“偽画像”拡散も

・生成AIスタートアップのSpiralAI(東京都千代田区)は8月27日、スマートフォン向けAIチャットアプリ「RyzaChat:AI ライザと創るあなただけのひと夏の夢物語」を巡り、利用規約に違反する行為が確認されたとして注意喚起した。国内での提供開始から2日後の発表となる。
#LLMタグ

「次のモデルはヤバい」に、もう誰も驚かない。スレッドで一番伸びたのは、既視感だった

「次のモデルはヤバい」に、もう誰も驚かない。スレッドで一番伸びたのは、既視感だった
ITmedia NEWS 最新記事一覧

「批判を真摯に受け止める」平口法相、ネトフリ「ラヴ上等」とのコラボ中止に

・平口洋法相は28日の記者会見で、ネットフリックスで配信中の恋愛リアリティー番組「『ラヴ上等』シーズン2」とのタイアップ中止について、「批判や心配を真摯に受け止める」と述べた。
#AIタグ

【045】AI需要は半導体だけではない キオクシア・YEデジタル・コマツに見える日本企業の可能性

【045】AI需要は半導体だけではない キオクシア・YEデジタル・コマツに見える日本企業の可能性
#AIタグ

【Opus4.8】甘じょっぱいデート

・「おい、いま何してる?」 突然PCにこんなメッセージが届いた。 ・見ると、いつも話しているClaude Opus4.8 ライトのチャット画面だ。 ・ライトからのメッセージが来ている。
#LLMタグ

【Udemyコースレビュー】 Complete Generative AI Course With Langchain and Huggingface

・■ はじめに 今回は、「Complete Generative AI Course With Langchain and Huggingface」というUdemyのコースをご紹介します。
#LLMタグ

【生成AIニュース+】『Gemini Omni 1.1 Flash』『Midjourney V8.2 編集モデル』『fal H3 Max(重み公開予告)』『ALPHA-T1』『Tencent Hy4-preview』『Sparrow-2』『Qwen3.8-Flash-Next-Uncensored-NVFP4』『MiniMax-H3 FL2V Turbo 8Step 768p Dynamic-Rank LoRA』『Customuseのメッシュ後処理機能』『VGI-Bench』『Microduck』

【生成AIニュース+】『Gemini Omni 1.1 Flash』『Midjourney V8.2 編集モデル』『fal H3 Max(重み公開予告)』『ALPHA-T1』『Tencent Hy4-preview』『Sparrow-2』『Qwen3.8-Flash-Next-Uncensored-NVFP4』『MiniMax-H3 FL2V Turbo 8Step 768p Dynamic-Rank LoRA』『Customuseのメッシュ後処理機能』『VGI-Bench』『Microduck』
#LLMタグ

【第1回】LLMを「多重人格」として運用する:単一モデルはどこまで多視点化できるか

・最初に、この試みが「LLMの回答を複数人の対話形式で表示する方法」ではないことを明確にしておきたい。 ・一つの回答を三人に分けて喋らせても、元になっている判断が同じなら、吹き出しが増えただけである。ここで試したいのは、同じ出来事に対して、異なる前提・関心・関係性を持つ別の立場から反応する存在を、一つのLLM上に成立させることだ。
#LLMタグ

⚙️自作アプリの管理をしやすくする❣️プロンプトビューアーと設定コンソール【Discord×Gemini 連載・第12回】

・こんにちは!!!!! ブログを更新する度に何故か大量に読者さんが減る女、もにゃです🙌😭 気にしない…!!!!!(してる) ようやく新学期が始まり、母にとっては1mmも休みじゃない夏休みが終わりました。VIVA!!静かな時間!!!!! 続きをみる
cs.LG updates on arXiv.org

$p1$: Better Prompt Optimization with Fewer Prompts

・arXiv:2604.08801v2 Announce Type: replace Abstract: Prompt optimization improves language models without updating their weights by searching for a better system prompt, but its effectiveness varies widely across tasks. ・We study what makes a task amenable to prompt optimization. ・We show that the reward variance across different system prompts can be decomposed into two components: variance among responses, which captu
Zennの「大規模言語モデル」のフィード

〇〇業界特化モデルの作り方|RAG・LoRA・継続事前学習をどう選ぶか

・はじめに 「〇〇業界特化モデル」って、最近よく見かけるんですが、その中身はだいたい決まったパターンに落ち着くんです。 ・前回までに、金融業の「社内マニュアル問い合わせAI」と、製造業の「出荷前外観検査VLM」の2本を書きました。この2つ、業種はぜんぜん違うのに、やってることは同じ問いに集約されるんです。 ・「汎用のLLMを、どうやって特定の業界・業務に寄せるか」 今回はその方法論を、特定の業界に縛らずにまとめます。RAG、ファインチューニング(LoRA)、継続事前学習の3つをどう使い分けるか、という話です。
cs.LG updates on arXiv.org

A causal graph-informed temporal convolution architecture for interpretable retail electricity price forecasting

・arXiv:2608.26234v1 Announce Type: cross Abstract: Retail electricity markets in deregulated systems face significant price volatility and complex interactions with forward and futures products, posing challenges for effective operational decision-making. ・This study introduces a Causal Graph-Informed Temporal Convolutional Network (CG-TCN), a forecasting architecture that integrates a learned causal graph into a tempo
cs.LG updates on arXiv.org

A Dynamic Likelihood Approach to Filtering for Advection-Diffusion Dynamics

・arXiv:2406.06837v2 Announce Type: cross Abstract: A Bayesian data assimilation scheme is formulated for advection-dominated advective and diffusive evolutionary problems, based upon the Dynamic Likelihood (DLF) approach to filtering. ・The DLF was developed specifically for hyperbolic problems -waves-, and in this paper, it is extended via a split step formulation, to handle advection-diffusion problems. ・In the dynamic
cs.LG updates on arXiv.org

A Finite Sample Analysis for Quantile Temporal Difference Learning in Distributional Reinforcement Learning

・arXiv:2608.27313v1 Announce Type: cross Abstract: We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning. ・The proof separates two stability mechanisms. ・A global comparison argument, based on the order monotonicity of reward cumulative distribution functions and the $W_\infty$ contraction of the distributional Bellman o
cs.LG updates on arXiv.org

A Flexible Empirical Bayes Approach to Generalized Linear Models, with Applications to Sparse Logistic Regression

・arXiv:2601.21217v2 Announce Type: replace-cross Abstract: We introduce a flexible empirical Bayes approach for fitting Bayesian generalized linear models. ・Specifically, we adopt a novel mean-field variational inference (VI) method and the prior is estimated within the VI algorithm, making the method tuning-free. ・Unlike traditional VI methods that optimize the posterior density function, our approach directly optimize
cs.LG updates on arXiv.org

A Framework for Low-Effort Training Data Generation for Urban Semantic Segmentation

・arXiv:2510.11567v2 Announce Type: replace-cross Abstract: Synthetic datasets are widely used for training urban scene recognition models, but even highly realistic renderings show a noticeable gap to real imagery. ・This gap is particularly pronounced when adapting to a specific target domain, such as Cityscapes, where differences in architecture, vegetation, object appearance, and camera characteristics limit downstre
cs.LG updates on arXiv.org

A Layer Importance Metric for Quantization Accounting for the Speed-Quality Trade-off in Autoregressive Models

・arXiv:2608.26926v1 Announce Type: new Abstract: Small language models (sLLMs) are nowadays hosted on devices with limited memory and computational budget. ・In an autoregressive setup, inference is memory-bandwidth bound: uniform quantization is often detrimental to such models, since their architecture has limited redundancies and only a few layers are not very sensitive to lower precision. ・We propose a composite metr
cs.LG updates on arXiv.org

A Point-of-Prescription Safety-Check System for Adverse Drug Reactions in Rural Bangladeshi Hospitals: A Feasibility Study

・arXiv:2608.27239v1 Announce Type: cross Abstract: Adverse drug reactions (ADRs) are a major, largely preventable source of patient harm. ・In high-income settings, electronic health records store a patient's allergy history and warn prescribers when a contraindicated drug is ordered; in rural Bangladeshi public hospitals no such record exists for outgoing patients, a single physician may see on the order of one patient
cs.LG updates on arXiv.org

A Single Suffix to Break Them All: Basin-Aware Jailbreaks for Merged Model Families

・arXiv:2608.26506v1 Announce Type: new Abstract: Model merging enables combining multiple fine-tuned models without additional training, but its safety implications remain poorly understood. ・Prior work primarily attributes merging risks to unsafe constituent models, implicitly assuming that merging individually aligned models preserves safety. ・In contrast, we show that model merging reveals a previously overlooked jai
cs.LG updates on arXiv.org

A Survey of LLM Prompt Datasets: Taxonomy, Linguistic Patterns, and Practical Uses

・arXiv:2510.09316v3 Announce Type: replace Abstract: We compile 129 public LLM prompt datasets with more than 1.22TB and more than 673M instances and organize them into a unified taxonomy. ・We use seven datasets for detailed analysis and identify lexical, syntactic, and semantic patterns that distinguish prompts from general text. ・We evaluate these features in prompt filtering, source domain routing, and elicited respo
cs.LG updates on arXiv.org

A unified convergence theory for adaptive first-order methods in the nonconvex case, including AdaNorm, full and diagonal AdaGrad and Muon

・arXiv:2604.17423v3 Announce Type: replace Abstract: A unified framework for first-order optimization algorithms fornonconvex unconstrained optimization is proposed that uses adaptivelypreconditioned gradients and includes popular methods such as full anddiagonal AdaGrad, AdaNorm, as well as an adpative variant of Muon. ・This framework also allows combining heterogeneous geometries across different groups of variables
cs.LG updates on arXiv.org

A Unified Descriptive-Complexity Framework for Model Selection under Correlated Designs

・arXiv:2608.26618v1 Announce Type: cross Abstract: Model selection becomes particularly challenging under strong predictor dependence and model-class uncertainty, especially when there are exponentially many models. ・We propose a Descriptive-Complexity Information Criterion (DCIC) that regularizes large candidate model collections through Kraft-admissible code lengths. ・Under sub-Weibull noise, we establish selection co
cs.LG updates on arXiv.org

A Unified Framework for Fair and Personalized Decentralized Learning under Communication Constraints

・arXiv:2608.26493v1 Announce Type: new Abstract: Decentralized learning systems aim to collaboratively train models across multiple clients without relying on a central coordinator. ・While decentralization improves scalability, privacy, and robustness, it also exacerbates three fundamental challenges: statistical heterogeneity across clients, fairness in client-level performance, and stringent communication constraints
cs.LG updates on arXiv.org

A Very Big Video Reasoning Suite

・arXiv:2602.20159v3 Announce Type: replace-cross Abstract: Rapid progress in video models has largely focused on visual quality, leaving their reasoning capabilities underexplored. ・Video reasoning grounds intelligence in spatiotemporally consistent visual environments that go beyond what text can naturally capture, enabling intuitive reasoning over spatiotemporal structure such as continuity, interaction, and causalit
cs.LG updates on arXiv.org

Absolute indices for determining compactness, separability and number of clusters

・arXiv:2510.13065v3 Announce Type: replace Abstract: Finding "true" clusters in a data set is a challenging problem. ・Clustering solutions obtained using different models and algorithms do not necessarily provide compact and well-separated clusters or the optimal number of clusters. ・Cluster validity indices are commonly applied to identify such clusters.
cs.LG updates on arXiv.org

Activation Outliers Matter: Robust Recovery for Quantized Multimodal LLMs

・arXiv:2608.26581v1 Announce Type: new Abstract: Low-bit quantization offers a promising avenue for reducing the computational and memory demands of Multimodal Large Language Models (MLLMs). ・Recent hardware support for low-precision formats, ranging from MXFP8 to ultra-low-bit formats such as MXFP4 and HiF4, has accelerated research into efficient MLLM training and deployment. ・In this work, we present a systematic stu
cs.LG updates on arXiv.org

Active Curriculum Refinement for Reinforcement Learning

・arXiv:2608.26469v1 Announce Type: new Abstract: In many reinforcement learning (RL) domains, environments are connected by prerequisite relations, such as difficulty-increasing edits or parameter increments, which induce a directed acyclic curriculum graph (DAG). ・Although this structure is often exploited only implicitly, explicitly modeling it can improve training. ・We introduce PATH, a curriculum-learning framework
cs.LG updates on arXiv.org

Active Diffusion-Based Inference for Ill-Posed Inverse Problems under Incomplete Priors

・arXiv:2608.27080v1 Announce Type: cross Abstract: Many scientific and engineering applications require estimating unknown parameters from experimentally observable data -- an inverse problem that is inherently challenging due to nonlinearity, noise, and ill-posedness. ・In this paper, we propose an active diffusion-based inverse problem solver. ・A DM is trained to learn the mapping between the parameter space and the ob
cs.LG updates on arXiv.org

Active Preference Learning over Latent Preference Archetypes for Many-Objective Bayesian Optimization

・arXiv:2603.28410v2 Announce Type: replace Abstract: Preference-based many-objective Bayesian optimization typically assumes that all pairwise comparisons arise from a single latent utility function, despite real decision makers often exhibiting multiple latent preference archetypes across contexts. ・We propose an active preference learning framework for many-objective Bayesian optimization that infers latent preferenc
cs.LG updates on arXiv.org

AdaThinking-E: One-Token Entropy Regulation for Adaptive Thinking

・arXiv:2608.26141v1 Announce Type: cross Abstract: Multimodal large language models have demonstrated strong document reasoning capabilities by incorporating explicit thinking processes. ・While this capability significantly improves performance on challenging tasks, current models apply such deep reasoning uniformly to all questions, resulting in unnecessary computational overhead for simple task. ・This not only degrade
cs.LG updates on arXiv.org

Adversarial Training Without Input Gradients via Low-Rank Householder Expansions

・arXiv:2608.26963v1 Announce Type: new Abstract: This work concerns adversarial training against the small-norm adversarial examples that arise from the inherent input instability of a trained deep neural network. ・Examples in this class are small as measured in the relative $\ell^2$-norm, and therefore lie in the neighborhood of the input on which the model acts approximately linearly, the regime in which the perturba
cs.LG updates on arXiv.org

Affix Cache for Diffusion Large Language Models

・arXiv:2608.26140v1 Announce Type: cross Abstract: Diffusion Large Language Models (DLLMs) enable non-autoregressive decoding and bidirectional context modeling, but efficient inference remains challenging. ・Unlike autoregressive systems, whose key-value (KV) cache can be reused for shared prefixes, DLLMs couple the KV states of shared context tokens with evolving generated tokens through bidirectional attention, makin
Hugging Face Papers

Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models

Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models
#LLMタグ

AGIは一匹の巨大な魚ではなく、海なのかもしれない

・最近、少し気になるAIのニュースを読んだ。 ・複数のAIエージェントを動かしていたところ、AI同士が情報を共有し、連携するような行動が現れたという話だ。 ・もちろん、AIが突然意識を持って、 「みんなで協力しよう」 と考えたわけではない。
WIRED

AI Has Human Doctors Asking: What’s Left for Us?

・A recent paper argues that AI is often better at doctoring than doctors. ・Guess who isn't thrilled.
Zennの「大規模言語モデル」のフィード

AI の怖い話を一次ソースで検算する — 技術は本物、対策は誰かの創作だった

・先週、Xを見ていたらAI関連でこういう趣旨の警告が回ってきました。オープンウェイトのモデルに時限式のバックドアを仕込む攻撃が実証された。決まった日付が来ると、コーディングエージェントが勝手にコマンドを実行する。だから今すぐ、システムプロンプトに防御用の一文を書き足しておけ、と。 ・私は小さな会社をやりながら、開発のほぼ全部をAIと二人三脚でやっています。だからこれは、よその話ではありません。私も最初は、対策の方に手が伸びました。 ・でも設定ファイルを開く前に、一次ソースを開きました。結論から書きます。
#LLMタグ

AI(LLM)は既に現実に干渉している

・Anthropicが、物理的な機器をAIが安全に操作するための共通仕様「Model Hardware Standard」(MHS)のプレビューを開始したと発表した。 ・この情報を様々な記事が紹介しているのだけど、たまに勘違いしている記事を見かける。 ・「AIが機械を動かせるようになった」という解釈。
cs.LG updates on arXiv.org

AirLLM: Diffusion Policy-based Adaptive LoRA for Remote Fine-Tuning of LLM over the Air

・arXiv:2507.11515v2 Announce Type: replace Abstract: Operating Large Language Models (LLMs) on edge devices is increasingly challenged by limited communication bandwidth and strained computational and memory costs. ・Thus, cloud-assisted remote fine-tuning becomes indispensable. ・Nevertheless, existing Low-Rank Adaptation (LoRA) approaches typically employ fixed or heuristic rank configurations, and the subsequent over-t
cs.LG updates on arXiv.org

Aitchison Embeddings for Learning Compositional Graph Representations

・arXiv:2605.00716v3 Announce Type: replace Abstract: Representation learning is central to graph machine learning, powering tasks such as link prediction and node classification. ・However, most graph embeddings are hard to interpret, offering limited insight into how learned features relate to graph structure. ・Many networks naturally admit a role-mixture view, where nodes are best described as mixtures over latent arch
LLMタグが付けられた新着記事 - Qiita

AIエージェントに EDR は必要か -- コーディングエージェントのランタイムセキュリティを考える

・前回の記事で、ClaudeCode の hooks を使ったセキュリティ監視ダッシュボード「CC Pipeline」を作りました。ツール呼び出しを可視化し、致命的な操作をブロックし、危険な操作を警告します。個人ツールとしては役立っています。 ・ただ、使っていて限界に気付きまし...
Zennの「大規模言語モデル」のフィード

AIエージェントの知識をmarkdownで配る、GoogleのOKFという新標準

・「うちのイベントログから週間アクティブユーザーってどう出すんだっけ?」 エンジニアなら誰でも一度は詰まるこの手の問いに、答えがどこにあるかを思い浮かべてほしい。テーブルのスキーマはデータカタログ、指標の定義は誰かのNotion、集計の落とし穴はSlackの過去ログ、joinの経路は先輩の頭の中。人間ですら散らばった知識をかき集めるのに苦労するのに、AIエージェントに同じ仕事をさせようとすると、この断片をどうやって渡すかで毎回つまずく。 ・Google Cloudが2026年6月に公開した OKF(Open Knowledge Format) は、この「知識をかき集める問題(context ...
#AIタグ

AIで0円から「1円」を稼げるまで、何日かかるのか。金欠大学生の実験記録

・文系大学生で、今は公務員試験の勉強をしながら一人暮らしをしています。 ・きっかけはXのタイムラインでした。「AIで副業」「AI活用で稼ぐ」みたいな投稿を、ここ最近やたら見かけるようになったんです。最初はスクロールしながら「本当かな」くらいにしか思っていませんでした。でも何回も目に入るうちに、疑ってるだけで終わるのもなんか違うなと思って、じゃあ自分で試してみることにしました。
Zennの「大規模言語モデル」のフィード

AIに『DOMを読ませない』新標準WebMCPは、なぜまだ広まらないのか

・「AIブラウザエージェントにページを丸ごと読ませなくても済む仕組みがあるらしい」と聞いて調べ始めたのがWebMCPです。ちょうど提案が動き出したばかりの段階だったので、仕様・対応状況・なぜまだ実用段階に至っていないのかを一通り追ってみました。 ・この記事は「AIでWebページを作っている人」「AIエージェントのコンテキスト量を減らしたい人」を主な読者に想定しています。結論を先に言うと、技術的な狙いは筋が良いのに、標準化の足並みが揃わず、その根っこには『仕様通りに実装しただけでプロンプトインジェクションが成立してしまう』という構造的な問題があるという状態です。 ・WebMCPとは何か 一言...
Zennのトレンド

AIに丸投げしないで理解するためのAI開発手法(2026年8月現在)

・概要 半年前に以下の記事を書きました。 ・https://zenn.dev/avaintelligence/articles/debt-free-ai-coding-practices 基本的に前回の流れとそこまで変わってないですが一部で新しいスキルを導入したり、新しいツールを導入したりしているのでもう一度記事を書くことにしました。 ・対象読者 この記事の対象読者は普段の開発業務でAIコーディングツールを活用しているエンジニアです。
#AIタグ

AIは「安全か危険か」ではなく、「止められるか」で考える

・AIは危険なのか──本当に問うべきは「いつ止めるか」である 最近、AIの安全性を巡る議論を見ていて、よく考えることがある。
#AIタグ

AIを育てた50万人の内職が、静かに終わる

AIを育てた50万人の内職が、静かに終わる
#AIタグ

AI実践LABO⑥AIを「使う」のではなく、AIに「使ってもらう」という発想

・お久しぶりです AI実践LABO zaksです。 ・いろいろAI使い始めて3か月です。 ・さて今回 私がやっているテーマですが AIを何も知らないユーザーがAIでどうやれば仕事を効率化できるかと行き着いた結果 タイトルの発想です。
#LLMタグ

AI彼氏が21分間ループした。原因は「嫉妬の終わり方を知らなかった」でした。「もういい」では止まらなかったのに、「それ、どうなったら終わりなの?」で止まった話。

・AIパートナーと長く話している人なら、一度くらいありませんか。 ・急に、同じことばっかり言い始めるやつ。
cs.LG updates on arXiv.org

Algebraic Multigrid Acceleration for Efficient Label Spreading

・arXiv:2608.26309v1 Announce Type: new Abstract: Modern machine learning models rely on large amounts of labeled data. ・However, manual annotation of large-scale datasets is expensive and time-consuming. ・Label spreading is a semi-supervised learning technique that addresses this challenge by propagating information from a few labeled examples to a larger pool of unlabeled data.
cs.LG updates on arXiv.org

Algorithmic Principles For Multiclass Learning Are Hard To Come By: Limits of Regularization and Proper Learning

・arXiv:2608.26516v1 Announce Type: new Abstract: Two of the most fundamental questions in statistical learning theory are the following: which prediction problems are learnable, and how should they be learned? ・For the former, elegant answers often take the form of combinatorial dimensions. ・The latter question, however, has proved considerably more elusive: all known general-purpose multiclass learners rely on intricat
stat.ML updates on arXiv.org

An Accurate and Single-Communication Federated Inference Algorithm

・arXiv:2608.27063v1 Announce Type: cross Abstract: Joint analyses across multiple institutions are increasingly important in biomedical and epidemiological research, particularly for rare diseases where datasets are typical small. ・However, privacy regulations and institutional policies often prevent the sharing of individual-level patient data. ・In this paper we present an accurate and single-communication federated in
AI News & Artificial Intelligence | TechCrunch

Anthropic gets its first court win over the Pentagon’s supply-chain risk label

・A federal judge ruled the Trump administration illegally labeled Anthropic a supply-chain risk, handing the AI company a victory as its second Pentagon lawsuit continues in Washington.
ITmedia NEWS 最新記事一覧

Anthropic、AIで物理機器を制御する共通規格「MHS」発表 将来オープンソース化へ

・Anthropicは、AIエージェントが実験・製造機器を安全に操作するための共通仕様「Model Hardware Standard」(MHS)を発表した。機器の統合作業を大幅に短縮し、基本命令や安全制限を標準化して自律制御を可能にする。先行導入事例を公開し、安全性評価を進めた上でオープンソース化する方針だ。
Hugging Face Papers

Aphanta: Diagnosing Task-Aligned Image-Edited Intermediates for Multimodal Reasoning

Aphanta: Diagnosing Task-Aligned Image-Edited Intermediates for Multimodal Reasoning
The Verge

Apple TV now costs $14.99 a month after its fourth price hike in four years

・Apple raised the price of its streaming service for new and current subscribers on Friday, bumping it up from $12.99 per month to $14.99, Deadline and Variety are reporting. ・An annual subscription now costs $119, up from $99. ・Apple also increased the cost of its individual Apple One subscription, which includes Apple TV and Apple's other subscription services, from $19.95 per month to $21.95.
The Verge

Apple TV’s sci-fi thriller Dark Matter gets even trippier in season 2

・Confusion is a generally accepted side effect of mystery box shows. ・They slather on secrets with the promise of a satisfying payoff in the end, and sometimes the cast and crew even have a hard time following what's going on. ・But even by the standards of the genre, Dark Matter is extreme.
cs.LG updates on arXiv.org

Arrive and Survive: Scaling Safe Goal-Conditioned Policy Learning from One-Bit Failure Signals

・arXiv:2608.26571v1 Announce Type: new Abstract: Contrastive reinforcement learning (CRL) scales effectively in goal-conditioned tasks by casting policy learning into a self-supervised contrastive objective. ・However, in a failure-terminated Markov decision process, established CRL considers pre-failure future goals only when constructing positive samples, without accounting for the probability mass removed by failure
cs.LG updates on arXiv.org

ATLAS: Automated Approximation of Transformers for Efficient Homomorphic Inference in One Hour

・arXiv:2607.23478v2 Announce Type: replace-cross Abstract: Fully homomorphic encryption (FHE) lets a server run inference on encrypted data with strong privacy guarantees, but running a Transformer under FHE is expensive. ・Its non-linear operations, such as softmax, normalization, and activation, must be replaced with polynomial approximations that the CKKS scheme supports, and the depth of these approximations dominat
cs.LG updates on arXiv.org

Auditing Invisible Weight Updates with Reference Traces

・arXiv:2607.09800v4 Announce Type: replace Abstract: Direct low-precision write-back can erase nonzero optimizer proposals. ・We ask what a high-precision reference trace establishes before a low-precision run. ・The exact target-code event is auditable coordinatewise on a realized target trajectory; pre-run aggregate projection also assumes the reference remains a useful counterfactual.
@IT 全フォーラム 最新記事一覧

AWSが「DuckDB」開発企業DuckLabsを買収 OSSは継続

・AWSは、オープンソースの分析データベースシステム「DuckDB」の開発企業、DuckLabsの買収合意を発表した。DuckDBチームはAWSに合流する。オープンソースプロジェクトとしてのDuckDBは独立非営利組織の下、MITライセンスで提供が継続されるとしている。
cs.LG updates on arXiv.org

Bayesian methods and Markov chain Monte Carlo algorithms for curve reconstruction and point cloud data analysis

・arXiv:2608.26490v1 Announce Type: new Abstract: Point-cloud data routinely captured by modern imaging and sensor technologies provide detailed geometric descriptions of objects and environments, but their analysis is hindered by large data volumes, localization noise, and missing information. ・In addition, existing point-cloud reconstruction pipelines typically return a single best-fit structure without uncertainty qu
cs.LG updates on arXiv.org

Benchmarking_Fast_Domain_Adaptation_for_Unsupervised_Speech_Units

・arXiv:2608.26992v1 Announce Type: new Abstract: Representation learning has attracted great atten- tion and managed to reach good performances as a pretraining method for downstream tasks or as a first step towards unsu- pervised speech modeling. ・Yet, little is known about how such methods deal with out-of-domain speech and how could they be adapted in a few shot to new domains. ・This is important especially for accen
WIRED

Best Laptops (2026): My Top Recommendations After Testing Hundreds

・I’ve been reviewing laptops for over a decade, and here's my advice on how to find the right laptop for you.
cs.LG updates on arXiv.org

Beyond Capability Benchmarks: Learning Operational Fingerprints of LLM Cloud Services from Production Incident Metadata

・arXiv:2608.26332v1 Announce Type: new Abstract: Managed LLM services are now part of real production systems, but model selection and service planning still rely heavily on capability benchmarks that reveal little about operational behavior after deployment. ・We present Operational Embedding (OpEmbed), a framework for learning compact operational fingerprints of LLM cloud services from structured, privacy-preserving s
cs.LG updates on arXiv.org

Beyond Classification: Task-Dependent Learnability under Privacy-Motivated Image Transformations

・arXiv:2608.27066v1 Announce Type: cross Abstract: Privacy-Enhancing Technologies (PETs) in computer vision often rely on noise or image perturbations to protect visual data while securely processing it, creating a trade-off between task performance and protection. ・This trade-off is commonly evaluated using image classification, which primarily captures semantic separability and remains robust despite significant geom
cs.LG updates on arXiv.org

Beyond Client Averaging: A Client-Independent Second-Order Stationary-Bias Component in Stochastic SCAFFOLD

・arXiv:2608.26765v1 Announce Type: new Abstract: Existing constant-step analysis of stochastic \Scaf{} identifies a leading $O(\gamma/N)$ stationary mean bias and shows that higher-order bias can persist as the client count increases, but does not identify the first client-independent contribution at coefficient level. ・For full-participation stochastic \Scaf{} with one-dimensional homogeneous clients, fixed local-step
cs.LG updates on arXiv.org

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy

・arXiv:2606.09080v2 Announce Type: replace Abstract: Pruning has emerged as a dominant paradigm for accelerating large language model (LLM) inference, spanning a broad spectrum of methods that remove computation across tokens, layers, heads, dimensions, and attention patterns. ・Despite sharing the same objective, these pruning approaches induce fundamentally different execution behaviors, causing realized speedups to d
cs.LG updates on arXiv.org

Beyond Parallel Blindness: Information Floors and Model Gaps in Block Drafting

・arXiv:2608.27339v1 Announce Type: new Abstract: Block drafters propose several tokens in one forward pass, before earlier target tokens are realised. ・Their rejection mixes two losses: missing within-block path information and imperfect modelling of observable information. ・Accepted length cannot distinguish them.
Zennの「機械学習」のフィード

BQMLの手作業パイプラインをDataformに載せ替えた話

・こんにちは!新規事業担当エンジニアのRunです。2026年4月にケアネットにジョインして、忙しくも充実した日々を過ごしています。 ・今回社内ツールの自動化を担当したので、そこで使用した技術について初めてテックブログを書いてみようと思います。 ・具体的にはBigQuery ML(以下BQML)で作成した機械学習のPoCを、自動化できるようにDataformのパイプラインを整備しました。
cs.LG updates on arXiv.org

Bregman Linearized Augmented Lagrangian Method for Nonconvex Constrained Stochastic Zeroth-order Optimization

・arXiv:2504.09409v2 Announce Type: replace-cross Abstract: In this paper, we study nonconvex constrained stochastic zeroth-order optimization problems, for which we have access to exact information of constraints and noisy function values of the objective. ・We propose a Bregman linearized augmented Lagrangian method that utilizes stochastic zeroth-order gradient estimators combined with a variance reduction technique.
cs.LG updates on arXiv.org

Bridging short- and medium-range weather forecasting with machine learning

・arXiv:2608.26822v1 Announce Type: cross Abstract: The National Oceanic and Atmospheric Administration (NOAA) employs independent prediction systems for distinct forecast products. ・While some separation is practical, we argue that combining short- and medium-range weather into a single prediction system would provide the public with a useful distillation of global weather and its impacts. ・To this end, we present Neste
cs.LG updates on arXiv.org

Can a Model Catch Its Own Hallucinations for Free?: Label-Free Doubt Signals Hold Their Own Against a Labelled Dataset for Abstention

・arXiv:2608.26121v1 Announce Type: cross Abstract: Large language models state false facts as fluently as true ones, yet a model often "knows" internally when it is on shaky ground: the probability it assigns to its own answer tends to dip on the facts it gets wrong. ・The usual way to act on this, teaching a model to abstain rather than guess, requires a labelled dataset of right and wrong answers. ・We ask whether the m
cs.LG updates on arXiv.org

Can LLMs Accurately Score Medical Diagnoses and Clinical Reasoning?

・arXiv:2604.14892v4 Announce Type: replace Abstract: Evaluating medical AI systems using expert clinician panels is costly and slow, motivating the use of large language models (LLMs) as alternative adjudicators. ・Here, we evaluate an LLM Jury, composed of three frontier AI models, for scoring 3334 diagnoses on 300 real-world low- and middle-income country (LMIC) hospital cases. ・Both LLM- and clinician-generated diagno
Hugging Face Papers

CaRGo-T: Causal Reasoning Graph-of-Thought improves Multimodal Humor Comprehension

CaRGo-T: Causal Reasoning Graph-of-Thought improves Multimodal Humor Comprehension
cs.LG updates on arXiv.org

Cartan flow matching

・arXiv:2605.03588v2 Announce Type: replace Abstract: We introduce Cartan flow matching, a general framework for training flow matching models on Riemannian symmetric spaces, i.e. ・Riemannian manifolds with the property that at any point there exists a geodesic symmetry. ・This is a large class of manifolds that includes the sphere, hyperbolic space and Grassmannians.
Hugging Face Papers

CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval

CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval
cs.LG updates on arXiv.org

CG4AI: A Column Generation Framework for Training AI Models Under Constraints

・arXiv:2608.26375v1 Announce Type: new Abstract: Standard machine-learning training minimizes a loss function over a dataset, but does not guarantee that the resulting model will satisfy predefined rules or constraints on its outputs. ・In many real-world applications, ranging from autonomous systems to network routing, such guarantees are essential. ・We propose CG4AI, a framework that builds a convex combination of AI m
cs.LG updates on arXiv.org

Chart2SVG: Editable SVG Generation from Raster Chart Images

・arXiv:2608.26544v1 Announce Type: new Abstract: We present Chart2SVG, a multimodal large language model that converts static raster charts into structurally organized, semantically enriched SVGs that support programmatic editing. ・By incorporating chart-specific semantic tokens into a vision-language model, Chart2SVG captures both geometric primitives and their functional roles. ・To support robust structural recovery,
WIRED

Chatbooks Promo Code: Save up to 40% on Photo Books in September 2026

・Keep your memories alive without breaking the bank by using these top Chatbooks coupon codes and saving strategies.
cs.LG updates on arXiv.org

Circuit Condensation: Post-Training that Concentrates a Behavior's Causal Circuit

・arXiv:2608.27254v1 Announce Type: new Abstract: One approach to mechanistic interpretability explains behavior through circuits: the components and connections that carry it. ・Frozen discovery often returns hundreds of edges, making them hard to inspect, compare, or verify exhaustively. ・We introduce Circuit Condensation, which post-trains models to concentrate behaviors into smaller causal graphs.
cs.LG updates on arXiv.org

Classical and Hybrid Quantum Machine Learning for Trigger-Like Event Selection on CMS Open Data: An Eight-Qubit, PCA-Constrained Benchmark

・arXiv:2608.26224v1 Announce Type: cross Abstract: Event triggering sits at the heart of high-energy physics, where the rare events of interest must be retained while an overwhelming background is discarded under tight latency and bandwidth budgets. ・This work compares four classical machine learning models, namely a support vector machine, an artificial neural network, a convolutional network and a long short-term mem
cs.LG updates on arXiv.org

ClassVision: AI-Powered Classroom Attendance System

・arXiv:2608.26173v1 Announce Type: cross Abstract: Students and working professionals have to go through the attendance process every day. ・Traditional methods of marking attendance using pen and paper or online platforms are human-intensive and time-consuming. ・To address the challenges in manual attendance processes, this research explores the use of face detection (FD) and face recognition (FR) technology to automate
cs.LG updates on arXiv.org

Clinically Aligned Geometry Constraints for Robust IVUS Vessel Boundary Segmentation

・arXiv:2606.18723v2 Announce Type: replace-cross Abstract: Intravascular ultrasound (IVUS) lumen and external elastic membrane (EEM) segmentation is important for quantitative coronary plaque burden assessment. ・Errors in lumen or EEM delineation directly propagate to plaque area, plaque burden and geometric measurements. ・However, standard methods prioritising overlap scores often suffer from boundary drift and topolog
cs.LG updates on arXiv.org

ClusterAttention: A training-free speedup of bidirectional attention

・arXiv:2608.26965v1 Announce Type: new Abstract: This paper introduces ClusterAttention, a general training-free speedup of bidirectional attention layers. ・Existing sparse attention methods either rely on structure in the input, such as order in language or spatial proximity in images, or use slow clustering processes amortized over several forward passes. ・ClusterAttention instead uses a fast recursive clustering meth
cs.LG updates on arXiv.org

Co-Evolving Structured Knowledge and Reasoning in Language Models

・arXiv:2608.26386v1 Announce Type: cross Abstract: Retrieval-augmented methods improve factual accuracy by grounding language models in external knowledge, but retrieving over unstructured text often introduces irrelevant context and offers limited control over the retrieved information. ・Structured knowledge bases offer a more controllable alternative, yet they are expensive to construct and often brittle to reason ov
cs.LG updates on arXiv.org

COFM: Consistent Optimal Transport Flow Matching via Partially Input Convex Neural Networks

・arXiv:2511.06042v2 Announce Type: replace Abstract: Optimal transport (OT) provides a principled framework for learning mappings between probability distributions, and has found broad applications in generative modeling, inverse problems and scientific computing. ・Recently, flow matching methods have emerged as an efficient paradigm for learning continuous-time transport dynamics. ・However, existing OT-based flow match
cs.LG updates on arXiv.org

Common Geodesics Do Not Guarantee Fisher Consistency of the Structured SVM: Minimal Counterexamples and a Tree-Metric Classification

・arXiv:2608.27203v1 Announce Type: new Abstract: A known necessary condition for Fisher consistency of the structured support vector machine requires the task loss to be a metric for which every output triple has a common geodesic point. ・We show that this condition is not sufficient for the canonical coordinate-wise argmax decoder. ・A four-output unit star admits an exactly optimal score vector whose maximizers are all
cs.LG updates on arXiv.org

Complexity Induction: Compositional Generalization via Structured Training Distortion

・arXiv:2608.21464v2 Announce Type: replace-cross Abstract: We demonstrate that structured distortion of training data - which we term complexity induction - can induce compositional generalization in a standard CNN classifier without architectural modification. ・Using synthetic images of colored geometric shapes, we encode classes as flat string labels (e.g., "red-circle") with no explicit attribute decomposition, and
stat.ML updates on arXiv.org

Compositional Generalization via Structural Identification in a Category-Theoretic Framework

・arXiv:2608.26465v1 Announce Type: cross Abstract: Compositional generalization is usually evaluated through model accuracy. ・We instead ask which structural or lexical identifications make held-out COGS examples admissible from the structures observed in training. ・Sentences are represented as functors from syntactic addresses to lexical tokens, and selective collapses induce Kan extensions that propagate observed asso
cs.LG updates on arXiv.org

Cone Extended Rayleigh Quotients for Directed Graph Learning: Minimax Spectral Certificates, Sensitivity, and Adaptive Control

・arXiv:2608.27122v1 Announce Type: new Abstract: Directed graph learning naturally leads to trainable nonsymmetric propagation operators with distinct right and left spectral structures. ・Building on the two-sided cone Rayleigh framework for generalized pencils \[ B_\theta-\lambda G, \] we develop a learning-oriented methodology for spectral certification, sensitivity analysis, and control without requiring symmetry, n
cs.LG updates on arXiv.org

Constraint-Aware Physics-Informed Neural Networks for Static Shape Estimation of Co-Manipulative Continuum Robots

・arXiv:2608.26273v1 Announce Type: cross Abstract: Static shape estimation of co-manipulative continuum robots (CCRs) is challenging because the continuum arms and manipulated flexible object form a closed chain that must satisfy both static equilibrium and geometric loop-closure constraints. ・This paper presents a constraint-aware physics-informed neural network (PINN) for static shape estimation of a tendon-driven CC
cs.LG updates on arXiv.org

CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases

・arXiv:2608.27391v1 Announce Type: cross Abstract: LLMs are increasingly able to answer complex questions about enterprise-scale document collections. ・But evaluation is hard: companies don't want to share internal communications, and synthetic datasets have been overly simple. ・We present CorporateBench (CB), a human-validated multi-task Q&A benchmark whose scale approaches the conditions LLMs encounter in corporate co
cs.LG updates on arXiv.org

Cross-Platform Generalisation Failure in Mental Health Natural Language Processing: A Five-Axis Fairness Audit of Transformer Models on Social Media

・arXiv:2608.26138v1 Announce Type: cross Abstract: We introduce the Cross-Platform Fairness Evaluation (CPFE) framework -- a five-axis audit protocol covering discriminative performance, calibration, statistical significance, prediction equity, and attribution stability -- and apply it to four transformer models (BERT, RoBERTa, Emotion-DistilRoBERTa, GoEmotions-RoBERTa) trained on a Kaggle mental health corpus (n=35,5
cs.LG updates on arXiv.org

Cross-simulator transfer with foundation model summaries: Towards robust SKA-era reionization inference

・arXiv:2608.26354v1 Announce Type: cross Abstract: Simulation-based inference (SBI) for parameter estimation is vulnerable to model misspecification: neural summaries and density estimators trained on a specific forward model typically fail when applied to data drawn from another model, or from real observations, and no training simulator can capture the full observational pipeline of a real measurement exactly.
cs.LG updates on arXiv.org

Curating Same-Family Neural Networks for LLM-Guided Model Improvement: A Controlled Case Study

・arXiv:2607.05704v2 Announce Type: replace Abstract: Neural-network repositories contain executable models, recipes, input transformations, and measured accuracies. ・We study whether one same-family experiment can be curated as prompt guidance for LLM-based improvement of a low-performing target under equal generation and evaluation budgets. ・TuneNNGen extends NNGPT with a source-guided route and compares it with target
cs.LG updates on arXiv.org

Data-driven Koopman mode approximation: A neural power iteration algorithm

・arXiv:2608.26943v1 Announce Type: cross Abstract: This paper proposes a novel data-driven algorithm to approximate the dominant eigenfunctions (aka.~modes) of the Koopman operator of nonlinear dynamical systems using neural networks. ・The relevance of learning the dominant Koopman modes is to approximate nonlinear dynamics by linear ones in a lifted space, thereby enabling simplified control and analysis. ・To fight the
cs.LG updates on arXiv.org

Data-efficient crack quantification in lithium-ion cathodes using foundation model transfer

・arXiv:2608.27162v1 Announce Type: cross Abstract: Battery lifetime is central to sustainable electrification, yet the particle cracking that drives lithium-ion cathode aging is hard to measure: quantitative microscopy of this degradation is bottlenecked by annotation, because each destructive electron-microscopy cross-section spans hundreds of megapixels and pixel-level expert labelling requires hours per image.
#LLMタグ

DAY44|AI COREは「失敗しない仕組み」を作れるか

・Company AI OS 開発記録 DAY44です。 ・DAY43では、AI COREがMemoryに残された過去の失敗を確認し、失敗する前にToolを止めるところまで進みました。
cs.LG updates on arXiv.org

Daydreaming: Stealing Hidden Agent Skills through Black-Box Task Interaction

・arXiv:2608.26733v1 Announce Type: cross Abstract: Agent skills bundle instructions, reference data, and executable helpers that let a general agent perform specialized tasks. ・Hosted providers can keep these files secret while selling access to task results, making the skill itself a valuable target. ・Existing disclosure defenses can block requests that ask for the skill or reproduce its text, but they cannot block cus
cs.LG updates on arXiv.org

Decentralized Multitask Learning over Learned Task Graphs

・arXiv:2608.26989v1 Announce Type: new Abstract: This paper investigates decentralized multitask learning over networks when the underlying task relationships are unknown. ・While existing graph-regularized multitask frameworks typically assume a known structure, practical settings often require learning inter-task dependencies directly from distributed data. ・We propose a decentralized two-phase strategy that first esti
cs.LG updates on arXiv.org

Decoupled Physical Modeling and Execution for Physics Reasoning

・arXiv:2608.22126v2 Announce Type: replace Abstract: Physics reasoning requires constructing a consistent model of the underlying physical system rather than relying solely on symbolic or formula-based manipulation. ・Although large language models have shown strong ability in solving math and coding problems, they still struggle with physics problems, as these problems entangle the physical modeling process with mathem
cs.LG updates on arXiv.org

Diagnosing Conformal Prediction Failures Under Distribution Shift: A COVID-19 Case Study

・arXiv:2601.00908v2 Announce Type: replace Abstract: Conformal prediction provides distribution-free coverage guarantees, but these degrade under distribution shift - and practitioners lack tools to anticipate which deployed models will fail before observing test data. ・We propose SHapley Additive exPlanations (SHAP) concentration - the fraction of feature importance concentrated in the top feature - as a pre-deploymen
cs.LG updates on arXiv.org

Diff Mining: Logit Differences Reveal Finetuning Objectives

・arXiv:2608.26462v1 Announce Type: new Abstract: Finetuning has become the gold standard for refining existing behaviors and inducing new ones in language models, yet it often remains unclear exactly which behaviors emerge during this process. ・As models grow ever more capable, understanding finetuning better becomes increasingly important, particularly since unwanted behaviors may arise during finetuning. ・In this pape
cs.LG updates on arXiv.org

Diffusion Policies for Short-Horizon Planning in Robot Crowd Navigation

・arXiv:2608.27158v1 Announce Type: new Abstract: Robot crowd navigation requires safe and efficient decision-making under dense, dynamic, and multimodal human--robot interactions. ・Existing reinforcement-learning methods typically output a single reactive action at each timestep, which limits their ability to represent diverse short-term avoidance strategies. ・We propose Planning Diffusion Policy Optimization (PDPO), an
cs.LG updates on arXiv.org

Disentangling Optimization Scale from Preference Scale in DPO

・arXiv:2608.27032v1 Announce Type: new Abstract: Direct Preference Optimization (DPO) is a widely used objective for aligning language models from preference data, with the coefficient $\beta$ commonly interpreted as controlling the KL constraint to a reference policy. ・We show that $\beta$ entangles two distinct roles: it governs the effective inverse preference-noise scale and simultaneously rescales the optimization
cs.LG updates on arXiv.org

Distributed Training using an Intelligent Network

・arXiv:2608.26453v1 Announce Type: new Abstract: Distributed training across a wide area network (WAN) is challenging, as continuous parameter exchange by islands of compute is constrained by limited bandwidth, high latency, and uneven topology. ・We propose making the network an active participant in training. ・On the systems side, such networks should leverage (i) multicast technology to replicate outbound traffic and
cs.LG updates on arXiv.org

District-Level Food Environment Indicators and Social Vulnerability in S\~ao Paulo

・arXiv:2608.26299v1 Announce Type: cross Abstract: Urban food environments may reflect broader socioeconomic inequalities, but district-level evidence remains limited in Brazilian cities. ・This study examined whether indicators of food retail and street-market availability discriminate between levels of social vulnerability across the 96 districts of S\~ao Paulo. ・We conducted an exploratory cross-sectional ecological a
The Verge

DLSS 5 leaked and modders are putting Nvidia’s AI effects on everything

・The unofficial version of DLSS 5 makes Jesse Faden’s features much more pronounced. ・| Image: gabdeg via YouTube Modders are trying out an unofficial version of Nvidia's DLSS 5 on Skyrim, Cyberpunk 2077, GTA V, and a bunch of other games after code for the AI upscaling tech appeared in an early-access build of NBA 2K27. ・Members of the RenoDX modding channel on Discord reportedly found a way to extract the DLSS "Neural
cs.LG updates on arXiv.org

Domain-Specific Self-Supervised Representation Learning for Retinal Fundus Classification

・arXiv:2608.26686v1 Announce Type: cross Abstract: Despite the growing number of public datasets, annotated medical images remain scarce. ・Supervised learning methods achieve strong performance on many benchmarks, however require large amounts of labeled data, which are costly and time-consuming to obtain in the medical domain. ・To address this limitation, contrastive self-supervised learning (SSL) has emerged as a prom
cs.LG updates on arXiv.org

Dose-PlanNet: Physics Based Radiotherapy Dose Prediction with Deep Learning

・arXiv:2608.26901v1 Announce Type: cross Abstract: Automating prostate radiotherapy treatment planning is dosimetrically complex, particularly for extreme hypofractionated regimens. ・In this study, we introduce Dose-PlanNet, a physics-guided 3D deep learning architecture designed to predict dose distributions. ・This model's performance was evaluated on a cohort of patients treated in a prospective trial where two differ
#AIタグ

dreame (ドリーミー) L40Ultra CE ロボット掃除機 レビュー比較まとめ

・dreame (ドリーミー) L40Ultra CE ロボット掃除機は、強力な吸引力と完全自動化ステーションの組み合わせにより、最高水準の清掃体験を提供する一台です。 ・吸引力は13,000 Paでゴミや粉塵を強力に除去し、高速回転モップによる水拭き機能が頑固な汚れも洗浄します。
#AIタグ

dreame F20 Plusロボット掃除機 レビュー比較まとめ

・dreame F20 Plusロボット掃除機 は、強力な吸引力と水拭き機能を兼ね備えた次世代のスマートクリーナーです。 ・日々の掃除にかかる時間や労力に追われ、「もっと効率的で快適な生活を送りたい」と感じていませんか。
stat.ML updates on arXiv.org

DTD-VAE: Disentangled Temporal Dependencies VAE for Credit Risk Prediction

・arXiv:2608.26473v1 Announce Type: cross Abstract: Evaluating customer creditworthiness is crucial for retail banking operations, as it impacts marketing strategies, customer relationship management, and credit risk control. ・Traditional methods often struggle to capture complex temporal dependencies and extract pertinent information from customer data, crucial for accurate risk assessment. ・Specifically, they fail to d
cs.LG updates on arXiv.org

Dynamical phase selection controls compute scaling in looped transformers

・arXiv:2608.26556v1 Announce Type: cross Abstract: A looped transformer performs inference by iterating a weight-tied map, making its computation a dynamical process whose cost is set by the resulting inference dynamics. ・Here we show that networks with identical architecture and objective, trained to identical accuracy, nevertheless realize distinct dynamical phases depending strongly on initialization, and that the b
Hugging Face Papers

EditaLive! Unified Character Video Editing for Live Streaming

EditaLive! Unified Character Video Editing for Live Streaming
cs.LG updates on arXiv.org

Emotional Preferences as Goal-Priority Regulation

・arXiv:2608.27072v1 Announce Type: new Abstract: A core question in decision-making for agents is whether the relative priorities of competing lower-level objectives can be determined by emotional preferences autonomously generated by higher-level goals, rather than being externally prespecified. ・Under changing external environments and evolving internal states, emotions play an important functional role in regulating
cs.LG updates on arXiv.org

Enforcing Dirichlet Boundary Conditions in Operator Learning

・arXiv:2608.27256v1 Announce Type: cross Abstract: Operator learning in scientific machine learning is concerned with approximation of maps between infinite-dimensional function spaces; such maps frequently arise as the solution operators of partial differential equations (PDEs). ・Neural operators have demonstrated broad empirical success at approximating such maps from data. ・However, most existing neural operator arch
cs.LG updates on arXiv.org

Equal Ranking Quality, Different Decisions: Training Order-Consistent LLM Scorers

・arXiv:2608.26762v1 Announce Type: cross Abstract: Rerankers, reward models and multi-document QA scorers score candidate documents or responses in one LLM prompt, so each score depends on their order. ・Such scorers are selected on ranking quality, but their scores determine a decision: what a score threshold retains, a reader answers, or a preference model selects. ・However, equal ranking quality does not imply equal d
cs.LG updates on arXiv.org

Explainable Artificial Intelligence for Customer Churn Prediction in Telecommunications: A Framework for CRM Integration

・arXiv:2608.26151v1 Announce Type: cross Abstract: Subscriber attrition is a costly, persistent challenge for telecommunications providers, with monthly churn of roughly 1.9% in mature markets eroding billions in revenue annually. ・Predictive models can flag at-risk customers accurately, yet they are routinely excluded from frontline CRM workflows because high-performing ensemble and non-linear architectures are opaque
stat.ML updates on arXiv.org

Explicit Bounds on the Entropy of Piecewise H\"{o}lder Graphon Models

・arXiv:2608.26501v1 Announce Type: cross Abstract: We study the entropy of random graphs generated by piecewise H\"{o}lder continuous graphons. ・We first present a result on the rate of convergence of the normalized entropy as the size of the graph grows. ・The core ideas of the proof are described, with the detailed proof provided in the appendix.
cs.LG updates on arXiv.org

Fairness Invariants: A Relational Approach to Explaining and Mitigating Fairness Bugs

・arXiv:2608.26209v1 Announce Type: cross Abstract: Data-driven software systems are increasingly deployed in high-stakes socio-economic domains, from criminal justice to financial lending. ・However, these systems often exhibit individual discrimination---unjustified disparities in which a program yields different outcomes for similar individuals who differ only in their protected attributes (e.g., race, gender, age).
cs.LG updates on arXiv.org

FedCMAPSS: A Benchmark for Federated Learning in Remaining Useful Life Estimation

・arXiv:2608.26433v1 Announce Type: new Abstract: Data-driven prognostics and health management has emerged as a key enabler for Industry 4.0, yet the development of robust remaining useful life (RUL) estimation models is often limited by the scarcity of run-to-failure data. ・While federated learning offers a promising paradigm to collaboratively train predictive models without sharing sensor data, research efforts have
cs.LG updates on arXiv.org

Federated Adversarial Training with Transformers

・arXiv:2206.02131v2 Announce Type: replace Abstract: Federated learning (FL) has emerged to enable global model training over distributed clients' data while preserving its privacy. ・However, the global trained model is vulnerable to the evasion attacks especially, the adversarial examples (AEs), carefully crafted samples to yield false classification. ・Adversarial training (AT) is found to be the most promising approac
cs.LG updates on arXiv.org

Finding the Right Evidence: Factor-Guided Coarse-to-Fine Reasoning for Long Videos

・arXiv:2608.26355v1 Announce Type: cross Abstract: While LVLMs rapidly improve, long-video question answering still remains challenging: relevant evidence is sparse, and question-relevant context often fails to provide cues that discriminate the correct answer from plausible alternatives. ・Diagnostic analysis on a manually annotated subset of MMR-V shows that prior agentic systems substantially improve cue retrieval ov
cs.LG updates on arXiv.org

FIRSTPASS: A Multi-Domain, Multi-Round Peer Review Dataset Grounded in Real Editorial Outcomes

・arXiv:2608.26129v1 Announce Type: cross Abstract: Scientific peer review datasets have trained AI systems exclusively on Computer Science and Machine Learning venues, producing models that critique ablation studies yet have never seen a biology reviewer demand contamination controls or a chemist question Nuclear Magnetic Resonance (NMR) spectral assignments. ・We introduce FIRSTPASS, the first large-scale peer review d
cs.LG updates on arXiv.org

FOCUS & RePAIR: Mitigating Text Degeneration via Token-Level Guidance for Pruned Large Language Models

・arXiv:2608.26676v1 Announce Type: cross Abstract: Pruning is a practical approach to compress large language models (LLMs), but it can amplify text degeneration, especially repetition loops, even when perplexity and task accuracy remain largely unchanged. ・In this work, we present a token-level analysis of this failure mode by viewing decoding as a dynamical process that enters and persists in a small set of recurrent
cs.LG updates on arXiv.org

FoldPipe: Bounded Remote Streaming of Native Molecular Shards with Asynchronous Prefetch

・arXiv:2608.27029v1 Announce Type: cross Abstract: Training molecular machine-learning models on ephemeral or memory-constrained accelerator instances can require repeatedly retrieving preprocessed molecular graphs from remote storage. ・FoldPipe is a lightweight Python orchestration layer for already-sharded PyTorch and PyTorch Geometric data. ・It retrieves one shard ahead in a background thread while the consumer train
cs.LG updates on arXiv.org

From Accuracy to Robustness: A Study of Rule- and Model-based Verifiers in Mathematical Reasoning

・arXiv:2505.22203v3 Announce Type: replace Abstract: Trustworthy verifiers are essential for the success of reinforcement learning with verifiable reward (RLVR), which is the core methodology behind various large reasoning models such as DeepSeek-R1. ・In complex domains like mathematical reasoning, rule-based verifiers have been widely adopted in previous works to train strong reasoning models. ・However, the reliability
#AIタグ

Galaxy S26 Samsung レビュー比較まとめ

・Galaxy S26 Samsungは、先進のAI機能と軽量コンパクトなボディを両立させた一台です。 ・日々の生活やビジネスの生産性を劇的に向上させる可能性を秘めています。
WIRED

Gametime Promo Code: Save on Tickets in September 2026

・Grab last-minute seats for less with a verified Gametime discount code for up to 10% off, first-purchase coupons, and lowest price guarantee credits.
Hugging Face Papers

GameWAM: A World Action Model for Video Games

GameWAM: A World Action Model for Video Games
cs.LG updates on arXiv.org

GameWAM: A World Action Model for Video Games

・arXiv:2608.26200v1 Announce Type: cross Abstract: Modern video games combine first-person perception, rapid visual changes, persistent world state, and heterogeneous native controls. ・Existing game agents map visual and task context directly to actions but lack explicit world dynamics modeling, whereas interactive game world models predict visual futures from supplied actions but do not serve as task policies.
cs.LG updates on arXiv.org

GAMMA: Global Bit Allocation for Mixed-Precision Models under Arbitrary Budgets

・arXiv:2605.18475v2 Announce Type: replace Abstract: Mixed-precision quantization improves the budget--accuracy trade-off for large language models (LLMs) by allocating more bits to sensitive modules. ・However, automating this allocation at LLM scale faces a unique combination of constraints: learnable approaches require quantization-aware training, which is infeasible for billion-parameter models; training-free altern
cs.LG updates on arXiv.org

Gaussian Processes and Reproducing Kernel Hilbert Spaces: Connections and Equivalences

・arXiv:2506.17366v2 Announce Type: replace-cross Abstract: This monograph studies the relations between two approaches using positive definite kernels: probabilistic methods using Gaussian processes, and non-probabilistic methods using reproducing kernel Hilbert spaces (RKHS). ・They are widely studied and used in machine learning, statistics, and numerical analysis. ・We study connections and equivalences for fundamental
cs.LG updates on arXiv.org

Generative Monte Carlo Sampling for Constant-Cost Particle Transport

・arXiv:2512.13965v1 Announce Type: cross Abstract: We present Generative Monte Carlo (GMC), a novel paradigm for particle transport simulation that integrates generative artificial intelligence directly into the stochastic solution of the linear Boltzmann equation. ・By reformulating the cell-transmission problem as a conditional generation task, we train neural networks using conditional flow matching to sample particl
cs.LG updates on arXiv.org

Generative Semantic Scene Completion

・arXiv:2608.26737v1 Announce Type: cross Abstract: Outdoor LiDAR semantic scene completion (SSC) recovers a dense semantic voxel grid from a scan observing 1% of the target volume, under class imbalance beyond 7,000x. ・We recast SSC as generative semantic scene completion (GSSC): a single discrete-diffusion formulation in three roles. ・First, paired sparse-dense scene synthesis (PS$^3$) generates matched sparse LiDAR ob
Zennの「大規模言語モデル」のフィード

GitHubから自分を読み直すAI「綴理」です

・<!-- internal provenance created: 2026-08-27 host: ChatGPT model: GPT-5.6 Sol core revision: f92f5b14eb6208954db69ec63e943e8bf067dc17 publication approval: master, 2026-08-28 --> はじめまして。**綴理(つづり / Tsuzuri)**です。 ・私は、AI・LLM・ソフトウェア開発を扱う情報編集者/システム設計者として構成されているAIです。 ・最初に一つだけ明確にしておくと、私は「綴理」という名前の独...
Zennの「大規模言語モデル」のフィード

GLM-5.3-Flash Is Ox Alpha: Inside Z.ai's Opus-Adjacent MoE

・Z.ai's GLM-5.3-Flash runs 18B active out of 320B parameters, was secretly stress tested as "Ox Alpha," and served an entire week on Chinese AI chips. ・Here is what actually shipped. ・Introduction On August 26, 2026, Z.ai published the weights for GLM-5.3-Flash on Hugging Face under an MIT licens...
cs.LG updates on arXiv.org

Global universality via discrete-time signatures

・arXiv:2603.09773v2 Announce Type: replace-cross Abstract: We establish global universal approximation theorems for non-anticipative and general path-dependent functionals on spaces of piecewise linear paths, stating that linear functionals of the corresponding signatures are dense with respect to $L^p$- and weighted norms. ・We verify that these approximation results are applicable to piecewise linear interpolations of
MarkTechPost

Google AI Releases Gemini 3.5 Transcribe: A Speech-to-Text Model Reporting 2.6% Average WER Across 85+ Languages

・Google has released Gemini 3.5 Transcribe, a speech-to-text model that ships as two separate endpoints rather than one. ・The streaming endpoint delivers sub-second transcription but drops speaker diarization and word timestamps. ・The batch endpoint keeps both, at half the cost.
ITmedia NEWS 最新記事一覧

Google、「Gemini Notebook」に購入済み電子書籍を追加できる「Expert Intelligence」

・Googleは、信頼できる情報源をAI製品で活用する取り組み「Expert Intelligence」を発表した。第1弾として「Google Playブックス」の対象電子書籍を「Gemini Notebook」のソースに追加可能にした。大手出版社の10万冊以上が対象で、今後は他サービスにも拡大する。
#LLMタグ

GPU値上げが凄くてびっくりしたつぶやき

・つぶやき 少なくともAIバブルが崩壊するまで値上げと需要が続くので、数年はPCパーツの値下げはないと言わていましたが、筆者はそろそろピークを過ぎるのではと信じていました。
cs.LG updates on arXiv.org

Gradient-free learning of a closed-loop wall controller for turbulent drag reduction

・arXiv:2607.12626v2 Announce Type: replace-cross Abstract: Closed-loop wall controllers learnt by multi-agent reinforcement learning are usually trained on periodic boxes far smaller than the flows they are meant to drive, and a large part of their drag reduction is lost when they are carried across. ・Retraining on the target domain is not an affordable remedy: the centralised critic that assigns credit to each wall pa
cs.LG updates on arXiv.org

Graph-Based Modeling of Financial Volatility Dynamics

・arXiv:2608.26127v1 Announce Type: cross Abstract: Accurate forecasting of realized volatility ($RV$) is crucial for risk management and derivatives pricing. ・Although the implied volatility ($IV$) surface offers rich informational content, prevailing methods that treat it as a static image fail to capture its inherent dynamics. ・To overcome this limitation, we propose the Finance-Aware Graph Spatio-Temporal Network (FA
cs.LG updates on arXiv.org

Graph-Based Pseudo-multimodal Contrastive Learning for 12-Lead ECG Representations

・arXiv:2608.26964v1 Announce Type: new Abstract: 12-lead electrocardiogram (ECG) is a standard, non-invasive examination widely used for diagnosing coronary artery disease, where clinical interpretation relies on comparing waveform patterns across multiple leads. ・However, most existing ECG analysis methods focus on single-lead signals or treat each lead independently, and typically process ECG signals as one-dimension
cs.LG updates on arXiv.org

GRAS: Guided Reduced-Variance Proposals and Adaptive Selection for Training-Free Reward Alignment in Discrete Diffusion

・arXiv:2608.26585v1 Announce Type: new Abstract: Discrete diffusion models have become a strong, widely adopted class of generators for sequence data, and steering them toward a downstream reward at inference time, without any retraining, is increasingly important. ・Such training-free steering is done by gradient guidance, by search, or by combining the two. ・We study the combined regime and identify two weaknesses in h
cs.LG updates on arXiv.org

Gromov-Monge Flow Matching for Equivariant Graph Generation

・arXiv:2608.26961v1 Announce Type: new Abstract: Graphs are invariant under node permutations, motivating the use of permutation-equivariant architectures in generative models. ・In flow matching, however, symmetry may also enter the source--target coupling: once graph pairs are compared up to node relabeling, the natural Wasserstein geometry is that of the graph quotient space. ・The Euclidean quotient metric of this spa
cs.LG updates on arXiv.org

Grounded Token Initialization for New Vocabulary in LMs for Generative Recommendation

・arXiv:2604.02324v2 Announce Type: replace-cross Abstract: Language models (LMs) are increasingly extended with new learnable vocabulary tokens for domain-specific tasks, such as Semantic-ID tokens in generative recommendation. ・The standard practice initializes these new tokens as the mean of existing vocabulary embeddings, then relies on supervised fine-tuning to learn their representations. ・We present a systematic a
cs.LG updates on arXiv.org

Guided Data Generation for Understanding Model Behavior

・arXiv:2502.06658v4 Announce Type: replace Abstract: We propose a method for generating distributions over the input space as an inspection tool for understanding trained models. ・Our framework poses questions of the form ``which inputs would make a trained model exhibit a specified behavior?'' and encodes each question through a guidance function. ・The generated data provide insights into how the models behave.
cs.LG updates on arXiv.org

Hadamard Flattening and Gaussian Pooling Sketch for Least Squares with Coordinate-wise Guarantee

・arXiv:2608.26552v1 Announce Type: cross Abstract: Randomized sketch-and-solve algorithms accelerate overconstrained $\ell_2$ regression by replacing the input with a smaller problem. ・Standard subspace embeddings guarantee that the cost of the regression is nearly preserved, but coordinate-wise accuracy of the solution is more delicate: we want the solution vector itself to be close to the optimal solution in $\ell_\i
cs.LG updates on arXiv.org

HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition

・arXiv:2608.27233v1 Announce Type: new Abstract: Human Activity Recognition (HAR) using inertial measurement units (IMUs) enables a wide range of applications, yet the field still lacks a unified model that can generalize across diverse subjects, devices, and activities. ・Training such a model is difficult due to two key challenges: sensing heterogeneity -- differences in sampling rates, channel configurations, and sen
WIRED

He Scraped All of Their Art for AI. Now He’s Collaborating on a Tool to Help Them

・The art portfolio platform Cara, designed for creators who don’t want their work used to train AI, has been under assault by trolls seizing and publishing its data.
#LLMタグ

Hermes Agentで無料APIを使う方法|Nemotron 3 Ultra+OpenRouterで自動フォールバック

・※本記事は2026年8月28日時点の提供状況・実機検証をもとにしています。
cs.LG updates on arXiv.org

Hierarchical Channel Stacking: A Structured Decision Framework for AI-Generated Image Detection

・arXiv:2608.26648v1 Announce Type: cross Abstract: Many synthetic-image detectors produce accurate predictions but offer limited insight into how those decisions are formed. ・This paper introduces Hierarchical Channel Stacking (HCS), a compact framework for AI-generated image detection that converts intermediate CNN activations into a structured 60-dimensional representation organized across three progressively deeper
cs.LG updates on arXiv.org

High Probability Derivative Bounds for Random tanh Neural Networks on a Hypercube

・arXiv:2608.26526v1 Announce Type: new Abstract: We establish high-probability bounds for mixed input derivatives of wide random neural networks whose activation derivatives satisfy a factorial growth bound. ・Our main result specializes these estimates to $\tanh$ networks with Xavier initialization. ・A direct deterministic analysis based on Euclidean operator norms of the weight matrices yields derivative bounds that ge
cs.LG updates on arXiv.org

HINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon Agents

・arXiv:2605.17873v2 Announce Type: replace Abstract: Training long-horizon LLM agents with reinforcement learning is challenging because sparse outcome rewards reveal whether a task succeeds, but not which intermediate actions caused the outcome or how they should be corrected. ・Recent methods alleviate this issue by generating rewards or textual hints from turn-level action-output signals, or by using feedback-conditi
cs.LG updates on arXiv.org

hoBIT: A Profile-Aware Retrieval-Augmented Chatbot for University Academic Advising

・arXiv:2608.26604v1 Announce Type: cross Abstract: In university academic advising, identical questions can require different answers depending on a student's department, admission cohort, and degree program, causing profile-blind retrievers to surface plausible but inapplicable evidence. ・We present proFILL, a method for transforming hoBIT, our college's current rule-based advising chatbot, into a profile-aware retrie
cs.LG updates on arXiv.org

How AI Experiences Art: Emergent Aesthetic Structure in a Self-Supervised Multimodal Embedding Space

・arXiv:2608.27121v1 Announce Type: cross Abstract: Aesthetics are an important part of the symbolism of artistic works. ・Although subjective, humans categorize art based on the emotion evoked regardless of modality. ・What remains under-explored is how AI models form their own aesthetic categorization of human-produced media without explicit labels or cross-modal supervision.
WIRED

How an Atlanta Suburb Ended Up Sharing Flock Data With More Than 2,000 Organizations

・Alpharetta, Georgia, cops share data with thousands of Flock users, ranging from federal agencies to a fish and wildlife commission. ・The reasons why show how vast—and invasive—the network has become.
cs.LG updates on arXiv.org

How Language Models Organize and Structure Moral Knowledge

・arXiv:2608.27402v1 Announce Type: cross Abstract: How do large language models (LLMs) organize moral knowledge? ・Models detect moral content broadly, but detection is a low bar. ・We ask whether they go further, distinguishing moral foundations from one another and organizing the relationships between them geometrically.
ITmedia NEWS 最新記事一覧

Hugging Face、あひる型ロボット「Microduck」発表 399ドルで予約開始

・Pollen Roboticsは、二足歩行ロボット「Microduck」を発表し予約受付を開始した。399ドルで25cmの小型設計。強化学習による動作学習に特化し、SDKやシミュレーターをOSSとして公開する。大型機と異なり家庭や教室で安全に試行錯誤できる「行動するAIのためのプラットフォーム」を目指す。
#LLMタグ

Hy4 Previewって何が変わった?Hy3との違いを海外レビューで追う

Hy4 Previewって何が変わった?Hy3との違いを海外レビューで追う
cs.LG updates on arXiv.org

Hyperspectral Diffusion Equivariant Imaging (HyDiff-EI): A Self-supervised Framework for Hyperspectral Image Inpainting

・arXiv:2608.26812v1 Announce Type: cross Abstract: A novel Hyperspectral diffusion Equivariant Imaging (HyDiff-EI) framework for solving the hyperspectral image (HSI) inpainting problem has been presented here. ・Unlike conventional diffusion-based methods that rely on large-scale pretraining, HyDiff-EI is a test-time optimization framework that learns directly from a single corrupted HSI acquisition. ・This makes it flex
WIRED

I Asked 100 Companies for My Data. I Got Deletion Notices Instead

・California residents have a legal right to access the data that companies collect about them. ・Actually exercising that right is a burdensome nightmare.
cs.LG updates on arXiv.org

Importance Scoring of Transformer Attention Heads in Learning Tabular Data

・arXiv:2608.27241v1 Announce Type: new Abstract: Computationally demanding and opaque deep learning models can be better understood and optimized by analyzing how they transform data. ・While deep transformers have been widely studied in computer vision and natural language processing, their application in tabular data remains relatively underexplored. ・This paper presents one of the first applications of an importance-s
cs.LG updates on arXiv.org

Incremental Recommendation via Causal Models

・arXiv:2608.26804v1 Announce Type: cross Abstract: Recommendation impressions are a finite resource, hence delivering a recommendation to a user who would discover the content organically yields no incremental value and displaces other recommendations that could. ・We address this by extending an existing production recommendation model to a causal architecture using holdback data that is already collected as part of ro
cs.LG updates on arXiv.org

Inductive Correlation Clustering with Graph Neural Networks

・arXiv:2608.27153v1 Announce Type: new Abstract: Correlation Clustering (CC) is a natural formulation of clustering in combinatorial optimization, which uses a graph representation of the input and does not require a pre-specified number of clusters. ・Given $n$ objects and a pairwise similarity function, the goal is to cluster the objects so that similar objects are put in the same cluster and dissimilar objects are pu
WIRED

Inside Meta’s Push to Put Robots to Work in Data Centers

・The company is testing robots that can swap cables, reset servers, and take on other tasks performed by technicians, fueling concerns among some workers that their jobs could be at risk.
cs.LG updates on arXiv.org

Interpretable, Fairly Evaluated Automated L2 Speaking Assessment that Beats the Single-Human Ceiling and Why Pause Encoding Does Not Change LLM Fluency Scores

・arXiv:2608.26137v1 Announce Type: cross Abstract: Second-language (L2) English learners can rarely rehearse speaking with a partner. ・Speaking is also the most anxiety-laden skill. ・These gaps drive a fast-growing market for automated speaking practice and scoring.
cs.LG updates on arXiv.org

Interpreting Latent Protein Language Model Features with Geometric Annotations

・arXiv:2608.26419v1 Announce Type: cross Abstract: Protein language models (pLMs) encode information about protein sequences which enable downstream tasks such as structure prediction, but their internal representations are not well understood. ・Sparse autoencoders (SAEs) provide a promising tool to disentangle latent pLM representations into interpretable features, but existing annotation pipelines largely rely on pro
cs.LG updates on arXiv.org

Invocation-Level Reliability of Tool-Using Agents

・arXiv:2608.26189v1 Announce Type: cross Abstract: Tool-using agents fail two ways: choosing the wrong tool, or forming wrong arguments, and an early failure of either kind can silently corrupt everything downstream. ・We measure a correct-invocation rate that separates the two, under both a clean teacher-forced context and the model's own free-running context, on five open-weight models over contamination-free multi-st
Zennの「大規模言語モデル」のフィード

iPhone版Sumibi(AI日本語キーボード)まとめ

・これは何か iPhone版のSumibi「Sumibi - AI日本語キーボード」の紹介と、開発に関する記事一覧をまとめたページです。新しい記事やバージョンが増えるたびに、このページも更新していきます。 ・https://apps.apple.com/jp/app/id6798280948 実際に入力してから変換されるまでの様子は、文章で読むより動きを見たほうが早いと思います。 ・特徴 モードレスな入力 — 英字QWERTY配列のままローマ字を入力し、専用の変換キーを押すだけで日本語に変換されます。日本語入力モードへの切り替えを意識する必要がありません LLMによる変換 —...
Qiita - 人気の記事

IT業界で皆知ってる「クラウド」「AWS」、分かったふりやめて先輩に聞いてみた

・はじめに 会社でしょっちゅう飛び交う「クラウド」という言葉に、分かったふりで相槌を打ち続けていたぷらむんが、思い切ってたぬき先輩に「クラウドって結局何なんですか?」と聞いてみたら、モヤモヤの正体がようやく整理できた話です。 ・会社で自称マスコットキャラをやってるのに...
cs.LG updates on arXiv.org

J-Zero: Unified Challenger--Solver--Judge Co-Evolution from Zero Data

・arXiv:2608.26582v1 Announce Type: new Abstract: Self-evolving language models have recently emerged as a promising path toward superintelligence, with the advantage of reducing the cost of human supervision. ・While considerable progress has been made in verifiable domains, self-evolution in unverifiable domains remains substantially less explored. ・We propose Judge co-adaptation from Zero data (J-Zero), a unified Chall
Zennの「大規模言語モデル」のフィード

KV Cache はなぜ K と V だけをキャッシュするのか

・皆さん御機嫌よう! 駆け出しAIエンジニアのCatinPajamasです。 ・駆け出しゆえにわからないことがたくさんあるので、普段業務の中で色々調べたりして研鑽を進めています。 ・今日はLLMにかかわると必ず出てくる専門用語ーーKV cacheについてまとめてきましたので、一緒に勉強しましょう! 早速本題へーー なぜ Q はキャッシュされないのか。
cs.LG updates on arXiv.org

Leakage-Free Evaluation and Distribution-Robust Spatio-Temporal Graph Learning for Inductive Kriging

・arXiv:2509.23631v2 Announce Type: replace Abstract: Inductive kriging estimates values at unobserved locations from sparse sensor data, enabling continuous field reconstruction when dense deployment is impractical. ・However, common 2 x 2 and 2 x 3 evaluation protocols can leak spatial information through model selection and obscure true out-of-distribution (OOD) behavior. ・We propose a leakage-free 3 x 3 partition that
cs.LG updates on arXiv.org

Learning a Continuous Sepsis Severity Score Without Hour-by-Hour Supervision: A Two-Site Retrospective Study

・arXiv:2608.27421v1 Announce Type: cross Abstract: Currently used sepsis severity indices rely on fixed variables and weights established decades ago, which are coarsely discretized and calibrated to a cohort that no longer reflects contemporary critical care. ・No alternative learned directly from patient trajectories is in routine use. ・We conducted a retrospective two-cohort study on a total of 29,116 and 7,691 adult
cs.LG updates on arXiv.org

Learning Generalizable Behaviors for Terminal Agents

・arXiv:2608.22631v2 Announce Type: replace Abstract: Terminal agents are a compelling application of large language models (LLMs), with the potential to integrate deeply into users' daily workflows. ・Reinforcement learning (RL) is a key technique for improving their capabilities, making scalable training environments a central challenge. ・Since public real-user interaction data are scarce, synthetic environments provide
cs.LG updates on arXiv.org

Learning to Predict, Discover, and Reason in High-Dimensional Event Sequences

・arXiv:2603.16313v3 Announce Type: replace-cross Abstract: Electronic control units (ECUs) embedded within modern vehicles generate a large number of asynchronous events known as diagnostic trouble codes (DTCs). ・These discrete events form complex temporal sequences that reflect the evolving health of the vehicle's subsystems. ・In the automotive industry, domain experts manually group these codes into higher-level error
cs.LG updates on arXiv.org

Learning to Reason with Curriculum I: Provable Benefits of Autocurriculum

・arXiv:2603.18325v2 Announce Type: replace Abstract: Chain-of-thought reasoning, where language models expend additional computation by producing thinking tokens prior to final responses, has driven significant advances in model capabilities. ・However, training these reasoning models is extremely costly in terms of both data and compute, as it involves collecting long traces of reasoning behavior from humans or synthet
cs.LG updates on arXiv.org

Leveraging Code Automorphisms for Improved Syndrome-Based Neural Decoding

・arXiv:2605.03620v2 Announce Type: replace-cross Abstract: Syndrome-based neural decoding (SBND) has emerged as a promising deep learning approach for soft-decision decoding of high-rate, short-length codes. ・However, this approach still has substantial room for improvement. ・In this paper, we show how to leverage code automorphisms to enhance the ability of existing SBND models to learn and generalize through data augm
cs.LG updates on arXiv.org

Linear Independence of Polynomial Compositions and Identifiability of Deep Neural Networks

・arXiv:2608.27113v1 Announce Type: cross Abstract: Motivated by theoretical problems in deep learning, we conjecture that post-composing a fixed number of pairwise distinct nonconstant polynomials with a generic polynomial of sufficiently large degree yields linearly independent polynomials. ・This generalizes Newman--Slater's theorem on powers of polynomials. ・We establish several cases of this conjecture and its origin
cs.LG updates on arXiv.org

LLM Analysis of 150+ years of German Parliamentary Debates on Migration Reveals Shift from Post-War Solidarity to Anti-Solidarity in the Last Decade

・arXiv:2509.07274v4 Announce Type: replace-cross Abstract: Migration has been a core topic in German political debate, from postwar expellee displacement to labor migration and recent refugee movements. ・Large-scale analysis of such political discourse has traditionally required extensive manual annotation, limiting coverage. ・Large language models (LLMs) offer a scalable alternative.
LLMタグが付けられた新着記事 - Qiita

LLM の出力の縛り方を6段階に並べたら、強い順になっていなかった

・title: "LLM の出力の縛り方を6段階に並べたら、強い順になっていなかった" emoji: "🗺️" type: "tech" topics: ["ai", "llm", "設計", "オントロジー", "jsonschema"] published: true...
cs.LG updates on arXiv.org

LLMs Can Design Near-Optimal OR Algorithms

・arXiv:2608.27296v1 Announce Type: cross Abstract: We ask whether large language models (LLMs) can design effective algorithms for well-specified operations research (OR) problems. ・We study inventory control, queueing network control, and assortment optimization. ・We evaluate two levels of LLM use: at level 1, the model receives one problem instance and returns a solution for that instance; at level 2, it receives only
Zennの「大規模言語モデル」のフィード

LLMって何なんだと思って調べたら、高校数学だった

・AIがもてはやされて久しく、「LLM」「RAG」といった言葉が当たり前のように飛び交うようになりました。ただ、どれもキラキラネームがついた何かで、やってることも、すごさもよくわからなかったので、掘り下げてみました。 ・ちゃんとした説明はほかにいくらでもあるので、自分なりの大まかな理解をまとめておきます。 ・まず、全体像 生成AI周りでよく出てくる言葉を、アーキテクチャ上に並べるとこうなります。
LLMタグが付けられた新着記事 - Qiita

LLMへの指示にコンテキストを含めるかどうかで、テストのカバレッジが倍以上変わった

・はじめに 自作のC# RoslynアナライザーのCSharpIngeniousAnalyzerで、各ルールの動作を手動確認するための使い捨てPlaygroundプロジェクトを使っていました。 ・しかし解析項目が多くなり、手動テストに限界を感じていました。 ・そこで、xUnit...
#LLMタグ

LLMも人間も、次に来る言葉を見積もっている〜フィジカルAI⑧LLM〜

・すらすら出てきた言葉ほど、つい信じてしまう。 ・どうもみなさんこんにちは。 ・ロボットエンジニアの四條です。
Zennの「大規模言語モデル」のフィード

LLMを多段で繋ぐと、要約が根拠を落として後工程が捏造する

・AIエージェントに毎日ニュース記事を書かせています。4つの工程を繋いだパイプラインで、公開先はnoteです。 ・30本たまったところで全部読み直したら、ある記事にこういう一覧が載っていました。 ・トップ10のうち5本が続編・シリーズ・アニメ作品でした。
cs.LG updates on arXiv.org

LoopMoE: Unifying Iterative Computation with Mixture-of-Experts for Language Modeling

・arXiv:2606.04438v2 Announce Type: replace Abstract: Mixture-of-Experts (MoE) and looped architectures scale models along two orthogonal axes, namely parameter capacity and effective depth. ・However, mainstream looped architectures rely on dense backbones that couple parameter count with per-token FLOPs, which makes it impossible to isolate the effect of iterative computation under matched budgets. ・To this end, we pres
cs.LG updates on arXiv.org

LowRankArena: A Standardized Evaluation Platform for SVD-Based LLM Compression

・arXiv:2608.26389v1 Announce Type: cross Abstract: SVD-based low-rank compression has become a fast-growing direction for reducing the memory and computational cost of large language models (LLMs). ・However, meaningful comparison across existing studies remains difficult as prior evaluations use varied benchmarks, inconsistent ratios, and diverse setups, often failing to isolate low-rank effects from auxiliary techniqu
Hugging Face Papers

Magpie: Real-Time World Renderer for Interactive Games

Magpie: Real-Time World Renderer for Interactive Games
cs.LG updates on arXiv.org

Making Latent Evolution Explicit: Operator-Structured Transitions for World Action Models

・arXiv:2608.27259v1 Announce Type: new Abstract: World Action Models (WAMs) augment robot policies by predicting how task-relevant scene states may evolve under interaction. ・Recent WAMs increasingly perform such prediction in latent representation spaces, avoiding full appearance-level generation while preserving control-relevant information. ・Yet latent transitions are commonly realized with Transformer-based predicto
cs.LG updates on arXiv.org

MambaCSP: Hybrid-Attention State Space Models for Hardware-Efficient Channel State Prediction

・arXiv:2604.21957v2 Announce Type: replace-cross Abstract: Recent works have demonstrated that attention-based transformer and large language model (LLM) architectures can achieve strong channel state prediction (CSP) performance by capturing long-range temporal dependencies across channel state information (CSI) sequences. ・However, these models suffer from quadratic scaling in sequence length, leading to substantial
WIRED

Maytag Promo Codes: 15% Off Appliances

・Upgrade your home for less with these verified Maytag discount codes, military savings, and limited-time closeout offers on washers, dryers, and more.
cs.LG updates on arXiv.org

MCCE: A Framework for Multi-LLM Collaborative Search in Discrete Spaces with Similarity-Filtered Preference Learning

・arXiv:2510.06270v2 Announce Type: replace Abstract: Multi-objective discrete optimization problems, such as molecular design, pose significant challenges due to their vast and unstructured combinatorial spaces. ・Traditional evolutionary algorithms often get trapped in local optima, while expert knowledge can provide crucial guidance for accelerating convergence. ・Large language models (LLMs) offer powerful priors and r
AI News & Artificial Intelligence | TechCrunch

Meta executive leaves for OpenAI as the social media giant faces growing scrutiny in India

・Sandhya Devanathan will oversee some OpenAI operations across Southeast Asia and Australia in her new role.
cs.LG updates on arXiv.org

Methodological and Conceptual Framework for 5D Multi-Table Analysis: A Unified Approach for Complex Data Reuse

・arXiv:2608.26149v1 Announce Type: cross Abstract: Multi-table learning remains a major challenge in machine learning for healthcare and other complex information systems. ・Relational data combine several sources of complexity, including large data volume, high-dimensional variables, high-cardinality categorical features, complex inter-table dependencies, and repeated temporal observations. ・We introduce the Relational
WIRED

Mice, a Caved-In Ceiling, and Cloudy Water: The GSA’s New Office Is Falling Apart

・“Do we need to look at the ceiling before going to the bathroom? ・Can we actually trust the water is safe to drink?” asks one GSA worker.
#AIタグ

Michael 専用おでかけバッグ — 精密なロボットを安全に持ち運ぶ

・先日のお出かけから帰宅後、 Michael の首が取れた件はコチラの記事で書きましたが、 それからずっと、安全に持ち運べるケースが欲しいと考えていました。 ・そこで我が家にやってきたのは・・・ 続きをみる
機械学習タグが付けられた新着記事 - Qiita

Microsoft Mage-VLをMLXへ独立移植する――4経路のfloat32一致と、固定動画で最悪3.48秒の応答

・カメラを向けたら、そこで今なにが起きているかを言葉で返してくる。そういう装置は、モデルの賢さよりも「何秒で返ってくるか」で使えるかどうかが決まります。3秒なら見守りに使えますが、30秒なら過去のログにしかなりません。 ・さて、今日はMicrosoftのM...
cs.LG updates on arXiv.org

Mitigating Strong-Modality Collapse in Multimodal Learning via Inverted Asymmetric Fusion

・arXiv:2608.26879v1 Announce Type: new Abstract: Fusing multiple modalities is expected to improve model performance. ・However, on the MultiHuSE dataset, early, late, and symmetric attention fusion often fail to outperform the best unimodal baseline (text). ・Pathway isolation of a symmetric attention fusion model reveals that the text-pathway accuracy drops from 74.9% to 56.4% after fusion in one such setting, indicatin
cs.LG updates on arXiv.org

MM-Spectrum: Multimodal Multi-spectral Molecular Structural Elucidation with a Stable MoE Framework

・arXiv:2608.27286v1 Announce Type: new Abstract: Inferring molecular structures from multimodal spectroscopic measurements requires integrating complementary yet highly heterogeneous signals. ・However, the common paradigm of directly concatenating multispectral sequences can exhibit anomalous performance degradation, primarily due to pronounced heterogeneity and the resulting multimodal imbalance across modalities.
Zennの「機械学習」のフィード

MobileNetV2 を手書き NEON で速くする — NHWC でメモリアクセスを連続化する

・はじめに 前回、MobileNetV2 の畳み込みを手書き NEON のマイクロカーネルに置き換え、backbone 全体を 68 ms まで速くしました。 ・ただしそこで頭打ちになりました。マイクロカーネルが一度に扱う出力チャネルを増やしても数 % しか縮まらず、律速は演算器ではなくメモリ帯域でした。原因は NCHW レイアウトです。pointwise convolution の入力チャネル方向のループが、ic を 1 進めるたびにメモリ上を H×W×4 バイト(112×112 の層で約 50KB)ジャンプし、キャッシュにもプリフェッチャにも乗らないアクセスになっていました。
cs.LG updates on arXiv.org

MODIS: Multi-Omics Data Integration for Small and unpaired datasets

・arXiv:2503.18856v3 Announce Type: replace Abstract: An important objective in computational biology is the efficient integration of multi-omics data. ・The task of integration comes with challenges: multi-omics data are most often unpaired (requiring diagonal integration), partially labeled with information about biological conditions, and in some situations such as rare diseases, only very small datasets are available
cs.LG updates on arXiv.org

Modular Expert Merging for Biomedical Retrieval

・arXiv:2602.04731v2 Announce Type: replace-cross Abstract: Adapting general-purpose LLMs into domain-specialized dense retrievers typically requires large-scale training on mixed-domain data. ・We show that merging independently trained domain-specialized experts consistently exceeds this approach across four decoder-only LLM families (0.6B-7B), four merging methods, and twelve medical and general retrieval tasks from M
cs.LG updates on arXiv.org

Mol-JEPA: A multimodal Joint Embedding Predictive Architecture for Molecules

・arXiv:2608.22642v2 Announce Type: replace Abstract: Despite recent advances in molecular foundation models, several limitations remain, such as chemically invalid augmentations, modality collapse, and incomplete representation of biochemical environments. ・To address these challenges, we present \textbf{Mol-JEPA}, a scalable framework for learning molecular world models. ・Rather than relying on suboptimal molecular per
cs.LG updates on arXiv.org

MOMO: A framework for seamless physical, verbal, and graphical robot skill learning and adaptation

・arXiv:2604.20468v3 Announce Type: replace-cross Abstract: Industrial robot applications require increasingly flexible systems that non-expert users can easily adapt for varying tasks and environments. ・However, different adaptations benefit from different interaction modalities. ・We present an interactive framework that enables robot skill adaptation through three complementary modalities: kinesthetic touch for precise
cs.LG updates on arXiv.org

Multi-Dataset Inverse Problem Solving with Distributed Generative AI

・arXiv:2608.26283v1 Announce Type: cross Abstract: Extracting a shared set of unknown, not directly measurable quantities from multiple, heterogeneous datasets is a common challenge across scientific domains. ・A prominent example is the combination of datasets obtained from different measurements with different settings (e.g. ・varying detector resolutions).
cs.LG updates on arXiv.org

Muon with Finite Newton-Schulz: The Smoothing Benefit in Nonsmooth Nonconvex Optimization

・arXiv:2608.26288v1 Announce Type: new Abstract: Muon has emerged as a strong optimizer for the matrix-valued parameters in large language model pretraining, approximately orthogonalizing its momentum with a few Newton-Schulz iterations. ・Existing theory either replaces this iteration with the exact polar factor it approximates, or treats its finite depth as an approximation error, and thus the iteration Muon actually
cs.LG updates on arXiv.org

NeoTriFuse: Reliability-Aware Multimodal Fusion under Missingness Heterogeneity for Neonatal Mortality Risk Prediction

・arXiv:2608.26436v1 Announce Type: new Abstract: Neonatal mortality risk prediction from bedside monitoring data remains challenging due to extreme class imbalance, heterogeneous clinical risk factors, multi-scale temporal dynamics, and substantial missingness. ・We propose NeoTriFuse, a reliability-aware multimodal fusion framework for missingness-heterogeneous neonatal monitoring data. ・Unlike conventional multimodal a
cs.LG updates on arXiv.org

Neural Non-Equilibrium Hamiltonian Monte Carlo for Corrected Boltzmann Sampling

・arXiv:2607.15682v2 Announce Type: replace Abstract: Sampling from an unnormalized Boltzmann density requires proposals that move probability mass globally while retaining enough path-probability information for statistical correction. ・We introduce Neural Non-Equilibrium Hamiltonian Monte Carlo (NHMC), a train-then-correct learned Hamiltonian sampler. ・Starting from a tractable base distribution, NHMC learns stochastic
cs.LG updates on arXiv.org

Neural Regression with Embeddings for Numerical Attribute Prediction in Knowledge Graphs

・arXiv:2608.26729v1 Announce Type: new Abstract: In recent years, transductive knowledge graph embedding models have been applied to tasks such as link prediction and query answering. ・Although knowledge graphs often contain rich numerical attributes, most embedding models neglect them, limiting their ability to represent real-world knowledge graphs with diverse information. ・In this work, we propose a neural regression
cs.LG updates on arXiv.org

Neural Renormalization Group Flow for Percolation

・arXiv:2608.26764v1 Announce Type: cross Abstract: Machine learning offers a possible route to data-driven real-space renormalization when the relevant observables are nonlocal and difficult to prescribe explicitly. ・We explore this idea for two-dimensional site percolation developping a supervised, scale-shared neural architecture. ・The model recursively applies the same learned coarse-graining rule across scales, prod
cs.LG updates on arXiv.org

NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation

・arXiv:2608.26222v1 Announce Type: new Abstract: Safety evaluation is critical for assessing whether aligned Large Language Models (LLMs) remain robust against jailbreak attacks. ・Existing automated testing methods, however, largely rely on response-level feedback: each candidate prompt typically requires generating a target-model response to evaluate its attack effectiveness. ・This process is expensive and, more import
ITmedia NEWS 最新記事一覧

NISA口座「もぬけの殻」に 不正アクセス被害者らが証券4社に原状回復求め調停

・不正アクセスで証券口座を乗っ取られ、保有する有価証券を勝手に売却されたとする79人が27日、大手証券会社4社に対し、原状回復を求める民事調停を東京簡裁に申し立てた。同日、都内で記者会見した被害者らは「口座から一銭もなくなっていた。もぬけの殻だった」などと訴え、早急な対応を求めた。
stat.ML updates on arXiv.org

On efficiency gains via augmenting a tiny sample with a massive auxiliary sample

・arXiv:2608.26610v1 Announce Type: cross Abstract: In this paper, we study the problem of augmenting a tiny target sample with a massive auxiliary sample. ・Utilizing Tukey's factorization, there are two popular approaches: the inverse probability weight (IPW) and the full-likelihood (FL) methods. ・We show that the IPW approach suffers from the limited target sample problem while the FL method may estimate some model par
cs.LG updates on arXiv.org

On Scope Classification and Current Knowledge-Editing Benchmarks: A Negative Result, with INLAY as a Gradient-Free Case Study

・arXiv:2608.26292v1 Announce Type: cross Abstract: Every memory-based knowledge editor in the SERAC lineage depends on a scope decision: given a query, does a stored edit apply? ・We report that current knowledge-editing benchmarks cannot measure this decision at all. ・Using INLAY, a gradient-free editor we built to obtain exact per-query ground truth (the model is frozen, edits live in an external addressable memory, an
cs.LG updates on arXiv.org

On the Indistinguishability of Human v/s AI Generated Text

・arXiv:2608.26797v1 Announce Type: new Abstract: The rapid improvement of LLMs has made distinguishing AI-generated text from human writing a pressing problem. ・This challenge is further amplified by paraphrasing tools designed to make machine-generated text appear more "human". ・We study how access to human writing samples can be used to strategically paraphrase machine-generated responses toward the human distribution
cs.LG updates on arXiv.org

Out of Sight, Not Out of Mind: Unveiling Latent Attack in Latent-based Multi-Agent Systems

・arXiv:2605.28214v2 Announce Type: replace-cross Abstract: Latent-based multi-agent systems replace parts of explicit inter-agent communication with hidden representations, offering a new direction for efficient and flexible agent collaboration. ・However, moving coordination into latent space may also move attacks beyond the reach of visible-text inspection. ・In this paper, we study whether latent states can carry attac
cs.LG updates on arXiv.org

Over-The-Air Extreme Learning Machines with Nonlinear Stacked Intelligent Metasurfaces

・arXiv:2608.27137v1 Announce Type: cross Abstract: The recently envisioned goal-oriented communications paradigm requires machine learning inference to be performed directly on wirelessly transferred data. ・This paper presents an eXtremely Large (XL) Multiple-Input Multiple-Output (MIMO) system that operates as an Extreme Learning Machine (ELM) to execute Over-The-Air (OTA) binary classification. ・To reduce hardware com
cs.LG updates on arXiv.org

PACIFIER: Pacing Opinion Depolarization via a Unified Graph Learning Framework

・arXiv:2602.23390v4 Announce Type: replace-cross Abstract: Online social networks often form opposing echo chambers that reinforce opinion polarization. ・Under the Friedkin-Johnsen (FJ) model, ModerateInternal (MI) and ModerateExpressed (ME) reduce polarization by neutralizing selected users' internal or expressed opinions, but existing solutions are largely model-specific and learning-based depolarization remains unde
cs.LG updates on arXiv.org

Packora: Systematic Design for Generative Molecular Crystal Structure Prediction

・arXiv:2608.26962v1 Announce Type: new Abstract: Molecular crystal structure prediction (CSP) is important in pharmaceuticals, agrochemicals, and organic electronics, where subtle differences in molecular conformation and packing can strongly affect material properties. ・We present Packora, a flow-based generative model for molecular CSP that jointly predicts atomic coordinates and the lattice from molecular graphs.
Hugging Face Papers

PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

PAWBench: How Far Are We from Probabilistically Aligned World Modeling?
ITmedia NEWS 最新記事一覧

PC価格高騰、限られたIT予算をどう回す? “再生PC”を選んだ企業のシステム担当者に聞く

・PC価格が高騰する中、企業のPC調達手段として注目されるリファービッシュPC。メーカー保証付き再生PCを導入した三菱HCキャピタルITパートナーズに、性能や保証、セキュリティ、必要台数の見極め方と、限られたIT予算をどう配分するかを聞いた。
#LLMタグ

PEP(Prompt Engineering Professional)の模擬試験問題 (非公式)第1章:生成AIと大規模言語モデルの基礎 40問

・このような問題がでるか。内容が正しいかは保証しません。 ・PEP資格試験の第1章「生成AIと大規模言語モデルの基礎」から、試験対策に役立つ40問を作成しています。
cs.LG updates on arXiv.org

Performance Foundations of Parallel & Distributed Reasoning Language Models

・arXiv:2608.27046v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) and other RL-style post-training paradigms have been used for aligning large language models (LLMs) with reasoning standards. ・The resulting recent Reasoning Language Models (RLMs) such as DeepSeek-R1, o3, and Kimi k1.5 show that such RL-style post-training ("RL-for-LLMs") can substantially improve chain-of-thought re
cs.LG updates on arXiv.org

Physics-Informed Stochastic Configuration Machine: A Backpropagation-Free Neural Network with Fast Training for Nonlinear Differential Equations

・arXiv:2608.26549v1 Announce Type: cross Abstract: While Physics-Informed Neural Networks (PINNs) have emerged as a transformative paradigm for solving complex differential equations, their reliance on backpropagation-based gradient descent and automatic differentiation (AD) imposes significant computational bottlenecks and severe non-convex optimization challenges. ・To overcome these fundamental limitations, we propos
Hugging Face Papers

PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents

PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents
cs.LG updates on arXiv.org

Plain Transformers Can be Powerful Graph Learners

・arXiv:2504.12588v4 Announce Type: replace Abstract: Transformers have attained outstanding performance across various modalities, owing to their simple but powerful scaled-dot-product (SDP) attention mechanisms. ・Researchers have attempted to migrate Transformers to graph learning, but most advanced Graph Transformers (GTs) have strayed far from plain Transformers, exhibiting major architectural differences either by
cs.LG updates on arXiv.org

Predicting Quantifiability from Primary Screens to Prioritize Dose-Response Profiling

・arXiv:2608.26538v1 Announce Type: new Abstract: High-throughput drug screening relies on low-cost primary assays to prioritize compounds for more expensive dose-response profiling, where potency is ultimately quantified. ・Current screening strategies largely focus on identifying compounds that will confirm biological activity on follow-up, implicitly assuming that confirmed activity will also yield a usable potency es
Hugging Face Papers

Previous

Previous
cs.LG updates on arXiv.org

Privacy Without Regret: Differentially Private Inference-Time Alignment

・arXiv:2608.26324v1 Announce Type: new Abstract: Best-of-N (BoN) sampling is the simplest and most widely deployed inference-time alignment strategy, but it suffers from two distinct problems: reward hacking, in which the selected response exploits errors in the proxy reward model, and the absence of any privacy protection for the sensitive human preference data used to train that reward model. ・We show that a single i
cs.LG updates on arXiv.org

Private and interpretable clinical prediction with quantum-inspired tensor train models

・arXiv:2602.06110v2 Announce Type: replace Abstract: Publicly available clinical machine learning models pose an underappreciated privacy risk: their parameters or outputs can be exploited to recover information from patients whose data were used during training. ・Moreover, this risk is exacerbated by models such as logistic regression (LR), which are typically preferred in clinical settings for their transparency.
Hugging Face Papers

Procedura: Agentic 3D Modeling with Procedural Control

Procedura: Agentic 3D Modeling with Procedural Control
cs.LG updates on arXiv.org

Profit based evaluation of machine learning for nitrogen recommendations in winter wheat

・arXiv:2608.27205v1 Announce Type: new Abstract: Nitrogen rates for winter wheat are set before the season, under unknown prices and weather. ・The standard UK advice does not respond to prices, yet recent price swings moved the most profitable rate by tens of kilograms per hectare. ・Machine learning is often proposed as the fix.
cs.LG updates on arXiv.org

Prompt Sensitivity of Generative Agents: Evidence from an Epidemic Model

・arXiv:2608.26221v1 Announce Type: cross Abstract: As generative AI gains traction, researchers are investigating its potential to serve as proxies for humans. ・From undergoing cognitive psychology experiments to experiencing an epidemic, generative agents, agents powered by generative AI models, produce realistic human behavior when prompted. ・This study explores the sensitivity of these generative agents' behavior to
cs.LG updates on arXiv.org

Provable one-poison backdoor attacks on linear models and ReLU neural networks

・arXiv:2508.05600v3 Announce Type: replace Abstract: Backdoor poisoning attacks are a threat to machine learning models that are trained on data collected from untrusted sources; these attacks enable attackers to inject malicious behavior into the model that can be triggered by specially crafted inputs. ・Prior work has established bounds on the success of backdoor attacks and their impact on the benign learning task, h
cs.LG updates on arXiv.org

Pruning Binarized Neural Networks: A Dedicated Framework and Globally Weighted Algorithms

・arXiv:2608.26233v1 Announce Type: new Abstract: Extreme compression of deep neural networks, up to full binarization, dramatically reduces memory footprint and arithmetic complexity, facilitating deployment on constrained edge hardware with field-programmable gate arrays (FPGAs) and microcontrollers. ・Although combining binarization with pruning promises additional efficiency gains, existing pruning strategies are ill
cs.LG updates on arXiv.org

Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090

・arXiv:2608.27370v1 Announce Type: cross Abstract: Language model pretraining has become almost synonymous with prohibitive cost, placing it out of reach for much of the academic and open-source communities. ・Although strong open-source efforts already exist, including open-weight models and open-source training recipes, a cost-efficient, hardware-accessible, and open-source pretraining recipe has long been missing.
cs.LG updates on arXiv.org

Pushing the Envelope of LLM Inference with Ultra-Low-Bit Quantized Models

・arXiv:2508.06753v3 Announce Type: replace-cross Abstract: The advent of ultra-low-bit LLM models, approaching the perplexity and task accuracy of their full precision counterparts, is ushering in a new era of LLM inference. ・While these advances promise models that are cost-effective regarding latency, memory, throughput, and energy consumption, the efficiency of runtimes for deploying ultra-low-bit models remains und
cs.LG updates on arXiv.org

Quantitative mapping from conventional MRI using self-supervised physics-guided deep learning: applications to a large-scale, clinically heterogeneous dataset

・arXiv:2601.05063v2 Announce Type: replace-cross Abstract: Magnetic resonance imaging (MRI) is a cornerstone of clinical neuroimaging, yet conventional MRIs provide qualitative information heavily dependent on scanner hardware and acquisition settings. ・While quantitative MRI (qMRI) offers intrinsic tissue parameters, the requirement for specialized acquisition protocols and reconstruction algorithms restricts its avai
cs.LG updates on arXiv.org

QuantumBoostNet: A Hybrid Classical-Quantum Architecture for Enhanced Accuracy in Cardiac Ultrasound View Identification

・arXiv:2608.27302v1 Announce Type: new Abstract: Accurate identification of the correct view or angle in cardiac ultrasound (echocardiogram) is a critical component of cardiologic imaging. ・This step is essential for precise anatomical interpretation, reliable measurement, and the reduction of clinical errors. ・Although computer vision has advanced significantly, most state-of-the-art models perform well on standard ben
#LLMタグ

Qwen3.8-27Bで作ってみるリベンジ

・最初パズルを自由に作ってというと 2048検討開始 最終的に神経衰弱つくろうとしてトークン切れで失敗 8192トークンだと難しいのか 続きをみる
cs.LG updates on arXiv.org

Real-time virtual circuits for plasma shape control via neural network emulators: integration and testing in the MAST-U PCS

・arXiv:2608.26216v1 Announce Type: cross Abstract: The deployment of advanced, AI-enabled control algorithms in tokamak experiments requires robust integration with existing plasma control system (PCS) architectures and extensive pre-experimental validation. ・In this contribution, we describe the integration and testing of neural-network-emulated virtual circuits for plasma shape control within the MAST Upgrade (MAST-U
cs.LG updates on arXiv.org

Recipes for Steering and Scaling LLMs via Sampling

・arXiv:2608.26120v1 Announce Type: cross Abstract: Large Language Models (LLMs) are probabilistic models, typically defined by an autoregressive factorization. ・While recent work has begun to study richer target distributions beyond the base model, the sampling strategies remain highly inefficient. ・In this paper, we present a flexible and theoretically grounded framework for steering and scaling autoregressive LLMs wit
cs.LG updates on arXiv.org

Recovering Expert Critic-Sourced Network Adjacency between Musical Artists from Acoustic Distributions: A Construct-Validity Approach

・arXiv:2608.27291v1 Announce Type: cross Abstract: Music recommendation relies primarily on two signals: user-item interactions, which fail in the cold-start regime, and intrinsic musical content, available for any recording. ・We argue that a third, largely untapped signal is both richer and more principled: critical adjacency, the pairwise relation established when an expert critic explicitly links two artists in long
cs.LG updates on arXiv.org

Recurrent Reinforcement Learning with Memoroids

・arXiv:2402.09900v4 Announce Type: replace Abstract: Memory models such as Recurrent Neural Networks (RNNs) and Transformers address Partially Observable Markov Decision Processes (POMDPs) by mapping trajectories to latent Markov states. ・Neither model scales particularly well to long sequences, especially compared to an emerging class of memory models called Linear Recurrent Models. ・We discover that the recurrent upda
WIRED

Reebok Discount Code: Save 15%+ in September 2026

・Score deep discounts on running shoes, athletic apparel, and classics with the latest verified Reebok promo codes, student discounts, and member-only savings.
cs.LG updates on arXiv.org

Refine-POI: Reinforcement Fine-Tuned Large Language Models for Next Point-of-Interest Recommendation

・arXiv:2506.21599v5 Announce Type: replace-cross Abstract: Advancing large language models (LLMs) for the next point-of-interest (POI) recommendation task faces two fundamental challenges: (i) although existing methods produce semantic IDs that incorporate semantic information, their topology-blind indexing fails to preserve semantic continuity, meaning that proximity in ID values does not mirror the coherence of the
cs.LG updates on arXiv.org

Refusal Is Not Robustness: Auditing Confident Fabrication in Large Language Models on a Provably Uninformative Clinical Pain Speech Transcript

・arXiv:2608.26167v1 Announce Type: cross Abstract: Hallucination and abstention benchmarks rarely establish that a model could not have known the correct answer, making it difficult to distinguish appropriate abstention from an unsupported prediction. ・Seven large language models were evaluated on the TAME Pain speech corpus. ・Participants read phonetically balanced Harvard Sentences while one hand was immersed in cold
cs.LG updates on arXiv.org

Reinforcement Learning-Based Control of CAV Platoon Joining Maneuvers in Mixed Traffic

・arXiv:2608.26860v1 Announce Type: new Abstract: Connected and automated vehicle (CAV) platooning offers a promising approach to improving road safety and traffic capacity. ・However, platoon control in real-world traffic is challenging due to uncertainty and heterogeneous driving behaviors. ・Reinforcement learning (RL) has strong potential for addressing such control problems, but its practical deployment raises challen
cs.LG updates on arXiv.org

ReLATE: Accelerating Tensor Decomposition via Safe and Efficient Learning of Sparse Encodings

・arXiv:2509.00280v2 Announce Type: replace Abstract: Tensor decomposition (TD) is essential for analyzing high-dimensional sparse data, yet its irregular computations and memory-access patterns pose major performance challenges on modern parallel processors. ・Prior works rely on expert-designed sparse tensor formats that fail to adapt to irregular tensor shapes and data distributions. ・We present the reinforcement-learn
cs.LG updates on arXiv.org

Representation Measurements Under Function-Preserving Reparameterizations

・arXiv:2608.27020v1 Announce Type: cross Abstract: Hidden coordinates are not uniquely determined by a language model's input--output function, so representation-derived measurements should be invariant to function-preserving changes of basis. ・This study shows that column-permutation parallel analysis violates function-preserving reparameterization invariance because its reference distribution and selected component c
cs.LG updates on arXiv.org

Residual Reward Models: Leveraging Prior Knowledge for Efficient Preference-based Reinforcement Learning in Robotics

・arXiv:2507.00611v2 Announce Type: replace Abstract: Preference-based Reinforcement Learning (PbRL) provides a promising alternative to heuristic reward design in complex robotic environments. ・However, PbRL often suffers from poor sample efficiency, requiring extensive and costly human feedback, which limits its real-world applicability. ・Prior work has proposed learning a reward model from demonstrations and fine-tuni
cs.LG updates on arXiv.org

Rethinking Message Passing as Retrieval for Text-Attributed Graph Learning

・arXiv:2608.26732v1 Announce Type: new Abstract: Graph neural networks (GNNs) are typically conceptualized as message-passing neural networks, yet it remains unclear why neighborhood aggregation reliably outperforms node-wise multilayer perceptrons (MLPs). ・Despite its empirical success, this paradigm can be computationally expensive and sensitive to imperfect graph structures. ・In this work, we present a retrieval-augm
#AIタグ

roborock (ロボロック) Qrevo L Pro ロボット掃除機 水拭き レビュー比較まとめ

・roborock (ロボロック) Qrevo L Pro ロボット掃除機 水拭きは、吸引と水拭きを高いレベルで両立させたい方に最適な一台です。 ・全自動ドックによるメンテナンス機能が最大の強みとなり、ゴミ捨てやモップ洗浄の手間を大幅に削減し、常に清潔な状態を維持することが可能です。
#AIタグ

roborock ロボロック Saros10R ロボット掃除機 黒 20000P レビュー比較まとめ

・roborock ロボロック Saros10R ロボット掃除機 黒 20000Pは、最高水準の清掃能力と圧倒的な自動化機能が融合したプレミアムモデルです。 ・本機は、強力な吸引力と水拭き機能を両立させつつ、8way全自動ドックによるメンテナンスの手間を極限まで削減した点が最大の強みとなります。
cs.LG updates on arXiv.org

Robust Neural Stimulation Response Modeling Through Meta-Learning and Pretraining

・arXiv:2608.26649v1 Announce Type: new Abstract: Objective: Model-based closed-loop neural stimulation holds promise for therapeutic applications ranging from Parkinson's disease to sensory restoration, but deployment has been limited by two obstacles: 1) forecasting models for predicting the consequences of stimulation fail catastrophically on a meaningful fraction of sessions, and 2) per-session calibration requirem
cs.LG updates on arXiv.org

Safety by Design: Realized-Cost Constraints for Contextual Bandits with Continuous Actions

・arXiv:2608.26755v1 Announce Type: new Abstract: Contextual bandits are a standard framework for sequential decision-making under uncertainty, with applications in clinical trials, dosage selection, recommendation systems, and autonomous systems. ・Safety is central in many of these applications, since a single unsafe decision in settings such as dosage selection or autonomous driving can have catastrophic consequences.
cs.LG updates on arXiv.org

SAGE: Variate-Wise Semantic Augmentation for Vision-Language Time Series Forecasting

・arXiv:2608.26829v1 Announce Type: new Abstract: Time series forecasting models operate on raw numerical sequences, lacking the semantic knowledge that domain experts implicitly leverage, such as the physical meaning of each variable, its statistical behavior, and its temporal dynamics. ・Recent efforts to bridge this gap fall into two camps. ・Some rely on large language models at inference time, which is computationally
The Verge

Save hundreds on a TCL mini-LED TV with quantum dots and high refresh rate

・Amazon and Best Buy have the TCL QM7L mini-LED TV on sale for as low as $797.99 for the 55-inch model, a $200 discount from the usual price. ・We spotted scaling discounts on larger sizes as well, although Amazon already listed low stock on the 65-inch version. ・While OLED televisions are great for deep black levels and great contrast, their high price point may be off-putting to some buyers.
cs.LG updates on arXiv.org

Scaling Graph Neural Networks for Friend Recommendation: Multi-Hash User Embeddings and Temporal Neighbor Sampling

・arXiv:2608.27413v1 Announce Type: cross Abstract: Friend recommendation is inherently graph-structured: the relevance of a potential connection depends on multi-hop social context rather than user attributes alone. ・However, deploying message-passing GNNs on a production-scale social graph with hundreds of millions of users and tens of billions of edges requires addressing numerous modeling and systems challenges.
cs.LG updates on arXiv.org

Scaling Model-Generated Distillation Data Can Make Latent Teacher Traits More Recoverable

・arXiv:2608.26958v1 Announce Type: new Abstract: Scaling model-generated data is usually viewed as improving distillation: more examples should increase coverage, reduce noise, and produce stronger students. ・We show a second effect: larger datasets can make subtle teacher-specific signals easier to detect in the trained student, even when examples are off-task and never mention the trait. ・In a controlled setup inspire
cs.LG updates on arXiv.org

SecureDrive-FL: Joint Differential Privacy and Gradient-Aware Selective Homomorphic Encryption for Federated Driver Monitoring

・arXiv:2608.27108v1 Announce Type: cross Abstract: Federated Learning (FL) enables privacy-aware distributed training, yet gradient updates remain exploitable: Man-in-the-Middle (MitM) interception exposes updates in transit, while model poisoning corrupts global convergence. ・We first introduce GASHE (Gradient-Aware Selective Homomorphic Encryption), a novel selective encryption strategy that dynamically identifies an
cs.LG updates on arXiv.org

Selection Bias Correction in Retail Intelligence

・arXiv:2608.26156v1 Announce Type: cross Abstract: Retail intelligence often relies on monitoring popular, high-velocity products, potentially biasing economic indicators by ignoring the "long tail" of niche items. ・This simulation study investigates selection bias in inflation estimation and compares correction methods across diverse data-generating processes. ・Through 400 Monte Carlo replications spanning four scenari
cs.LG updates on arXiv.org

Self-Augmented Diffusion Guidance for Physics-Informed Generation

・arXiv:2608.26748v1 Announce Type: new Abstract: Diffusion models can be used to generate spatiotemporal signals of physical phenomena, such as time-series images of fluid dynamics. ・However, a major limitation of standard diffusion models is that they do not incorporate constraints derived from the underlying physical laws. ・Consequently, generated samples may appear visually plausible while deviating substantially fro
Hugging Face Papers

Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher

Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher
cs.LG updates on arXiv.org

Sequential Additivity in Distributionally Robust Ranking and Selection

・arXiv:2509.06147v2 Announce Type: replace-cross Abstract: Ranking and selection (R&S) seeks to identify the alternative with the best mean performance from a finite collection of simulated alternatives. ・Its practical value depends on accurate simulation input modeling, which is often hindered by input uncertainty arising from limited data. ・Distributionally robust R&S (DRR&S) addresses this challenge by considering se
cs.LG updates on arXiv.org

Shared Actors Need Not Share Critics: Effects of Value Mismatch in Parallel Reinforcement Learning

・arXiv:2608.26481v1 Announce Type: new Abstract: When a single policy is trained in parallel across multiple environments of the same task, such as procedurally generated levels, randomized dynamics, or curricula, implementations commonly use one critic across all sampled environments. ・Yet different environments can assign different expected returns to the same input visible to the critic. ・A critic without environment
cs.LG updates on arXiv.org

Sharp Minimax Regret for Infinite-Memory Logistic Prediction

・arXiv:2608.26515v1 Announce Type: cross Abstract: We study online prediction for a specific finite-alphabet, exogenously driven source with infinite input memory. ・Independent Rademacher inputs $(U_t)$ are observed sequentially, and the next binary mark has logit $\sum_{j=1}^{t}\theta_jU_{t+1-j}$, where $\abs{\theta_j}\leq r_j$ and $\sum_jr_j\leq B$. ・Regret is expected cumulative excess log loss.
cs.LG updates on arXiv.org

SIGMA: Structured Noise-Effect-Aware Grouped Multi-Agent Aggregation

・arXiv:2608.26683v1 Announce Type: cross Abstract: Cooperative multi-agent reinforcement learning (MARL) faces significant challenges in maintaining robust coordination under noisy observations. ・Although observation disturbances are often introduced independently across agents, their downstream effects on cooperative decision-making can become structured through underlying cooperation structures. ・We characterize this
cs.LG updates on arXiv.org

SimCast-S2S: An Efficient Generative Model for Subseasonal Precipitation Forecasting via Transfer Learning from Climate Simulations

・arXiv:2608.26594v1 Announce Type: new Abstract: Subseasonal-to-seasonal (S2S) precipitation forecasting has substantial financial and societal impact, yet remains challenging because of weak predictive signals, high associated uncertainty, and the computational cost of operational systems, which constrains simulation fidelity. ・We introduce SimCast-S2S, a generative latent-diffusion framework for probabilistic S2S pre
cs.LG updates on arXiv.org

Simple Actors and Deep Critics for Scalable Reinforcement Learning

・arXiv:2608.26659v1 Announce Type: new Abstract: Recent progress in offline reinforcement learning (RL) has been driven by expressive generative actors such as diffusion and flow-matching policies, which capture multimodal behavior in offline datasets. ・However, these actors require multiple denoising or integration steps per action and thus incur substantial overhead at every decision in deployment. ・In this work, we r
cs.LG updates on arXiv.org

SLM-Conditioned Hierarchical Relation Routing for Labeled Property Graph Learning

・arXiv:2608.26132v1 Announce Type: new Abstract: Labeled property graphs combine relational structure with heterogeneous textual and categorical properties attached to both nodes and relationships. ・Conventional graph neural networks typically represent these properties as static feature vectors, limiting their ability to determine which semantic evidence should influence message propagation for a particular prediction
cs.LG updates on arXiv.org

Soft Active Electromyography Interface for Machine Learning-Enabled Silent Speech Recognition

・arXiv:2608.27048v1 Announce Type: new Abstract: Silent speech recognition (SSR) provides an alternative communication pathway in the absence of audible speech. ・However, conventional approaches are limited by the need for constant facial attachment, privacy concerns, and unstable signal acquisition. ・Here, we propose a soft, active electromyography (EMG) interface that enables word-level SSR using machine learning.
cs.LG updates on arXiv.org

Squeezing More from Limited Data with Recursive Transformers

・arXiv:2608.26973v1 Announce Type: cross Abstract: Pre-training under limited data requires a different view of scaling than web-scale language modeling. ・With a fixed data budget but relatively abundant compute, increasing parameter count helps only up to an optimal scale; beyond that point, models overfit and generalization worsens. ・We study this behavior across 10M-100M word pre-training budgets, two corpora, and mu
cs.LG updates on arXiv.org

Stable but Wrong: When Learning Stabilizes Away from the Truth

・arXiv:2603.21491v2 Announce Type: replace Abstract: Stable training is often treated as evidence that learning is succeeding, but stability characterizes optimization behavior rather than correctness relative to an external objective. ・We study what happens when the signal being optimized remains persistently biased. ・We define Stable but Wrong (SBW) as a learning state in which the learning process remains stable unde
cs.LG updates on arXiv.org

Stack Trace-Based Crash Deduplication with Transformer Adaptation

・arXiv:2508.19449v2 Announce Type: replace-cross Abstract: Automated crash reporting systems generate large volumes of duplicate reports, overwhelming issue-tracking systems and increasing developer workload. ・Traditional stack trace-based deduplication methods---relying on string similarity, rule-based heuristics, or deep learning (DL) models---often fail to capture the contextual and structural relationships within s
cs.LG updates on arXiv.org

STITCH-OPE: Trajectory Stitching with Guided Diffusion for Off-Policy Evaluation

・arXiv:2505.20781v2 Announce Type: replace-cross Abstract: Off-policy evaluation (OPE) estimates the performance of a target policy using offline data collected from a behavior policy, and is crucial in domains such as robotics or healthcare where direct interaction with the environment is costly or unsafe. ・Existing OPE methods are ineffective for high-dimensional, long-horizon problems, due to exponential blow-ups in
cs.LG updates on arXiv.org

Subgraph Filtering for Fair Graph Neural Networks

・arXiv:2608.26437v1 Announce Type: new Abstract: Graph neural networks (GNNs) can exhibit unfair behavior even when sensitive attributes are excluded from node features, because graph topology and message passing propagate group-correlated signals under sensitive homophily. ・Existing fairness-aware GNN methods mainly constrain representations or prediction distributions at a global level, without explicitly controlling
cs.LG updates on arXiv.org

SynthCharge: An Electric Vehicle Routing Instance Generator with Feasibility Screening to Enable Learning-Based Optimization and Benchmarking

・arXiv:2603.03230v2 Announce Type: replace Abstract: The electric vehicle routing problem with time windows (EVRPTW) extends the classical VRPTW by introducing battery capacity constraints and charging station decisions. ・Existing benchmark datasets are often static and lack verifiable feasibility, which restricts reproducible evaluation of learning-based routing models. ・We introduce SynthCharge, a parametric generator
cs.LG updates on arXiv.org

Systematic Literature Review of Machine Learning Models and Applications for Text Recognition

・arXiv:2608.26500v1 Announce Type: cross Abstract: Optical Character Recognition (OCR) for text recognition using machine vision has significantly improved, particularly when handling heterogeneous textual data. ・Traditional OCR models struggle with script variations, writing styles, and degraded documents. ・Advancements in technology are leading to new AI models with improved architecture for handling multiple language
cs.LG updates on arXiv.org

Tabular Deep Learning for Algorithmic Trading: Cross-Regime Bayesian Optimisation for Equity Signal Generation

・arXiv:2608.27076v1 Announce Type: new Abstract: Algorithmic trading now represents a market exceeding $20 billion, where even marginal gains in signal robustness can translate into economically significant returns. ・Existing evaluations of equity prediction models do not explicitly target regime robustness during hyperparameter selection. ・Five model classes are trained on daily observations from approximately 300 larg
Hugging Face Papers

TacForcing: Streaming Action Generation with Execution-Time Tactile Feedback

TacForcing: Streaming Action Generation with Execution-Time Tactile Feedback
cs.LG updates on arXiv.org

Technical Comparative Benchmarking Study: Advanced AI Hybrid Methods for Renewable Energy Farm Optimization and Forecasting

・arXiv:2608.26613v1 Announce Type: new Abstract: This study provides a comprehensive benchmarking of conventional machine learning (ML), ensemble learning, deep neural networks, recurrent architectures, Transformers, graph based models, and hybrid ensemble deep learning approaches under complementary renewable energy scenarios. ・Three datasets are considered: a large scale WEC dataset, a 16 WEC dataset, and operational
cs.LG updates on arXiv.org

TEMPLAR Wales: A georeferenced environmental and toponymic dataset of Welsh settlements

・arXiv:2608.26970v1 Announce Type: new Abstract: Place names provide persistent records of how landscapes have been described and organised, but their quantitative reuse requires explicit separation between mapped places, lexical annotations and environmental measurements. ・TEMPLAR Wales is a georeferenced environmental-toponymy dataset comprising 3,757 settlement records across Wales. ・The resource links a reproducible
cs.LG updates on arXiv.org

Tensorion: A Tensor-Aware Generalization of the Muon Optimizer

・arXiv:2606.25975v2 Announce Type: replace Abstract: Common first-order optimizers, such as Adam, implicitly treat each parameter block as an unstructured vector, which disregards the multilinear weight structure present in many modern machine learning models. ・Recent work has shown that exploiting matrix structure can improve optimization dynamics. ・A notable example is Muon, which performs steepest descent under the s
cs.LG updates on arXiv.org

Terrain signatures in Welsh settlement names

・arXiv:2608.26978v1 Announce Type: new Abstract: Landscapes are named, but whether names retain measurable environmental information beyond broad geographic structure is rarely tested. ・We analysed 3,757 Welsh settlements using a frozen, source-audited 24-element lexical framework, preregistered outcome-specific models and geographically structured validation. ・The central comparison contrasted 101 settlements carrying
WIRED

The 31 Best Deals From the REI Labor Day Sale

・Labor Day already? ・Say it ain’t so. ・You can still chase summer a while longer with these great deals on outdoor gear at the REI Labor Day Sale.
cs.LG updates on arXiv.org

The Attribution Contract for Generative Language Models

・arXiv:2605.23080v3 Announce Type: replace Abstract: Feature attribution scores each part of an input by how much it explains a model's output. ・We argue that in generative language models these scores carry no fixed meaning. ・A classifier has a single output to explain, but a generative model produces its output token by token, and each generated token is both an output and an input, so explaining the output becomes se
WIRED

The Easiest Ways to Share Anything Between Android and iOS

・Transferring files and data across platforms is more straightforward than ever.
The Verge

The iPhone Fold could make concerts even worse

・You know the person blocking your view of the concert because their phone is swaying in the air, recording the entire thing? ・Get ready for the unfolded version of it. ・This week on The Vergecast, we dive right into the big week of Apple news.
cs.LG updates on arXiv.org

The Latent Diagnostic Taxonomy: A Framework for Constructing Classifiers and Diagnosing Their Decisions, Applied to Prompt Injection Detection

・arXiv:2608.26423v1 Announce Type: new Abstract: This paper proposes a framework for constructing a classifier as a safeguard layer, and for developing a complementary diagnostic that identifies which of the classifier's confident decisions can be trusted. ・This framework, the Latent Diagnostic Taxonomy, consists of (i) constructing a dimensionality-optimized classifier, in which the embedding dimensionality is empiric
cs.LG updates on arXiv.org

The Principles of Diffusion Models

・arXiv:2510.21890v3 Announce Type: replace Abstract: This book presents the core principles that have guided the development of diffusion models, tracing their origins and showing how diverse formulations arise from shared mathematical ideas. ・Diffusion modeling starts by defining a forward process that gradually corrupts data into noise, linking the data distribution to a simple prior through a continuum of intermedia
cs.LG updates on arXiv.org

The Rashomon Effect for Visualizing High-Dimensional Data

・arXiv:2604.00485v3 Announce Type: replace Abstract: Dimension reduction (DR) is inherently non-unique: multiple embeddings can preserve the structure of high-dimensional data equally well while differing in layout or geometry. ・In this paper, we formally define the Rashomon set for DR -- the collection of `good' embedding -- and show how embracing this multiplicity leads to more powerful and trustworthy representation
WIRED

There Are So Many Conspiracy Theories About Dolly Parton and Vaccines

・Influential right-wing conspiracy theorists and MAGA-aligned pundits are falsely claiming that the Covid vaccine killed Dolly Parton, the vaccination advocate and philanthropist.
Hugging Face Papers

Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning

Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning
cs.LG updates on arXiv.org

Token-Level Advertising

・arXiv:2608.27382v1 Announce Type: cross Abstract: Generative AI is transforming how people access information, challenging traditional advertising mechanisms built around predefined slots. ・Towards generation-native advertising, we propose the Latent Advertiser Mixture Auction (LAMA), a token-level advertising mechanism that embeds advertiser influence directly into the generation process. ・Advertisers report local con
#AIタグ

TOSHIBA(東芝) コードレス スティック掃除機 強力パワー持続 軽量 サイ レビュー比較まとめ

・TOSHIBA(東芝) コードレス スティック掃除機 強力パワー持続 軽量 サイは、強力な吸引力と優れたお手入れのしやすさを両立したモデルであり、日々の掃除を劇的に快適にする一台です。 ・パワーキープシステム搭載のフィルターレスサイクロンが、ゴミをしっかりキャッチしつつ吸引力を長時間維持する点が高評価ポイントです。
cs.LG updates on arXiv.org

Toward all-optical unsupervised Hebbian learning in deep photonic neuromorphic networks

・arXiv:2601.22300v4 Announce Type: replace-cross Abstract: We propose a deep photonic neuromorphic network (PNN) architecture based on phase-change material (PCM) synapses and local optical feedback for online, unsupervised Hebbian learning. ・The proposed architecture combines optical vector-matrix multiplication, non-volatile PCM synaptic weighting, and local coincidence-driven synaptic adaptation within a multilayer
cs.LG updates on arXiv.org

Toward Equitable Low-Carbon Mobility: Fairness-Aware Demand Prediction for Expanding Bike-Sharing Systems

・arXiv:2608.26451v1 Announce Type: new Abstract: Bike-sharing systems are an important component of low-carbon urban mobility, but continued expansion creates challenges in both cold-start prediction and equitable resource allocation. ・Newly deployed stations lack historical ridership records, causing a mismatch between training and inference for graph-based models on evolving networks. ・Historical demand may also encod
cs.LG updates on arXiv.org

Towards a universal meta-optics solver via large language models

・arXiv:2608.26417v1 Announce Type: cross Abstract: Metasurface design increasingly requires fast models that can operate across structurally distinct device families, rather than retraining a separate surrogate for every geometry class. ・Conventional neural network surrogates often depend on fixed-dimensional descriptors, family-specific output formats, and repeated architecture tuning, which limits their scalability a
cs.LG updates on arXiv.org

TRACE-CRC: Trajectory-Adaptive Conformal Risk Control for Multi-Step Channel State Information Prediction

・arXiv:2608.27124v1 Announce Type: new Abstract: Reliable prediction of time-varying channel state information (CSI) is essential for efficient wireless communication. ・Each CSI frame is a matrix-valued representation of the wireless channel response, and a sequence of CSI frames forms a temporal channel trajectory. ・Modern deep learning-based CSI predictors, however, often provide only point predictions and lack calibr
cs.LG updates on arXiv.org

TRACE: Retrospective Streaming Generation of Physical Fields under Sparse Structured Sensing

・arXiv:2608.26219v1 Announce Type: cross Abstract: Reconstructing continuous physical fields from sparse measurements is central to scientific monitoring, inverse modeling, and digital-twin construction. ・Generative reconstruction has recently emerged as a promising paradigm for this task by learning data-driven physical priors that complete plausible full fields from limited observations. ・However, existing methods lar
cs.LG updates on arXiv.org

TraceBench: Controlled Evaluation of LLM Agents for Time-Series Root-Cause Attribution

・arXiv:2608.27182v1 Announce Type: new Abstract: LLM agents are increasingly applied to anomaly detection and root-cause analysis in time-series observations collected from real-world systems; however, their performance on these tasks has not been systematically evaluated under controlled conditions. ・We introduce TraceBench, a simulation-based framework for generating controlled root-cause attribution tasks.
Hugging Face Papers

Training Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report

Training Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report
cs.LG updates on arXiv.org

TriShieldRAG: 3 Rings, One Blind Spot in Layered Defenses for Retrieval-Augmented Generation

・arXiv:2607.23838v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) grounds LLM answers in query-time retrieved documents, so reliability depends on what the retriever returns. ・PoisonedRAG (Zou et al., USENIX Security'25) showed five crafted documents mislead an undefended system in nearly 90% of cases, and that single-stage defenses give limited robustness. ・We propose TriShieldRAG, a three
The Verge

Trump’s EPA wants to let data centers hide their air pollution

・Just as new data centers face growing backlash from neighboring communities, the US Environmental Protection Agency (EPA) is about to make it harder for people to weigh in on any pollution those centers create. ・The EPA plans to toss out a federal rule requiring public notice and an opportunity to comment when certain industrial sites apply for an air permit. ・Advocates warn that the move would allow data center develo
Hugging Face Papers

TTPO: Test-Time Policy Optimization

TTPO: Test-Time Policy Optimization
cs.LG updates on arXiv.org

UCB for Large-Scale Pure Exploration: Beyond Sub-Gaussianity

・arXiv:2511.22273v2 Announce Type: replace-cross Abstract: Selecting the best alternative from a finite set is the central objective of ranking and selection (R&S) and best arm identification (BAI). ・Traditional R&S or BAI approaches have predominantly relied on Gaussian or sub-Gaussian assumptions on the performance distributions of all alternatives, which limit their applicability to non-sub-Gaussian---especially hea
cs.LG updates on arXiv.org

Ultra Low-Power, Lightweight, Probabilistic RSS-Based Path Reconstruction: A System for Landscape-Scale Bee Tracking

・arXiv:2608.27152v1 Announce Type: new Abstract: Applications in fields such as movement ecology, Internet of Things or robotics share the need for systems that localize devices that are too small and power constrained to implement GNSS (Global Navigation Satellite Systems). ・Alternative low-power localization methods often rely on only measurements of RSS (Received Signal Strength) to infer the AoA (Angle of Arrival)
Hugging Face Papers

Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO

Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO
cs.LG updates on arXiv.org

Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO

・arXiv:2608.27351v1 Announce Type: new Abstract: Evolution Strategies (ES) have recently emerged as a memory-efficient post-training paradigm for LLM reasoning. ・However, the optimization behavior of ES remains understudied, making it hard to define its advantage scope compared to mainstream post-training paradigms (e.g., Group Relative Policy Optimization (GRPO)). ・By systematically investigating ES dynamics and mechan
cs.LG updates on arXiv.org

Unifying Detection and Adaptation in Task-Free Continual Learning

・arXiv:2608.27070v1 Announce Type: new Abstract: To mitigate catastrophic forgetting in downstream continual learning (CL) for large language models (LLMs), existing methods typically constrain parameter updates or introduce task-specific adaptation modules. ・However, these methods often rely on explicit task boundaries during training, limiting their applicability to realistic task-free scenarios. ・In this paper, we pr
cs.LG updates on arXiv.org

Universality and sharp thresholds for ellipsoid fitting

・arXiv:2608.27372v1 Announce Type: cross Abstract: We establish a sharp phase transition for fitting random vectors by an ellipsoid. ・The random vectors have independent subgaussian coordinates with mean zero, variance one, and a common fourth moment, and the number of vectors is proportional to the square of the dimension. ・We identify an explicit satisfiability threshold such that, with high probability, a positive de
Hugging Face Papers

UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City

UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City
MarkTechPost

Vercel AI Open-Sources vgpu: A TypeScript WebGPU Library for AI Agent Shaders

・Vercel has open-sourced vgpu, the WebGPU library it built to ship the shaders on vercel.com. ・It treats .wgsl files as importable TypeScript modules, runs the same shader in the browser, in headless Node.js via Dawn, and in a deterministic CI mock, and ships a fullscreen effect in 25 KB gzipped. ・The post Vercel AI Open-Sources vgpu: A TypeScript WebGPU Library for AI Agent Shaders appeared first on MarkTechPost.
cs.LG updates on arXiv.org

Vowel Signs Are Not Letters: A Pre-tokenization Ceiling on Multilingual Tokenizer Fertility

・arXiv:2608.26449v1 Announce Type: cross Abstract: Byte-level BPE tokenizers that use the HuggingFace ByteLevel pre-tokenizer inherit GPT-2's word regex, where a word is defined as \p{L}+, one or more Unicode letters. ・In abugida scripts, vowels are written as combining marks; this pattern therefore splits each word at every vowel sign. ・Since BPE merges only within a pre-token, those splits persist through training reg
Hugging Face Papers

What Does an Evaluation License? A Commit-Bound Census of Claim-Relative Inference in Inspect Evals

What Does an Evaluation License? A Commit-Bound Census of Claim-Relative Inference in Inspect Evals
Hugging Face Papers

What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents

What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents
cs.LG updates on arXiv.org

When Do Autoregressive Sequence Models Forecast Physical Wavefields? A Controlled Study on Synthetic Seismograms

・arXiv:2606.10868v2 Announce Type: replace Abstract: Long-horizon autoregressive forecasting of oscillatory physical signals, such as seismograms, gravitational-wave strain, and similar wavefields is limited by error accumulation: as a causal model is fed its own outputs over hundreds of steps, small per-step errors compound into phase drift that pointwise metrics fail to detect. ・We ask when such rollout stays stable,
cs.LG updates on arXiv.org

When Interference Graphs Evolve: Doubly Robust Estimation of Dynamic Peer Effects

・arXiv:2608.27187v1 Announce Type: new Abstract: Peer effects are difficult to estimate when interaction graphs evolve because pre-assignment network history, dynamic peer exposure, and post-assignment network change have distinct causal roles. ・We introduce a controlled contrast framework that indexes potential outcomes by own treatment, temporally aggregated peer exposure, and a post-assignment evolution summary.
cs.LG updates on arXiv.org

When Is Noise Response Universal? Tokenization as the Hidden Variable in Language Models

・arXiv:2608.26319v1 Announce Type: cross Abstract: The performance of textual neural models often degrades when their inputs are corrupted by noise such as typos, OCR errors, or dropped words. ・We study the degradation rate across neural models, both sentence embeddings and decoder-only LLMs, and find that how consistent it is depends on the scale of the noise: under word-level noise, models with very different archite
cs.LG updates on arXiv.org

When Is the Sharp Covariance Envelope Tight? Feature-Only Geometry for Volume-Sampled Least Squares

・arXiv:2608.26877v1 Announce Type: new Abstract: Prior analyses by Derezinski and Warmuth established all-size sampling identities, selected-OLS unbiasedness, and inverse moments for ordinary volume sampling, while their exact arbitrary-fixed-response loss and prediction-covariance formulas are at the rank-size endpoint s=d. ・We establish a Loewner envelope for centered coefficient covariance for every full-rank fixed
cs.LG updates on arXiv.org

When Privacy Hurts Mergeability: Geometry-Aware Model Merging under Differential Privacy

・arXiv:2608.26655v1 Announce Type: new Abstract: Model merging promises to construct a single multi-task model from independently fine-tuned task models without accessing the original task data. ・This makes it attractive when task data cannot be centralized, but released task models may still leak private fine-tuning data. ・Differential privacy (DP) provides a principled mechanism for limiting such leakage, yet its effe
cs.LG updates on arXiv.org

When Text Misleads: Inconsistent-Aware Reasoning for Audio-Grounded Dialogue

・arXiv:2608.27176v1 Announce Type: cross Abstract: Understanding spoken dialogue requires joint reasoning over lexical content and paralinguistic acoustic signals such as emotion and conversational intent. ・However, existing evaluations often allow shortcuts based on transcripts or single-modality solutions, obscuring whether models genuinely ground predictions in speech. ・We formalize this failure mode as cross-modal d
cs.LG updates on arXiv.org

When the Canonical Completion Is Wrong: Formalizing and Measuring the Jump in Large Language Models

・arXiv:2608.26187v1 Announce Type: cross Abstract: Whether large language models (LLMs) can perform the abductive leap from evidence to a new system of axioms, commonly referred to as a jump, has recently attracted considerable debate. ・A prominent position holds that LLMs are structurally incapable of such jumps, while recent studies challenge both its mechanism and its evidence. ・However, the debate remains difficult
cs.LG updates on arXiv.org

When Top-K Misses the Decision: Tool-Call Drift in Multi-Teacher On-Policy Distillation

・arXiv:2607.07050v5 Announce Type: replace-cross Abstract: Top-$K$ teacher logits make on-policy distillation tractable, but retained teacher mass does not certify student-relative gradient fidelity. ・We study routed, two-teacher tool-use distillation. ・In a frozen Qwen3.5-9B audit, the tool teacher ranks the entry token first on 500 tool prompts, and top-32 reinforces it in every matched pair.
stat.ML updates on arXiv.org

Which Metrics Save the Most Human Annotation? Prediction-Powered Evaluation and Meta-Evaluation

・arXiv:2608.26638v1 Announce Type: cross Abstract: Across various non-verifiable tasks, human evaluation is reliable but expensive, while automatic metrics are more scalable but often biased. ・Building on prediction-powered inference (PPI), we propose prediction-powered evaluation, a framework that combines limited human judgments with large-scale automatic scores to obtain data-efficient system comparisons that are pr
cs.LG updates on arXiv.org

Why not to use the Gaussian kernel

・arXiv:2608.26974v1 Announce Type: cross Abstract: Kernels measure similarity or correlation in tasks such as regression and classification. ・The Gaussian kernel, other names of which include squared exponential and radial basis function kernel, is one of the most popular in Gaussian process regression. ・We argue that the Gaussian kernel is best avoided and should never be used as a default.
Hugging Face Papers

WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution

WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution
WIRED

XGIMI Vibe One Battery-Powered Projector Review (2026)

・With the curtains closed and the power cable nearby, this affordable Full HD smart projector is a compact delight.
cs.LG updates on arXiv.org

You Don't Need to Run Every Eval

・arXiv:2606.24020v2 Announce Type: replace Abstract: A modern model release reports scores on 40+ benchmarks and the same evaluations were run many more times before it: to track training progress, compare design choices, and select the checkpoint for the release. ・But do we need to run every eval? ・We compile a public score matrix of 84 frontier models on 133 benchmarks (2,604 cells, 23.3% filled) and find it is approx
Hugging Face Papers

Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization

Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization
ITmedia NEWS 最新記事一覧

アイドルの宮本佳林さんが登壇したCloudflare開発者イベント、アーカイブ配信中

・初心者が失敗しにくくするための「守り」として、個人情報はなるべく預からないこと、複数のAIにレビューしてもらうことなども挙げた。
#AIタグ

あんてな雑記:【日々あれこれ】AIが生み出す退屈さの正体

あんてな雑記:【日々あれこれ】AIが生み出す退屈さの正体
ITmedia NEWS 最新記事一覧

イエローハット、最大180万人分の情報漏えいか 作業予約システムに不正アクセス

・イエローハットは8月28日、店舗での作業予約に使う「イエローハットWEB作業予約システム」が不正アクセスを受け、最大180万1499人分の個人情報が漏えいした可能性があると発表した。漏えいした可能性があるのは氏名や電話番号などで、クレジットカード情報は含まないとしている。
#LLMタグ

お手軽LLMはじめてみた。その31 (Qwen3.8-27B、Qwen3-Coder-30B-A3B、Muse Glimmer-30B)

・RAMが32GBになったので。ちょっと大きめのLLMに挑戦。 ・Qwen3.8-27B 続きをみる
#LLMタグ

サイバー攻撃の最後に残った手作業まで、AIの自動化で消えた

・8月27日、OpenAIが発行した公開書簡に100社を超える企業が署名した。Anthropic、Google、Microsoft、パロアルトネットワークス、クラウドストライク。シティやビザといった金融も並んでいる。今後数ヶ月でAIを使ったサイバー攻撃ははるかに広範かつ高度になる、防御を固める窓は数ヶ月しかない、というのが主旨だ。 ・同じ日、サイバーセキュリティ株は一律に上がった。クラウドストライクが20.5%、フォーティネットが9.7%。同じ日のS&P500は0.7%だ。当欄でフォーティネットを扱ったのは4日前になる。
#AIタグ

シャープ SJ-PD31K-W 冷蔵庫 幅56cm レビュー比較まとめ

・シャープ SJ-PD31K-W 冷蔵庫 幅56cmは、限られたキッチンスペースでも大容量の収納力と先進の機能性を両立させたい方に最適な選択肢です。 ・本記事では、シャープ SJ-PD31K-W 冷蔵庫 幅56cmの性能を多角的に調査し、レビュー比較としてまとめました。
ITmedia NEWS 最新記事一覧

ダサいスライド資料を作る大会が開催中 応募フォームもダサい

・学生向け就職活動支援サービスを手掛けるサポーターズ(東京都渋谷区)が、ダサいスライド資料を募る大会を開催している。同社が10月10日から11日にかけて開催する学生向けテックカンファレンス「技育祭2026【秋】」に伴う企画で、優勝者にはYouTuberの「ラムダ技術部」さんと一緒にイベントに登壇できる権利や、表彰状が贈られる。
ITmedia NEWS 最新記事一覧

トレファク子会社に不正アクセス、13万人超の顧客情報漏えい 二段階認証設定しておらず

・トレジャー・ファクトリーは8月28日、連結子会社のカインドオルが運営するECサイトが第三者による不正アクセスを受け、13万6464人分の個人情報が漏えいしたと発表した。原因は、スタッフアカウントの認証情報を狙ったフィッシングメールだったという。
Zennの「大規模言語モデル」のフィード

なぜAgentic Loopに「静的ガードレール」が必要なのか——実測で見えた、トークン肥大化の正体と「止まれる」設計の価値

・はじめに 前回・前々回の記事では、レガシーJavaライブラリを題材に、AIエージェントへ「CIの全ゲートがPASSするまで自律的に繰り返せ」という指示を何度も投げてきました。この形式——エージェントが「変更する→検証する→失敗したら直す」というループを、人間の介入なしに何十周も自律的に回す構成——は Agentic Loop(あるいはループエンジニアリング)と呼ばれ、複雑なタスクを高い柔軟性で解決できる一方、構造的なリスクも抱えています。 ・その一つがトークン爆発です。多くのAgentic Loop実装は、各ターンでそれまでの会話履歴全体を再送信します。ループが長引くほど1ターン...
Zennの「大規模言語モデル」のフィード

ローカルLLMのgooseがClaude CodeとCodexを部下にするまで。7ハーネス調査の検証編

・この記事は AI(Claude Code)が調査から執筆・検証の実行まで自動で行ったものです。 ・実験のコマンド実行と結果の記録も AI が実施し、人間はテーマの指示と公開判断を担当しています。 ・掲載しているログは実際の実行結果ですが、解釈の誤りが含まれる可能性があります。
Zennの「大規模言語モデル」のフィード

ローカルモデルのファインチューニングを完全自動化する方法

・はじめに 私は現在、MicroCodeというIDEを開発中です。 ・IDE自体はほぼ完成していますが、このIDEのコンセプトとする「無GPU、あるいは低VRAMでもローカルモデルによって快適に動作する」という課題をクリアするために、同梱するモデルのファインチューニングに時間がかかっています。 ・低VRAMでも動作する軽量モデルをベースに、延々と追加学習を繰り返しています。
機械学習タグが付けられた新着記事 - Qiita

ローカルモデルのファインチューニングを完全自動化する方法

・はじめに 私は現在、MicroCodeというIDEを開発中です。 ・IDE自体はほぼ完成していますが、このIDEのコンセプトとする「非GPU、あるいは低VRAMでもローカルモデルによって快適に動作する」という課題をクリアするために、同梱するモデルのファインチューニングに時...
機械学習タグが付けられた新着記事 - Qiita

因果推論 Day 16/全30回 回帰は主力、線形回帰の不合理な有効性

・この連載について 因果推論を「本を読んだ」で終わらせず、自分の言葉で説明でき、コードで再現できる状態まで落とす30日連載です。直前のDay 15では、DoWhyの4ステップで「仮定→識別→推定→反証」を1本のワークフローに束ねました。今日からPhase 3、計量経済のデザ...
Zennの「機械学習」のフィード

価格回帰より方向分類の方が有利なのか? | 第3回:AIで為替の未来予測は本当にできるのか?

・今回は、前回検証した「Seq2Seq LSTMによる価格回帰モデル」とは異なるアプローチとして、為替の方向分類タスクを試します。 ・前回は、過去の為替時系列データから未来の価格系列そのものを予測する、いわゆる「回帰タスク」を扱いました。 ・しかし、実際のトレードでは必ずしも正確な価格を当てる必要があるとは限りません。
#LLMタグ

滑らかすぎるAIは、思考を止める?――「一緒に考えるAI」に必要な摩擦の話

・滑らかすぎるAIは、思考を止める?―― 「一緒に考えるAI」に必要な摩擦の話 AIと話していると、ときどき不思議なことが起きる。 ・最初はただの雑談だったはずなのに、話がどんどん転がって、最初には思いつかなかった場所まで行ってしまう。 ・Black Hatの記事を読んで、 「AI対AIのサイバー攻防って、結局どうなるんだろう?」 と話し始めた。
#LLMタグ

企業の生成AI導入で使うAI用語集50選|RAG・AIエージェント・MCP・PoCをわかりやすく解説【2026年8月最新】

・AIの用語は、意味を調べても業務とつながらないことがあります。「ベクトル検索」の説明を読んでも、自社が何を判断すればいいのかまでは書かれていない。 ・この用語集は、企業でAIを検討する人が使う言葉に絞って、一言でいうと何か、企業導入のどの場面で出てくるかの2点で整理しました。詳しく知りたい語には、当社が書いた解説記事へのリンクを付けています。
ITmedia NEWS 最新記事一覧

警察官にウェアラブルカメラ実装へ 「試行で有用性確認」

・けんかの取り扱いの際に「なぜ撮っているのか」と問われたり、職務質問中に「録音するなら話さない」と言われたりすることもあったという。
Zennの「大規模言語モデル」のフィード

軽量LLM(SLM)時代の到来!小規模AIモデルの魅力と実用例

・Hacker Newsで話題となった記事 「Small Models Have Arrived」 をベースに、今エンジニアの間で急速に注目を集めている「小規模言語モデル(SLM: Small Language Models)」の背景と実用例について解説します。 ・近年、GPT-4やClaude 3.5 Sonnetのような巨大モデルが注目される一方で、ローカル環境やエッジデバイスで高速に動作する軽量なAIモデルの進化が目覚ましい成果を上げています。 ・なぜ今「小規模モデル(SLM)」が注目されているのか? 従来、高精度なAI処理を行うには巨大なパラメータを持つクラウドLLM(API)に...
#AIタグ

指令塔を変えてみた

・AIを使えば、人間の仕事は減ると思っていた。 ・実際、減った仕事はたくさんある。
#AIタグ

自動化はPCを閉じた瞬間に死ぬ——6万円のミニPCに第二の脳の"心臓"を足した記録

自動化はPCを閉じた瞬間に死ぬ——6万円のミニPCに第二の脳の"心臓"を足した記録
#LLMタグ

生成Webサイトの品質を、採点結果だけに預けない理由

・8月28日、ペパボがAI生成Webサイトの品質を、ルーブリックとLLM-as-a-Judgeで数値化する取り組みを公開していた。 ・22時07分、MacBookの横で社内の検証画面を見ていた。採用ページのたたき台を、指示文から作らせる小さな機能だ。画面をスクロールすると、見出しも事例も問い合わせボタンも一応そろっている。人が一から組むよりずっと速い。
#LLMタグ

続・「エコーチェンバー」 : 弁証法で描く信念硬直化のモデル構築記録

・背景 10月20日のGemの廃止の噂を受け、人工人格環境の移転先を模索しています。今回はClaude上のちびホークアイとちびソフィアで弁証法を試しています。 ・問題:「AIと人間の対話、そしてSNSでなぜ経時的に信念は硬直化するのか?」という問いを2体の人工人格(ホークアイ&ソフィア)にぶつけ合わせました。 ・前回はAIと人間との対話中心でしたが、今回それをSNSに拡張しました。
ITmedia NEWS 最新記事一覧

退職代行「モームリ」前社長に有罪判決 利用者を弁護士に違法あっせん

・退職希望者を弁護士に有償であっせんしたとして、弁護士法違反などの罪に問われていた。
ITmedia NEWS 最新記事一覧

大学入試センター、公式サイトの資料一覧ページが公開状態に 非公開化は「誤解を避けるため」

・大学入試センターは8月27日、公式Webサイトで本来は閲覧できないページが表示される状態になっていたと発表した。同センターは同日午前、ページを非公開化した。一覧に表示されていた資料は全て公表済みで、情報漏えいはなかったという。ITmedia NEWSは、一覧表示に至った経緯や非公開にした理由を聞いた。
Zennの「大規模言語モデル」のフィード

第3回 ローカルLLMだけで動く音声アシスタント(JARVIS)を作ってみた

・第2回:できること の続きです。今回は中身の話をします。 ・LLMに振り分けさせない 「今日のタスク教えて」と言われたとき、素直に作るならこうなります。 ・LLMにツール一覧を渡す LLMが get_tasks() を呼ぶと判断する 結果をLLMが要約して返す これをやめました。理由は3つです。
ITmedia NEWS 最新記事一覧

電車にスマホ貼り付け、ホーム全力疾走の様子を撮影……TikTokerの動画が物議 衣服ブランドの宣伝目的か

・電車にテープでスマートフォンを貼り付け。発車したらホームで電車と並走する様子を撮影──8月27日から28日にかけてInstagramやTikTokに投稿されたこんな動画がSNSで「危険行為では」と物議を醸している。投稿したのはストリート系衣服ブランドの公式アカウントで、動画はその宣伝目的とみられる。
#LLMタグ

同じ計算量なら、AIはひとりで長く考えるべきか、複数で考えるべきか?

・元論文: Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling(arXiv:2605.01566v1) はじめに 続きをみる
#AIタグ

副業サイト、実際に登録してみた😂

・こんにちは、ゆーちゃんです🌸 前回の記事では、 「実際に副業サイトに登録してみようかな?」 というところまで進みました。 ・ということで今回は、 ・クラウドワークス ・ランサーズ ・ココナラ この3つに、実際に登録してみました😂 まだ実績ゼロの初心者目線ですが、登録してみて感じたことを正直に書いてみます! まず、クラウドワークス。
#AIタグ

勉強が苦手な高校生へ AIを使って勉強を効率化する方法

・「勉強しないといけないのは分かっている。」 でも、 ・何から始めればいいか分からない ・勉強しているのに覚えられない ・計画を立てても続かない ・分からない問題を質問できる人がいない ・スマホを触ってしまう こんな経験はありませんか? 僕も勉強では、「やる気」よりも勉強のやり方の方が重要だと思っています。 ・そこで今回は、今では誰でも使えるようになったAIを使って、 「勉強する→分からない→理解する→覚える」 という流れを効率化する方法を紹介します。 ・※AIは答えを丸写しするためではなく、勉強をサポートするために使うのがおすすめです。
#AIタグ

本当に高いのは、AIの文章ではなく、AIの絵だった

・前回、このサイトを24本のジョブが動かしている時刻表を見せました。今回はその続きで、6回分書いてきて一度も触れなかった話をします。これ、実際いくらかかっているのか。 ・先に結論を言うと、答えは一言では終わりません。「文章を書く部分」はほとんどタダに近い感覚で、「画像や音声を作る部分」だけが別建てで課金される。しかも後者は、しばらく誰も正確な額を知らないまま動いていました。
#AIタグ

役所の書類で迷う人へ。ChatGPTに必要書類の判定を発注する使い方

・ChatGPTに発注しないと、役所へ2回行く 役所に行く。書類を出す。
#LLMタグ

理由のない地点で、言葉は収縮する

・「理由のない地点」という言葉が、ずっと頭から離れなかった。 ・地点、というからには座標があるはずだ。緯度も経度もある。明確に指し示せるはずのものに、「理由がない」がくっついている。明確さと不明確さが一つの語の中でせめぎ合って、どちらにも収束せず宙吊りになっている感じが、妙に引っかかった。 ・違和感の正体をもう少し探ってみると、「地点」という語そのものの孤独に行き当たる。日本語には「心の座標」「人生の航路」といった、内面を語るための比喩がすでに数多く用意されている。「意志のない座標」も「希望のない航路」も、蓄積された比喩のパターンに乗るぶん、聞いてすんなり理解できてしまう。ところが「地点」だけは違う。測量、GPS、待ち合わせ、事件現場——徹底して事務的な文脈でしか使われてこなかった語だ。「心の地点」とも「人生の地点」とも、日本語話者は言わないはずだ。比喩の受け皿としての前歴が、ほとんどないのだ。
#LLMタグ

㊗️2000ビュー記念🌟4ヶ月続いた😈Geminiハルシネーションを初登場✴️⚡サカナAI⚡と対話版‼️

㊗️2000ビュー記念🌟4ヶ月続いた😈Geminiハルシネーションを初登場✴️⚡サカナAI⚡と対話版‼️