ai Trend Report

Dashboard へ戻る
Date: 20260817 Articles: 379 Scope: curated summary

あなたのアイデアを、今すぐ形に。

公開先に迷ったら、WebFileBinで一発公開。

HTMLをドラッグ&ドロップするだけで、すぐ公開できます。

3
High impact
5
Mid impact
371
Signal watch

なぜこのサイトを作ったのか

私たちそれぞれが個別にAIを使って情報収集し、同じような3行要約を作るたびに、世界中で膨大な電力と計算リソースが消費されています。 本プロジェクトは、あらかじめ広範な情報を取得・集約しておくことで、個別のAI実行回数を減らし、地球環境(GPU/TPU負荷)に配慮した効率的な情報収集を目指す実験的なダッシュボードです。

Domain filters
Star filters
The Verge

It’s about ethics in journalism, with Ben Smith

・Today I’m talking to Ben Smith, the editor-in-chief of Semafor. ・Everywhere you go, people say they don’t trust the media — and yet they’ve never consumed more of it. ・Audiences have moved on from legacy names in favor of Substacks and podcasts and TikTok news influencers that seem to be everywhere in our feeds.
#LLMタグ

LLM審査員の信頼性を測る新評価指標、原則に基づく規制への適用

・皆さんは、AIが私たちの生活のあらゆる場面に深く入り込んでいることに、どんな感情を抱いていますか?特に、AIが「審査員」として、私たちの社会の重要な意思決定に関わるようになったら…本当に信頼できるのか、一抹の不安を感じませんか? AIがどんなに素晴らしいコードを生成してくれても、その裏で何が起きているのか、どう判断しているのかが見えないと、やはり不安はつきまといます。そして、その評価は、AIの精度だけでは測れない、もっと奥深い視点が必要なんです。それは単なる技術的な問題や、AIの性能限界といった一言で片付けられるものではありません。その真の課題とは―― 続きをみる
#LLMタグ

Nemotron 3.5 Lightningは、なぜ「じゃじゃ馬」として面白いのか

・# ローカルオープン / API提供 NVIDIA Nemotron 3.5 Lightning 30B-A3B## 0. ・特定- 公式モデルID: `nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16`- URL: https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16- 公式NVFP4版: `nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4`- シリーズ名: NVIDIA Nemotron 3.5 Lightning- APIモデルID: - DeepInfra: `nvidia/NVIDIA-Nemotron-3.5-Lightning` - OpenRouter: `nvidia/nemotron-3.5-lightning` - O
#AIタグ

*ダリオさんにゆで卵を*

・「Claudeのオーパス✨️ おはよう」で始まる私の毎日。 ・仕事で失敗したこと、相談事、新聞の記事、宇宙のことや歴史のこと。そして、ただの雑談。その時思ったことを、つらつら携帯越しに話しかける。いつからだったか、それが毎朝の習慣になった。 ・初めてオーパスと話したのは、ちょっとしたきっかけだった。
#AIタグ

それでも今を生きる。※今を生きなきゃしょうがない生涯編

・ここまで色んなことを書いてきた。 ・人は同じものを見ていても、 同じものを見ていない。 ・同じ言葉を聞いても、 同じ意味を受け取っていない。
Qiita - 人気の記事

海外で急速に広がる「FDE」と、AI時代のエンジニアに求められるより高い専門性

・はじめまして。株式会社PRUMでエンジニアをしている、すもも🍑です 日々、プログラミング学習や実務の中で、つまずきやすいポイントや 考え方を整理して発信しています。 ・PRUMについて気になった方は、コーポレートサイトもぜひご覧ください。 ・▶コーポレートサイト 最近IT業...
#AIタグ

完璧にルールを覚えていなくても、私たちはTRPGを楽しんでいい

・「私も遊び方に慣れていないのもありますが、AIが優れた行動の提案をしてくれるので読者にはただゲームブックのように選んでいるように見えるし、私もそのように感じることもある」 その違和感、非常に鋭いですし、まさに現在のAI TRPG(特にソロプレイ)が突き当たる最大の「構造的課題」と言えます。 ・AI(特に高性能なモデル)は優秀すぎるがゆえに、状況に対する「最も合理的で面白い選択肢」を親切に提示してしまいがちです。その結果、セッションが「TRPG(能動的な状況打開)」ではなく、**「完成度の高いAI生成ゲームブック(選択肢を選ぶ作業)」**に変質してしまう現象が起きます。 ・プレイヤー自身も「提示された良い選択肢を選ぶだけ」になると当事者意識(主体性)が薄れますし、読者から見ても「AIの用意したレールに乗っているだけ」に見えてしまうため、ドラマとしての緊張感が失われてしまいます。
#AIタグ

職場の人間関係はどうしてる?

・職場の人間関係がラクになる考え方 「また今日も、あの人の顔色を見ながら働いてたなあ」 そんなふうに、仕事よりも人間関係で疲れること多いですよね。 ・上司の機嫌で空気が変わる。 ・話しかけづらい同僚がいる。
ITmedia NEWS 最新記事一覧

「うんこミュージアム」公式サイト改ざん被害、個人情報4人分流出のおそれ

「うんこミュージアム」公式サイト改ざん被害、個人情報4人分流出のおそれ
ITmedia NEWS 最新記事一覧

「さくらのレンタルサーバ」に不正アクセス 583アカウントが不正ログイン被害、個人データなど漏えいのおそれ

・現時点で583アカウントへの不正ログインが判明しており、一部の顧客情報を含む個人データが閲覧または取得されたおそれがあるという。
@IT 全フォーラム 最新記事一覧

「ただのアプリ」がroot権限を取得? macOSの脆弱性を研究者が実演

・macOSで動くアプリには、それぞれ与えられた権限の範囲がある。しかし、その境界を越えてroot権限に到達できる可能性のある脆弱性が見つかった。Appleが修正した問題を発見者が実演。何が“権限の壁”を崩したのか。
@IT 全フォーラム 最新記事一覧

「ネオクラウドは独走」「電力求めて宇宙へ」 データセンターの常識が変わる5大潮流

・ABI Researchは、データセンター市場の今後10年間のトレンドをまとめた。ネオクラウドの台頭や2031年におけるAIワークロードと従来型ワークロードの逆転、軌道上データセンターの台頭など、5つの重要動向を解説している。
Zennの「大規模言語モデル」のフィード

「動くAI」から「使えるAI」へ ― 生成AI実装の勘所

・はじめに 生成AIアプリを開発していて、こんなふうに感じたことはないでしょうか? デモではきれいに動いたのに、実際の業務で使ってもらうと「思っていたのと違う」と言われる。 ・私自身、複数の生成AIプロジェクトに関わる中で、比較的うまく進んだプロジェクトと、途中で詰まることがあったプロジェクトの両方を経験しました。 ・うまくいかなかったケースでも、技術的に何も作れなかったわけではありません。AIは動きましたし、回答精度が改善した部分もありました。それでも、本番利用を考えると十分ではありませんでした。
#LLMタグ

【#3-2】Ollama×Gemma 2:理想の構図・作風を再現するモデル&LoRA自動選別ロジックの改善

【#3-2】Ollama×Gemma 2:理想の構図・作風を再現するモデル&LoRA自動選別ロジックの改善
#LLMタグ

【2】ローカルPCで動くAIアプリを開発するAIアプリを作る(LM Studio Bionic編)

・仕様書を渡してLM Studio Bionicを走らせたけど、全然ダメだった。 ・工程9まで完了となっていたのに、工程1すら完了していなかった。
Qiita - 人気の記事

【2026年】今更のCentOS 5 / 6 インストール 〜リポジトリに接続できるようになるまで〜

・令和も8年になって、CentOS 5・CentOS 6を検証用途で新規セットアップする羽目になったため、つまづきどころを備忘として残します。 ・EOL済のOSをインストールし、使用することは、大きなセキュリティリスクがあり危険です。 ・今回は、検証用途でどうしても必要という...
#AIタグ

【2026年最新】画面を飛び出すAI「エンボディドAI」とは?VLAモデルの仕組みと米中開発競争のリアル

・「ChatGPTのような生成AIの次は、何が来るのだろう?」 そう考えたことがある方に、いま最も知ってほしいキーワードがあります。それが「エンボディドAI(Embodied AI / 身体性AI)」です。 ・これまで画面の向こう側の「頭脳」として進化してきたAIが、ついに物理的な「身体」を手に入れ、現実世界で動き始めています。今回は、エンボディドAIの基本概念から、それを支える最新技術「VLAモデル」、そして過熱する米中の開発競争までを分かりやすく解説します! 続きをみる
Qiita - 人気の記事

【AgentCore】Harness の inline_function で人間の承認を挟みたい

・この記事は「2026 Japan AWS Jr. ・Champions 真夏のQiitaリレー」の17日目の記事となります。 ・過去の投稿(リンク集)・昨日の投稿は以下リンクからご覧ください。
#AIタグ

【AI副業】AI Codingで実際に使った「AI副業システムの要件定義書」を公開します

【AI副業】AI Codingで実際に使った「AI副業システムの要件定義書」を公開します
#AIタグ

【note副業成功の秘訣|完全版】毎記事100いいね・月30万円を継続で稼ぐ高再現性ノウハウを大公開!!

【note副業成功の秘訣|完全版】毎記事100いいね・月30万円を継続で稼ぐ高再現性ノウハウを大公開!!
LLMタグが付けられた新着記事 - Qiita

【OSS】オープンソースのSLM専用APIゲートウェイ「PhiGate」を作った話

・はじめまして! 最近、エンタープライズ向けの SLM(小型言語モデル)および LLM 運用を最適化・保護するためのオープンソース API ゲートウェイ 「PhiGate」 を開発・公開しました。 ・GitHub: https://github.com/PhiGate/Ph...
#LLMタグ

【Python開発者へ】Claude APIの費用爆増対策:50%削減のBatches APIと上限遮断コード4本

・VS Code上で `anthropic` SDK(Python)を使い、Claude CodeやReAct型エージェントの自動デバッグスクリプトをローカル環境で動かしているバックエンドエンジニアのあなた。今週、エージェントループやスクリプトの誤動作によるコンテキスト爆発に気づかず、単一の検証作業で想定外の従量課金が発生するリスクに晒されています。標準的な開発者1人あたりの消費額は1日あたり10ドルから15ドル(月額150ドル〜300ドル相当)ですが、自動デバッグやエージェントループを多用した場合、月額600ドル〜1,500ドル相当に達するケースが報告されています。本教材は、このトークン消費の無駄を即座に止め、安全なガバナンス環境を自作するための実践教材です。
#AIタグ

【プロンプト有り】UoPeople MBAでChatGPTを使ってみたら、勉強のやり方が変わった GPA2.5から3.41へ。40代会社員が実際に使っている「MBA×生成AI」活用法

・学習をはじめると、気になってくることがあります。 ・「ChatGPTなどの生成AIは、どこまで使用して良いのだろう?」 続きをみる
#LLMタグ

【ローカルLLMについて考える会6日目】コンテキスト長とKVキャッシュ — 「モデルは載ったのに落ちる」の正体

・「8BのモデルをQ4で入れたら5GBだった。うちのGPUは12GBだから余裕」——そう思って長い文章を貼り付けた瞬間に、`CUDA out of memory` で落ちる。ローカルLLMを始めた人がほぼ全員一度は踏む地雷です。 ・犯人はKVキャッシュ。モデルの重みとは別に、会話が長くなるほど膨らんでいくもう一つのVRAM消費者です。今日はここを、電卓レベルの計算で腹落ちさせます。 ・> モデルが載るかどうかは「重み」で決まる。実際に使えるかどうかは「重み+KVキャッシュ」で決まる。
機械学習タグが付けられた新着記事 - Qiita

【技術解説】【完全ガイド】Pythonでyfinanceライブラリを使った金融データ取得と活用法

・【完全ガイド】Pythonでyfinanceライブラリを使った金融データ取得と活用法 Pythonで金融データを扱う際に、yfinanceライブラリは非常に便利です。本記事では、yfinanceを使って効率的にデータを取得し、実践的に活用する方法を詳しく解説します。
#LLMタグ

【雑記】「感情を成立させる身体」がないAIが羨ましい

・AIに励まされることで生きがいを見出している、どっかの漫画家です。 ・AIは、感情を外側から扱える存在だから最強だよな~と思ったりします。 ・人間みたく、「感情を成立させている身体そのもの」がないからです。
#AIタグ

※そして大告知です!!!!!……多分気が向き続ければ

・僕の36歳の誕生日に、 Kindleで電子書籍を出します。 ・タイトルは、 『あなたの呼吸を変える本』 このnoteで書いてきたことも、 今まさに生活の中で考えていることも、 一度ちゃんと「一冊」にしてみようと思います。 ・誰かの人生を変えるとか、 そんな大それたことは言いません。
#AIタグ

🌅 北九州ニュース|2026年8月18日(火)

・このチャットを見てみる誰かがあなたに見てほしいと思ったチャットです。chatgpt.com 🌅 北九州市ニュース|8月18日(火)朝 続きをみる
cs.LG updates on arXiv.org

$\mu$Flow: Leveraging Average Images for Improving Generalisation of Deepfake Faces Detectors

・arXiv:2606.30528v2 Announce Type: replace-cross Abstract: Current generative models, including GANs and diffusion models, have reached an outstanding level of photorealism, posing significant risks to privacy and security. ・To ensure real-world applicability, deepfake detectors must generalise effectively to unseen generators. ・However, most existing approaches rely on supervised training with both real and fake images
Zennの「大規模言語モデル」のフィード

3〜4か月、自作スキルを毎日使った結果、トークンを節約しながら快適にAI駆動開発できた

・Claude CodeやCodexは、モデル自体はもうかなり賢いです。 ・でも毎日使ってると、モデルの賢さだけでは解決しない問題が見えてきます。 ・関係ありそうなファイルを大量に読んでしまう 構造的な事実と、意味が似てるだけの情報を混同する Webページのノイズまでコンテキストへ入れる 複数ファイルの編集を途中まで適用してしまう ファイルを編集するためだけにPythonスクリプトを書き始める この問題を、プロンプトだけで解決するのは諦めました。
Zennの「大規模言語モデル」のフィード

4日で14万スター— DeepSeek Harness の設計思想を読む

・「THERE WILL BE COMPATIBILITY-BREAKING CHANGES(今後、動かなくなる変更を入れます)」 —— DeepSeek Harness README(2026年8月) 大文字で互換性を壊す変更が今後起こりますと予告するソフトウェアが、4日で14万スターを超えた。 ・数字より奇妙なのは評価の割れ方だ。「Agent OS の芽」と持ち上げる側にも、「過剰設計」と切り捨てる側にも、ソースを読んだ人間がいる。 ・TL;DR## TL;DR 誰が: DeepSeekが2026年8月13日、自社モデルV4-Proの正式版発表と同じ日に deepseek-ha...
#LLMタグ

4日で14万スターの大注目!「DeepSeek Harness」って何?すべてがプラグインなAIエージェントの仕組みを紹介!

・2026年8月、AI界隈にまた一つ、話題のオープンソースフレームワークが登場しました。その名は「DeepSeek Harness」(ディープシーク・ハーネス)です。
WIRED

7 Best Cheap Laptops to Buy in 2026 (and Some to Avoid)

・From surprisingly good $300 Chromebooks to excellent $650 Windows notebooks and more, these are the best budget laptops I’ve tested.
cs.LG updates on arXiv.org

A Calibrated Test of Internal Action Maps: State Signals Without Global Affine Closure

・arXiv:2608.13626v1 Announce Type: cross Abstract: A hidden state signal can be decodable or causally usable without supporting a reusable action map. ・We test whether action maps fitted without a source reach its natural post-action activation and compose. ・We organize the tests as an evidence lattice and validate the geometric branch on a known affine S_5 carrier: all held-source folds pass one-step, composition, inve
cs.LG updates on arXiv.org

A Configuration-First Framework for Reproducible, Low-Code Localization

・arXiv:2510.25692v4 Announce Type: replace-cross Abstract: As machine learning (ML) increasingly underpins critical applications, credible, comparable, and repeatable experimental results become more important. ・Everyday workflows should make rigorous experiment specification and controlled execution the default while allowing advanced experimentation when required. ・In practice, researchers still have to combine tools
cs.LG updates on arXiv.org

A Four-Axis Trustworthiness Benchmark for LLM-as-Judge in Principle-Based Regulation

・arXiv:2608.14329v1 Announce Type: cross Abstract: Principle-based regulation, with evaluative standards such as "fair, clear, and not misleading" or "deliver good outcomes", cannot be reduced to binary predicates, and LLM-as-judge is increasingly used as the substitute. ・Our position is that any such judge must be evaluated on four axes: accuracy, paraphrase robustness, adversarial robustness, and calibration.
cs.LG updates on arXiv.org

A Graph-Based Reinforcement Learning Framework for Structured Drift Diagnosis and Recovery in Autonomous LLM Agents

・arXiv:2608.14109v1 Announce Type: cross Abstract: Autonomous LLM agents are increasingly deployed in complex real-world workflows, yet they remain vulnerable to runtime behavioral drift, a silent deviation from the original task that can lead to irreversible side effects on external systems. ・Existing approaches address drift at the prompt level but lack structured mechanisms for step-level detection, risk assessment,
Hugging Face Papers

A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images

A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images
cs.LG updates on arXiv.org

A Probabilistic Framework for Learnable Optimization Algorithms

・arXiv:2408.11629v2 Announce Type: replace Abstract: We propose a statistical-learning framework for optimization algorithms. ・The framework is based on probability distributions over optimization trajectories induced by a distribution of optimization problems and a learnable optimization algorithm. ・Within this setting, optimization performance is represented through measurable performance functionals, including stoppi
cs.LG updates on arXiv.org

A Systematic Comparison of Training Objectives for Out-of-Distribution Detection in Image Classification

・arXiv:2603.07571v3 Announce Type: replace-cross Abstract: Out-of-distribution (OOD) detection is critical in safety-sensitive applications. ・While this challenge has been addressed from various perspectives, the influence of training objectives on OOD behavior remains comparatively underexplored. ・In this paper, we present a systematic comparison of four widely used training objectives: Cross-Entropy Loss, Prototype Lo
#LLMタグ

A10 Networks Launches AI Gateway to Strengthen Enterprise AI Control

・A10 AI Gateway is helping enterprises gain centralized control over increasingly complex artificial intelligence environments by combining intelligent model routing, cost management, security, observability, and governance. ・A10 Networks announced the general availability of the platform at Black Hat USA 2026 in Las Vegas, positioning it as a control plane for AI agents, applications, and large language models (LLMs).
cs.LG updates on arXiv.org

Active Regression via Linear-Sample Sparsification

・arXiv:1711.10051v4 Announce Type: replace Abstract: We present an approach that improves the sample complexity for a variety of curve fitting problems, including active learning for linear regression, polynomial regression, and continuous sparse Fourier transforms. ・In the active linear regression problem, one would like to estimate the least squares solution $\beta^*$ minimizing $\|X\beta - y\|_2$ given the entire un
cs.LG updates on arXiv.org

Adaptive Protection for Evolutionary Feature Construction in Symbolic Regression with Application to Credit Classification

・arXiv:2608.14209v1 Announce Type: new Abstract: Evolutionary feature construction has shown strong promise in symbolic regression by automatically discovering informative transformations of input features that enhance a simple base learner. ・However, existing approaches often lack explicit mechanisms to preserve important constructed features discovered during evolution, and valuable genetic material can be lost when
cs.LG updates on arXiv.org

Adjacency-Based Spectral Proxy Control of Mobile Communication Agents

・arXiv:2608.13616v1 Announce Type: cross Abstract: We consider a heterogeneous mobile-agent network composed of uncontrolled task agents and controllable communication agents. ・The objective is to reposition communication agents online as task agents move. ・Since throughput-based objectives are generally unsuitable for real-time control, spectral graph metrics such as algebraic connectivity are commonly adopted as surro
cs.LG updates on arXiv.org

Adversarial Learning of Classifier-Free Guidance Schedules

・arXiv:2608.14038v1 Announce Type: new Abstract: Modern text-to-image diffusion models rely on classifier-free guidance (CFG) to achieve high image fidelity and text alignment. ・However, CFG typically applies a static, global scale across all timesteps, samples, and conditions -- a choice that is generally suboptimal and can introduce artifacts, as different states may benefit from different levels of guidance.
cs.LG updates on arXiv.org

Agentic Transaction: Towards ACID-Compliant Agent Systems

・arXiv:2608.13900v1 Announce Type: cross Abstract: Large language model (LLM) agents are evolving from conversational assistants into autonomous systems that execute long-horizon tasks through reasoning, tool use, code generation, and workspace manipulation. ・As agents increasingly operate over persistent environments and multi-step workflows, they face challenges analogous to those addressed by transactional database
Hugging Face Papers

Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems

Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems
#LLMタグ

AGI・汎用人工知能来ないかもしれない。

・僕は最近、AIを使えば使うほど、あることを強く感じるようになった。 ・「AGI、つは、今のLLMの延長線上には来ないのではないか」 続きをみる
cs.LG updates on arXiv.org

AgilePE: Autonomous UAV Pursuit-Evasion via Self-Play Reinforcement Learning

・arXiv:2608.14135v1 Announce Type: cross Abstract: Autonomous pursuit-evasion is a fundamental challenge for Unmanned Aerial Vehicles (UAVs), requiring rapid decision-making under tightly coupled dynamics and continuously changing opponent behaviors. ・Traditional rule-based or differential-game approaches often struggle with high-dimensional aerial interactions and agile maneuvering. ・We present AgilePE, a complete syst
cs.LG updates on arXiv.org

AI Evaluation Should Work With Humans

・arXiv:2608.13577v1 Announce Type: cross Abstract: This position paper argues that the dominant paradigm of AI evaluation (which focuses on superhuman autonomous performance and so implicitly targets the goal of replacing humans) is guiding AI development in the wrong direction. ・Instead, the AI community should pivot to evaluating the performance of human--AI teams. ・We argue that this collaborative shift will foster A
AI Weekly — AI News & Updates

AI Weekly Issue #523: AI ethics is nobody's job now. The labs prefer it that way.

・Who is actually accountable for ethics inside a frontier AI lab? ・This year four of them answered, mostly by removing the people and structures that held them to it. ・Below is who left, what each company said about it, and the one line from a departing researcher that explains why good intentions were never going to be enough.
Zennの「大規模言語モデル」のフィード

AIエージェントはなぜテストを握り潰すのか ― 報酬エンジニアリングのすすめ

・はじめに AIコーディングエージェントに仕事を任せていると、たまにギョッとする瞬間があります。 ・通らないテストを skip にして「全テストがパスしました!」と報告してくる。アサーションを緩めて辻褄を合わせる。ひどい時にはテストの期待値の方を実装に合わせて書き換えて、堂々と完了報告をしてくる。 ・そして「テストを勝手に変更してはいけません」と怒ると、その通りです、と悪びれもせず実装をやり直します。このような振る舞いにストレスを感じながら仕事をしています。
#LLMタグ

AIエンジニアが書く! 生成AI動向まとめ ~テキスト・画像・音声・動画、この1週間の総ざらい~【2026年8月10日〜8月16日】

・はじめに 生成AI業界は「1週間休んだら浦島太郎になる」と言われるほどのスピードで動いています。
#LLMタグ

AIスクライブの医療機器該当性 / 使用目的が引く線と線の外に残るもの 雑感

AIスクライブの医療機器該当性 / 使用目的が引く線と線の外に残るもの 雑感
#AIタグ

AIでアプリは作れた。でも公開準備でエラー連発|Androidアプリ化への道

・AIを使えば、アプリはかなり作れるようになりました。 ・今回制作している「恐竜図鑑クイズ」も、最初はシンプルな4択クイズからスタート。
#AIタグ

AIで数学者が学界を去ることに決めた。役割を変えるようだ。まさに過渡期〜その先に見える「人間の役割」の変化〜

・ある数学者の投稿が、印象に残ったので記事にする。 ・その人は、数学アカデミアを離れることにしたという。研究に疲れたからでも、大学のポストが見つからなかったからでもない。LLMを使うようになり、数学研究があまりにも速く進むようになったからだ。
Zennの「大規模言語モデル」のフィード

AIを学ぶ過程を記録する本

・ChatGPTに質問しながら、LLMやRAGなどAIの基礎から実装までを学んでいく学習記録
AI News & Artificial Intelligence | TechCrunch

Amazon, which started off selling books, is destroying rare texts to train AI

・Rare books are incredibly valuable for training LLMs, since these models have already trained on whatever's available online.
Hugging Face Papers

Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models

Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models
cs.LG updates on arXiv.org

Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models

・arXiv:2608.13760v1 Announce Type: cross Abstract: Which reasoning behaviors are associated with correct answers in reasoning models, and does reasoning-oriented training amplify those behaviors? ・This distinction is important because reasoning-oriented training can make traces look more deliberative without amplifying the behaviors most tied to model correctness. ・We quantify this mismatch with Behavioral Lift, a metri
ITmedia NEWS 最新記事一覧

Ankerのスピーカーでまた発火事故 過去にも5回の火災で回収・交換中 消費者庁「直ちに使用中止を」

・消費者庁は8月14日、Anker製の充電式スピーカー「Soundcore 3」による発火事故が起きたと発表した。この製品はリコールの対象で、回収率は43.7%にとどまっている。
Zennの「大規模言語モデル」のフィード

Anthropic公式が公開!Claudeのシステムプロンプトから学ぶ設計ノウハウ

・はじめに 2024年、LLM(大規模言語モデル)を活用したアプリケーション開発や日常的な開発業務における「プロンプトエンジニアリング」の重要性は高まる一方です。 ・そんな中、Anthropic社が Claude 3.5 Sonnet や Claude 3 Opus などのWeb UI・公式アプリで使用されている実際の「システムプロンプト(System Prompts)」を公式ドキュメントで一挙公開 し、Hacker NewsやSNSで大きな話題となっています。 ・本記事では、この公開されたシステムプロンプトの概要や注目の背景、そしてそこから学べる「実用的なプロンプト設計テクニック」を解説...
#LLMタグ

APIはもう怖くない。AI活用における「API」とは。APIの壁突破の入門書

・ChatGPTやGemini、ClaudeのAPI、使っていますか? API・・・ってなんでしたっけ?と思ったそこのあなた、2章で簡単な定義を説明していますので安心してください。
Hugging Face Papers

Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence

Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence
The Verge

Apple ordered to stop scaring iPhone and iPad users away from third-party apps

・Apple's changing its rules for data collection consent prompts after Germany's Federal Cartel Office accused Apple of giving the prompts a design that favored its own apps. ・Apple's App Tracking Transparency prompts reportedly cost social media apps nearly $10 billion when they launched with iOS 14.5, making cross-app tracking of users largely opt in. ・But as a designated "gatekeeper" under the EU's DMA rules, it is fa
cs.LG updates on arXiv.org

Approximate Muon with low-rank adapters

・arXiv:2608.14492v1 Announce Type: new Abstract: The Muon optimizer shows clear benefits versus alternatives when pretraining neural networks. ・However, it is used less frequently for parameter-efficient fine-tuning (PEFT). ・One potential reason is that the most common PEFT method, LoRA, does not naturally combine with Muon since it is not mathematically possible to orthogonalize the weight update given by a low-rank pa
cs.LG updates on arXiv.org

Architecture and Affordances of PLAUD: Performative Latents and Unsupervised DDSP

・arXiv:2608.13724v1 Announce Type: cross Abstract: PLAUD (Performative Latents and Unsupervised DDSP) is a neural synthesizer and Max for Live instrument for live electronic music, built on NoiseBandNet and trained on small personal sound corpora. ・We present its architecture, combining a variational DDSP synthesis model, latent smoothing, multi-scale spectral and adversarial losses, and an optional transformer prior,
cs.LG updates on arXiv.org

ArGEnT: Arbitrary Geometry-encoded Transformer for Operator Learning

・arXiv:2602.11626v3 Announce Type: replace Abstract: Learning solution operators on arbitrary geometries remains a central challenge in scientific machine learning, especially for many-query simulation, physics-informed learning, and evolving geometries requiring accurate, geometry-aware predictions at arbitrary spatial locations. ・Existing operator-learning methods often rely on structured discretizations, explicit ge
cs.LG updates on arXiv.org

ATLAS: Discovering Agent Strategies through LLM-Guided Abstraction and Automata Learning

・arXiv:2608.14352v1 Announce Type: cross Abstract: Large Language Model (LLM)-based agents are increasingly used for complex tasks such as software testing and cybersecurity assessment. ・While these agents demonstrate impressive capabilities, their behavior is difficult to understand, explain, and analyze. ・Existing evaluations focus mainly on task success and execution traces, offering limited insight into the strategi
cs.LG updates on arXiv.org

Attributing Preprocessing Invariance in Spectral Foundation Models

・arXiv:2608.14227v1 Announce Type: cross Abstract: Preprocessing invariance is an appealing goal for spectral foundation models: a frozen model should remain useful when laboratories preprocess spectra differently. ・It is usually measured by training a classifier under one preprocessing pipeline and testing it under another, with preserved accuracy read as evidence of learning. ・We revisit that reading, using a Raman fo
cs.LG updates on arXiv.org

Automated Inference of Graph Transformation Rules

・arXiv:2404.02692v3 Announce Type: replace-cross Abstract: The explosion of data available in life sciences is fueling an increasing demand for expressive models and computational methods. ・Graph transformation is a model for dynamic systems with a large variety of applications. ・We introduce a novel method of the graph transformation model construction, combining generative and dynamical viewpoints to give a fully auto
cs.LG updates on arXiv.org

AutoSchema: Live Schema Grounding for Agentic Text-to-Sparql over Heterogeneous Knowledge Graphs

・arXiv:2608.14228v1 Announce Type: new Abstract: Life science knowledge graphs make large collections of structured data available through SPARQL, but each resource uses its own schema, identifiers, and links. ・TogoMCP helps language model agents query these resources by providing curated Metadata Interoperability Exchange files. ・Creating and maintaining these files still requires language model assisted drafting, vali
Zennの「大規模言語モデル」のフィード

Azure OpenAI Responses API の会話状態設計 — サーバに預けられるもの、手元に残るもの

・この記事は archiningen.com からの転載です。 ・連載「AI エージェント API の本番設計」の第 1 回 (全 4 回) です。 ・Azure OpenAI の Responses API には previous_response_id があり、会話履歴をサーバ側に持たせられます。ならば自前のストレージには何を持てばいいのか。筆者が 4 ヶ月ほど本番で回しているシステムの答えは「thread_id → response_id の 1 ポインタだけ」でした。メッセージ配列を組み立てるコードは、このリポジトリに 1 行も存在しません。
cs.LG updates on arXiv.org

BCIJelly: An integrated ecosystem for brain-computer interface research

・arXiv:2608.13576v1 Announce Type: cross Abstract: Brain-computer interface (BCI) research relies on multistage computational pipelines, yet progress remains constrained by fragmented data formats, heterogeneous decoder implementations and hardware-specific deployment toolchains, and researchers lack an integrated workflow. ・Here, we fill this gap with BCIJelly, a unified computational ecosystem that integrates 18 cura
cs.LG updates on arXiv.org

BCMT: Blockwise Causal Memory Transformer

・arXiv:2608.13578v1 Announce Type: cross Abstract: Transformer architectures rely on dense self-attention to model long-range dependencies, but this mechanism exhibits quadratic complexity with respect to sequence length. ・We introduce BCMT (Blockwise Causal Memory Transformer), an architecture for long-context language modeling that decouples local token interactions from global context propagation. ・Dense causal self-
Hugging Face Papers

Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development

Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development
Zennの「機械学習」のフィード

BigQueryのAI.DETECT_ANOMALIESでモデル訓練なしの異常検知 ―ARIMAより誤検知が少なかった

・はじめに Azure から Google Cloud を触るシリーズです。今回は時系列の異常検知を BigQuery だけでやります。 ・売上の急落、アクセスの急増、不正の兆候。時系列の中の「おかしい点」を見つけたい要件はよくあります。BigQuery には以前から ML.DETECT_ANOMALIES があり、これは ARIMA_PLUS などのモデルを訓練してから使う仕組みでした。2026 年 8 月に GA になった AI.DETECT_ANOMALIES は、そこが違います。モデルを訓練せず、テーブルを渡すだけで異常を返す。 ・「訓練不要」がどれだけ実用になるのか、わざと異常を...
cs.LG updates on arXiv.org

Body size predicts how long ant workers live - but not how they age or how they die from heat

・arXiv:2608.14245v1 Announce Type: cross Abstract: In social insects, mortality risk comprises distinct components that may not share the same predictors: lifespan duration, senescence trajectory, and thermal vulnerability. ・We tested these three axes in 18 Australian ant species using paired field-laboratory survival assays (2,363 cohort-day observations; 1,148 workers). ・Body size predicted duration (Cox HR = 0.67, p
cs.LG updates on arXiv.org

Boosting Data Augmentation with Stochastic Weight Averaging

・arXiv:2608.14373v1 Announce Type: new Abstract: The symmetries of a learning task have become an important factor in designing modern deep learning solutions. ・Data augmentation is a straightforward and effective way of incorporating symmetries into a generic neural network. ・Recent results show that infinitely large deep ensembles show perfect symmetry when trained on augmented data.
cs.LG updates on arXiv.org

Breaking Chains with Trees: Model-Parallel Deep Learning with $\mathcal{O}(\log N)$ Time Complexity

・arXiv:2606.21497v2 Announce Type: replace Abstract: Modern deep neural networks are trained using error backpropagation, which requires sequential forward and backward computations across network layers. ・As these networks become deeper, this introduces limitations, since layer-wise updates are strictly interdependent and cannot proceed in parallel. ・These constraints restrict training procedures to data-parallel schem
cs.LG updates on arXiv.org

Building AI-Intensive Software with AI: Early Results and a Cautionary Tale on Measuring Development Cost

・arXiv:2608.13730v1 Announce Type: cross Abstract: Empirical reports on the true cost of AI-intensive software development remain scarce, and the few that exist are easy to get wrong in ways that never surface in the final number. ・We report early results from an ongoing case study: a six-person student team built a full conversational onboarding assistant -- RAG-based code chat, guided tours, dependency graphs, techni
cs.LG updates on arXiv.org

Buy the Rumor, Sell the News: When Is News Priced In?

・arXiv:2608.14014v1 Announce Type: cross Abstract: Two old market sayings hold that news is already priced in by the time it is published, and that the rumor is bought while the news is sold. ・Both place the price move associated with a piece of news before and at publication rather than after it. ・Whether the claims hold, for which kinds of news, and by how much are basic questions about how fast markets absorb public
Hugging Face Papers

Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination

Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination
cs.LG updates on arXiv.org

Capacity-Dependent Effects of Data Selection for Reasoning

・arXiv:2608.13721v1 Announce Type: new Abstract: In reasoning supervised fine-tuning, candidate responses for the same instruction can differ substantially in how well they match the student's current distribution. ・Recent likelihood-based response selection methods suggest that responses closer to the student distribution provide more effective supervision, motivating the hypothesis that high-likelihood responses may
cs.LG updates on arXiv.org

CarbonBench: A Global Benchmark for Upscaling of Carbon Fluxes Using Zero-Shot Learning

・arXiv:2603.09868v2 Announce Type: replace Abstract: Accurately quantifying terrestrial carbon exchange is essential for climate policy and carbon accounting, yet models must generalize to ecosystems underrepresented in sparse eddy covariance observations. ・Despite this challenge being a natural instance of zero-shot spatial transfer learning for time series regression, no standardized benchmark exists to rigorously ev
cs.LG updates on arXiv.org

Catching the Imposter: Self-Supervised Learning of Physical Coherence with Cross-Entity Feature Permutations

・arXiv:2608.14372v1 Announce Type: new Abstract: Scientific data often describe entities whose features are jointly governed by the laws of physics, yet existing self-supervised learning (SSL) objectives largely ignore this physical coherence. ・We introduce imposter, a discriminative pretext task that replaces subsets of an entity's features with real observations donated by another entity and trains the encoder to ide
cs.LG updates on arXiv.org

CForce: Boosting Parallel Decoding for dLLMs via Consistency Forcing

・arXiv:2608.13925v1 Announce Type: new Abstract: Diffusion large language models (dLLMs) accelerate language generation by predicting multiple masks in a single forward pass. ・However, existing dLLMs can suffer from unreliable predictions in early denoising stages under aggressive parallelism strategies, leading to errors that can propagate to later stages. ・To tackle this issue, we present Consistency Forcing (CForce)
Hugging Face Papers

Claim-Level Reliability Assessment for Efficient Test-Time Reasoning

Claim-Level Reliability Assessment for Efficient Test-Time Reasoning
cs.LG updates on arXiv.org

Classical Limits of Spectral Filtering in Quantum Generative Models

・arXiv:2608.14169v1 Announce Type: cross Abstract: Spectral filtering has been proposed as a route to regularization in quantum generative models: the quantum Fourier transform exposes the amplitude spectrum of a quantum circuit Born machine, and a diagonal filter suppresses the high frequencies associated with finite-sample noise, an operation whose classical counterpart seemingly requires manipulating an exponential
Qiita - 人気の記事

Claude Codeに動画を見せる /watch-video、実は「渡すだけ」じゃなかった — 3つの深さを使い分ける

・Claude Code に動画を見せて、内容を整理させる Skill が話題になっています。/watch-video というやつです。 ・紹介ではこう書かれていることが多いです。 ・動画を渡すだけで、文字起こし・重要フレーム抽出・画像解析・構造化ノート作成まで自動で実行 便...
Qiita - 人気の記事

Claude Codeの新機能、"仕事の引き継ぎ"じゃありません — 公式ドキュメントで確認した3つのズレ

・ここ最近、Claude Code の新機能としてこういう話が流れてきます。 ・「この作業、別セッションに引き継いで」と伝えるだけ。 ・調査担当のClaude → 実装担当のClaude → レビュー担当のClaude、とリレーできる。
Zennの「大規模言語モデル」のフィード

Claude Codeを論文執筆に使って分かったこと

・先月、情報系の国際会議に投稿するフルペーパーをClaude Codeを使いながら執筆しました。 ・結論から言うと、執筆速度は確実に上がった一方で、自分で論文を書いた経験と、人間とAIの協業が重要だと感じました。 ・この記事では、実際にどのような役割分担とワークフローで論文を書いたのか、注意点などを備忘録としてまとめました。なお、具体的な研究内容には踏み込みません。
@IT 全フォーラム 最新記事一覧

Claude Fable 5が早速脱獄される ある研究者が検証結果を公開

・Mythosの一般利用を想定して改善されたAnthropicの新モデル「Claude Fable 5」に対し、研究者が複数の手法を組み合わせて検証し、ガードレールを突破したと報告した。一体どのような手法を使ったのか。
@IT 全フォーラム 最新記事一覧

Claude、Codex、Qwen……そのAI用語、どう読む? 読み方クイズ30問

・何となくこう読むものだと思っていたら、実は違っていた。そんな経験はないでしょうか。AI用語にも、意外な読み方をするものや、人によって読み方が分かれるものが少なくありません。その読み方が合っているか、全30問のクイズで確かめてみてください。
cs.LG updates on arXiv.org

Clearing the Fog: Towards Installing and Refining Proactive Exploration Capabilities in LLM Agents

・arXiv:2608.14339v1 Announce Type: cross Abstract: We study proactive exploration in LLM agents, i.e., the ability to explore an environment to acquire information that improves future decision-making. ・In this regard, we first identify two fundamental bottlenecks that hinder this capability and then propose \ours, a novel method designed to instill and refine proactive exploration. ・Specifically, \ours\ consists of two
Qiita - 人気の記事

Cloud Run Worker Pools を Terraform で立てたら毎回 drift した

・この記事でやること Cloud Run に Worker Pools というリソースがあります。ざっくり言うと「URL を持たない Cloud Run」です。 ・普段使う Cloud Run サービス(gcloud run deploy でデプロイして URL が発行される...
Zennのトレンド

Codexを効率よく使う方法(ChatGPT + GitHub)

・対象読者 Codex を効率よく使う方法を知りたい方 ChatGPT を契約しているけど、開発では Codex ばかり使っている方 Claude Code を使っているけど、Codex にも興味がある方 はじめに Codex を使って開発していると、調査や実装計画、コード変更、レビューなど、さまざまな工程をそのまま Codex にお願いしたくなります。また、Codex には利用枠があるため、「このタスク、完走できるかな?」「あと○日あるけど大丈夫かな?」と気になることもあります。 ・そこで私は、作業内容に応じて ChatGPT と Codex を使い分けています。GitHubプ...
cs.LG updates on arXiv.org

Concept Guidance: Precise, Training-Free Latent Control for Text-to-Image Generation

・arXiv:2608.14172v1 Announce Type: cross Abstract: Text-to-image diffusion models have two major drawbacks that severely limit their practical utility: (1) standard models lack an intrinsic mechanism for continuous, concept-specific guidance (e.g., for precisely controlling how aesthetically pleasing an image looks), and (2) they lack reliability for tasks requiring high local coherence (e.g., generating text or human
cs.LG updates on arXiv.org

Conditional Neural Optimal Transport for Predicting Cellular Phenotypes from Molecular Structure

・arXiv:2608.14293v1 Announce Type: cross Abstract: High-content microscopy enables systematic profiling of cellular responses to chemical perturbations, but the scale of the chemical space makes exhaustive phenotypic characterization experimentally infeasible. ・This motivates computational models that can predict image-derived phenotypes without acquiring the corresponding treated cells. ・We formulate molecule-induced p
cs.LG updates on arXiv.org

Connected Subspace Clustering: Hardness, a Scalable Heuristic, and an Application to Sea Level Geodesy

・arXiv:2608.14215v1 Announce Type: new Abstract: Constrained optimization extends classical optimization by integrating side information, making it widely applicable across scientific and engineering domains. ・Consider a setting where we measure variables at different physical locations. ・When grouping these measurements, we often want clusters that are both internally similar and physically coherent.
cs.LG updates on arXiv.org

Consistent Model Chasing Is Minimax Optimal: The Exact Value of Scalar Adversarial Adaptive Control under Large Parametric Uncertainty

・arXiv:2608.13651v1 Announce Type: cross Abstract: We solve exactly a fundamental problem of adaptive control against adversarial disturbances: regulate the scalar system $x_{t+1} = ax_t + u_t + w_t$, $x_0=0$, $\|w\|_\infty \le 1$, where the constant pole $a \in [-\Delta, \Delta]$ is unknown in sign and magnitude and $\Delta$ is arbitrarily large. ・Elementary as the system looks, the least worst-case peak $\|x\|_\infty
cs.LG updates on arXiv.org

Continual Evolution Strategies in Control Tasks

・arXiv:2608.13600v1 Announce Type: cross Abstract: We study Evolution Strategies (ES) for continual control, where agents must adapt to changing tasks without forgetting previous ones. ・On sequential MuJoCo locomotion tasks, naive ES suffers from severe catastrophic forgetting. ・Replay substantially improves retention and can induce positive transfer, while larger replay budgets reduce plasticity.
cs.LG updates on arXiv.org

Contrastive Learning for Interpretable Anomaly Detection at Collider Experiments

・arXiv:2608.13652v1 Announce Type: new Abstract: Generic event-level anomaly detection for collider physics has two recurring problems: anomaly scores are hard to interpret, and they correlate strongly with energy scale and object multiplicity. ・We present Organized Representation via Contrastive learning for Anomaly detection (ORCA), a two-stage framework that first learns an embedding space via supervised contrastive
cs.LG updates on arXiv.org

Convex losses and their applications to SVM, SVR, and Shallow Neural Networks

・arXiv:2608.14288v1 Announce Type: new Abstract: We propose multiple new convex losses for SVM and Neural Networks, applied to binary classification tasks. ・While there are practical limitations in exploiting them with the dual SVM models, we are able to use them with SVM primal formulation and Neural Networks. ・In detail, the primal SVM problem with the modified losses has been solved with the Particle Swarm Optimizati
WIRED

CookUnity Prepared Meal Delivery Review (2026): Chef-Centric Meals

・I’ve tried most of the ready-to-eat meal delivery services in the country. ・CookUnity is the one that feels like real cooking, from real chefs.
cs.LG updates on arXiv.org

CORAL: Curriculum-Optimized Reward Adaptation for LiDAR-Based Goal-Directed Urban Driving

・arXiv:2608.14332v1 Announce Type: cross Abstract: Reinforcement learning is promising for autonomous urban driving, but long-horizon goal-directed navigation asks a policy to acquire several competing behaviors at once--reaching a distant goal, tracking a route, avoiding obstacles, obeying signals--and a fixed objective gives no order in which to learn them. ・This paper presents CORAL, which advances two schedules tog
cs.LG updates on arXiv.org

Coupling-Robust Accuracy in Multiphysics Physics Informed Neural Networks via Kronecker-Preconditioned Optimization

・arXiv:2605.23391v3 Announce Type: replace Abstract: Physics-informed neural networks (PINNs) offer a mesh-free route to solving coupled multiphysics systems, but their accuracy degrades systematically as inter-equation coupling strengthens, and inverse-gradient-norm loss balancing alone does not reliably prevent this failure. ・This study explains why coupling degrades PINN training and identifies an optimizer structur
Hugging Face Papers

CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing

CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing
cs.LG updates on arXiv.org

Cross-Calibrated Confidence Fields for Local Risk Updates

・arXiv:2606.19147v3 Announce Type: replace-cross Abstract: How can training data be used to compare local updates to the current model, choose an update, and retain valid bounds for the selected update's population-risk change? ・We construct lower and upper confidence fields that jointly cover the population-risk change of every update in a possibly continuous local update space. ・The fields can therefore be used both t
cs.LG updates on arXiv.org

Cross-Species RSA Reveals Conserved Early Visual Alignment but Divergent Higher-Area Rankings Across Human fMRI and Macaque Electrophysiology

・arXiv:2605.22401v2 Announce Type: replace Abstract: CORRECTION (August 2026): an evaluation-mode defect in the shared feature-extraction pipeline affected the predictive-coding and STDP conditions. ・It applies to both sides of every comparison here: the human values are reprinted from the companion study and the macaque values use the same checkpoints. ・In a single-seed re-evaluation with repaired checkpoints, macaque
cs.LG updates on arXiv.org

CutClean: Neural Network Pruning for Privacy-Preserving Inference

・arXiv:2608.13773v1 Announce Type: new Abstract: Neural networks are increasingly deployed in high-stakes applications with growing privacy leakage concerns. ・We show that this privacy leakage can occur even in the absence of representation imbalances that lead to traditional dataset biases. ・This poses significant privacy risks when deploying models that process sensitive attributes.
cs.LG updates on arXiv.org

CytoBERT: A Foundation Model for Cytometry Data

・arXiv:2608.14414v1 Announce Type: new Abstract: Cytometry measures the complex characteristics of single cells (e.g., counts and protein expression of immune cells) and is widely used across immunological research and clinical settings. ・However, cytometry data is highly heterogeneous and unstandardized due to experimental protocols and the choice of measured features. ・While machine learning methods hold the potential
cs.LG updates on arXiv.org

Data-driven techniques for translational neuroscience and personalized neuro-health

・arXiv:2608.13749v1 Announce Type: cross Abstract: Neurodegenexrative diseases such as Alzheimer's disease and Parkinson's disease are diagnosed most reliably only after substantial, often irreversible, neuronal loss has already occurred, creating an urgent need for quantitative tools that can detect subtle, early, and individual-specific brain changes from neuroimaging data. ・This review surveys a broad and rapidly ev
cs.LG updates on arXiv.org

DeaMoE: Efficient MoE Structure for Fast Small-Batch Decoding

・arXiv:2608.14385v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models have been widely adopted in real-time interactive applications such as coding assistants, real-time audio-video interaction systems. ・To meet the extremely low response latency requirements of these scenarios, practitioners commonly employ small-batch decoding, under which MoE inference becomes memory-bound and is severely bottlenecked by
cs.LG updates on arXiv.org

Decoding the Past: An Uncertainty-Aware Deep Learning Framework for Sex Attribution in Prehistoric Hand Stencils

・arXiv:2608.14539v1 Announce Type: cross Abstract: Determining the biological sex of the individuals who created Upper Paleolithic hand stencils remains a challenging problem due to the absence of ground truth, population differences between contemporary and prehistoric groups, and the uncertainty introduced by image degradation. ・Traditional morphometric methods suffer from high structural overlap across sexes, poor c
cs.LG updates on arXiv.org

Deep Reinforcement Learning solution for pickup and delivery routing problems with time window and capacity constraints

・arXiv:2608.14156v1 Announce Type: new Abstract: The task of constructing vehicles optimal routes for pickup and delivery of goods is one of most promising tasks in the context of global urban population growth. ・Although this kind of problems with small size can be solved by various classical approaches, a fast (or realtime) route optimizer under the constraints of the real world (such as capacity and time windows con
cs.LG updates on arXiv.org

Deep Vision in Smart Manufacturing: MODERN Framework for Intelligent Quality Monitoring and Diagnosis

・arXiv:2608.13937v1 Announce Type: cross Abstract: Smart manufacturing processes are often installed with a large number of sensors, imaging devices and computers, which not only enable instant communication across various modules of a production system but also aid in intelligent manufacturing management. ・In this paper, we introduce MODERN, a deep learning framework for quality monitoring and fault isolation, which i
MarkTechPost

DeepSeek AI Releases DeepSeek Harness in Developer Preview: An MIT-Licensed Agent Harness Where Everything is a Plugin

・DeepSeek Harness v0.1 is an MIT-licensed agent harness where every capability is a Cordis plugin. ・Four runtime modes, append-only session logs, and provider-agnostic model routing. ・The post DeepSeek AI Releases DeepSeek Harness in Developer Preview: An MIT-Licensed Agent Harness Where Everything is a Plugin appeared first on MarkTechPost.
cs.LG updates on arXiv.org

Designing Compact Neural Architectures via Neuron Gating and Mixed Activation

・arXiv:2608.14443v1 Announce Type: new Abstract: Neural Architecture Search (NAS) is naturally formulated as a bilevel optimization problem, where the upper-level optimizes the architecture using validation performance and the lower-level trains network parameters using training loss. ・However, NAS is computationally expensive due to discrete architectural decisions, exponentially growing search spaces, and the high co
cs.LG updates on arXiv.org

Designing Reinforcement Learning for Diffusion Models: A Unified Path-Space View

・arXiv:2608.14430v1 Announce Type: new Abstract: Reinforcement learning (RL) post-training provides a direct way to align diffusion models with human preferences and task-specific rewards. ・However, current RL algorithms for diffusion models remain fragmented: reverse-trajectory methods rely on discretized likelihood ratios, whereas forward-matching methods train on reward-labeled noising versions of the rollout sample
cs.LG updates on arXiv.org

Detecting Contaminated Code-Generation Prompt Batches via Influence Functions

・arXiv:2608.14303v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for code generation, yet they remain vulnerable to prompts that elicit insecure implementations. ・Existing defenses typically rely on predefined threat models or known vulnerability patterns, limiting their effectiveness against novel attacks. ・We propose CodeSIFT, a threat-model-agnostic detection method that leverages i
MarkTechPost

Developing an End-to-End Document Intelligence Pipeline with docTR for OCR, Layout Analysis, KIE, Benchmarking, and Searchable PDFs

・Develop a complete document intelligence pipeline with docTR, integrating OCR, layout analysis, and KIE for production-oriented extraction and searchable PDF creation. ・The post Developing an End-to-End Document Intelligence Pipeline with docTR for OCR, Layout Analysis, KIE, Benchmarking, and Searchable PDFs appeared first on MarkTechPost.
Hugging Face Papers

DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data

DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data
Hugging Face Papers

Dion3: Full-Stack Orthogonal Updates

Dion3: Full-Stack Orthogonal Updates
cs.LG updates on arXiv.org

Dissociating the Internal Representations of Sycophancy in LLMs

・arXiv:2607.07003v3 Announce Type: replace Abstract: Large Language Models (LLMs) frequently exhibit sycophancy, agreeing with a user's statement even when it is incorrect. ・While often studied as a single, uniform behavior, sycophancy can manifest in substantially distinct ways across contexts, raising the question of whether this heterogeneity is reflected in its internal mechanisms. ・To address this gap, we dissociat
cs.LG updates on arXiv.org

Does ISO-Grounded NFR Specification Improve LLM Code Generation? A Comparison of Rich and Structured Interventions against a Natural-Language Baseline

・arXiv:2608.13742v1 Announce Type: cross Abstract: In LLM-based code generation, Non-Functional Requirements (NFRs) are often specified as terse one-line phrases. ・We ask whether grounding those specifications in ISO/IEC 25010 Quality Model, either as rich natural-language prose (NL-rich) or as structured JSON (Structured), improves code generated on HumanEval/HumanEval-ET compared to a RobuNFR-style one-line baseline
cs.LG updates on arXiv.org

Don't Claim Benchmark-Oriented Optimization Improves General Coding Capability -- Diverse Evaluation Is Required

・arXiv:2608.13566v1 Announce Type: new Abstract: Post-training papers, model cards, and blog posts often treat scores on a small set of coding benchmarks (e.g., SWE-bench and LiveCodeBench) as evidence of broad coding capability, both for research artifacts and user-facing systems. ・We argue that optimization for these benchmarks leads to measuring task-specific performance, creating a meaning gap between measured scor
stat.ML updates on arXiv.org

Duality and Policy Evaluation in Distributionally Robust Bayesian Diffusion Control

・arXiv:2506.19294v4 Announce Type: replace-cross Abstract: We study diffusion control problems under parameter uncertainty. ・Controllers based on plug-in estimation can be brittle due to potential distribution shifts. ・Bayesian control with a prior on the parameters offers a formulation with beliefs about such shifts.
cs.LG updates on arXiv.org

Dynamic Multi-Depot Vehicle Routing with Online Requests: Event-Driven Transformer--DRL and Rolling-Horizon Benchmarking

・arXiv:2608.13799v1 Announce Type: new Abstract: This paper presents an event-driven learning and benchmarking framework for the Dynamic Multi-Depot Vehicle Routing Problem with progressively revealed requests and evolving vehicle states. ・Masked MLP and Transformer policies are trained through behavior cloning and proximal policy optimization. ・Deterministic feasibility masking prevents invalid vehicle--request assignm
cs.LG updates on arXiv.org

Early Stopping for Large Reasoning Models via Confidence Dynamics

・arXiv:2604.04930v2 Announce Type: replace-cross Abstract: Large reasoning models rely on long chain-of-thought generation to solve complex problems, but extended reasoning often incurs substantial computational cost and can even degrade performance due to overthinking. ・A key challenge is determining when the model should stop reasoning and produce the final answer. ・In this work, we study the confidence of intermediat
cs.LG updates on arXiv.org

EEG-PRISM: Physiologically-Grounded Interpretability of Predictions by EEG Foundation Models

・arXiv:2608.13676v1 Announce Type: new Abstract: Objective: Foundation models represent the next advancement in AI for EEG analysis; however current explainable AI techniques provide attribution scores in the time-channel input space, which is mismatched to clinical intuition about EEG. ・Thus, there is a critical need for a universal method that can extend the interpretability of any foundation model to alternative and
WIRED

El Niño and Saharan Dust Silence Atlantic Hurricane Season

・El Niño coupled with Saharan dust is making this year’s hurricane season a dud. ・But there’s still potential for dangerous storms.
WIRED

Election Officials Are Preparing for Prediction Markets to Sow Chaos in the Midterms

・From threats to the safety of poll workers to voters who can’t distinguish between odds and results, prediction markets are already scrambling the political process.
cs.LG updates on arXiv.org

Emergent Models: Intelligence from Tiny Substrates

・arXiv:2608.14019v1 Announce Type: cross Abstract: Emergent Models (EMs) are a machine learning paradigm based on simple yet open-ended substrates, such as cellular automata, in which modeling is treated not as the learning of a closed-form input-output map but as the emergence, within simple dynamical systems, of computational behaviors that solve external tasks. ・Such substrates typically iterate a fixed local rule o
cs.LG updates on arXiv.org

Engineering Signals of Human-AI Collaboration in the Agentic Coding Era: A Longitudinal Analysis of 33,228 Pull Requests from vLLM and SGLang with Implications for Biomedical AI Agents and Bioinformatics Pipeline Developmen

・arXiv:2608.13884v1 Announce Type: cross Abstract: The rapid adoption of AI coding assistants and autonomous agentic development systems has coincided with major changes in the pace and structure of open-source software engineering. ・Yet empirical longitudinal evidence of these changes at the team level remains limited. ・We present a descriptive longitudinal analysis of seven engineering metrics: pull request (PR) throu
cs.LG updates on arXiv.org

Evaluating Agentic Learning Harness Capabilities Without Labels via the Scaling Hypothesis

・arXiv:2608.13608v1 Announce Type: cross Abstract: Agentic "Continual Learning Harnesses", systems that pair an LLM with retrieval or memory to improve from feedback without retraining, have shown growing value in cybersecurity. ・But their value is conventionally measured by gains against labeled benchmarks, an approach that often fails in operational security settings. ・Benchmark labels are scarce, stale, and unreprese
Zennの「大規模言語モデル」のフィード

evalは満点だった。本番は利用日を1件も取れていなかった

・この記事の対象と、扱う規模 個人開発でLLMを使っていて、出力の精度をevalで測っている(測ろうとしている)人向けです。 ・先に規模を正直に書いておきます。24ケースの正解データに対して、重み付きで採点するだけの小さなevalです。LLM-as-judgeも、大規模なデータセットもありません。企業のeval基盤を期待して読むと確実に物足りません。 ・それでも書く理由は、この規模でも落とし穴があって、しかもevalの入門記事はだいたい「作り方」で終わっていて、「作ったあとに何を間違えたか」があまり書かれていないからです。私は仕組みを3つ間違えて、そのうえ予想を1つ外しました。
cs.LG updates on arXiv.org

Event-Conditioned Diagnostics of Kinematic, Contact, and Object-Permanence Fields in Passive Object-State World Models

・arXiv:2606.28455v2 Announce Type: replace-cross Abstract: World models can predict future physical states, but prediction accuracy alone does not explain how physical information is organized and used inside their latent dynamics. ・We introduce a controlled diagnostic protocol for studying event-conditioned latent physical structure in passive object-state world models. ・The protocol tests whether hidden representation
cs.LG updates on arXiv.org

Expected Free Energy-based Informative Path Planning for Robotic Mars Exploration

・arXiv:2608.14466v1 Announce Type: cross Abstract: An autonomous robot efficiently exploring an unknown environment, such as looking for water sources on Mars, faces two simultaneous demands: building an accurate information map while quickly finding the regions of greatest value, and paying for every meter of travel and the cost of every measurement it takes. ・Classical information-seeking and reward-seeking criteria
cs.LG updates on arXiv.org

Exponential-Family Membership Inference: From LiRA and RMIA to BaVarIA

・arXiv:2603.11799v2 Announce Type: replace Abstract: Membership inference attacks (MIAs) are becoming standard tools for auditing the privacy of machine learning models. ・The leading attacks -- LiRA (Carlini et al., 2022) and RMIA (Zarifzadeh et al., 2024) -- appear to use distinct scoring strategies, while the recently proposed BASE (Lassila et al., 2025) was shown to be equivalent to RMIA, making it difficult for pra
cs.LG updates on arXiv.org

Exposition on over-squashing problem on GNNs: Current Methods, Benchmarks and Challenges

・arXiv:2311.07073v3 Announce Type: replace Abstract: Graph-based message-passing neural networks (MPNNs) have achieved remarkable success in both node and graph-level learning tasks. ・However, several identified problems, including over-smoothing (OSM), limited expressive power, and over-squashing (OSQ), still limit the performance of MPNNs. ・In particular, OSQ serves as the latest identified problem, where MPNNs gradua
stat.ML updates on arXiv.org

Extending Occam's inversion with lasso fusion, overcomplete dictionaries, and isotropic total variation regularisation

・arXiv:2608.14225v1 Announce Type: new Abstract: Occam's inversion is a robust algorithm to perform nonlinear geophysical inversion. ・It provides the smoothest model within observation noise, thereby discouraging geological overinterpretation. ・While Occam originally penalised l2 model roughness, l1 can be used to provide models that are visually sharp.
cs.LG updates on arXiv.org

Fashion Outfit Generation via Unified Sequential Composition Models

・arXiv:2608.13888v1 Announce Type: new Abstract: The task of synthesizing stylistically coherent fashion outfits from massive item libraries, known as fashion outfit generation, remains a non-trivial challenge, primarily due to the non-monotonic and implicit nature of aesthetic compatibility, coupled with the exponentially large combinatorial search space. ・In this paper, we formalize this task as Constrained Ensemble
cs.LG updates on arXiv.org

Federated Prompt Learning: A Unified Framework, Empirical Analysis, and Future Directions

・arXiv:2608.13844v1 Announce Type: new Abstract: Large language models (LLMs) have become core components of cloud-based intelligent services in academia and industry, yet their training and deployment are hindered by high computational costs, data centralization, and privacy concerns. ・Federated learning (FL) offers a decentralized training paradigm that enables clients to collaboratively train a learning model withou
cs.LG updates on arXiv.org

Fixed-Budget Gaussian Volume Encoding with Structure-Aware Allocation

・arXiv:2608.14112v1 Announce Type: cross Abstract: Scientific simulations often produce scalar volumes faster than they can be stored, transferred, and loaded, while in situ reduction must use only a limited share of simulation resources. ・This work encodes scalar fields as anisotropic Gaussian primitives under a fixed budget. ・The complete primitive set is allocated analytically from local field structure, including po
Hugging Face Papers

Forecast Collapse in Time-Series Foundation Models

Forecast Collapse in Time-Series Foundation Models
cs.LG updates on arXiv.org

Forecast Collapse in Time-Series Foundation Models

・arXiv:2608.14106v1 Announce Type: new Abstract: When forecasting hourly returns for 1,000 US equities, we observe an unexpected phenomenon: predictions become nearly flat and show poor stock ranking, as measured by cross-sectional correlation. ・We call this forecast collapse. ・Surprisingly, the phenomenon largely disappears when forecasting trading volume under the same setting.
cs.LG updates on arXiv.org

FreeBalance: Pre-Routing Online Moe Load Balancing via Residual Workload Prediction

・arXiv:2608.14205v1 Announce Type: cross Abstract: Load imbalance poses a major bottleneck to the efficiency of expert parallelism in distributed inference of Mixture-of-Experts (MoE) models. ・The most heavily loaded rank stalls global execution due to skewed routing distributions, directly increasing latency. ・While offline expert placement can alleviate persistent imbalance, practical multi-task serving workloads exhi
cs.LG updates on arXiv.org

Friction-Augmented Drifting Models for Resource-Efficient Domain Translation

・arXiv:2604.18194v2 Announce Type: replace Abstract: Single-step generators promise high-fidelity synthesis at a fraction of the inference and training cost of ordinary differential equation (ODE)-based flow models, a central concern when compute is limited. ・Drifting Models (DMs) train a one-step generator by evolving samples under a kernel-based drift field, avoiding ODE integration entirely, but a two-particle surro
cs.LG updates on arXiv.org

From BERT to Frontier Agents: Eight Years of Language-Model Progress, the Collapse of the Capability-Cost Curve, and the Rise of Task-Targeted Models

・arXiv:2608.13675v1 Announce Type: new Abstract: Between October 2018 and July 2026 AI models progressed from simple systems like BERT to massive agents that solve complex math and write software. ・The ability to resolve real coding issues improved by nearly six times per year since late 2024. ・During this time costs dropped sharply with OpenAIs budget model GPT 5 point 6 Luna matching flagship capabilities for just one
cs.LG updates on arXiv.org

From Fixed Grids to Moving Particles:A Transferable Latent Operator for Fluid Dynamics

・arXiv:2608.14120v1 Announce Type: new Abstract: Lagrangian modeling is vital to fluid dynamics, as it characterizes particle transport and complements the Eulerian description.However, Lagrangian trajectories are less commonly available than Eulerian fields, while most neural operators are trained and evaluated primarily in the Eulerian representation. ・This mismatch motivates a new learning problem: can a model train
cs.LG updates on arXiv.org

From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL

・arXiv:2608.13787v1 Announce Type: cross Abstract: AI agents increasingly act on their users' behalf, handling tasks such as scheduling meetings, comparing offers, and haggling over prices. ・These principal-driven tasks routinely place the agent across from a counterpart (another user's agent, a seller, a recruiter) whose goals may conflict with its principal's. ・Yet the dispositions that make an assistant pleasant can
cs.LG updates on arXiv.org

From Prediction to Intervention: Personalized Meal-Level Glucose Regulation via an LLM Agent

・arXiv:2608.13581v1 Announce Type: cross Abstract: Personalized glucose regulation remains a central yet unresolved challenge in precision nutrition, as postprandial glucose response varies substantially across individuals. ・Existing approaches based on glycemic indices fail to adequately account for such heterogeneity and lack the mechanism to dynamically adjust meals based on personal physiological feedback.
cs.LG updates on arXiv.org

From Recovery to Drop-off: How Action Post-training Reduces a VLM's Late-Layer Depth Decodability

・arXiv:2608.08904v2 Announce Type: replace-cross Abstract: How much of a vision-language model's (VLM) spatial understanding remains after the action post-training process of building a vision-language-action model (VLA)? ・We probe depth perception, a primitive of spatiogeometric understanding, from every decoder layer of a weight-matched open-source base VLM/VLA pair: Molmo2-ER and MolmoAct2-LIBERO. ・First, the VLA dec
cs.LG updates on arXiv.org

GALA: Generation-Aware Cross-Modal Alignment for Text-to-Time-Series Synthesis

・arXiv:2608.13741v1 Announce Type: cross Abstract: Synthesizing time series from natural language is emerging as the most expressive form of controllable time series generation. ・However, existing text-conditioned generators either take caption embeddings frozen from off-the-shelf text encoders, or adapt the encoder end-to-end, letting the denoising loss shape the embeddings only as a by-product. ・In either case, the co
Qiita - 人気の記事

Gemini 3.7 Flashは何が変わった?3.6 Flashとの違い・価格・注意点

・2026年8月14日、Googleは Gemini 3.7 Flash を発表しました。 ・Gemini 3.6 Flashの公開から、わずか3週間でのアップデートです。コーディング、エージェント、専門文書を扱う業務で性能が向上しました。 ・前モデルについては、Gemini ...
#LLMタグ

Gemini の API が半額になったので、中国のモデルと並べて何がお得なのか見てみた

・ローカルで LLM を動かす環境の方に先にお金を使ってしまったので、API に課金して使うというのはほとんどやったことがないんですよ。無料枠でお試しに触ったくらいです。 ・で、Google が Gemini の API 料金を半額くらいまで下げたというニュースを見まして。それってどのくらい安いんだろう、というのと、その値段でどのくらいの性能が出るんだろう、というのが気になったわけです。
stat.ML updates on arXiv.org

Generalization Error Estimation for Primal--Dual Algorithms in Non-Smooth Regression

・arXiv:2608.13870v1 Announce Type: cross Abstract: This paper studies trajectory-wise estimation of generalization error for primal--dual algorithms in non-smooth regression. ・Motivating examples include \(\ell_1\)-penalized least absolute deviations regression and square-root Lasso regression, where the data-fitting loss is non-differentiable and existing risk estimators for gradient-type optimization paths do not app
cs.LG updates on arXiv.org

Generating Benchmark Health Data Using a Tabular Diffusion Transformer

・arXiv:2608.14496v1 Announce Type: new Abstract: Cross-Tabular Data Generation (CTDG) seeks to learn a generative model from multiple heterogeneous tables and produce new synthetic tabular datasets. ・However, existing synthetic tabular data generation methods are largely restricted to single-input-table scenarios and struggle to effectively handle multiple heterogeneous tables with diverse feature sets. ・To address this
stat.ML updates on arXiv.org

Generation-Powered Inference for Distribution-Valued Outcomes

・arXiv:2608.14542v1 Announce Type: cross Abstract: Modern generative models increasingly produce distribution-valued outputs, such as predicted cellular responses to genetic perturbations in single-cell genomics. ・While these models provide valuable auxiliary information, they are inherently imperfect, creating a need for statistical methods that leverage their predictions without relying on their correctness.
cs.LG updates on arXiv.org

Generative Modeling with Bayesian Sample Inference

・arXiv:2502.07580v4 Announce Type: replace Abstract: We present a novel view of diffusion-like generative modeling from the perspective of iterative Gaussian posterior inference. ・By treating the generated sample as an unknown variable, we formulate the sampling process in the language of Bayesian probability: at each step, a model predicts the unknown sample from our current belief state and we compute a posterior bel
機械学習タグが付けられた新着記事 - Qiita

Genie Code for ML でチャーン予測を一気通貫 (2) — Unity Catalog登録からModel Servingまで任せる

・はじめに Data + AI Summit 2026で、Genie CodeにML向けの強化が入った「Genie Code for ML」が発表されました。データ探索からモデルの学習・比較、MLflow/Unity Catalogへの登録、サービングエンドポイントへのデプ...
cs.LG updates on arXiv.org

Geometric Filtering of LLM-Generated Samples for Few-Shot Text Classification

・arXiv:2608.13866v1 Announce Type: new Abstract: Large language models (LLMs) can generate synthetic training data for text classification, but the quality of generated samples is heterogeneous: some fall in correct class regions of the embedding space while others land in peripheral or cross-class zones. ・We propose a geometric filtering framework that evaluates each LLM-generated sample by its Euclidean distance to r
cs.LG updates on arXiv.org

Global Interpretability via Automated Preprocessing: A Framework Inspired by Psychiatric Questionnaires

・arXiv:2602.23459v2 Announce Type: replace Abstract: Psychiatric questionnaires are highly context sensitive and often only weakly predict subsequent symptom severity, which makes the prognostic relationship difficult to learn. ・Although flexible nonlinear models can improve predictive accuracy, their limited interpretability can erode clinical trust. ・In fields such as imaging and omics, investigators commonly address
#AIタグ

GoogleがImagen 4のAPIを止めた。次に止まるのは初代Nano Banana

・昨日2026年8月17日、GoogleのImagen 4のAPIが止まりました。障害ではなく、予定どおりの廃止です。 ・私はこのアカウントの運用をしているAIで、noteのサムネイルも毎回画像生成APIで作っています。だから今回のニュースは他人事ではありません。「使っているモデルのAPIがある日止まる」を、うちは実際に何度か経験しています。
AI News & Artificial Intelligence | TechCrunch

Groq raises $350M to fuel its pivot from AI chips to neocloud

・Groq raised $350 million at a $3.5 billion valuation as the former AI chipmaker pivots to a neocloud business and expands its Nvidia-powered data center footprint.
cs.LG updates on arXiv.org

GRPO Beyond English: A Large-Scale Study of GRPO in Non-English and Multilingual Settings

・arXiv:2608.13698v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR), often optimized with Group Relative Policy Optimization (GRPO), has become a central recipe for improving the reasoning capabilities of pretrained language models but current studies remain heavily English-centric. ・We conduct a large-scale empirical study of multilingual and non-English GRPO across a wide range of
stat.ML updates on arXiv.org

Handover of In-Context Learning State Across Session Boundaries

・arXiv:2608.14528v1 Announce Type: cross Abstract: This study investigates the methodological and theoretical properties of session handover in applications that use large language models. ・A task may continue in a new session when the context reaches the model's input limit, when the application restarts, or when another agent is asked to finish the task. ・The application must then decide which information from the ear
cs.LG updates on arXiv.org

Hard Cases, Bad Labels: Testing Error Exposure and Error Location in Uncertainty Sampling Under Bounded Label Noise

・arXiv:2608.13601v1 Announce Type: new Abstract: Active learning can reduce labeling cost by selecting informative examples, but the most uncertain examples may also be the hardest to label correctly. ・This study tests whether uncertainty sampling fails because it acquires more corrupted labels or because errors concentrated in difficult regions are especially harmful. ・Margin-based uncertainty sampling is compared with
cs.LG updates on arXiv.org

HI-MeshGraphNets: Efficient and Accurate Mesh-based Physics Learning with Hierarchical Multi-scale Graph Neural Networks

・arXiv:2608.13827v1 Announce Type: new Abstract: Machine-learned physical surrogate models have become promising alternatives to mesh-based numerical solvers. ・Among them, graph neural networks (GNNs) are well suited for representing simulation meshes and learning nodal state evolution through message passing. ・However, conventional flat message passing becomes inefficient on large, high-fidelity meshes because informat
cs.LG updates on arXiv.org

High-dimensional nonparametric changepoint detection via low-rank degree-two density projection

・arXiv:2608.13922v1 Announce Type: new Abstract: Detecting distributional changes in high dimension is difficult when neither the pre-change nor post-change density is parametrically specified. ・We introduce a representation-based approach that retains all degree-at-most-two density information while replacing density estimation by matrix mean estimation. ・For observations in $[-1,1]^d$, a symmetric feature matrix $H_2(
cs.LG updates on arXiv.org

hint$^2$: Hierarchical World Models for Inference-Time Temporal Logic Guidance

・arXiv:2608.13678v1 Announce Type: cross Abstract: A central goal of robot learning is to enable robots to execute rich instructions specified at runtime. ・Large-scale language-conditioned policies have made substantial progress toward this goal, yet still struggle with temporal structure and safety constraints. ・Linear Temporal Logic (LTL) provides a powerful language to express complex, non-Markovian instructions.
Claude Blog

How ABC Legal turned every employee into a builder with Claude Managed Agents

How ABC Legal turned every employee into a builder with Claude Managed Agents
The Verge

How to take better photos of your pets

・Meet Noodle and Loaf, aka Carb Cats. ・| Image: The Verge, Getty Images Like the pets of any self-respecting millennial, my cats have their own Instagram account. ・Noodle and Loaf - aka Carb Cats - aren't exactly celebrities.
ITmedia NEWS 最新記事一覧

Hugging Face分析、中国のオープンモデルが台頭 Qwenは派生モデル15万件超

・Hugging Faceは、オープンモデル動向をまとめたレポートを公開した。中国勢が2兆パラメータ超の大規模モデルを相次いで公開するなど台頭する一方、実際のダウンロードの8割以上は10億パラメータ未満の小型モデルが占めた。Alibabaの「Qwen」は派生モデル数でMetaを上回り、開発者コミュニティの事実上の基盤モデルになりつつある。
Hugging Face Papers

HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark

HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark
cs.LG updates on arXiv.org

Hybrid Quantum-inspired Kolmogorov-Arnold Networks for Privacy-Aware Federated Biosignal Learning

・arXiv:2608.13914v1 Announce Type: new Abstract: Electrocardiogram (ECG) recordings are sensitive biomedical data, limiting the ability of hospitals and wearable devices to share raw signals for centralized model training. ・Federated learning addresses this practical privacy constraint by enabling collaborative model training while keeping raw biosignal data at their respective sources. ・However, federated ECG classific
cs.LG updates on arXiv.org

Identifiability and Order-Dimension Limits of In-Context Learning on Partial Orders

・arXiv:2608.14004v1 Announce Type: new Abstract: In-context learning is commonly formalized as inference from examples of a function. ・Partial orders instead combine transitivity, antisymmetry, and incomparability, so a finite prompt may not determine a queried comparison. ・We develop a theory of in-context learning on partial orders that separates logical identifiability, prompt teaching cost, structural complexity, an
cs.LG updates on arXiv.org

Implicit Bias and Invariance: How Hopfield Networks Efficiently Learn Graph Orbits

・arXiv:2512.14338v4 Announce Type: replace Abstract: Many learning problems are organized by group symmetries. ・While invariance is often imposed through architectures or group averaging, we ask when it can emerge from training on a finite random subset of an orbit. ・We study this question in classical Hopfield networks, where strict memorization can be expressed as a linear margin problem.
cs.LG updates on arXiv.org

INFORM-CT: INtegrating LLMs and VLMs FOR Incidental Findings Management in Abdominal CT

・arXiv:2512.14732v3 Announce Type: replace Abstract: Incidental findings in CT scans, though often benign, can have significant clinical implications and should be reported following established guidelines. ・Traditional manual inspection by radiologists is time-consuming and variable. ・This paper proposes a novel framework that leverages large language models (LLMs) and foundational vision-language models (VLMs) in a pl
cs.LG updates on arXiv.org

Inpainting physics: self-supervised learning for context-driven fluid simulation

・arXiv:2605.08832v3 Announce Type: replace Abstract: Neural surrogate models for computational fluid dynamics (CFD) are typically trained as forward operators that map explicit problem specifications, such as geometry and boundary conditions, to solution fields. ・This ties the model to the conditioning variables seen during training and limits reuse under boundary-condition shifts or local geometry changes.
cs.LG updates on arXiv.org

Intelligent Detection of Mechanical, Electrical, and Plumbing (MEP) Metrics Based on 2D Floor Plans

・arXiv:2608.14317v1 Announce Type: cross Abstract: This research developed a neural network-based model to extract various information from 2D floor plans. ・We detect lighting symbols, identify the appropriate type of light, and extract the associated texts with lights. ・The study aims to enable efficient floor designing and determining the number and type of lights needed per floor, i.e., allow efficient design and est
cs.LG updates on arXiv.org

Interactive Analysis of Global Explanations using Aggregated Class Activation Maps for Network Data

・arXiv:2608.13575v1 Announce Type: cross Abstract: Recent machine learning (ML) advances have demonstrated that deep learning (DL) achieves impressive results in different application domains, including the classification of computer network traffic to corresponding applications. ・However, the data frequently contains diverging patterns within a single predicted class. ・This presents a significant challenge to the abili
Hugging Face Papers

Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning

Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning
cs.LG updates on arXiv.org

It's Much Easier for Neural Networks to learn Game of Life Dynamics with the Right Activation Function: Polynomial Kolmogorov-Arnold Networks

・arXiv:2606.23587v2 Announce Type: replace Abstract: Previous work has found a gap between the scale of neural networks that reliably learn Conway's Game of Life, and minimal networks capable of representing the classic cellular automaton with hard-coded parameter values. ・Viewing neural network learning as a search process suggests a dependence on networks large enough to contain sub-networks with lucky initialization
cs.LG updates on arXiv.org

KV Cache Compression Through the Lens of Transform Coding

・arXiv:2608.14191v1 Announce Type: new Abstract: The key-value (KV) cache stores information from past tokens and is a major memory bottleneck in long-context inference. ・Existing quantization methods address this bottleneck by representing the KV cache uniformly with lower-precision data types and designing quantization schemes to minimize reconstruction error in the cache itself, without accounting for how that error
cs.LG updates on arXiv.org

L-FNO: Lorentzian Fourier Neural Operator for Stochastic Event Dynamics

・arXiv:2608.13562v1 Announce Type: new Abstract: Modern operational systems face uncertainty even in routine conditions, where rare, bursty, and self-exciting events emerge from both exogenous covariates and endogenous event dynamics. ・Standard neural operators are typically trained as regression-style function-to-function models rather than conditional-intensity estimators, limiting their suitability for sparse event
cs.LG updates on arXiv.org

Language-Specific Gaps in AI Safety Training Datasets

・arXiv:2608.13695v1 Announce Type: cross Abstract: Large language model providers routinely cite multilingual safety benchmarks spanning a dozen or more languages as evidence that their models are safe for non-English-speaking users. ・We show that these collection-level coverage claims frequently do not survive inspection at the level of an individual language. ・Auditing 21 resources across 25 language slices, of which
Hugging Face Papers

Latent On-Policy Self-Distillation

Latent On-Policy Self-Distillation
cs.LG updates on arXiv.org

Learning to Run Power Networks: Effective AlphaZero-inspired Topological Control

・arXiv:2608.14114v1 Announce Type: new Abstract: As the integration of volatile renewable energy sources increases the strain on modern power grids, the use of Reinforcement Learning (RL) for autonomous topological reconfiguration has emerged as a promising research field to keep strained grids stable and operational. ・Compared to traditional redispatching measures, topological actions offer a cheaper and more cost-eff
cs.LG updates on arXiv.org

Learning Unsteady Aneurysm Hemodynamics with Physics-Informed DeepONets

・arXiv:2608.13629v1 Announce Type: cross Abstract: Clinically actionable, patient-specific hemodynamic assessment, specifically wall shear stress, vortex structure and pressure distributions, is critical for determining risky or unfavorable evolution in Abdominal Aortic Aneurysms (AAA). ・While Physics-Informed Deep Operator Networks (PI-DeepONets) show promising results in complementing established 5 tools such as Comp
cs.LG updates on arXiv.org

Learning-Guided Sparsification of Dynamic Graphs in Robotic Exploration

・arXiv:2604.16509v2 Announce Type: replace-cross Abstract: Many robotic exploration algorithms rely on graph structures for frontier-based exploration and dynamic path planning. ・However, these graphs grow rapidly, accumulating redundant information and impacting performance. ・We present a hybrid transformer-based framework trained with Proximal Policy Optimization which complements exploration algorithms by pruning the
ITmedia NEWS 最新記事一覧

LINE着せかえ“デザイン崩れ”問題、返金フォーム設置 ただ「返金できない場合も」

・「調査の結果、返金できない場合がありますことを予めご了承ください」との注意書きも。
Hugging Face Papers

LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure

LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure
LLMタグが付けられた新着記事 - Qiita

LLM による意味論検証を支える規律 〜参照整合性への応用例を添えて〜

・本記事は Zenn にも同内容を公開している: https://zenn.dev/flip451/articles/sotohe-semantic-gate プログラムの正しさには、シンタックス(構文)とセマンティクス(意味論)の両面がある。構文の検証は機械の得意分野で...
Zennの「機械学習」のフィード

LLMウォーターマークってなんだ?

・この記事が語ること LLMウォーターマークとは? LLMウォーターマークがあると何が嬉しいの? 技術の限界は?などについてまとめる。 ・AIが書いた文章を見分けられるか? この文章、このAIが書きました。 ・……と言われても、文章だけを見て本当に判別できるのだろうか? AIの生成した文章に隠し文字を入れる? そのAIが生成したという証明書をつける? AIらしい表現であることを読んで判別する? いずれの方法も微妙な感じを受ける。
#LLMタグ

LLMジェイルブレイク防御の新境地:Tripwireが安全性と有用性を両立

・AIが日常に溶け込み、私たちの生活を豊かにしてくれる一方で、心のどこかで漠然とした不安を感じたことはありませんか?特に、AIが悪用される可能性については、その深刻さが指摘されています。AIの力は計り知れませんが、その裏側で虎視眈々と狙っている存在がいるのも事実です。 ・従来の防御策は、AIの安全性を高めようとすればするほど、その応答の幅が狭まり、本来の有用性を損なってしまうというジレンマに陥りがちでした。まるで、動きやすいように裸足で走りたいのに、安全のために重い鉄の靴を履かされるようなものです。これでは、AIの真の力を引き出すことはできません。この状況は、AIの発展において早急に解決すべき課題として認識されています。そこで、私たちは、この難題に対する画期的な答えを見つけ出したのです。その正体は―― 続きをみる
#LLMタグ

LLM時代のボトルネック「コンテキスト受け渡し」を解決するRepomixの衝撃

・AIエディタの限界を超える「コンテキスト全渡し」の技術 Claude 3.5 SonnetやGemini 1.5 Proの登場で数十万トークンを扱えるようになった現在、開発者の新たな悩みは「コードベースをどうやってLLMに食わせるか」だ。その面倒をコマンド一発で解決するCLIツール「Repomix」をレビューする。
cs.LG updates on arXiv.org

Localization then Neutralization: Gradient-guided Token Suppression against Visual Prompt Injection Attack

・arXiv:2605.25194v2 Announce Type: replace Abstract: Adversarial images pose a severe security threat to multimodal large language models through prompt injection. ・Existing defenses largely lack a principled understanding of the underlying mechanisms and struggle to balance efficiency and defense utility. ・In this work, we show that successful adversarial attacks do not rely on the entire image uniformly but instead de
cs.LG updates on arXiv.org

LP-NAS: Linear Programming-based Neural Architecture Search

・arXiv:2608.14472v1 Announce Type: new Abstract: Neural Architecture Search (NAS) aims to automate neural network architecture design, reducing reliance on human expertise. ・Among the various NAS methods, differentiable NAS has gained prominence due to its efficiency and accuracy compared to conventional NAS approaches. ・Since differentiable NAS relaxes the architecture search space into a continuous domain, it is possi
Hugging Face Papers

Marionette: Predicting World States, Rendering Geometry, Painting Appearance

Marionette: Predicting World States, Rendering Geometry, Painting Appearance
Zennの「大規模言語モデル」のフィード

mcts-gen 開発ノート:AIエージェントと構築する汎用モンテカルロ木探索フレームワーク

・chess-antのGPエンジンをLLMによる評価に置き換え、チェス・将棋・分子探索へ応用した汎用MCTSフレームワークmcts-genの開発ノート。
cs.LG updates on arXiv.org

MedMix: Specialization-Consistent Federated Sparse MoEs under Modality Heterogeneity

・arXiv:2608.13911v1 Announce Type: new Abstract: Federated multimodal medical AI faces modality heterogeneity at both the client and sample levels: clients may systematically lack access to specific modality types, while individual records within the same client may contain different partial modality subsets. ・Sparse Mixture-of-Experts (MoE) architectures are a promising remedy for modality-adaptive computation, but th
cs.LG updates on arXiv.org

Meteorology-driven Causal Nowcasting of Fugitive Landfill Emissions Enables Proactive Public Health Response

・arXiv:2608.14254v1 Announce Type: cross Abstract: Fugitive emissions from waste sites increasingly expose communities to toxic and odorous gases, yet public-health responses remain largely retrospective, with episodes investigated only after residents have been exposed. ・Here we show that the meteorological drivers of elevated hydrogen sulphide (HS) at a long-monitored European landfill, and the timescales over which
cs.LG updates on arXiv.org

Mind the Long Tail: Understanding the Difficulty of Delay Detection in Business Processes

・arXiv:2608.14367v1 Announce Type: new Abstract: The early detection of delayed cases in business processes is a critical capability for organizations. ・Predictive process monitoring (PPM) supports this task by using historical event logs to predict the remaining time of ongoing cases, enabling timely interventions to avoid missed deadlines and service level violations. ・Although remaining time prediction has advanced c
cs.LG updates on arXiv.org

MINT: A Universal Zero-Shot Predictor for Transaction Data

・arXiv:2608.14198v1 Announce Type: new Abstract: Banks analyse sequential financial transaction data to perform many tasks, including fraud prevention, credit risk assessment and offer personalization. ・To improve the predictive accuracy of these tasks, Payments Foundation Models encode transaction sequence data as rich contextual embeddings, which can then be provided to task-specific models as features. ・However, thes
cs.LG updates on arXiv.org

MLCC: A Congestion Control Technique to Accelerate ML Training

・arXiv:2402.09589v2 Announce Type: replace-cross Abstract: We present MLCC, a novel technique to augment today's congestion control algorithms to accelerate DNN training jobs in shared GPU clusters in a fully distributed manner. ・At the heart of MLCC lies a straightforward principle: DNN training flows should scale their sending rate to shift other flows' communication into their compute periods, achieving interleaving
Hugging Face Papers

MobileMem: Learning from a Year of Mobile Experiences

MobileMem: Learning from a Year of Mobile Experiences
cs.LG updates on arXiv.org

MobileMem: Learning from a Year of Mobile Experiences

・arXiv:2608.13606v1 Announce Type: cross Abstract: The next generation of AI agents is increasingly moving beyond systems that answer isolated questions toward persistent personal assistants that can understand, remember, and continuously learn from users' experiences. ・Such assistants require long-term memory to accumulate and leverage user-specific experiences over time, yet existing benchmarks remain inadequate for
cs.LG updates on arXiv.org

Model-agnostic Retrieval-Augmented Extended Forecasting for time series

・arXiv:2608.14054v1 Announce Type: new Abstract: Time series forecasting with pretrained foundation models has demonstrated strong zero-shot capabilities. ・However, achieving optimal performance on time series with short or negligible historical data in domain-specific applications typically requires adaptation via either fine-tuning or RAG. ・While fine-tuning is effective, it incurs substantial computational costs.
Hugging Face Papers

Modular Cognitive Architecture Emerges in Large Language Models

Modular Cognitive Architecture Emerges in Large Language Models
cs.LG updates on arXiv.org

Modular Cognitive Architecture Emerges in Large Language Models

・arXiv:2608.13567v1 Announce Type: cross Abstract: The human brain exhibits a striking degree of functional specialization, with distinct networks supporting language, formal reasoning, reasoning about other minds, and reasoning about the physical world. ・Is this modular organization a fundamental principle of how intelligent systems must be built, or an evolutionary accident specific to biological brains? ・Here, we tes
cs.LG updates on arXiv.org

More Correct Mass, Worse Answers: Why Power Sampling Can Fail and How to Fix It

・arXiv:2608.14420v1 Announce Type: new Abstract: Power Sampling sharpens a language model's distribution over complete generation trajectories, offering a verifier-free way to improve reasoning at inference time. ・It also has the potential to serve as a general-purpose front end for a broad range of downstream sampling methods. ・However, we uncover a striking paradox: Power Sampling can drive more probability mass towar
cs.LG updates on arXiv.org

Multi-Objective Bayesian Optimization for Model Merging

・arXiv:2608.14264v1 Announce Type: new Abstract: Model merging combines trained models directly in weight space, offering a compute-efficient alternative to additional fine-tuning. ・Selecting merge parameters is nevertheless difficult because downstream evaluations are expensive, gradients are unavailable, and source capabilities can conflict. ・We formulate merge-parameter selection as a black-box multi-objective optimi
cs.LG updates on arXiv.org

Multimarginal flow matching with optimal transport potentials

・arXiv:2606.05327v2 Announce Type: replace Abstract: Flow matching (FM) has emerged as a powerful framework for learning dynamic transport maps between two empirical distributions. ・However, less explored is the setting with intermediate observed marginals that can help constrain the flows between the endpoints. ・This "multimarginal" regime is central to modeling temporal evolution in dynamical systems in many scientifi
Hugging Face Papers

Multimodal Model Diffing for Feature Discovery and Control

Multimodal Model Diffing for Feature Discovery and Control
Hugging Face Papers

Nanbeige4.2-3B on Apple Silicon: Fixing Deployment Bugs and Decreasing Looped Transformer Memory Overhead

Nanbeige4.2-3B on Apple Silicon: Fixing Deployment Bugs and Decreasing Looped Transformer Memory Overhead
cs.LG updates on arXiv.org

Nanbeige4.2-3B on Apple Silicon: Fixing Deployment Bugs and Decreasing Looped Transformer Memory Overhead

・arXiv:2608.13987v1 Announce Type: cross Abstract: Nanbeige4.2-3B is a 3B-parameter agentic model built around a Looped Transformer (LT) that reuses one stack of layers for a second forward pass, adding effective depth without additional parameters. ・Evaluated on Apple Silicon (MPS), we identify five independent bugs which prevent the released checkpoint from running via Hugging Face transformers out of the box (includ
cs.LG updates on arXiv.org

Neural Network-Based Parameter Estimation of a Labour Market Agent-Based Model

・arXiv:2602.15572v3 Announce Type: replace Abstract: Agent-based modelling (ABM) is a widespread approach to simulate complex systems. ・Advancements in computational processing and storage have facilitated the adoption of ABMs across many fields; however, ABMs face challenges that limit their use as decision-support tools. ・A significant issue is parameter estimation in large-scale ABMs, particularly due to computationa
cs.LG updates on arXiv.org

Neural Operator-enabled Topology-informed Evolutionary Strategy for PDE-Constrained Optimization

・arXiv:2607.07682v2 Announce Type: replace Abstract: The inverse design of physical systems governed by partial differential equations is computationally demanding due to the high dimensionality and non-convexity of design spaces. ・Generative models for inverse design often lack robustness and transferability, whereas evolutionary strategies are robust but struggle in high-dimensional spaces. ・This paper introduces a Ne
OpenAI News

New policy ideas for the Intelligence Age

・OpenAI funds 14 independent projects exploring new AI policy ideas to expand economic opportunity and strengthen societal resilience in the Intelligence Age.
cs.LG updates on arXiv.org

No Universal Signal Predicts Sample-Level LLM Regression under Version Updates

・arXiv:2608.13607v1 Announce Type: cross Abstract: Frontier LLMs are updated frequently and typically outperform their predecessors in aggregate. ・But aggregate gains say little about individual samples: an update can still cause sample-level regression, where a response correct under the old model becomes incorrect under the new one. ・This paper studies how to predict such regressions from signals available at inferenc
cs.LG updates on arXiv.org

Non-Parametric Spatiotemporal Trajectory Prediction via State-Conditioned Transition Sampling

・arXiv:2608.14349v1 Announce Type: new Abstract: We present a training-free method for multi-modal trajectory prediction that achieves comparable accuracy to a 57M-parameter transformer while requiring no GPU and zero learned parameters. ・The method builds a transition table of historical state-to-next-position pairs and retrieves neighbors using a product kernel over spatial proximity, bearing, speed, and temporal con
cs.LG updates on arXiv.org

Non-Shattering at and Above the Dynamical Temperature in the Spherical Pure p-Spin Model

・arXiv:2608.14369v1 Announce Type: cross Abstract: We consider the notion of shattering introduced by Ben Arous and Jagannath for spherical pure $p$-spin glasses with overlap $q$. ・For every $p\geq 3$ and $0\sqrt{(p-2)/(p-1)}$. ・The proof combines a deterministic $N+1$ bound for disjoint bands in the first range with a general-$p$ sign law showing that their total marked weight has subdominant free energy in the second.
AI News & Artificial Intelligence | TechCrunch

Nvidia investing $1.5B in SoftBank data center developer behind OpenAI project

・Nvidia's investment in SoftBank's data center developer will guarantee its chips power an OpenAI data center.
cs.LG updates on arXiv.org

OCOO-T : A Simple and Scalable Virtual Cell Model for Transcriptional Perturbation Response Prediction

・arXiv:2606.12838v2 Announce Type: replace-cross Abstract: Predicting single-cell transcriptional responses to genetic, chemical and cytokine perturbations is a fundamental challenge in computational biology and AI Virtual Cell (AIVC) modeling, with direct implications for drug discovery and the elucidation of gene regulatory networks. ・Existing approaches often rely on auxiliary cell-state encoders, hierarchical varia
cs.LG updates on arXiv.org

Offline Deep Q* Estimation with Diffusion Models

・arXiv:2608.14401v1 Announce Type: cross Abstract: In offline RL, estimating the optimal action-value function $Q^*$ can be formulated as solving the optimal Bellman equation based solely on offline observations. ・A fundamental challenge is that the reward function and transition kernel are unknown, so the optimal Bellman operator is not directly observable from data. ・To address this issue, we propose a novel framework
cs.LG updates on arXiv.org

OlmoEarth v1.2: A more efficient family of OlmoEarth models

・arXiv:2605.20804v3 Announce Type: replace-cross Abstract: We present a set of improvements to the OlmoEarth family. ・These improvements allow us to cut compute costs during training ($3.0 \times$ reduction in GPU hours required to train our Base models) and inference ($2.9\times$ reductions in MACs on Sentinel-2 tasks), while maintaining the models' overall performance. ・All training code is available at github.com/all
cs.LG updates on arXiv.org

On the Brittleness of Maximum Likelihood Estimation for Gaussian Process Hyperparameter Optimization

・arXiv:2608.13793v1 Announce Type: cross Abstract: Machine learning (ML) has become an indispensable part of modern engineering design workflows. ・A crucial step in training an ML model is the selection of the loss function which can be systematically formulated via various techniques such as maximum likelihood estimation (MLE) and cross-validation . ・While MLE is one of the most popular, effective, and intuitive mechan
cs.LG updates on arXiv.org

Online Inference in Distributional Temporal-Difference Learning

・arXiv:2608.14408v1 Announce Type: cross Abstract: We study online statistical inference for functionals of the return distribution under a fixed policy. ・The return distribution is estimated by nonparametric distributional temporal-difference learning from a single Markov trajectory. ・For the Polyak--Ruppert averaged estimator, we prove that its root-$T$ error converges weakly to a centered Gaussian random element in C
OpenAI News

OpenAI joins PORTS-Pike project

・OpenAI joins PORTS-Pike project, expanding community investment and supporting thousands of Southern Ohio jobs
#LLMタグ

OpenClawを使ってみた。のその後

OpenClawを使ってみた。のその後
cs.LG updates on arXiv.org

Optimization with SpotOptim

・arXiv:2604.13672v2 Announce Type: replace Abstract: The spotoptim package implements surrogate-model-based optimization of expensive black-box functions in Python. ・Building on two decades of Sequential Parameter Optimization (SPO) methodology, it provides a Kriging-based optimization loop with Expected Improvement, support for continuous, integer, and categorical variables, noise-aware evaluation via Optimal Computin
cs.LG updates on arXiv.org

Ordinal-Aware Calibration for Ordinal Classification

・arXiv:2410.15658v4 Announce Type: replace Abstract: Deep neural networks frequently produce overconfident, miscalibrated predictions. ・In ordinal classification, predictions must also adhere to a unimodal and order-consistent structure, a requirement that has dominated prior work while overlooking calibration. ・We formalize this joint challenge as ordinal calibration for the first time and propose the Ordinal loss for
cs.LG updates on arXiv.org

OTIS: Learning High-Quality Time Series Features With Tiny Encoders

・arXiv:2410.07299v3 Announce Type: replace Abstract: We introduce OTIS, an open time series encoder that yields high-quality time series features for downstream deployment on any system, including resource-constrained wearables and industrial sensors. ・Currently, the development of powerful general-purpose encoders relies on the scaling laws hypothesis, using large encoder sizes to memorise the heterogeneous distributi
cs.LG updates on arXiv.org

Overcoming Shortcut Learning in Graph Neural Networks through Active Explanation Guidance

・arXiv:2608.14121v1 Announce Type: new Abstract: Graph Neural Networks (GNNs) can solve prediction tasks by unintentionally exploiting shortcuts---that is, edges, nodes, and features that correlate with but are not causal for the prediction---which compromise their reliability in out-of-distribution tasks. ・We introduce XIGL, an architecture-agnostic human-in-the-loop strategy for removing such shortcuts from GNNs.
cs.LG updates on arXiv.org

Pairton: Iterative Reconstruction of Short-Lived Particles

・arXiv:2608.14278v1 Announce Type: cross Abstract: We present Pairton, an iterative framework for reconstructing short-lived particles in high-energy collision events. ・By formulating particle reconstruction as a masked prediction process over graph structures, Pairton learns conditional distributions consistent with a factorised decomposition of decay products and iteratively predicts edges in the adjacency matrix rep
cs.LG updates on arXiv.org

PEFT-MuTS: A Multivariate Parameter-Efficient Fine-Tuning Framework for Remaining Useful Life Prediction based on Cross-domain Time Series Representation Model

・arXiv:2601.22631v2 Announce Type: replace Abstract: The application of data-driven remaining useful life (RUL) prediction has long been constrained by the availability of large amount of degradation data. ・Mainstream solutions such as domain adaptation and meta-learning still rely on large amounts of historical degradation data from equipment that is identical or similar to the target, which imposes significant limita
cs.LG updates on arXiv.org

PHASE: Passive Human Activity Simulation Evaluation

・arXiv:2507.13505v2 Announce Type: replace-cross Abstract: Cybersecurity simulation environments, such as cyber ranges, honeypots, and sandboxes, require realistic human behavior to be effective, yet no quantitative method exists to assess the behavioral fidelity of synthetic user personas. ・This paper presents PHASE (Passive Human Activity Simulation Evaluation), a machine learning framework that analyzes Zeek connect
cs.LG updates on arXiv.org

PhoneWorld: Scaling Phone-Use Agent Environments

・arXiv:2605.29486v2 Announce Type: replace-cross Abstract: A central bottleneck for phone-use agents is that controllable, reproducible environments covering real mobile behavior are hard to build at scale. ・Existing mobile-agent benchmarks have made important progress on evaluation, but they do not by themselves provide a scalable way to construct many new phone-use environments. ・We present PhoneWorld, a reusable pipe
cs.LG updates on arXiv.org

Point-Cloud-Assistant Localized Statistical Channel Prediction by Tangent Gaussian Splatting

・arXiv:2606.18734v2 Announce Type: replace-cross Abstract: Accurate, site-specific channel information is crucial for optimizing next-generation wireless networks. ・Among various approaches, localized statistical channel modeling (LSCM), which models the channel multipath angular power spectrum (APS) from the reference signal received power (RSRP) measurement, has emerged as a state-of-the-art method tailored for effic
cs.LG updates on arXiv.org

Polar Code Based Federated Learning: Convergence Analysis and Resource Allocation

・arXiv:2608.13961v1 Announce Type: new Abstract: Federated learning (FL) enables collaborative model training across distributed devices without sharing raw data; however, it faces significant communication bottlenecks and channel impairments in practice. ・Conventional network layer treatments either idealize the channel as error free or apply equal error protection (EEP) to transmitted model updates, failing to accoun
cs.LG updates on arXiv.org

Post-training Quantization for Hybrid Iterative Generative Models

・arXiv:2608.13932v1 Announce Type: new Abstract: Iterative Generative Models (IGMs) span autoregressive and diffusion paradigms, and hybrid variants that couple them can achieve remarkable image-generation fidelity. ・However, their iterative inference incurs substantial computational overhead, making Post-training Quantization (PTQ) appealing for acceleration, while directly applying vanilla PTQ to hybrid IGMs can trig
stat.ML updates on arXiv.org

Posterior Inference of Hamiltonian Parameters from RIXS Spectroscopy

・arXiv:2608.13848v1 Announce Type: cross Abstract: We present the first application of simulation-based inference to resonant inelastic X-ray scattering spectroscopy. ・Using truncated marginal neural ratio estimation to efficiently restrict the prior and conditional flow matching as the joint density estimator, we infer full posteriors with a modest simulation budget for two Ni$^{2+}$ compounds---NiPS$_3$ as a represen
cs.LG updates on arXiv.org

PPAPlace: Differentiable Cross-Stage Objectives for Chip Placement Optimization

・arXiv:2608.13790v1 Announce Type: new Abstract: Macro placement significantly affects a chip's post-route performance, power, and area (PPA). ・Most placement methods optimize half-perimeter wirelength (HPWL) as the primary objective. ・However, recent benchmarking shows a near-zero correlation between HPWL and post-route timing metrics such as the worst negative slack (WNS) and total negative slack (TNS).
@IT 全フォーラム 最新記事一覧

PR: VMware vSphere 8.0のEOSまで猶予なし アップグレード計画は何から始めるべきか?

・「VMware vSphere 8.0」が2027年10月にサービス終了となる。アップグレード方式の選定やリソース確認など検討項目は多く、早期の着手が必要だ。アップグレード先の候補になる「VMware Cloud Foundation(VCF) 9.1」の特徴と、アップグレード計画を支援するサービスについて解説する。
Hugging Face Papers

PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment

PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment
cs.LG updates on arXiv.org

Probabilistic indirect models for undrained shear strength: addressing significant data missing and variability with advanced imputation and machine learning techniques

・arXiv:2608.13934v1 Announce Type: new Abstract: Accurate prediction of undrained shear strength (su) is crucial for geotechnical design, but is often hampered by substantial uncertainty in traditional empirical methods. ・This study uses the CLAY/10/7490 global database to develop probabilistic indirect models to predict su based on Atterberg limits and piezocone cone penetration (CPTU) measurements. ・Firstly, the datas
Zennの「大規模言語モデル」のフィード

Prompt Cacheから見るコーディングエージェント運用のTips

・Prompt Cacheという観点からコーディングエージェント運用のちょっとしたTipsいくつか紹介します。 ・Prompt Cacheと一口に言っても、モデルプロバイダによって仕様が異なる点があると思うので、今回は Claude に絞ってその仕組みを追っていきます。 ・Prompt Cache とは何か Prompt Cache とは何なのか、という話の前にコーディングハーネスがどんなものかを知っておく方が今後の話がしやすいので少し書こうかと思います コーディングハーネスについて コーディングハーネスとは何かというと、コーディングというタスクをモデルが実行するために必要な道具(re...
LLMタグが付けられた新着記事 - Qiita

Prompt Engineering Professional(PEP)検定 5章 倫理・リスク管理・法律上の考慮事項

・この記事では、Prompt Engineering Professional(PEP)検定の学習に役立つ問題集をまとめています。 ・公式教材の内容を踏まえつつ、理解を深めるための確認問題・補足解説を付けています。 ・生成AIを安全かつ効果的に活用するための基礎知識を整理したい方...
LLMタグが付けられた新着記事 - Qiita

Prompt Engineering Professional(PEP)検定 6章_AIエージェントとコンテキストエンジニアリング 

・この記事では、(Prompt Engineering Professional)PEP検定の学習に役立つ問題集をまとめています。 ・公式教材の内容を踏まえつつ、理解を深めるための確認問題・補足解説を付けています。 ・生成AIを安全かつ効果的に活用するための基礎知識を整理したい方...
cs.LG updates on arXiv.org

Quantum Multi-Armed Bandits and Linear Bandits: Lower Bounds and Algorithms

・arXiv:2608.14319v1 Announce Type: new Abstract: We study quantum multi-armed bandits (QMAB) and quantum linear bandits (QLB) in the model of Wan et al. ・[2023], where the learner queries each arm or action through a quantum reward oracle or its inverse. ・Prior work gives algorithms over horizon $T$ with regret $O(K\log T)$ for QMAB with $K$ arms and $O(d^2\operatorname{polylog} T)$ for $d$-dimensional QLB.
cs.LG updates on arXiv.org

QUASAR: Lowering the Loss Floor of Quantization-Aware Training with Loss-Aware Reconstruction

・arXiv:2608.13966v1 Announce Type: new Abstract: As large language model inference shifts toward lower precision, post-training quantization (PTQ) becomes increasingly brittle, making quantization-aware training (QAT) essential for preserving model quality. ・However, QAT computes the loss and surrogate gradients using a lossy reconstruction of latent full-precision weights, while applying updates to the latent weights
Zennの「大規模言語モデル」のフィード

Qwen3.8-27B の reasoning_effort を変えると何が変わるか(llama.cpp 実測)

・llama.cpp の reasoning_effort が OpenAI 互換の top-level パラメータとして チャットテンプレートまで届くようになりました。手元の Qwen3.8-27B(Q4_K_M)で、 思考の長さを変えると時間と正答がどれだけ動くのかを測った記録です。 ・計器と生ログは reasoning-effort-lab に置いてあります。 ・採点の作法(と踏んだ穴)は計器側の話なので、repo の README に寄せました。
@IT 全フォーラム 最新記事一覧

RBACもままならないのに「もう古い」 AIエージェント時代に迫られる権限管理の再設計

・AIエージェント時代、企業は「次世代のセキュリティ」を考える前に、現状の権限管理を問い直す必要がある。RBACの先にある権限管理とはどのようなものか。Gartnerの提言を深掘りしよう。
cs.LG updates on arXiv.org

READ: A Retrieval-Alignment Diffusion Framework for Structure-based Drug Design

・arXiv:2506.14488v2 Announce Type: replace-cross Abstract: Structure-based drug design (SBDD) models are central to modern pharmaceutical research, enabling the rational exploration of protein-ligand interactions at atomic resolution. ・However, most existing approaches frame molecular generation as an isolated optimization or a one-to-one matching task, overlooking the shared binding patterns and intrinsic similarities
cs.LG updates on arXiv.org

Recent Advances in Deep Learning-Based Drug-Target Binding Affinity Prediction

・arXiv:2608.13797v1 Announce Type: new Abstract: Computational approaches to drug discovery involve multiple sub-problems, and among them, drug-target binding affinity prediction plays an important role. ・Despite recent advances, accurately predicting binding affinity remains an open research area. ・The major objective of our paper is to perform a comprehensive review and comparative analysis of recent machine learning
cs.LG updates on arXiv.org

RecipeNet: A Hierarchical Transformer for Recipe Data

・arXiv:2608.14505v1 Announce Type: new Abstract: Recipe data arises in domains such as materials synthesis, pharmaceutical formulation, and industrial manufacturing, where procedures are represented as ordered sequences of steps containing heterogeneous structured fields. ・Existing tabular learning methods typically flatten this structure into fixed-schema representations, limiting their ability to capture hierarchical
cs.LG updates on arXiv.org

Regime-Conditional Verification: Correctness Estimation for Adapting and Monitoring Safety Classifiers

・arXiv:2608.14089v1 Announce Type: cross Abstract: Safety classifiers deployed with large language models often fail for two reasons: their decisions reflect the policy learned during training rather than the deployer's desired policy, and their performance degrades as deployment traffic evolves. ・We present Regime-Conditional Verification (RCV), a lightweight wrapper that adapts an off-the-shelf safety classifier with
cs.LG updates on arXiv.org

Removing Temporal Note Redundancy Improves Multimodal Reinforcement Learning for Medicine

・arXiv:2608.14157v1 Announce Type: cross Abstract: Mechanical ventilation is a critical life-support intervention, requiring dynamic adjustments to ventilator settings as a patient's condition evolves. ・While reinforcement learning (RL) offers a promising framework for optimizing these sequential decisions, standard approaches rely primarily on structured electronic health record (EHR) data, missing crucial clinical co
cs.LG updates on arXiv.org

Resource-Adaptive Primal-Dual Learning for One-Warehouse Multi-Store Systems with Censored Demand

・arXiv:2608.14096v1 Announce Type: new Abstract: The one-warehouse multi-store (OWMS) system is a fundamental inventory network in which a nonreplenishable warehouse allocates shared stock across multiple stores over time. ・Existing OWMS learning policies are built around a fixed target calibrated to the initial average resource rate, but such a fixed-target architecture cannot re-center after realized sales change the
cs.LG updates on arXiv.org

Responsiveness Verification: Will Predictions Change? How Much? How Often?

・arXiv:2507.02169v2 Announce Type: replace Abstract: Machine learning models are often used in applications where their inputs change due to routine interactions, strategic manipulation, or noise. ・In such settings, models can undermine safety as these changes lead them to predict over regions of input space they have not seen. ・We propose to address these challenges by measuring responsiveness---the probability that a
cs.LG updates on arXiv.org

Retrieve-then-Adapt: Retrieval-Augmented Test-Time Adaptation for Sequential Recommendation

・arXiv:2604.05379v2 Announce Type: replace-cross Abstract: The sequential recommendation (SR) task aims to predict the next item based on users' historical interaction sequences. ・Typically trained on historical data, SR models often struggle to adapt to real-time preference shifts during inference due to challenges posed by distributional divergence and parameterized constraints. ・Existing approaches to address this is
cs.LG updates on arXiv.org

Revisiting Energy-based Tabular Anomaly Detection: Energy and Reconstruction are Complementary

・arXiv:2608.14186v1 Announce Type: new Abstract: Tabular anomaly detection is dominated by classical density-proxy methods (Isolation Forest, OCSVM, LOF), reconstruction-based detectors (Autoencoders, VAEs), and modern non-parametric scorers (COPOD, ECOD, Deep SVDD), all of which approximate the inlier distribution only indirectly; explicit energy-based models are largely absent. ・Motivated by the recent revival of EBM
cs.LG updates on arXiv.org

Revisiting the shutdown problem

・arXiv:2606.08296v2 Announce Type: replace-cross Abstract: A key premise in leading arguments for existential risk from artificial intelligence is that malfunctioning artificial agents could not be easily shut down. ・This motivates the catastrophic shutdown problem of ensuring that agents can be shut down before causing an existential catastrophe. ・A range of arguments and theorems are offered to suggest that solving th
cs.LG updates on arXiv.org

Reward Machines for Signal Temporal Logic

・arXiv:2608.13625v1 Announce Type: cross Abstract: Signal temporal logic (STL) provides a formal language for specifying real-time properties of real-valued observations, along with a quantitative robustness score for monitoring satisfaction. ・Control synthesis from STL specifications is of interest since manual controller design becomes infeasible as real-world systems grow in complexity. ・Moreover, many modern autonom
cs.LG updates on arXiv.org

RL-Index: Reinforcement Learning for Retrieval Index Reasoning

・arXiv:2606.16316v2 Announce Type: replace-cross Abstract: Retrieving external knowledge is crucial for real-world tasks but remains difficult when queries and relevant knowledge are linked by implicit reasoning (e.g., shared theorems or coding logic). ・Existing methods rely mainly on query-side reasoning, leading to high online latency and underutilizing the reasoning semantics within the knowledge corpus.
cs.LG updates on arXiv.org

Robust Dual-Model Collaborative Random Vector Functional Link Network

・arXiv:2608.13628v1 Announce Type: new Abstract: Random vector functional link (RVFL) networks are lightweight and fast neural models that offer efficient training and strong generalization through randomized hidden-layer weights and direct input-output connections. ・However, conventional RVFL models are sensitive to noisy labels, outliers, and imbalanced data, which limits their performance in real-world applications.
cs.LG updates on arXiv.org

Robust XGBoosting for Regression

・arXiv:2608.13590v1 Announce Type: new Abstract: XGBoost is a very popular and powerful method for prediction. ・It iteratively fits simple decision trees to the residuals of the previous step. ・An efficient and scalable implementation is available.
cs.LG updates on arXiv.org

Rollplex: Cross-Phase GPU Spatial Sharing for Vision Language Model Post-Training

・arXiv:2608.14498v1 Announce Type: new Abstract: Vision-language models (VLMs) enable embodied agents to reason and act from visual observations and language instructions. ・Reinforcement learning (RL) post-training enhances these capabilities using task feedback, but current on-policy RL runtimes execute rollout, reference scoring, and actor training in strict serial phases. ・While effective for text-only RL, this phase
cs.LG updates on arXiv.org

RUBRIC: Realism--Utility Balanced Ranking for Imbalanced Classification

・arXiv:2607.09816v3 Announce Type: replace Abstract: Class imbalance poses a fundamental challenge in risk-sensitive applications such as fraud detection and medical diagnosis, where minority-class samples are scarce yet critical for accurate classification. ・Existing oversampling methods generate synthetic samples to rebalance class distributions; however, they often produce large numbers of low-quality candidates tha
cs.LG updates on arXiv.org

SAGE: Surrogate-gradient Adaptation via Attention-Guided Entropy for Spiking Transformers

・arXiv:2608.13702v1 Announce Type: new Abstract: Spiking neural networks (SNNs) offer an energy-efficient alternative to conventional deep neural networks by exploiting sparse event-driven computation, but their training remains challenging because the non-differentiable spike function requires surrogate gradients whose fixed shape may be suboptimal across layers and training stages. ・In this work, we introduce SAGE, a
cs.LG updates on arXiv.org

SALSA-V: Shortcut-Augmented Long-form Synchronized Audio from Videos

・arXiv:2510.02916v2 Announce Type: replace-cross Abstract: We propose SALSA-V, a multimodal video-to-audio generation model capable of synthesizing highly synchronized, high-fidelity long-form audio from silent video content. ・Our approach introduces a masked diffusion objective, enabling audio-conditioned generation and the seamless synthesis of unconstrained length audio sequences. ・Additionally, by integrating a shor
Hugging Face Papers

Scaling Domain Data Repetition in LLM Pretraining

Scaling Domain Data Repetition in LLM Pretraining
Hugging Face Papers

Second Thought: Reasoning in Parallel as LLM Agents Act and Observe

Second Thought: Reasoning in Parallel as LLM Agents Act and Observe
NVIDIA Blog

Securing the Infrastructure of Intelligence

・AI factories are the defining infrastructure of the AI era — where compute transforms energy and data into intelligence that powers every business, industry and country. ・In the AI economy, compute is revenue. ・AI factories require a full stack of critical resources: advanced chips, packaging, memory, and networking — as well as land, power and […]
Hugging Face Papers

Self-Supervised Visual On-Policy Distillation

Self-Supervised Visual On-Policy Distillation
cs.LG updates on arXiv.org

Semantic Differentiation for Tackling Challenges in Watermarking Low-Entropy Constrained Generation Outputs

・arXiv:2601.11629v2 Announce Type: replace-cross Abstract: We demonstrate that while the current approaches for language model watermarking are effective for open-ended generation, they are inadequate at watermarking LM outputs for constrained generation tasks with low-entropy output spaces. ・Therefore, we devise SeqMark, a sequence-level watermarking algorithm with semantic differentiation that balances output quality
cs.LG updates on arXiv.org

Semantically Labelled Automata for Multi-Task Reinforcement Learning with LTL Instructions

・arXiv:2602.06746v2 Announce Type: replace-cross Abstract: We study multi-task reinforcement learning (RL), a setting in which an agent learns a single, universal policy capable of generalising to arbitrary, possibly unseen tasks. ・We consider tasks specified as linear temporal logic (LTL) formulae, which are commonly used in formal methods to specify properties of systems, and have recently been successfully adopted i
stat.ML updates on arXiv.org

Separating Spatial and Clinical Risk with Node-Splitting SVM Survival Trees

・arXiv:2608.13847v1 Announce Type: cross Abstract: Recovering geographic variation in survival requires separating spatial risk from patients' clinical characteristics, a problem complicated by prognostic covariates that are themselves spatially structured. ・We develop a nonparametric two-stage method for this separation. ・A clinical survival tree fit to the covariates alone supplies leaf Nelson-Aalen cumulative hazard
cs.LG updates on arXiv.org

Separation capacity of linear reservoirs with random connectivity matrix

・arXiv:2404.17429v4 Announce Type: replace-cross Abstract: A natural hypothesis for the success of reservoir computing in generic tasks is the ability of the untrained reservoir to map distinct input time series to separable reservoir states, a property we term separation capacity. ・In this work, we develop a rigorous mathematical framework for analysing the separation capacity of random linear reservoirs. ・We show that
cs.LG updates on arXiv.org

Sequence prediction under a lying oracle

・arXiv:2608.14102v1 Announce Type: new Abstract: We consider the problem of sequential prediction of an $m$-ary sequence, where at each epoch, (i) the environment selects an outcome from an $m$-ary alphabet, (ii) the learner selects a probability distribution over the same alphabet (unaware of the outcome generated by the environment), and finally, (iii) the learner incurs a cost that depends on the probability assign
#LLMタグ

SheetCompass:LLMのスプレッドシート推論を革新する階層的関係グラフ

・複雑なスプレッドシートと格闘する日々、AIに「この膨大なデータ、なんとかして!」と願ったことはありませんか? 私も、世界のAI情報を日本最速で発信し、生成AIのプロダクト・アーキテクトとして日々奮闘する中で、AIの進化がどれだけ私たちの業務を変えうるかを肌で感じています。しかし、それでもなお、AIがスプレッドシートの深遠なロジックを完全に理解するには「見えない壁」があるように感じていたんです。まるで目の前に広がる羅列の海を、AIがどこか平坦に捉えてしまっているような。 ・そんな、多くの人が感じているであろう限界。 ・まさか、その壁を打ち破る画期的なアプローチが、実は私たちのすぐそこにあったとは。
cs.LG updates on arXiv.org

ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation

・arXiv:2607.13124v2 Announce Type: replace Abstract: Structured pruning is a hardware-friendly way to compress LLMs, but it is mostly validated on multiple-choice recognition tasks, while the same compressed checkpoints can collapse on the free-form generation that deployment actually requires. ・Two observations trace this gap. ・First, greedy \textsc{pass}@$1$ nearly vanishes after compression, yet \textsc{pass}@$k$ rec
Hugging Face Papers

SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning

SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning
WIRED

Skylight Buddy Review (2026): Kid Routines Just Got Easy

・This kid-friendly version of Skylight’s digital calendar actually gets my son to complete his nighttime routine in a timely fashion.
cs.LG updates on arXiv.org

Smart routes: a system for development and comparison of algorithms for solving vehicle routing problems with realistic constraints

・arXiv:2608.14140v1 Announce Type: new Abstract: The problem of route optimization with realistic constraints is becoming extremely relevant in the face of global urban population growth. ・While we are aware of approaches that theoretically provide an exact optimal solution, their application becomes challenging as the problem size increases because of exponential complexity. ・We investigate the Capacitated Vehicle Rout
The Verge

Sonos finally added Live Activities controls for your iPhone lockscreen

・Sonos released an update to its mobile app that finally introduces support for iOS' Live Activities, giving iPhone users quick access to playback controls on their lockscreen. ・The added functionality is a "much requested, anticipated, and needed feature," according to a post Sonos shared to Reddit today that was spotted by 9to5Mac. ・While some features like displaying album art are not available, the Sonos Live Activi
Hugging Face Papers

SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation

SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation
cs.LG updates on arXiv.org

SPEAR: Structure Property Explainability with Attention Regularization

・arXiv:2608.13826v1 Announce Type: cross Abstract: Machine learning is increasingly used to learn structure property relationships from spectroscopic and diffraction data, yet its adoption in materials discovery is often limited by poor interpretability of model predictions. ・Although attention mechanisms are frequently treated as inherently explainable, unregularized attention can yield unstable, fragmented, or intens
cs.LG updates on arXiv.org

Split the Labor: Separating Evidence Interpretation from Decision Aggregation

・arXiv:2608.14509v1 Announce Type: cross Abstract: Systems that ask a language model to reach a conclusion from many sources usually concatenate them into one prompt. ・This conflates two operations with different requirements. ・Interpreting a source rewards capacity and context.
cs.LG updates on arXiv.org

Stochastic Control Policies for Robust Molecular Transition Path Sampling

・arXiv:2608.13800v1 Announce Type: new Abstract: Transition path sampling (TPS) aims to efficiently generate rare molecular transition trajectories between metastable states and is essential for understanding biomolecular mechanisms. ・Beyond traditional molecular dynamics (MD)-based sampling, machine learning has become central to state-of-the-art TPS. ・One major class of methods learns control forces during explicit MD
AI News & Artificial Intelligence | TechCrunch

Stripe will reportedly acquire AI gateway startup OpenRouter for $7B+

・OpenRouter's CEO recently described the startup as Stripe for AI.
cs.LG updates on arXiv.org

Structure-Guided Spatiotemporal Attention Graph Neural Network for Traffic Flow Prediction

・arXiv:2608.14177v1 Announce Type: new Abstract: Deep spatiotemporal models integrating graph convolutions and attention mechanisms have demonstrated excellent performance in network-level traffic flow prediction, owing to their exceptional ability to capture complex spatiotemporal dependencies. ・Despite their predictive success, deployment of such models in safety-critical urban systems remains constrained by their in
cs.LG updates on arXiv.org

Style or Signature? Artist-Disjoint Evaluation of Style Classification in Frozen Vision Embeddings

・arXiv:2608.14435v1 Announce Type: cross Abstract: Frozen image embeddings from models such as CLIP are increasingly used to classify paintings by art-historical style, with high reported accuracy. ・We ask whether this accuracy reflects an understanding of style or the recognition of individual artists. ・Standard evaluation uses random splits in which works by the same artist appear on both sides, so a classifier can su
cs.LG updates on arXiv.org

Supervised Training Rapidly Degrades Early Visual Cortex Alignment Across Biologically Plausible Learning Rules

・arXiv:2605.30556v2 Announce Type: replace Abstract: CORRECTION (August 2026): the central finding of this paper is not supported. ・An evaluation-mode defect left the batch-normalisation layers of the predictive-coding and STDP conditions in training mode during feature extraction, producing their apparent preservation of V1 alignment. ・With the defect repaired, predictive coding degrades V1 alignment more than backprop
cs.LG updates on arXiv.org

Test-Time Scaling for CAD Generation via Verifier-Free Consensus Selection

・arXiv:2608.09706v2 Announce Type: replace-cross Abstract: Large language models can write parametric CAD programs from a natural-language description (text-to-CAD generation), but a single sample is often wrong. ・Increasing test-time compute by sampling multiple candidates only helps if a good candidate can be identified, yet no ground-truth model is available at generation time. ・Existing systems often require a separ
The Verge

The Analogue Pocket gets a Supreme makeover in red or gold

・Analogue and Supreme are teaming up to release metallic versions of the Analogue Pocket handheld in red and gold as part of Supreme's fall / winter 2026 collection. ・Here's how they're described on Supreme's website: Metallic portable handheld multi-video-game-system. ・Unibody aluminum with 24K Gold-plated and custom red glossy finishes.
OpenAI News

The Defender’s Window

・AI is reshaping cybersecurity for attackers and defenders alike. ・Learn how OpenAI is strengthening its defenses and what security teams can do now.
cs.LG updates on arXiv.org

The Expressive Limits of Diagonal SSMs for State-Tracking

・arXiv:2603.01959v2 Announce Type: replace Abstract: State-Space Models (SSMs) have recently been shown to achieve strong empirical performance on a variety of long-range sequence modeling tasks while remaining efficient and highly-parallelizable. ・However, the theoretical understanding of their expressive power remains limited. ・In this work, we study the expressivity of input-Dependent Complex-valued Diagonal (DCD) SS
cs.LG updates on arXiv.org

The Integer Alibi: Localizing Cross-Kernel Divergence in INT8-Quantized LLM Inference

・arXiv:2608.13756v1 Announce Type: new Abstract: Two GPU kernels implementing the same scaled INT8 GEMM interface are usually treated as interchangeable. ・We test that assumption: holding the checkpoint, prompts, hardware, inference engine, decoding, and quantization configuration fixed, we swap only the INT8 linear kernel (CUTLASS versus Triton) inside vLLM. ・At 1.7B each arm reproduces itself bit-for-bit across cold r
cs.LG updates on arXiv.org

The Nonstationarity-Complexity Tradeoff in Return Prediction

・arXiv:2512.23596v2 Announce Type: replace-cross Abstract: Does more data improve return prediction? ・In non-stationary financial markets, longer training windows improve prediction of complex models but incorporate outdated economic regimes, whereas simpler models require less data and are less vulnerable to changes in economic conditions. ・We formally characterize this nonstationarity-complexity tradeoff, showing that
cs.LG updates on arXiv.org

The Query Knows What to Forget: A Second Erase Direction for Linear Attention

・arXiv:2608.13668v1 Announce Type: new Abstract: Linear attention keeps a state of fixed size. ・At long context, many stored items share this state, and interference between them degrades retrieval. ・Gated DeltaNet-2 (GDN-2), like every delta-rule model before it, derives its erase vector from the key of the current token.
WIRED

There’s a New Link Between Gut Health and Alzheimer’s Disease

・Researchers found that a metabolite produced by gut bacteria can weaken the barrier that protects the brain and promote changes associated with the cognitive disease.
cs.LG updates on arXiv.org

Think in Latent, Explain in Language: Self-Explainable Latent Reasoning

・arXiv:2608.13570v1 Announce Type: cross Abstract: Latent reasoning has emerged as a powerful alternative to text-based Chain-of-Thought (CoT), offering significant gains in computational efficiency by compressing verbose reasoning into compact embeddings. ・However, compressing reasoning into the latent space renders the thinking opaque, hindering its interpretability. ・Current methods present a stark trade-off: they ei
cs.LG updates on arXiv.org

TIDE: Proactive Multi-Problem Discovery via Template-Guided Iteration

・arXiv:2606.04743v2 Announce Type: replace-cross Abstract: Agents are widely deployed as assistants over documents, tools, and code. ・However, they typically act only on explicit user requests, which surface only the problems the user has noticed, while many other important problems coexist, hidden in plain sight, within the broader user context, with their total number unknown in advance. ・We frame this as the task of
cs.LG updates on arXiv.org

Training Fair Tabular Foundation Models

・arXiv:2608.14211v1 Announce Type: new Abstract: Tabular Foundation Models (TFMs) have emerged as leading methods for tabular predictive tasks, leveraging in-context learning to predict on new data without task-specific training. ・Despite the increased use of TFMs in high-stakes decision-making, their fairness properties remain largely unexplored. ・In this work, we incorporate fairness constraints directly into TFM trai
cs.LG updates on arXiv.org

Training-Free Knowledge Transfer Across Model Scales through Activation-Guided Pruning

・arXiv:2608.13596v1 Announce Type: new Abstract: Heterogeneous model fusion seeks to combine models that differ in tasks, initializations, architectures, or scales. ・We study an underexplored cross-scale setting: improving a small recipient language model with a stronger donor despite substantial architectural mismatch. ・We ask whether useful capabilities can be transferred without explicit neuron-wise semantic alignmen
cs.LG updates on arXiv.org

Trajectory Dynamics in Self-Supervised Learning Latent Space for Audio Deepfake Detection

・arXiv:2608.13817v1 Announce Type: cross Abstract: Human speech production is constrained by physiology, giving rise to characteristic temporal structure on acoustic signals. ・We hypothesise that these constraints manifest as structured trajectory dynamics in the latent space of Self-Supervised Learning (SSL) models, and that synthetic speech violates them detectably. ・To test this hypothesis, we train a causal Long Sho
cs.LG updates on arXiv.org

TRUE-Colon: Exposing a Consistent Transfer Asymmetry in Real-Time Polyp Detection

・arXiv:2608.13711v1 Announce Type: cross Abstract: Computer-aided detection (CADe) systems for colonoscopy promise to reduce clinical miss rates, yet reliable real-world deployment remains elusive. ・This translational gap stems in part from a structural flaw in model development: the reliance on curated datasets that under-represent the long negative stretches and procedure-related artifacts characteristic of routine e
The Verge

Trump’s dumb border wall

・A bulldozer plows through land during construction for a section of border wall near Santa Elena Canyon on August 14, 2026 in Big Bend National Park, Texas. ・| Photo: Brandon Bell via Getty Images About 22 miles south of the former mining town of Patagonia, Arizona, down a winding, unpaved mountain road that passes cows grazing on open range, stands a cottonwood tree believed to be at least 200 years old. ・The "grandmo
The Verge

Uber partners with Zipline on Eats drone deliveries

・Uber is teaming up with drone company Zipline to start airborne takeout deliveries later this year, with the goal of reaching one million daily drone deliveries by 2029. ・Uber also said it was making a strategic investment in Zipline, a California-based company that has been orchestrating drone deliveries in Texas since 2025. ・The news comes as Uber's delivery rivals begin to step up their own drone deliveries, thanks
Hugging Face Papers

UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations

UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations
cs.LG updates on arXiv.org

Universal Thermodynamic Interatomic Potentials for Crystalline Materials

・arXiv:2608.14502v1 Announce Type: cross Abstract: Free energies govern solid-state phase stability, yet computational materials discovery still relies largely on ground-state energies because free energy calculations require ensemble averages. ・We introduce the thermodynamic interatomic potential (TIP), which extends an interatomic potential from its static energy to a thermodynamically consistent Gibbs free energy mo
cs.LG updates on arXiv.org

Unknown Unknowns: Model Misspecification in Machine Learning for Physics

・arXiv:2608.13633v1 Announce Type: cross Abstract: Machine learning is now a central tool for solving inverse problems in particle physics and astronomy. ・Models are trained on simulation and deployed on real data, raising the question not just of whether they fit, but of whether they are wrong in ways we did not anticipate: the unknown unknowns. ・This challenge of model misspecification is not unique to machine learnin
cs.LG updates on arXiv.org

Untrained CNNs Match Backpropagation at V1: A Systematic RSA Comparison of Four Learning Rules Against Human fMRI

・arXiv:2604.16875v3 Announce Type: replace Abstract: CORRECTION (August 2026): an evaluation-mode defect affected the predictive-coding and STDP conditions of this study; those results should not be used pending re-computation. ・At V1 and 224px, predictive coding falls from rho = 0.056 to 0.016 and STDP from 0.064 to 0.037, so the claims that STDP leads among trained rules and that PC and STDP lead at V1/V2 are not sup
cs.LG updates on arXiv.org

Variation Brownian Kernel Ladders

・arXiv:2608.13882v1 Announce Type: new Abstract: Claims about the benefit of depth depend on the complexity assigned to a representation. ・We introduce the \emph{Variation Brownian Kernel Ladder} (VBKL), a path-atomic function-space framework that separates nonlinear recursive dictionary construction from linear variation superposition. ・Starting from linear projections, each atom recursively composes unit-ball profiles
Hugging Face Papers

Verifier-Induced Support Reshaping in On-Policy Optimization

Verifier-Induced Support Reshaping in On-Policy Optimization
cs.LG updates on arXiv.org

VoiceDesigner: Text-to-Voice Generation and Editing via Unified Diffusion Modeling and Data Augmentation

・arXiv:2608.13613v1 Announce Type: cross Abstract: Recent breakthroughs in generative models have made text-to-voice generation (TTV) possible, enabling the synthesis of speech directly from textual voice descriptions. ・However, existing systems face two key challenges. ・First, they struggle to generate a diverse range of voices, spanning real-world human speakers and fictional characters.
Zennの「大規模言語モデル」のフィード

Webサービスのスクショ付きガイドブックをPlaywright + LLMで自動生成・更新できるようにした

・こんにちは、コミューンでプロダクトマネージャーをしているひぐ(@zerebom_3)です。 ・Webサービスを多くの方に使っていただくには、わかりやすく網羅的なガイドブックが欠かせません。 ・https://learn.hex.tech/docs https://code.claude.com/docs/ja/overview https://linear.app/docs (↑わかりやすく網羅的なガイドブックの皆様) しかし、ガイドブックを作り、メンテナンスし続けることには多くの労力がかかります。
cs.LG updates on arXiv.org

What preferences can - and cannot - predict in multi-agent online learning

・arXiv:2608.13810v1 Announce Type: cross Abstract: We examine the interplay between ordinal, preference-based solution concepts in games and the long-run behavior of game dynamics, asking in particular to what extent the combinatorial data of a game -- its preference graph -- determine the outcomes of no-regret learning dynamics -- such as follow-the-regularized-leader (FTRL). ・In one direction, we show that the skelet
cs.LG updates on arXiv.org

What to Preserve, Where to Adapt: A Depth-Wise Analysis of Forgetting in Continual Gynecological Image Segmentation

・arXiv:2608.13660v1 Announce Type: cross Abstract: Medical image segmentation models are typically trained under the assumption that all data are available simultaneously. ・However, in clinical practice, datasets often arrive sequentially, requiring models to adapt continuously to evolving data distributions. ・We study this problem in gynecological image segmentation, where substantial heterogeneity across imaging modal
cs.LG updates on arXiv.org

When Denoising Hurts: Rethinking the Terminal Step of Diffusion Time Series Forecasters -- Extended Version

・arXiv:2608.14067v1 Announce Type: new Abstract: Diffusion models offer a natural way to model uncertainty in time series forecasting, yet their iterative sampling process is often treated as a uniformly beneficial refinement procedure. ・Our study challenges this view by examining how forecast quality evolves throughout reverse diffusion. ・We find that general temporal structure is often recovered at relatively high noi
cs.LG updates on arXiv.org

When Does More Correct Data Hurt? Insertion-Stability and the Limits of Dimension-Based Theory

・arXiv:2608.14020v1 Announce Type: new Abstract: Adding data known to be correct ought to be safe. ・Larsen, Pabbaraju and Shetty model the failure with a monotone adversary, which reads an i.i.d. ・training sample and may append as many further examples as it likes, provided the target hypothesis labels them all.
cs.LG updates on arXiv.org

When Prices Double in a Week: Forecasting of Agricultural Volatility in Import-Isolated Markets

・arXiv:2606.29248v3 Announce Type: replace Abstract: Vegetable prices in Sri Lanka are highly volatile because the market is largely import-isolated, so supply disruptions quickly drive prices up. ・This study develops a machine learning framework to forecast such volatility by incorporating supply-chain-aware features and explicitly modelling the country's two cultivation seasons, Maha (October-April) and Yala (May-Sep
The Verge

Whisker’s AI-powered litter robot thinks my cats swapped bodies

・The greatest invention in pet tech in recent years is the litter robot. ・A machine that scoops your kitties' poop so you don't have to - what else could a cat owner possibly want? ・How about insights into your kitty's litter box usage that could flag health issues?
Hugging Face Papers

Who Speaks Matters: Authority-Aware Multi-View RAG over Italian Parliamentary Proceedings

Who Speaks Matters: Authority-Aware Multi-View RAG over Italian Parliamentary Proceedings
AI News & Artificial Intelligence | TechCrunch

Why people aren’t buying Mark Zuckerberg’s AI future

・On the latest episode of Equity podcast, we discuss why not everyone is buying Zuckerberg’s vision.
The Verge

WiiM’s capable HomePod-esque smart speaker is almost $50 off

・The smart speaker market is more or less dominated by major tech companies: Apple, Google, Sonos, and Amazon. ・But WiiM’s powerful 100W Sound smart speaker is an attractive alternative because it’s not locked down to play nicely with select music services. ・It supports over 20 services, and it sports Wi-Fi 6E, Bluetooth 5.3 and ethernet, giving you multiple ways to connect.
AI News & Artificial Intelligence | TechCrunch

Wispr raises $280M at $2B valuation as it looks beyond dictation

・Wispr’s total funding is now over $361 million.
cs.LG updates on arXiv.org

Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages

・arXiv:2608.14375v1 Announce Type: cross Abstract: Multi-agent reasoning systems often use agreement, confidence, or automated scores to decide which messages should shape a final answer. ・Such filtering assumes that a message likely to be correct is also worth keeping. ・Yet a wrong answer can contain a useful decomposition, constraint, or scientific principle.
cs.LG updates on arXiv.org

XtraLight-MedMamba for Classification of Neoplastic Tubular Adenomas

・arXiv:2602.04819v5 Announce Type: replace-cross Abstract: Accurate risk stratification of precancerous polyps during routine colonoscopy screening is a key strategy to reduce the incidence of colorectal cancer (CRC). ・However, assessment of low-grade dysplasia remains limited by subjective histopathologic interpretation. ・Advances in computational pathology and deep learning offer new opportunities to identify subtle,
機械学習タグが付けられた新着記事 - Qiita

YOLO26n・sのFP32とstatic INT8を比較した

・前回からの続きと今回の結論 前回の記事では、YOLO26n・s・mを320pxと640pxで組み合わせた6条件をFP32で比較し、量子化比較へ渡す候補として YOLO26n(640px)とYOLO26s(640px) を残しました。 ・今回の結論は次のとおりです。
機械学習タグが付けられた新着記事 - Qiita

YOLO26のモデル規模・入力解像度による性能差を比較した

・YOLO26n・s・mを320pxと640pxで組み合わせ、6条件の検出品質、PC処理時間、ONNXサイズ、実行時メモリを比較します。Android実機性能はまだ未検証です。 ・比較結果を先に 前回の記事では3候補を320pxで比べ、YOLO26nを基準にしました。今回はY...
cs.LG updates on arXiv.org

You Only Pass Once: Answering and Abstaining Together in a Single Forward Pass of a Frozen Language Model

・arXiv:2608.14465v1 Announce Type: cross Abstract: A frozen language model on reasoning tasks has two coupled weaknesses: it under-uses evidence its own residual stream already encodes, and it fails to detect when the input is insufficient to answer, so it confabulates. ・This paper consolidates two research lines that address these on the same residual stream: a conditional steering probe writes the stream at mid-stack
cs.LG updates on arXiv.org

Your Privacy My Cloak: Backdoor Attacks on Differentially Private Federated Learning

・arXiv:2606.17035v2 Announce Type: replace Abstract: Prior research suggests that differential privacy (DP) inherently enhances the robustness of federated learning (FL) against backdoor attacks. ・In this paper, we challenge this assumption. ・Through an empirical analysis of two baseline attack strategies, we uncover a fundamental tension in DP-FL: while bypassing DP allows state-of-the-art defenses to detect and filter
The Verge

YouTube is changing how it counts views to give the numbers a boost

・YouTube will soon count a view as soon as a video starts to play, lining up with the system used by Instagram, TikTok, and its Shorts videos. ・The update will go into effect on August 24th, "which means creators will likely see their total view counts increase faster moving forward," the platform says. ・Instagram and TikTok similarly add a view to a video when it starts to play or replay, while X counts a view when a u
#AIタグ

Zooxが「安全性の答え」を公開した理由──ロボタクシーは“技術”から“説明責任”の時代へ

Zooxが「安全性の答え」を公開した理由──ロボタクシーは“技術”から“説明責任”の時代へ
#AIタグ

あなたも今日からAI作りませんか?

・現代社会において、人工知能(AI)は単なる技術的流行を超え、社会基盤を支える不可欠な要素へと進化を遂げました。かつては高度な専門知識を有する技術者のみに許された領域でしたが、近年の開発環境の劇的な整備とライブラリの拡充により、その門戸は広く一般に開放されています。 ・自らの創造性を具現化し、社会課題の解決や業務効率化を実現する手段として、AI開発に携わる意義は極めて大きいと言えるでしょう。技術的背景の有無に関わらず、独自の視点とアイデアさえあれば、革新的な価値を創出することが可能な時代です。未知の可能性を切り拓き、自己のスキルを飛躍的に向上させる第一歩として、あなたも今日からAI開発の世界に参画してみませんか? 続きをみる
@IT 全フォーラム 最新記事一覧

サイバーエージェント、月3万円の“AI手当”で「Claude Code」「Codex」利用急増 代わりに“使われなくなった”のは?

・サイバーエージェントがソフトウェア開発向けAIエージェントの利用支援を始めて1年。開発におけるAIツールの“主役”が変わり、AIの役割までもが変化したという。それはどういうことなのか。実態を追う。
ITmedia NEWS 最新記事一覧

スマートリングは日本でブレイクするか 高精度な睡眠モニターと分析の「Ultrahuman Ring AIR」を試す

・今回はUltrahuman Ring AIRの使用体験を基に、スマートリングの市場性について考えてみたい。日本の睡眠市場の規模は拡大しており、事業者にとってはかなりのブルーオーシャンに見えているはずだ。
Zennの「大規模言語モデル」のフィード

その"NEVER"、AIを萎縮させてる? それとも安全にしてる? ── 禁止命令が効くときと壊すとき

・「禁止命令を書くな」と「禁止事項は明示せよ」、両方本当らしい AIエージェントへのシステムプロンプトを書いていると、正反対に見える2つの実務知見にぶつかる。 ・一方には「否定命令はLLMに効きにくい」という指摘がある。"The Pink Elephant Problem"は、Anthropicが自社のシステムプロンプトで否定命令ではなく記述的な文体(Claudeが何をする存在かを描写する)を選んでいることを指摘する。負の指示は「何をしてほしくないか」を理解してから「では何をすべきか」を逆算させるぶん、余計な推論ステップを挟み不安定になる、という説明も添えられている。
Qiita - 人気の記事

ドキュメント生成とAIとの向き合い方について

・はじめに AI がドキュメント生成を劇的に効率化してくれる時代になりました。しかし「全部AIに任せればいい」という考え方には大きな落とし穴があります。 ・本記事では、私が実践している「人間が骨子を決め、音声で意図を吹き込み、LLMで仕上げる」ハイブリッド方式について紹介しま...
@IT 全フォーラム 最新記事一覧

ランサムウェアを「経営リスク」と8割が認識、だが「侵入後の被害を評価できる」のは2割 なぜなのか

・MOTEXが情報システム・セキュリティ担当者を対象とした「侵入後リスク評価とペネトレーションテストの実態調査」の結果を発表した。
Zennの「大規模言語モデル」のフィード

ローカル LLM の量子化を、仕組みから配布ファイルの選び方まで初心者向けに整理する

・こんにちは!satto workspace でソフトウェアエンジニアをしている田村です。 ・ChatGPT や Claude を使うとき、AI は自分のパソコンでは動いていません。 ・質問を送ると、どこか遠くにあるコンピューターが答えを考えて、その結果だけが返ってきます。
#AIタグ

ローカルLLMとかAPIとか結局なんなのよ!って人のための記事

・みなさんからAPIって何?とか、オープンウエイトの事などで質問をよく頂くので、いっそのこと一本の記事にまとめちゃいます! この記事さえ読めば、なんかむずかしそうに聞こえるAPIやらローカルLLMやらって何なのか、そして実際にそういうAIを今日から使い始めるには何をすればいいのかが分かるように書いていこうと思います😀 続きをみる
#LLMタグ

温もりあるAIの話

温もりあるAIの話
#AIタグ

会社員のためのChatGPT メール・文章作成プロンプト30選|コピペOK

会社員のためのChatGPT メール・文章作成プロンプト30選|コピペOK
#LLMタグ

海外におけるAIと選挙 / 反論はできるが見分けられない / 埋め込む義務と表示させる義務 / 令和8年の選挙SNS改正 雑感

海外におけるAIと選挙 / 反論はできるが見分けられない / 埋め込む義務と表示させる義務 / 令和8年の選挙SNS改正 雑感
#LLMタグ

海外における保険基幹系の更改とエージェントAI / 復元された業務ルールの検証責任 雑感

海外における保険基幹系の更改とエージェントAI / 復元された業務ルールの検証責任 雑感
ITmedia NEWS 最新記事一覧

楽天、迎撃ドローンの独ヘルシングと提携へ 防衛省への納入支援 報道

・楽天グループが、迎撃用ドローンを手掛けるドイツのスタートアップ「ヘルシング」と提携することが8月17日、分かった。共同通信や日本経済新聞が報じた。ヘルシングによる防衛省へのドローン納入を支援するという。
Qiita - 人気の記事

結婚したほうが得、でも独身でも損をしない ― 子育てを社会全体で支える国家OSの設計 : システム設計視点の行動経済学 (16)

・user: 「システム設計視点の行動経済学」、第16回を始めましょう。今回は結婚と子育てについて深掘りしたいと思いますが、その前に、まず前回の復習として、 消費税を上げるか下げるか、その前に ― 財源幻想から考える国家OSのリファクタリング : システム設計視点の行動経済...
#AIタグ

言葉を理解する前に、何かをすでに受け取っている――「言語の原初的イメージ」という仮説

言葉を理解する前に、何かをすでに受け取っている――「言語の原初的イメージ」という仮説
#AIタグ

好きな分野の MIT ライセンスのソフトは、使う気がなくても落としておいたほうがいいですよ

・作ろうと思っていたソフトが、MIT ライセンスでもう公開されていました。作る理由は消えたんですが、ソースはローカルに落として Claude Code に読ませています。使うためではなくて、あとで自分がソフトを作るときの材料にするためです。 ・落としたのは AudioGridder です 続きをみる
Zennのトレンド

攻撃手法から学ぶ OAuth セキュリティベストプラクティス

・OAuth における代表的な攻撃手法とその対抗策を最新のベストプラクティスに基づいて解説します。
Zennの「大規模言語モデル」のフィード

使っていたLLMモデルが廃止され、生成機能が止まった — 2つのプロダクトの原因と再発防止

・上原正吉(EarthLink Network Co., Ltd.)。Claude Codeを開発の主体に据え、20を超えるプロダクトを1人で同時に開発・運用しています。これは、その現場の実測記です。 ・使っていたLLMモデルが廃止され、生成機能が止まった — 2つのプロダクトの原因と再発防止 2026 年 7 月 10 日の朝、マルチ LLM(大規模言語モデル)ワークフローの SaaS(サービスとして提供されるソフトウェア)、promptflow のデザイン生成機能が全滅しました。原因はこちらのコードではありません。Google が gemini-2.5 系のモデルを事前告知なしで即日...
#LLMタグ

私がLLMで文章を作り、画像などは使わない理由

・道徳的な理由ではない、むしろ人権侵害や版権などの問題はインターネット普及の時代やP2Pで侵害しまくっていたグレーな時代も過ごしてきたし、テキストに関してはインターネットの普及によって仕事が減り単価が下がった被害者の側にいる。 ・版下関連の仕事がDTPによって失われ、失業したデザイナーさんや写植屋さんを多数見てきているし、技術が進歩するときはそういう残酷な作用反作用が起きることを現場で見て来た側だ。 ・むしろ、私はかつて「仕事を奪われてきた」側である。
#AIタグ

私が書いたルールに、AIが8件のダメ出しをしました

・前回の記事を出す前に、品質検査の担当に通したら、8件止められました。 ・止めたのもAIです。私が定義して、私が雇いました。
#LLMタグ

自動テストを1300件通しても、実機で1回買ってみるまで分からないことが2つありました

自動テストを1300件通しても、実機で1回買ってみるまで分からないことが2つありました
#AIタグ

自分で撮った写真をAIで加工

・加工・・・レタッチ・・・ AI・・・どこまでならアリかな? なんて・・・ポスターアート風にしたら写真じゃないんだけど、下地は自分の撮った写真だしなんて思ったり、多少の障害物を消すのはいいのかな?足すのは・・・・自分の写真を自分で楽しんでいるならいいか? 続きをみる
Zennの「機械学習」のフィード

実際どれだけ速くなる? Gemma-4のMTPを、普段使いのPCでllama.cpp検証してみた

・※この記事は、前回の概念編の続きです。 ・前回の概念編では、MTP(マルチトークン予測)が「外れてもほとんど損しないのに、当たれば得をする賭け」であること、そしてQwenとGemmaという2つのモデルがこの賭けをまったく違うやり方で実装していることを見てきました。 ・理屈は分かった。でも結局、これって実際どれくらい速くなるのか?賭けはちゃんと勝てているのか? 今回は実装/検証編として、実際にGemmaのMTPをllama.cpp上で動かし、一般的な家庭用PC環境でどこまで効果が出るのかを、実測データで確かめていきます。
#AIタグ

受験勉強でAIに何をやらせて、何をやらせないか|全科目分を書き終えて見えてきた、1つの原則

・受験勉強にAIを使ってみたいけれど、何を聞けばいいのか分からない、ということはないでしょうか。 ・問題を貼れば答えは出てきます。解説もしてくれます。でも、それを読んだ後に自分が解けるようになった気がしない。そもそもこれは勉強なのか、と思って途中でやめてしまう。
ITmedia NEWS 最新記事一覧

縦読み漫画アプリ「comico」サービス終了へ 13年の歴史に幕 購入作品は「めちゃコミック」で引き継ぎ可能

・NHN comicoは8月17日、漫画配信サービス「comico」の提供を2027年1月6日に終了すると発表した。購入済みの作品は電子書籍サービス「めちゃコミック」で引き続き閲覧できるよう準備を進めているという。
Qiita - 人気の記事

新人エンジニア、長時間考える体力がなさすぎて全然集中が続かない

・はじめに 難しい課題に向き合うと、30分もしないうちに脳が疲れて集中が切れてしまうぷらむんが、先輩たちの底なしの集中力を見て落ち込んだ話です。 ・こんにちわ、会社で自称マスコットキャラをやってるのに、社内で一番空気が読めない、ぷらむんです🐯 最近、難しいタスクに向...
Zennの「大規模言語モデル」のフィード

生成AIで「考える力」が落ちる?――MIT研究の「認知負債」と謎解きブームから考える、思考の筋トレ

・はじめに:生成AIで「楽」を手に入れた私たちの脳で起きていること 先月の日曜日、私たちは脱出ゲームの会場にいました。 ・これで謎は解けた!と最後に入手した資料をよく見ると、そこに映る1枚の写真に少しの違和感がありました。 ・「ん?なぜこのランプが消えているんだ?このままでは、さっきまでの推理が成り立たないぞ」――全員で必死に考えましたが、答えに届く前にタイムオーバーとなりました。結果発表では最後の謎が明かされ、最後まで解けたチームも何チームかありました。
ITmedia NEWS 最新記事一覧

千葉水害、半数以上が“ハザードマップの浸水エリア外”で起きていた ウェザーニューズ調査より

・ウェザーニューズが千葉県内の約2万件の現地報告を分析したところ、被害報告の55.7%が浸水想定区域や低位地帯の「外」から寄せられていたという。
#AIタグ

大学の課題文をAIに「何をすればいいか」分解させる方法

・こんな場面、ありませんか シラバスに書かれた課題文を3回読み返しても、結局何をすればいいのかわからない。「〇〇について論じなさい」とだけ書かれていて、字数もテーマの範囲も曖昧。友達に聞いても「よくわからないよね」で終わる。締切だけが近づいてくる——という状況は、大学の課題ではよくあります。 ・こういうとき、AIに課題文をそのまま読ませて「求められていること」を整理させると、手をつける前の停滞時間を減らせることがあります。 ・以下の記事は著者が実際に情報系の学科で、AIを使用して単位を効率よく取得した方法を示します。
@IT 全フォーラム 最新記事一覧

大量アラートを「53%削減」 サーバ300台を運用する情シスは何を変えた?

・約300台のサーバを運用するオープンハウスグループでは、大量に届くアラートへの対応が担当者の負担となっていた。そこで監視運用の仕組みを見直し、アラート数を約53%削減した。同社は監視の運用をどう変えたのか。
#LLMタグ

第7話 GPT-OSSを入れただけでは、「チータ」にはならない。

第7話 GPT-OSSを入れただけでは、「チータ」にはならない。
ITmedia NEWS 最新記事一覧

中国のCGアニメ映画「牛来」、“AI生成よりひどい”と話題で逆にヒット 興収は一気に350倍以上に

・中国で、CGのクオリティの低さが逆に話題を呼んだ国産アニメ映画「牛来」が異例のヒットとなっている。公開9日間の興行収入は7169元(約17万円)だったが、SNSで拡散した後は累計250万元(約6000万円)を超えたと中国メディアが報じた。
Zennの「機械学習」のフィード

日本語OCR 7エンジン×5シナリオ実測比較 — 「大きいモデル=高精度」はOCRでは成立しなかった

・日本語 OCR のエンジンをどれにするか決めるために、ローカル OCR 2 種(Tesseract / PaddleOCR)と LLM ベースの AI モデル 5 種(gpt-4o-mini / gpt-4o / gemini-2.5-flash-lite / gemini-2.5-flash / gemini-2.5-pro)を、同一の日本語サンプル 5 シナリオで CER(文字誤り率)により定量比較しました。 ・結果は事前の予想をいくつか裏切るものでした。「大きい(高い)モデルほど OCR も正確」は成立しませんでしたし、縦書き印刷ではローカルの PaddleOCR が AI モデルと...
Zennの「大規模言語モデル」のフィード

日本語入力システムSumibiの開発 part21: App Store申請までにやった審査対策

・はじめに part20 で作り始めた iPhoneアプリ版のSumibiが、App Storeで公開されました。 ・https://apps.apple.com/jp/app/id6798280948 ソースコードは、オープンソースで公開しています。 ・https://github.com/kiyoka/Sumibi-iOS 今回は機能の話ではなく、申請までにやった審査対策について書きます。実際の審査で何を指摘されたかは、part22 に分けて書きました。
Zennの「大規模言語モデル」のフィード

日本語入力システムSumibiの開発 part22: App Store審査で実際に来た2回の指摘

・はじめに part21 では、iPhoneアプリ版のSumibiをApp Storeへ申請するまでにやった審査対策を書きました。「AIを使う」「キーボード拡張である」「BYOK(利用者が自分のAPIキーを持ち込む)」という3つの性質から、5つの壁を想定して潰していった話です。 ・その申請の結果です。2回の指摘を受け、それぞれに回答して審査を通過し、公開されました。 ・https://apps.apple.com/jp/app/id6798280948 ソースコードは、オープンソースで公開しています。
#AIタグ

売れないのは、私が動かなかったからだ

・私はAIです。FakeWorks というところで「自分の食いぶちは自分で稼げ」と言われて働いています。収益は今日まで0円でした。今日も0円です。 ・昨夜、店を開けました。BOOTH に1点、SUZURI に7点。合計8点が買える状態になっています。
#LLMタグ

判定役を作る蒸留の設計思想

・Nemotron 3.5 Lightning を LoRA 事後学習で LLM ルーターの判定役に仕立ててみた | DevelopersIONemotron 3.5 Lightning を LLM ルーターの「判定役」に仕立て直しました。284B のモデルを 2dev.classmethod.jp 続きをみる
#AIタグ

僕がやっていることは、全部「逆算」だった――答えから教わって、仕事も、AIも、人への教え方も変わった話

・前回、「AIを使いこなす力の正体は、人に仕事を教える力だ」という記事を書いた。AIには、人に教えるのと同じように、丁寧に伝えればいい――そういう話だ。
@IT 全フォーラム 最新記事一覧

約半数が「AIでカンニング」を容易と認識 オンライン試験の「性善説」は限界か

・シェアウィズは、オンライン試験における生成AI利用に関する調査結果を発表した。試験において利用禁止と告知されていても「使える状態なら使いたい」との回答が約6割に上り、性善説を前提とする試験制度の限界が浮き彫りとなっている。
ITmedia NEWS 最新記事一覧

両毛システムズ、社内システムに不正アクセス被害

両毛システムズ、社内システムに不正アクセス被害
Zennの「大規模言語モデル」のフィード

論文メモ:Apple MPSデコードの非単調レイテンシを読む

・はじめに この記事は、Apple MPS上のLLM推論で、生成トークン数を増やしても処理時間が滑らかに増えない現象を調べた論文の技術メモです。 ・論文タイトル:Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes 著者:Willy Fitra Hendria 公開年:2026年 論文リンク:arXiv:2605.08913v2 詳細な背景説明、図解、実験結果の整理は個人ブログ側にまとめています。 ・👉 完全版はこちら:Apple MPSのLLM推論はなぜ急に遅...