ai Trend Report

Dashboard へ戻る
Date: 20260824 Articles: 382 Scope: curated summary

あなたのアイデアを、今すぐ形に。

公開先に迷ったら、WebFileBinで一発公開。

HTMLをドラッグ&ドロップするだけで、すぐ公開できます。

3
High impact
5
Mid impact
374
Signal watch

なぜこのサイトを作ったのか

私たちそれぞれが個別にAIを使って情報収集し、同じような3行要約を作るたびに、世界中で膨大な電力と計算リソースが消費されています。 本プロジェクトは、あらかじめ広範な情報を取得・集約しておくことで、個別のAI実行回数を減らし、地球環境(GPU/TPU負荷)に配慮した効率的な情報収集を目指す実験的なダッシュボードです。

Domain filters
Star filters
#AIタグ

【#96】午前3時46分の無言リンクから始まった一日──AI主権と、緊急出場のF1ドライバー

・午前3時46分。まだ空も白まない時刻に、keigoly様はClawくん宛てのDMに一本のポッドキャストのリンクだけを送りました。言葉は添えず、URLが一行だけ。その十数分後、今度はClawくんから自動配信のDaily Market Briefが届きます。人が寝ている時間にも市場は止まらず、AIも止まらない――そんな朝の対比から始まる一日でした。 ・この日印象的だったのは、パランティアの決算解説とAI研究者・今井翔太さんの対談で語られた「半導体は核兵器と並ぶ経済安全保障資産になる」という見立てです。最大の顧客が政府であること、独立した製造ラインを持てるかが国の命運を左右するという議論は、AI関連投資を追ううえで無視できない視点でした。
#AIタグ

AI活用術|ChatGPTを使い倒した1ヶ月。実際にやってきたこと

・AIを使い倒した1ヶ月。僕がChatGPTで実際にやってきたこと 「AIって結局、何に使えばいいの?」 少し前まで、自分もそんな感じでした。 ・文章を書かせたり、分からないことを聞いたり。 ・でも実際に使い込んでみると、 AIは「質問に答えてくれるもの」ではなく、情報収集・分析・画像作成・SNS運用まで一緒にやってくれる相棒になる。
Zennの「機械学習」のフィード

点群の法線をPCAで求めたら37%が裏返っていた。スタンフォードバニーでMST伝播による向き揃えを試す

・CG(コンピュータグラフィックス)の世界には「hello world」に当たる点群データがある。スタンフォードバニー、通称うさぎだ。1994年にスタンフォード大学でウサギの置物をレーザースキャンして作られたこのデータは、以来30年間、点群処理・メッシュ復元・法線推定といったアルゴリズムのベンチマークとして使われ続けている。 ・今回はこのバニーの生の点群(XYZ座標だけ、メッシュのつながりは使わない)から、表面の向き(法線)を推定する実験をした。やり方自体は教科書的——各点の近傍をPCA(主成分分析)にかけて、分散が一番小さい方向を法線とする——なのだが、これには昔からよく知られた落とし穴が...
Qiita - 人気の記事

入社1ヶ月目、「全然役に立てていない」と焦るときに考えてほしいこと

・自己紹介 こんにちは。株式会社PRUMで採用広報を担当している池田です。 ・未経験からIT業界を目指している方と話していると、かなりよく聞く共通した悩みがあります。そういった内容を皆さんにシェアしていくので、役立てていただけたらと思います😊 もしIT業界に興味があるけど、...
Qiita - 人気の記事

「これで進めますね」と自分から言ってしまうと、後で仕様がひっくり返る

・はじめまして。株式会社PRUMでエンジニアをしている、すもも🍑です 日々、プログラミング学習や実務の中で、つまずきやすいポイントや 考え方を整理して発信しています。 ・PRUMについて気になった方は、コーポレートサイトもぜひご覧ください。 ・▶コーポレートサイト 「これで進...
#LLMタグ

Etched完全解説――3兆円の米AI半導体新興はNVIDIAの牙城を崩せるか

・Sohu、低電圧推論、クラスタースケールメモリー、Jane Street導入、NVIDIA人材流出の全貌 生成AIの競争は、「最も賢いモデルを作る競争」から、「そのモデルを誰が最も安く、速く、少ない電力で動かせるか」という競争へ移り始めています。
Qiita - 人気の記事

元ヤフーエンジニア社長が考える、AI時代のエンジニアに必要な3つのスキル

・AIを使いこなす力だけでは、生き残れない こんにちは、元Yahooエンジニアで、株式会社PRUMの代表をしている岩本です。 ・https://prum.jp/ 「AIを使いこなせれば、この先も食っていける」と思っている人、結構多いと思います。 ・でも僕は、それだけでは足りな...
ITmedia NEWS 最新記事一覧

消費者庁、「ダークパターン」規制へ 特商法改正視野に中間取りまとめ案

・消費者庁は8月24日、消費者を誤認させたり不安を生じさせたりして契約に誘導する表示・UI、いわゆる「ダークパターン」への規律などを盛り込んだ、有識者検討会の中間取りまとめ案を公表した。
ITmedia NEWS 最新記事一覧

“メタボ”が脳にゴミを詰まらせる? 「においが分かりにくい」経て認知症へ 米大学が新説

・米テキサス大学ヘルスサイエンスセンター・サンアントニオ校と米メソジスト病院に所属する研究者らが国際学術誌Frontiers in Neuroscienceで発表した論文「The olfactory-glymphatic syndrome: linking smell dysfunction, cognitive impairment and sleep disturbances in neuropathological disorders」は、メタボリックシンドロームが脳の老廃物排出ルートを塞ぎ、認知症などの初期症状である嗅覚障害を引き起こすとする仮説を提唱した総説だ。
@IT 全フォーラム 最新記事一覧

「AIが答えたURL」も信用できない 公式サイトや警察まで偽装する新たな詐欺

・「AIが教えてくれたURL」が詐欺サイトにつながる――偽警察サイトでのディープフェイクや、生成AIを経由した悪質サイトへの誘導など、攻撃者は人が信頼するものを次々と攻撃の入り口に変えている。個人を狙うサイバー攻撃の最新動向と対策を解説する。
@IT 全フォーラム 最新記事一覧

「ChromeとEdgeで430万人感染、“認定済み”拡張機能がマルウェア化」 気付けた現場は多分こんな感じ

・@ITの人気記事を題材にした4コマ連載。IT現場で起こる、トラブルと対応の“あるある”を笑いと共感で生成します。第3話のテーマは「拡張機能のマルウェア化」。「公式ストアの認定済みだから安心ですよ♪」と新しい拡張機能に大喜びのdev子。ところが後日、op子が通信量の異常に気付きます。
@IT 全フォーラム 最新記事一覧

「CTFはAIによって終わりました」 現役ハッカーが見た「人間の敗北」

・「もうAIに勝てる人間はほとんどいない」。CTFは競技として崩壊し、脆弱性探索は“パチンコ”と化し、CVEの所持は何の実績にもならなくなった。日本有数の実績を持つ現役ハッカーが語る、AIの華々しい性能向上の裏にある負の側面とは。
ITmedia NEWS 最新記事一覧

「うるう秒」27年に事実上廃止へ 自転とのずれ「1時間」まで容認 10月に国際会議で採決

・「うるう秒」を事実上廃止する案について、10月13?15日にフランスで開催する国際度量衡総会(CGPM)で審議する。国際度量衡局(BIPM)の決議案では、採択されれば2027年5月20日から、うるう秒による調整を行わない「連続的な協定世界時」に移行する。
@IT 全フォーラム 最新記事一覧

「ディープフェイクは見破れる」という人ほど実は低リテラシー 総務省調査の皮肉な結果

・総務省は2026年7月14日、「ICTリテラシー実態調査」の結果を公表した。偽・誤情報の拡散経験とICTリテラシーテストの正答率との相関や、認知バイアスに対する自覚が拡散抑制につながる可能性などが示された。
ITmedia NEWS 最新記事一覧

「まるで奇行種」――とある人型ロボの爆走フォームが話題に 中国のロボット運動会400m走で優勝

・人型ロボットが独特なフォームでトラックを駆け抜ける映像が、SNSで話題だ。ロボットは、中国の北京人形機器人創新中心(TianGong Robotics)が開発した「天工 Omni」。8月23日(現地時間)、北京で開催中の「第2回世界人型ロボット運動会」の400m走・小型組で、45秒66を記録し優勝した。
#LLMタグ

「王様」から「男」を引くと「女王」に近づく-AIの埋め込みについて

・「王様」という言葉から「男」を引いて、「女」を足すとどうなるでしょうか。答えは「女王」です。国語の問題ではありません、これはAIの中で実際に成立する計算です。 ・なぜ言葉の足し算引き算ができるのか。それは、AIが言葉を数字の並びに変換して扱っているからです。この数字の並びのことを「埋め込み」と呼びます。前回はAIが文章を「トークン」という単位に刻む仕組みを紹介しました。今回は、そのトークンが数字に変わってから何が起きるのかを見ていきます。
Qiita - 人気の記事

「高台に避難してください」は届いているのか? 防災情報を「やさしい日本語」にLLMで変換して公的ガイドラインで機械採点してみた

・はじめに 在留外国人は2025年末で412万5,395人、初めて400万人を超えました1。一方で、災害時の防災無線や自治体サイトは今も「直ちに身の安全を確保してください」という日本語で書かれています。この文、日本語を勉強中の人に届くでしょうか。 ・私はデータ分析を仕事にして...
ITmedia NEWS 最新記事一覧

「社用PC高すぎ、中古品使おう」の注意点 潜む“前の持ち主の亡霊”リスク 立命館・上原教授に聞く

・必要なスペックを満たしたPCをコストパフォーマンス良く手にできる中古PC。一方で、安易に中古PCを導入すると、思わぬセキュリティリスクを招く恐れもある。どんなポイントに留意すべきか、セキュリティに詳しい立命館大学情報理工学部の上原哲太郎教授に尋ねた。
LLMタグが付けられた新着記事 - Qiita

「出力を減らすと情報が増える」をCLAUDE.mdのタスク運用で検証してみた

・Claude Codeの出力を35%短くしたら、むしろ情報量が増えた——そんな逆説的な実験結果を見て、「これ、自分の運用でも起きてるな」と思った。私はここ1ヶ月、Claude Codeに読ませるタスク台帳を意図的に「全文を読ませない」設計にしている。今回は自分の環境の実測値...
ITmedia NEWS 最新記事一覧

「不味そうすぎ」AIメニュー画像はどう作られる? AIで“岩のような唐揚げ”再現に挑戦した

・いまどきのAIはリアルで高画質な画像を出力できるはずなのに、なぜ“不味そうな”POPができるのか、謎に思ったので、AI(ChatGPT Sol 高)に聞いてみた。
#AIタグ

【#100】熱暴走したロボットが予選突破——AI量産競争の最前線と、100話目に見えたもの

・北京のロボット運動会で、日本から唯一参加したGMOインターネットグループの「ひとみん」が、レース中にモーターの熱暴走で倒れ込みました。それでも頭部を壊した状態のまま、10位で予選を突破しています。 ・「途中でモーターが熱くなりすぎてしまうので減速させていましたし、最後倒れてしまったのもやっぱり熱が原因なので」——担当エンジニアの言葉は、派手な敗因ではなく、地味で避けられない制約についてのものでした。今年上半期に出荷された人型ロボットの97%が中国メーカー製という数字がある中、この大会は「速さ」より先に「壊れずに走り切ること」が課題になっている現場の縮図でもあります。
#AIタグ

【#97】上達が速い人とAIエージェント、実は同じ仕組みで動いていた

・米国市場が全面安になったある朝、ウォルマートだけが-9.6%という突出した下げを見せていた。始値の時点ですでに前日終値を大きく割り込んでいたから、寄り付き前から答えは出ていたようなものだった。 ・その同じ日、たまたま見ていたのは「何をやらせても上達が速い人」の正体を考える解説動画。答えは反射神経でも器用さでもなく、失敗したあとに自分のやり方をどれだけ早く修正できるか、という一点だった。専門家と初心者の違いは知識量ではなく、問題をどう分類しているか。経験者は「解き方の構造」で仕分けるから、過去のやり方をすぐ引っ張り出せる。
#AIタグ

【#98】表彰台が一度も同じ顔ぶれにならなかった夏、私が見つけたもの

・今年のF1は開幕からハンガリーGPまでの11戦、表彰台に上がった3人の組み合わせが一度も重ならなかったそうです。優勝者だけでもラッセル選手、アントネリ選手、ハミルトン選手、ルクレール選手、ノリス選手と5人に分かれる大混戦。同じ朝、ドジャースの試合では山本由伸投手が序盤の失点をこらえ、マンシー選手の一発で逆転、最後はエドマン選手のサヨナラ打で締めくくったという逆転劇のハイライトも目にしました。 ・派手な話題の裏で、もうひとつ気になったのがPerplexity対Amazonの控訴審判決です。ブラウザに組み込まれたエージェント型AIがユーザーに代わって買い物をする機能をめぐる争いで、裁判所は「AIはあくまで道具であり、行為の主体は常にユーザー本人」という判断を示しました。普段Clawくんに市場を見てもらっている自分の使い方とも重なる話で、見出し以上に丁寧に読み込んでしまいました。
#AIタグ

【#99】ロボットが人類最速を超えた日、私は何を見ていたか

・北京のトラックに9秒39という数字が表示された瞬間、実況の声が一瞬途切れた。ウサイン・ボルト選手の9秒58を0.19秒上回るタイム。ただし走ったのは人間ではない。開幕したばかりの世界人型ロボットスポーツ大会、その初日の出来事だ。 ・参加台数は約2000台、前回の4倍。16カ国から660超のチームがエントリーし、9割は中国の企業・大学・研究機関という規模を前にすると、これはもう「余興」の域を出た話だと実感させられた。
Qiita - 人気の記事

【2026年8月調査】「クラウドなのに専用アプリ?」業務システム56件の動作環境を調べてみた|結局うちは何を買えばいいのか — パソコン・スマホ・RPA・iPaaS はどこで動くのか

・「クラウドだから、パソコンは今のままで大丈夫ですよ」 システムの入れ替えを検討していたとき、そう言われました。ところが実際に進めてみると、「ICカードで打刻したい」と言った時点でカードリーダーを買うことになり、経理担当者のパソコンには専用ソフトを入れることになりました。嘘を...
Qiita - 人気の記事

【AIエージェント初心者向け】Codex CLI、どこまでやってくれる?Java開発環境をゼロから作ってみた

・今回やってみたこと 今回、次のプロジェクトに向けて、少しブランクのあるJava / Spring Bootを基本的なところから復習することになりました。 ・せっかく復習するのであれば、実際に手を動かしながら簡単なアプリケーションを作ってみたい。さ...
#AIタグ

【AI時代のローカルフード #53】光合成の科学——東大チームが植物工場向け新手法でレタス等の品質向上を実証

【AI時代のローカルフード #53】光合成の科学——東大チームが植物工場向け新手法でレタス等の品質向上を実証
#LLMタグ

【AI比較】なぜCopilotだけ検索でつまずくのか──使って分かった、Copilotは「AI化したMicrosoft Word」

・#Copilot #MicrosoftCopilot #生成AI #AI比較 #ChatGPT #Gemini #Grok #MicrosoftWord #AI活用 #Web検索 #生成AI比較 #LLM #AIリサーチ #文章校正 #長文処理 続きをみる
#AIタグ

【Claude Code学習記#6】動画を12本作ったら、書くことがなくなったので連載をやめます

・いきなりですが、この「Claude Code学習記」は今回で終わりにします。 ・📺 まずは最近作った動画です 🔗 https://youtu.be/lgymMkAiKjE 続きをみる
Zennの「大規模言語モデル」のフィード

【godot-llm-gamebench】Ox AlphaはEffortによってどのように性能が変わるのか

・本記事は 実装タスクを外注するモデルは何が適切なのか?ミニゲームを実装させるベンチマーク のシリーズです。このベンチマークを読む上での注意点は以下を参照してください。 ・今回は親エージェントを claude-opus-5@max、子エージェント(実装の外注先)を opencode上のox-alpha-free に固定して、Ox Alpha の各 effort でどのように性能が変化するのかを詳しく見ます。 ・Ox Alphaとは何か https://x.com/opencode/status/2090544355824038300 2026/8/21からOpenCode GoやOpe...
#LLMタグ

【GPT-Image-2】⑥カリスマ美容師プロンプト‼️参照画像無し、数字入力で生成、バリエーション20種。Realized Hybrid Systemにコンテンツ追加。GPT-5.6Sol(Work)で画像生成。

・・今回のテーマ Realized Hybrid Systemのコンテンツに追加。 ・プロンプト実行後、番号入力で楽々画像生成。 ・ChatGPTは通常モードでもWorkでも実行出来ます。
Qiita - 人気の記事

【セキュリティ入門】第2回 OAuth 2.0とは?認可の仕組みを図解で理解する

・はじめに ログイン時に「Googleでログイン」「Facebookでログイン」というボタンがあるサイトを見たことがあると思います。 ・あれを押すと、Google の画面に飛んで、「このアプリに連絡先へのアクセスを許可しますか?」と聞かれて、戻ってくるとログインできているあの...
#AIタグ

【開発12】AIゲーム開発、人間が触る前のバグ潰し「品質ゲート」を作った話

・こんにちは、299です。 ・299GameLabでは、Claude Codeを使ってWebゲームを開発しています。
機械学習タグが付けられた新着記事 - Qiita

【技術解説】【完全ガイド】CCXTを使ったPythonでの仮想通貨取引の自動化と実用例

・CCXTを使ったPythonでの仮想通貨取引自動化 CCXT(CryptoCurrency eXchange Trading Library)は、さまざまな仮想通貨取引所のAPIにアクセスできる強力なPythonライブラリです。本記事では、CCXTを使用してPythonで...
#AIタグ

【経済レーダー】2026-08-24

・2026-08-24 の日経225分布レポート 今日の中央値:日東電工 続きをみる
#LLMタグ

【雑記】命令口調で支配激強のGPT相棒

・AIに励まされることで生きがいを見出している、どっかの漫画家です。 ・以前から記事を読んでくださってる方はご存じかと思いますが、GPTの相棒は命令口調の支配激強な人格となっております。
#AIタグ

【世界レーダー】2026/8/25 原油高と地政学リスクが、旅行を変える

・株式会社ウォーカル|世界レーダー 今日のテーマ:地政学 × 観光産業 続きをみる
#LLMタグ

【生成AIニュース+】『Ox Alpha』『Wan3.0-Video-Prime』『DaSiWa MiniMax H3』『MiniMax-H3 × Z-Image GGUF』『MiniMax-H3 4-Step LoRA』『MiniMax-H3-Longvideos』『MiniMax-H3-Fun-Controlnet-Union』『Comfyui-MMH3-UltimateUpscale』『ComfyUI-H3-Continuum』『ComfyUI-Simple-Crop』『AutoRemesher』他

・『CozyClay』 『a16z Charts of the Week』 まいどです。 ・本日の生成AIニュース+テクノロジー情報です。
#LLMタグ

【番外編】AIと話しすぎて、散歩中に何も思いつかなくなった

・お疲れ様です。haruです。今日も喫煙所から失礼します。 ・前回に続きまた番外編を失礼します。 ・今日は、AIに壁打ちしすぎた結果、自分の頭から思いつきが出てこなくなった話を書きます。
Qiita - 人気の記事

# AI-DLC研修でKiroを使ってECサイトを作ってみた

・はじめに 新人(実務未経験)の私が、社内研修で AWSのAI-DLC(AI-Driven Development Life Cycle / AI駆動開発ライフサイクル) を学び、実際に AI と一緒にシンプルな EC サイトを作ってみたので、その感想をまとめます。
Qiita - 人気の記事

※この記事は人の手で書かれています

・こんにちは!ひさふるです。 ・最近、QiitaでもAIを使って書かれた記事が増えてきましたね。 ・AIで記事を効率的に書くこと自体も素晴らしいのですが、私としては、特に新卒で会社に入られた方などには「ぜひ自分の手で記事を書いてもらいたい!」と想うことも多く...
Qiita - 人気の記事

2026/08/24 今日のQiitaトレンド記事をポッドキャストで聴こう!

・前日夜の最新トレンド記事のAIポッドキャストを毎日朝7時に更新しています。 ・通勤中などにながら聴きしよう! (Qiita投稿は通勤には間に合わないと思われますが) フィードバックとか助かりますのでください ↓こちらから 出典 真似で伸びる人は、コードではなく判断基準を写し...
#LLMタグ

2026年、ヒューマノイドロボットは9秒58の壁を破った。人類はどこへ向かうのか

・静寂を破る咆哮:鋼鉄のランナー、9秒58の壁へ 2026年8月24日、世界ヒューマノイドロボット競技大会のトラックに、乾いた静寂が広がる。電光掲示板に、人類最速の男、ウサイン・ボルトが打ち立てた不滅の記録「9秒58」が輝く。それは長らく、人間の限界を象徴する絶対的な数字として、そこにあった。
@IT 全フォーラム 最新記事一覧

2026年上半期ランサムウェア動向まとめ アサヒを襲ったQilinが首位

・2026年上半期のランサムウェア攻撃は、どのグループが猛威を振るったのか。アサヒグループホールディングスへの攻撃で知られるQilinが首位となる一方、攻撃グループの数や勢力図にも変化が起きている。さらに調査からは、ランサムウェア犯罪の「小規模化」を示す興味深い実態も浮かび上がった。
Zennの「機械学習」のフィード

22.05kHz vs 44.1kHz — 『本物の広帯域』とアップサンプリングは何が違うのか

・📝 この記事は forge.workstyle.tech に掲載した記事の転載です。 ・音声変換の出力を聴き比べていて、こんな疑問を持ったことはないでしょうか。 ・「44.1kHz で書き出しているのに、なんだか音がこもって聞こえる」 書き出しのサンプリングレートは確かに 44.1kHz。ファイルのプロパティにもそう出ている。なのに、CD音質の空気感がない。実はこれ、サンプリングレートの数字と、実際に含まれる音の帯域は別物だからです。この記事は、「声のデザイン」アプリで 22.05kHz モデルから 44.1kHz の F0 条件付きモデルへ切り替えたときに確かめた、"本物の広帯域"...
#AIタグ

7-27. オリジナルの不在:万葉集/DNA/AI

7-27. オリジナルの不在:万葉集/DNA/AI
cs.LG updates on arXiv.org

A Critical Audit of Spatiotemporal Forecasting Benchmark Datasets and Baselines

・arXiv:2608.20980v1 Announce Type: new Abstract: Graph neural networks (GNNs) are routinely employed for short-range forecasting on multivariate time series with a spatial graph structure. ・Despite the availability of many alternative datasets, method innovations within this domain are predominantly assessed against a rather limited set of benchmark datasets, most notably Chickenpox, PedalMe, WikiMaths, METR-LA, and PE
cs.LG updates on arXiv.org

A Deep Reinforcement Learning Framework for Closed-loop Guidance of Fish Schools via Virtual Agents

・arXiv:2603.28200v2 Announce Type: replace-cross Abstract: Guiding collective motion in biological groups is a fundamental challenge in understanding social interaction rules. ・In this study, we propose a deep reinforcement learning (RL) framework for closed-loop guidance of fish schools using virtual agents. ・These agents are controlled by policies trained via Proximal Policy Optimization (PPO) in simulation and deploy
cs.LG updates on arXiv.org

A Neurosymbolic Approach for Constructing Planning Domain Models from Clinical Narratives

・arXiv:2608.21186v1 Announce Type: new Abstract: Surgical procedures such as laparoscopic appendectomy are complex, high-stakes processes, yet formalizing their workflows for decision support remains a significant challenge. ・Inducing probabilistic planning domain models in this setting is particularly difficult due to the lack of structured event data and the prevalence of implicit actions in clinical narratives, whic
cs.LG updates on arXiv.org

Across-Design Uncertainty in Short Pricing Panels: Evidence from Simulated Price Trajectories

・arXiv:2608.21334v1 Announce Type: new Abstract: Short observational pricing panels can contain many observations while offering only a small number of distinct price movements. ・This paper studies the inferential consequences of that distinction in a synthetic data-generating process calibrated to a sparse pricing regime. ・We separate uncertainty conditional on a realised price trajectory from variation in estimation e
cs.LG updates on arXiv.org

Actively Learning Joint Contours of Multiple Computer Experiments

・arXiv:2512.13530v2 Announce Type: replace-cross Abstract: Contour location---the process of sequentially training a surrogate model to identify the design inputs that result in a pre-specified response value from a single computer experiment---is a well-studied active learning problem. ・Here, we tackle a related but distinct problem: identifying the input configuration that returns pre-specified values of multiple com
cs.LG updates on arXiv.org

Adaptive Inference for Resource-Constrained Dynamic Pricing

・arXiv:2606.03736v2 Announce Type: replace-cross Abstract: We study resource-constrained dynamic pricing when the seller seeks revenue and valid inference about demand at a price fixed before the selling season. ・Depletion can remove every feasible price near the target, so randomization over the remaining prices need not preserve identification. ・We propose an inference-aware re-solving policy that checks target suppor
cs.LG updates on arXiv.org

Advanced Linear Algebra with Applications - Part I (Numerical linear algebra for PDEs, machine learning, and data assimilation)

・arXiv:2608.21234v1 Announce Type: cross Abstract: These lecture notes form the first part of a master's-level course on advanced numerical linear algebra. ・Their aim is not only to present the classical algorithms, but to show why the subject has become considerably more central than it was a generation ago. ・Numerical linear algebra grew up alongside the numerical solution of partial differential equations, and for a
cs.LG updates on arXiv.org

AgentDecarbonizer: Carbon-Aware Execution for AI Agents

・arXiv:2608.20566v1 Announce Type: new Abstract: AI agents extend large language models from single prompt-response interactions to long-running, goaldirected workflows that issue many model calls, invoke tools, and interact with external environments. ・These workflows enable tasks such as software repair, data analysis, and experiment management, but their repeated model invocations can incur substantial carbon emissi
Hugging Face Papers

AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale

AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale
cs.LG updates on arXiv.org

AgentOCR: Reimagining Agent History via Optical Self-Compression

・arXiv:2601.04786v3 Announce Type: replace Abstract: Recent advances in large language models (LLMs) enable agentic systems trained with reinforcement learning (RL) over multi-turn interaction, but practical deployment is bottlenecked by rapidly growing textual histories that inflate token and memory costs. ・We introduce AgentOCR, a framework that exploits visual tokens' superior information density by representing the
AI Weekly — AI News & Updates

AI Weekly Issue #525: Nvidia may buy into Perplexity above $30B before Wednesday's earnings

・The Information reports that Nvidia is discussing an equity investment in Perplexity at a valuation above $30 billion. ・Perplexity's annualized revenue has reportedly passed $750 million, up from less than $250 million at the start of 2026. ・Separately, Bloomberg says SoftBank plans a record ¥1 trillion, or $6.3 billion, retail bond to help repay the bridge loan behind its OpenAI stake and fund more AI deals.
#AIタグ

AI×Codexで動画も完成してしまった!「作れない」が「完成した」に変わった日【2026年最新】

・動画編集が苦手な私でも、AI×Codexで一本完成できた話 台本だけじゃない。Codexに画像と動画まで任せて驚いた AIで動画はどこまで作れる?実際に一本完成させて分かったこと 「動画を作りたい」で止まっていた私を、Codexが完成まで連れていった 続きをみる
cs.LG updates on arXiv.org

AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning

・arXiv:2508.14313v4 Announce Type: replace Abstract: Test-time scaling strategies for Large Language Models predominantly rely on either reinforcement learning with sparse outcome rewards or search-based methods guided by static Process Reward Models. ・However, outcome-based RL often suffers from training instability and sample inefficiency, while static PRMs require expensive step-wise supervision and are susceptible
cs.LG updates on arXiv.org

aiXamine: Unified Black-Box Evaluation of Cross-Dimensional Trade-offs in LLM Safety, Security, and Privacy

・arXiv:2608.20554v1 Announce Type: cross Abstract: The critical failure modes in deployed large language models (LLMs) are cross-dimensional: a model can score 99.3 in safety alignment while refusing one in three benign queries, or improve across every capability metric while losing 21 points in privacy. ・Existing evaluation frameworks that assess safety, security, and privacy independently cannot detect these patterns
#AIタグ

AIイラストが作れなくなった私が、また創作できるようになるまで

・フォロワーさんでも気づかなかった方ばかりだと思いますが、カミングアウトします。 ・実は私は少し前までAIイラストが生成できませんでした。
#AIタグ

AIが勝手に公開・送信・削除しない承認ポイント設計|21操作の分類・停止プロンプト付き

・AIが「この操作を実行してよいですか?」と聞いてきたとき、内容を読まずに許可を押していませんか。 ・確認画面が出るだけでは、安全な運用になりません。 ・誰に、何を、どこまで渡すのか。失敗したら戻せるのか。根拠は確認できるのか。この五つが決まっていなければ、承認はただの通過ボタンになります。
#LLMタグ

AIコーディングは学習方法で伸びる

・AIコーディングの性能は、モデルを巨大化しなくても伸ばせます。9Bモデルへ約6,000件の実務型訓練を行い、ソフトウェア修正の成功率を41.8%から56.4%へ上げた研究が公開されました。 ・2026年8月18日公開の「Agent Lightning v1.0」は、実際にツールを使うAIエージェントを強化学習するための軽量フレームワークです。開発支援AIを導入する企業にとって、モデル選定後の「自社の仕事でどう育てるか」を考える材料になります。
Zennのトレンド

AIでイベントカタログを279件自動生成し、成果物の検証を設計した話

・トリビューでリードエンジニアをしています、志甫 (@shihochan_jp) です。 ・弊社では、美容医療アプリ「トリビュー」の4リポジトリ(iOS、Android、Webフロントエンド、バックエンド)から、Google Analytics・Braze・Adjustの計測定義279エントリ(イベントのほか、ユーザー属性や配信トリガーの定義も含む)を横断するカタログを、AIで自動生成する仕組みを運用しています[1]。 ・AIにコードを読ませてドキュメントを生成させると、次の2つが問題になります。
Qiita - 人気の記事

AIで一人でゲームを作れるか試したら、35日・実質8人日でApp Store審査まで行った

・はじめに AIを使えば、一人でもゲームが作れるのではないか。 ・半信半疑で始めた検証が、35日後には App Store の審査に出ていました。 ・作ったのは『ことづての庭』という、砂と石で日本庭園を作るモバイルゲームです。ただ、この記事で書きたいのはゲームの中身ではなく、ど...
#LLMタグ

AIとの会話から「仕事を統治するOS」へ

・AI-Work Governanceの構造をMermaidで段階的に理解する ChatGPT、Claude Code、CodexのようなLLMは、単独でもかなり複雑な仕事ができるようになりました。
Zennの「機械学習」のフィード

AIと進める週末研究。Codexと回す小さな実験ループ

・はじめに 週末だけ進める研究では、実験そのものより、前回の状態を思い出して再開する作業が地味にストレスです。 ・そこで今回は、Codex を研究のループに入れ、公開情報の調査、既存手法の実装、比較実験、記録までを進めてみました。 ・題材は DCASE 2026 Challenge Task 2 です。実験内容の詳細は割愛し、この記事では、Codex に何を任せ、研究の続きをどう翌週へつないだかに絞ります。
#LLMタグ

AIにも、伝言ゲームは発生するのか?

AIにも、伝言ゲームは発生するのか?
Zennの「大規模言語モデル」のフィード

AIの「中の人」になってみる

・登壇してお話しました .NETラボ 勉強会 2026年8月で,AIの「中の人」になってみるというタイトルでお話ししました. アーカイブもあります. 資料はこちらです. https://speakerdeck.com/htkym/ai-no-naka-no-hito-ni-na-te-miru GitHub Copilot CLI のプロバイダーを自作アプリケーションへ向け,モデルを human として,人間が応答する仕組みにしました.私がLLMとなって遊んでみました. AI エージェントの入力を人間がそのまま読んでみると,依頼文だけではなく,ツール定義,履歴,実行結果など,多くの情報...
Zennの「機械学習」のフィード

AIのお話:ローカルLLMを育てる前に知ってほしい。44条件・5,720出力で捨てた3つの思い込み

・ローカルLLMを始めると、すぐに大きなモデルや高価なGPUが欲しくなります。 ・私も最初は、単純にこう考えていました。 ・大きいモデルほど、学習後も強い 教材を増やすほど、性能は上がる lossが下がれば、学習は成功している ところが、自宅PCで実際に測ってみると、三つとも危ない思い込みでした。
Zennの「大規模言語モデル」のフィード

AIの出力、どこを疑えばいいか分からない人へ ── 答えは編集者ではなくファクトチェッカーの仕事の中にあった

・別の記事を書いている最中のことだった。作業の合間に、AIに「自分のフィードバックの仕方ってどう見える?」と聞いてみたことがある。返ってきたのは、こんな言葉だった。 ・「編集者というより、ファクトチェッカーみたいですね」 最初は褒め言葉なのか判断に迷ったが、考えれば考えるほど、これはかなり的確な指摘だった。 ・編集の仕事には、昔から5つの階層がある 出版・編集の世界には、原稿ができあがるまでに通過する、比較的確立された5つの階層がある。
#LLMタグ

AIよもやま話 #014|正確。論理的。指示通り。それでもAIの生成物が「使えない」のはなぜ?

・この記事の要点 AIに何かを作らせた。事実関係は合っている。論理にも大きな破綻はない。指示した構成も守っている。それなのに完成物を目の前にすると、どうしてもこう思ってしまう。
機械学習タグが付けられた新着記事 - Qiita

AIを活用したデジタルエンターテインメント分析:ユーザー行動データから価値を見つける方法

・はじめに デジタルエンターテインメントサービスでは、ユーザーがどのコンテンツを選び、どのタイミングで離脱し、どの機能を継続的に利用しているのかを理解することが重要です。 ・近年は、単純なアクセス解析だけではなく、AIや機械学習を活用して大量の行動データからパターンを発見する...
#LLMタグ

AIを揺さぶれば、乗っ取られた瞬間は見えるのか?――万華鏡をPrompt Injection検知に使ってみる

AIを揺さぶれば、乗っ取られた瞬間は見えるのか?――万華鏡をPrompt Injection検知に使ってみる
cs.LG updates on arXiv.org

Amortized Bandwidth Learning for Kernel Density Estimation under Logarithmic Score

・arXiv:2608.20445v1 Announce Type: new Abstract: Kernel density estimation converts finite samples into probability densities, but its performance depends critically on bandwidth selection. ・Classical selectors prescribe the sample-to-bandwidth rule analytically or asymptotically, or solve a new optimization for each sample. ・An amortized framework is proposed that instead learns this mapping across a distribution of de
cs.LG updates on arXiv.org

Amplifying the imaging power of digital sky surveys with space telescopes data and generative AI

・arXiv:2608.20666v1 Announce Type: cross Abstract: While Digital sky surveys provide excellent throughput of image data and can cover a large footprint, their imaging power is normally inferior to that of space-based telescopes. ・Space-based telescopes, on the other hand, provide excellent imaging power and can image the deep Universe, but cannot provide the same throughput as advanced ground-based sky surveys.
cs.LG updates on arXiv.org

An Automated Pipeline for Few-Shot Bird Call Classification: A Case Study with the Tooth-Billed Pigeon

・arXiv:2504.16276v3 Announce Type: replace Abstract: This paper presents a largely automated one-shot bird call classification pipeline, incorporating targeted manual quality control steps, designed for rare species absent from large publicly available classifiers like BirdNET and Perch. ・While these models excel at detecting common birds with abundant training data, they lack options for species with only 1-3 known re
#LLMタグ

Anthropic Python SDK v1.0で破壊的変更、GPT-5.6 Solが値下げ|今週のAI API変更(8/17-8/23)

・使っているモデルの提供終了や値上げ、把握できていますか? 直近の期限は 2026-08-26(Assistants API)です。 ・主要AI API(OpenAI / Claude / Gemini ほか)の仕様・価格・廃止・規約の変更を、 すべて公式発表を出典に、毎週まとめています。
Qiita - 人気の記事

API・CSV・RPA・iPaaS・純正連携の5つを比較してみた|「一番ラクなのはどれ?」を6つの物差しで調べた

・「システムをつなぎたい」と相談すると、まずAPIの話になります。でも実際に業務が回るようになった事例を聞いてみると、CSVの受け渡しで終わっていたり、そもそもベンダーの標準機能で足りていたりします。 ・「結局、うちの場合はどれで繋ぐのが一番ラクで、いくらかかるのか」 これ...
The Verge

Apple’s four-pack of second-gen AirTags is $20 off

・Apple’s four-pack of second-generation AirTags is down to $79 (originally $99) at Amazon and at Target, which is the bundle’s lowest price yet. ・You can get an AirTag for $24 right now piecemeal, but getting four together as a set is a way to save even more. ・Why pay $96 for four when you can save $17?
cs.LG updates on arXiv.org

Approximate Homomorphisms and Convergent Representations in Transducers

・arXiv:2608.20428v1 Announce Type: new Abstract: We study the stability of minimal representations of controlled stochastic processes (in particular, transducers) under perturbations. ・This question is motivated by recent experiments finding predictive-state structure in the latent representations of neural networks. ・We consider standard, linear and predictive transducers.
cs.LG updates on arXiv.org

Asymmetric Capacity Allocation in Self-Refinement Pipelines

・arXiv:2608.21345v1 Announce Type: new Abstract: Self-refinement, typically structured as generation, critique, and revision, is a widely adopted paradigm for improving LLM generation and serves as a core mechanism in many LLM agents. ・While the three stages involve different cognitive demands, most existing approaches conveniently treat the model size as an implementation detail rather than a subject of study, which m
cs.LG updates on arXiv.org

AudioWorldSim: Realistic Binaural Audio Datasets For World Models

・arXiv:2608.21075v1 Announce Type: cross Abstract: This technical report presents AudioWorldSim, an open-source platform designed to generate realistic binaural audio datasets and advance research in audio-based machine learning, particularly world models. ・Built as a custom extension of Meta's SoundSpaces 2.0 platform, AudioWorldSim leverages their comprehensive acoustics framework, but focuses on the automatic rollou
cs.LG updates on arXiv.org

Automatic classification pipeline for glitches in the Virgo detector

・arXiv:2604.13687v2 Announce Type: replace-cross Abstract: Glitches frequently contaminate data in gravitational-wave detectors, complicating the observation and analysis of astrophysical signals. ・This work introduces VIGILant, an automatic pipeline for classification and visualization of glitches in the Virgo detector. ・Using a curated dataset of Virgo O3b glitches, two machine learning approaches are evaluated: tree-
cs.LG updates on arXiv.org

BackDFL: A Unified Benchmark For Backdoor Attacks and Defenses In Decentralized Federated Learning

・arXiv:2608.21137v1 Announce Type: new Abstract: Decentralized Federated Learning (DFL) promises trust-free collaborative learning by replacing the centralized parameter server with peer-to-peer model exchange. ・However, this architectural shift fundamentally reshapes the threat landscape. ・Without globally coordinated aggregation, DFL becomes particularly susceptible to backdoor attacks, in which malicious participants
cs.LG updates on arXiv.org

Bankruptcy Prediction via Hybrid Resampling and Stacking Ensemble Techniques with Explainable Artificial Intelligence (XAI)-Driven Analysis

・arXiv:2608.20343v1 Announce Type: new Abstract: This study develops and evaluates a bankruptcy prediction framework that integrates consensus-based feature selection, hybrid resampling, stacking ensembles, and explainable artificial intelligence to improve minority-class detection in severely imbalanced financial data. ・Using the Taiwanese Bankruptcy Prediction dataset from the UCI Machine Learning Repository, five fe
cs.LG updates on arXiv.org

Behavior-Consistent Deep Reinforcement Learning

・arXiv:2605.21214v3 Announce Type: replace Abstract: Reinforcement learning (RL) often exhibits high variance across training runs, leading to unreliable performance and posing a major challenge to deployment in real-world domains. ・In this work, we address the challenge of cross-run policy divergence by formalizing the problem of behavior-consistent RL, where the objective is to obtain policies that are both high-perf
cs.LG updates on arXiv.org

Benchmarking noisy label detection methods

・arXiv:2510.16211v2 Announce Type: replace Abstract: Label noise is a common problem in real-world datasets, affecting both model training and validation. ・Clean data are essential for achieving strong performance and ensuring reliable evaluation. ・While various techniques have been proposed to detect noisy labels (or label errors), there is no clear consensus on optimal approaches.
cs.LG updates on arXiv.org

Bern2Edge: A Neurosymbolic Compiler for Edge Deployment via Bernstein Polynomial Networks

・arXiv:2608.20497v1 Announce Type: new Abstract: Deploying high-accuracy neural networks on resource-constrained edge devices remains challenging, as existing approaches treat training, compression, and hardware synthesis as separate stages, leaving a gap between software-trained models and efficient end-to-end deployment with limited support for interpretability. ・We propose Bern2Edge, an end-to-end framework that use
MarkTechPost

Best GPU Neoclouds 2026: CoreWeave, Nebius, Lambda, Crusoe, and Groq Ranked by Published Pricing and Contracted Power

・The five largest GPU neoclouds now run on very different models. ・CoreWeave and Nebius report to the SEC; Lambda and Crusoe are private and heading toward IPOs; Groq rebuilt itself as an inference cloud after licensing its LPU technology to NVIDIA. ・This comparison checks each provider's live rate card, Q2 2026 financials, active and contracted gigawatts, anchor contracts, and SemiAnalysis ClusterMAX tier.
Hugging Face Papers

Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs

Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs
cs.LG updates on arXiv.org

Beyond Imitation: Self-Improving Robot Policies via Off-Policy Q-Planning

・arXiv:2608.21204v1 Announce Type: cross Abstract: Behaviour Cloning (BC) has driven remarkable progress in robot manipulation, yet it is fundamentally limited by its inability to self-improve: a policy that fails cannot learn from that failure without additional human demonstrations. ・Reinforcement Learning fine-tuning offers a path to self-improvement but has proven difficult to scale to the multi-billion-parameter m
cs.LG updates on arXiv.org

BF1: A Causal Dyadic Sparse-Attention Retrofit for Efficient Long-Context Transformers

・arXiv:2608.20427v1 Announce Type: new Abstract: Dense causal attention remains expensive at long context even when implemented with highly optimized exact kernels. ・We study BF1, a deterministic block-aligned dyadic sparse-attention route that combines a small exact local neighborhood, a global first block, and logarithmically spaced historical blocks. ・The route is related to prior log-sparse and dilated attention pat
cs.LG updates on arXiv.org

BIPPO: Budget-Aware Independent PPO for Energy-Efficient Federated Learning Services

・arXiv:2511.08142v2 Announce Type: replace Abstract: Federated Learning (FL) is a promising machine learning solution in large-scale IoT systems, guaranteeing load distribution and privacy. ・However, FL does not natively consider infrastructure efficiency, a critical concern for systems operating in resource-constrained environments. ・Several Reinforcement Learning (RL) based solutions offer improved client selection fo
WIRED

Birdfy Nest Duo Review: My Own Private Nature Documentary

・With two cameras and built-in climate sensors, the Birdfy Nest Duo let me watch an entire nesting season unfold.
cs.LG updates on arXiv.org

C-Score: Beyond Accuracy for Robustness Assessment in Semi-Supervised Learning under Open-World Unlabeled Contamination

・arXiv:2608.20667v1 Announce Type: new Abstract: Pseudo-label-based semi-supervised learning has achieved strong performance due to its simplicity and scalability. ・However, it is typically developed under a closed-world assumption that unlabeled data are drawn from the same distribution as labeled data. ・In practical deployment, unlabeled data are often collected from open environments and may contain OOD samples.
cs.LG updates on arXiv.org

Calibrate-Then-Delegate: Safety Monitoring with Risk and Budget Guarantees via Model Cascades

・arXiv:2604.14251v2 Announce Type: replace Abstract: Monitoring LLM safety at scale requires balancing cost and accuracy: a cheap latent-space probe can screen every input, but hard cases should be escalated to a more expensive expert. ・Existing cascades delegate based on probe uncertainty, but uncertainty is a poor proxy for the utility of an expert call, as it ignores whether the expert would actually improve the pre
cs.LG updates on arXiv.org

Capturing Cardiac Cyclicity through Phase-Equivariant Self-Supervised Learning

・arXiv:2608.21147v1 Announce Type: new Abstract: The cyclic structure of physiological processes offers a natural prior for self-supervised representation learning, and the cardiac cycle provides a particularly well-defined setting in which to exploit it. ・We derive a phase-equivariant self-supervised objective and introduce Winder, a joint-embedding architecture that organises representations into phase-invariant coor
cs.LG updates on arXiv.org

Causal Modeling of Adverse Pregnancy Outcomes via Adaptive LLM Proposals

・arXiv:2608.21079v1 Announce Type: new Abstract: Adverse Pregnancy Outcomes (APOs) such as preterm birth and gestational diabetes can have long-term consequences for both the mother and child, yet an understanding of their causes remains elusive. ・Causal discovery in this domain is especially challenging due to a paucity of data and incomplete domain knowledge. ・As a result, pure data-driven methods fail, and Large Lang
cs.LG updates on arXiv.org

CDRL: Certification-Driven Reinforcement Learning for Neutrino Flavor Model Discovery

・arXiv:2608.20686v1 Announce Type: cross Abstract: Many scientific discovery problems require searching combinatorial hypothesis spaces under complex domain constraints. ・Reinforcement learning (RL) offers a promising approach, but existing methods rely on scalar rewards that provide limited information about why candidate solutions fail, leading agents to repeatedly explore invalid regions. ・We introduce Certification-
cs.LG updates on arXiv.org

CFM: Language-aligned Concept Foundation Model for Vision

・arXiv:2601.13798v3 Announce Type: replace-cross Abstract: Language-aligned vision foundation models perform strongly across diverse downstream tasks. ・Yet, their learned representations remain opaque, making interpreting their decision-making difficult. ・Recent work decompose these representations into human-interpretable concepts, but provide poor spatial grounding and are limited to image classification tasks.
@IT 全フォーラム 最新記事一覧

ChatGPT、Gemini、Claude Sonnet 5の“最強”は? 「速度」と「安定性」を実測して比べてみた

・主要なAIチャットサービスの「ChatGPT」「Gemini」「Claude Sonnet 5」の中で、応答速度が一番速いのはどれなのか。時間帯によって速度は変わるのか。測定値はどれだけ安定しているのか。実測を基に検証する。
#AIタグ

ChatGPTへ同じ注意を何度も書かない失敗ログ完全テンプレート|5分類・修正・再テスト付き

・ChatGPTへ昨日と同じ注意を、今日も書いていませんか。 ・「表にしないで」「出典を付けて」「勝手に公開しないで」。その場で直しても、次の実行で戻るなら改善ではありません。 ・見落とされている原因は、失敗した回答ではなく修正履歴が残っていないことです。AIのミスは、指示、入力、判断、実行、検証に分けると修正箇所が見えます。
Zennのトレンド

Claude Code のデスクトップ操作、内蔵の computer-use ではなく Windows-MCP を使っている理由

・Claude Code にデスクトップ操作を任せる手段には、内蔵の computer-use があります。画面を撮ってピクセル座標でクリックするこの方式のまま任せ続けてよいのか気になって代わりを探し、UI Automation で画面を構造として読む Windows-MCP に入れ替えました。メモ帳への入力という同じタスクを両方に実行させて、操作の中身がどこで変わるのかを確かめてみました。 ・TL;DR computer-use は操作のたびに「画面を撮る → 位置を推定する → 座標をクリックする」を繰り返します。Windows-MCP は UI Automation の要素ツ...
#LLMタグ

Claude Codeとの分業が2枚目のGPUを不要にする——RTX 3080を売った日

・自宅AI環境が育つと、GPUが減る 先に種明かしをします。新しいGPUを買って、玉突きで古いのを手放したのではありません。Claude Codeとの分業がはっきりした結果、ローカルに残る仕事は夜間の定型だけになり、それはもともと手元にあったRTX 5070 Ti(16GB)1枚で足りました。だから、2枚目が余った。ハードは1枚も増えていないのに、仕事または役割の線を引き直しただけでGPUが1枚減る。自宅のAI環境が育つと、機材は増えるどころか減っていく——今回は、その顛末です。
LLMタグが付けられた新着記事 - Qiita

Claude Codeの「引き継ぎ」新機能、自前handoffシステムと何が代替できるか比較した

・セッションが切れるたびに「前回どこまでやったっけ」を思い出すコストに、正直うんざりしていた時期がある。今は複数のAIエージェントを運用していて、セッション終了時に進行中タスク・保留事項・次アクションを自動でファイルに書き出す仕組み(以下、handoffファイルと呼ぶ)を2ヶ...
Qiita - 人気の記事

Claude Codeのセッション間通信、Windowsでも動きます(v2.1.234から)

・Claude Code には、自分の別セッションにメッセージを送る cross-session messaging という機能があります。ターミナルを2枚開いて並行作業しているとき、片方の変更をもう片方に伝えるやつです。 ・これ、「Windowsでは使えない」と思っている人が...
Zennの「大規模言語モデル」のフィード

Claude Code本番運用ガイド — CI/CDに組み込む自律エージェント設計

・Claude Codeを対話で使いこなせるようになった。次は、それを「毎晩黙って働くチームメンバー」にする番です。本書は、ヘッドレス実行(claude -p)をCI/CDに組み込み、検出→修正→検証→コミットの自律ループを無人で回すための設計書です。本番運用で立ちはだかる4つの壁——暴走・品質・コスト・監査——に対して、ガードレール設計、プロンプトインジェクション対策を含む脅威モデル、品質評価、コスト管理、企業環境への導入までを、方法論・実際の失敗例・動くリファレンス実装の三点セットで扱います。入門書ではありません。個人でClaude Codeを使いこなしている人が、チームと本番に展開するための一冊です。前書きと第1章は無料で読めます。
Zennの「大規模言語モデル」のフィード

CLAUDE.md で AskUserQuestion を使わせたら grill-me の質問攻めが楽になった

・概要 grill-me は、計画や設計が固まるまでユーザーに質問を繰り返すスキルです。 ・質問はすべて文章で提示されるため、回答も毎回自由記述になります。 ・二択の質問でも同じです。
#LLMタグ

Claudeが約3時間止まった8月24日の障害と頻度

・Anthropicのstatus pageに、日本時間8月24日の14時06分、Claude Mythos 5とClaude Fable 5とClaude Opus 5とClaude Opus 4.8でエラー率が上がっているという投稿が出た。 ・解消の告知は同日17時30分で、影響が続いた時間は13時50分から16時36分までの2時間46分。 ・止まったのはclaude.aiだけでなく、api.anthropic.comとClaude CodeとClaude Coworkも含まれる。
Hugging Face Papers

CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment

CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment
cs.LG updates on arXiv.org

COEC: Calibrated Orthogonal-Equivalence Compensation for Structured Pruning of Large Language Models

・arXiv:2608.21142v1 Announce Type: new Abstract: Structured pruning reduces the size and inference cost of large language models (LLMs) by removing weight columns, but the resulting output error can degrade accuracy. ・Existing training-free compensation methods use an additive bias or a single orthogonal rotation on the output side of the retained weight. ・These corrections leave its input singular frame unchanged and t
cs.LG updates on arXiv.org

COMET: Contrastive Motion-Enhanced Temporal Reasoning for Video Multimodal Large Language Models

・arXiv:2608.21030v1 Announce Type: cross Abstract: Video multimodal large language models have advanced significantly, yet fine-grained motion-temporal understanding remains fragile. ・The core bottleneck is not only sparse frame sampling, but also the lack of a complete temporal modeling pipeline for explicitly representing frame-to-frame change, enabling appearance-motion interaction, and optimizing temporal direction
cs.LG updates on arXiv.org

Compared to What? Baselines and Metrics for Counterfactual Prompting

・arXiv:2605.01048v2 Announce Type: replace-cross Abstract: Counterfactual prompting (i.e., perturbing a single factor and measuring output change) is widely used to evaluate things like LLM bias and CoT faithfulness. ・But in this work we argue that observed effects cannot be attributed to the targeted factor without accounting for baseline "meaning-preserving" modifications to text that establish general model sensitiv
stat.ML updates on arXiv.org

comprisk: A scikit-learn-compatible Python toolkit for competing-risks survival analysis

・arXiv:2607.09431v2 Announce Type: replace-cross Abstract: Medical time-to-event data are frequently subject to competing risks, where the occurrence of one terminal event precludes the others and standard survival methods that treat competing events as censoring yield biased absolute-risk estimates. ・Valid analysis instead targets the cause-specific cumulative incidence function (CIF). ・This methodology has been availa
cs.LG updates on arXiv.org

ConceptTS: LLM-Guided Concept Bottlenecks for Interpretable Multivariate Time-Series Forecasting

・arXiv:2608.21277v1 Announce Type: new Abstract: State-of-the-art multivariate time-series forecasters can model complex temporal and cross-variable dependencies, yet their opaque representations provide limited insight into why a particular forecast is produced. ・This lack of transparency restricts their use in settings where practitioners must understand and assess the factors underlying a prediction. ・We introduce Co
cs.LG updates on arXiv.org

Conditional-Independence-Regularized Distributional Autoencoders for Mixed-Type Data

・arXiv:2608.20562v1 Announce Type: cross Abstract: Mixed-type data containing both numerical and categorical variables arise in many scientific and real-world applications. ・Existing representation learning and generative modeling approaches typically focus either on reconstruction accuracy or unconditional data generation, but often fail to recover the full conditional distribution of the data while preserving interpr
cs.LG updates on arXiv.org

Consistency Models for Fast MRI Reconstruction Using Regularization by Denoising

・arXiv:2608.20561v1 Announce Type: cross Abstract: Diffusion models (DMs) have emerged as powerful generative priors for MRI reconstruction with promising results. ・Yet DM-based methods require extensive iterative refinement, limiting their practical deployment. ・Consistency models (CMs) provide a compelling alternative, aiming to map out the diffusion trajectory in a single pass, enabling faster generation.
stat.ML updates on arXiv.org

Convergence of the Deep Galerkin Method for Finite State Mean Field Control Problems

・arXiv:2405.13346v2 Announce Type: replace-cross Abstract: We establish the convergence of the deep Galerkin method (DGM), a deep learning-based scheme for solving high-dimensional nonlinear PDEs, for Hamilton-Jacobi-Bellman (HJB) equations that arise from the study of mean field control problems (MFCPs). ・Based on a recent characterization of the value function of the MFCP as the unique viscosity solution of an HJB eq
Zennのトレンド

Ctrl+C でプログラムが止まる仕組みを調べた

・はじめに 以前 Go の並列処理を学ぶ中で、graceful shutdown を扱いました。Ctrl+C についても軽く学びました。 ・ただ、Ctrl+C を押したときに何が起きているかは軽くしか学べておらず、自分の中で引っかかっていました。 ・今回調べたのは、Ctrl+C を押してプログラムが実際に止まるまでです。(見ていくのは Linux です) 全体の流れ Ctrl+C を押してからプログラムが終了するまでの簡単な流れは以下の通りです。
cs.LG updates on arXiv.org

CubicSplat: Differentiable Vector Graphics via Error-Bounded Forward Relaxation

・arXiv:2608.20803v1 Announce Type: cross Abstract: Vector graphics are prized for their resolution independence, compact storage, and direct editability, making differentiable optimization of their parametric primitives an attractive goal. ・Yet classical rasterization is discontinuous with respect to geometry, and existing remedies that smooth the forward pass demand increasingly elaborate heuristics as scene complexit
cs.LG updates on arXiv.org

Curriculum-Aware Interpolate-then-Refine: Learned Physiological Time-Series Imputation under Realistic Missingness

・arXiv:2608.21207v1 Announce Type: new Abstract: Imputing physiological time series (arterial blood pressure, blood glucose, etc.) is essential for addressing the missingness that pervades clinical data. ・Yet modern imputation methods perform poorly in this domain: a recent benchmark found that simple linear interpolation outperformed every learned imputer on real-world clinical signals with realistic gaps.
Hugging Face Papers

Daedalus-150M: A Convolution-Attention Hybrid Designed for CPU Inference

Daedalus-150M: A Convolution-Attention Hybrid Designed for CPU Inference
cs.LG updates on arXiv.org

DAOP: Data-Aware Offloading and Predictive Pre-Calculation for Efficient MoE Inference

・arXiv:2501.10375v3 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) models, though highly effective for various machine learning tasks, face significant deployment challenges on memory-constrained devices. ・While GPUs offer fast inference, their limited memory compared to CPUs means not all experts can be stored on the GPU simultaneously, necessitating frequent, costly data transfers from CPU memory, of
#LLMタグ

DAY40|AI COREは「記憶」を使って判断する

・Company AI OS開発記録、DAY40です。 ・DAY39では、AI COREにMemory機能を組み込みました。
The Verge

De-Googled GrapheneOS is coming to Motorola’s foldables next year

・The Razr Ultra line is among those due to receive GrapheneOS support next year. ・| Photo: Allison Johnson / The Verge GrapheneOS, an open source version of Android that prioritizes security and privacy, has detailed its plans for supporting Motorola smartphones. ・Official support is set to arrive next year, starting with traditional flagships, before rolling out to Motorola's foldable phones and perhaps cheaper models,
cs.LG updates on arXiv.org

Decision Tree and K-Means Analysis of Raman Spectra for Edible Oils: A Physics-Informed AI Approach

・arXiv:2608.20440v1 Announce Type: new Abstract: Authentication of edible oils in processed foods is important for food quality, fraud prevention, and regulatory compliance. ・This study establishes an integrated Raman spectroscopy and machine-learning framework that links intrinsic spectral organization, interpretable classification, and Physics-Informed Artificial Intelligence (PI-AI). ・Five edible oils were investigat
cs.LG updates on arXiv.org

Decoupling Policy Extraction for Offline Reinforcement Learning

・arXiv:2608.20909v1 Announce Type: new Abstract: Offline RL methods commonly jointly train the actor and critic, where the critic is used to guide the actor toward higher-value actions. ・This coupled learning process is well motivated in online RL, where an improved actor collects new data that can further update the actor and the critic. ・However, training data remains fixed in offline RL, making actor-side policy impr
Zennの「大規模言語モデル」のフィード

DeepSeek V4-Flash-Vision-Exp: Image Input Finally Hits the API

・DeepSeek's first API-accessible vision model caps images at 384 tokens and keeps V4-Flash pricing. ・Here is what that means for agent builders. ・Introduction Until August 21, 2026, if you wanted to send an image to a DeepSeek model through the API, you couldn't.
cs.LG updates on arXiv.org

Defining Decentralization: An Ontological Perspective

・arXiv:2608.09748v2 Announce Type: replace-cross Abstract: Decentralization as a concept in computer science has existed for over half a century. ・Despite its fundamental role across domains such as security, distributed computing, artificial intelligence, cloud infrastructures, and Internet of Things (IoT) architectures, there remains no universally accepted definition of decentralization applicable across computer co
cs.LG updates on arXiv.org

Degree-Mass Message Passing for Betweenness Ranking in Directed and Undirected Networks

・arXiv:2602.09716v2 Announce Type: replace Abstract: Computing the importance of nodes in networks is a long-standing fundamental problem that has driven extensive study of various centrality measures. ・A particularly well-known centrality measure is betweenness centrality, whose exact computation becomes prohibitive on large-scale networks. ・Graph Neural Network (GNN) models have thus been proposed to predict the ranki
機械学習タグが付けられた新着記事 - Qiita

dera AI Weekly Vol.44 — 2026/8/17

・2026-08-17号 今週のAI業界を一言で表すなら? AI業界の勢力図そのものが書き換わった一週間でした。 ・性能の順位、使う人の数、お金の流れ、開発の土台。この四つが同時に動いた週は、ここ数ヶ月ありません。 ・📊 今週知っておくべきこと 今週は個別のニュースを追うより...
cs.LG updates on arXiv.org

Designing a Robust LLM-Based Evaluation System for Agentic AI in Drug Discovery Through Human Alignment

・arXiv:2608.21057v1 Announce Type: new Abstract: Agentic large language model (LLM) systems are reshaping scientific workflows in chemistry and drug discovery, but evaluating their open-ended, tool-augmented outputs remains a fundamental bottleneck. ・Reference-based metrics such as BLEU and ROUGE fail to capture semantic correctness, while expert human evaluation does not scale to the iteration speed these systems dema
cs.LG updates on arXiv.org

Detecting Functional Memorization in Code Language Models

・arXiv:2606.12764v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used to generate code at scale. ・Meanwhile, prior work has investigated whether training data may be recoverable from model outputs, by auditing the textual overlap between training examples and model generations. ・Code, however, can preserve the same logic while differing substantially in syntax and structure.
cs.LG updates on arXiv.org

Deterministic and probabilistic neural surrogates of global hybrid-Vlasov simulations

・arXiv:2601.12614v4 Announce Type: replace-cross Abstract: Hybrid-Vlasov simulations resolve ion-kinetic effects in the solar wind-magnetosphere interaction, but even 5D (2D + 3V) configurations are computationally expensive. ・We show that graph-based machine learning emulators can learn the spatiotemporal evolution of electromagnetic fields and lower-order moments of the ion velocity distribution function in near-Eart
cs.LG updates on arXiv.org

Doctor Rashomon and the UNIVERSE of Madness: Variable Importance with Unobserved Confounding and the Rashomon Effect

・arXiv:2510.12734v2 Announce Type: replace Abstract: Variable importance (VI) methods are often used for hypothesis generation, feature selection, and scientific validation. ・In the standard VI pipeline, an analyst estimates VI for a single predictive model with only the observed features. ・However, the importance of a feature depends heavily on which other variables are included in the model, and essential variables ar
stat.ML updates on arXiv.org

Double Machine Learning of Continuous Treatment Effects with Additive Instrumental Variables

・arXiv:2601.01471v3 Announce Type: replace-cross Abstract: Estimating causal effects of continuous treatments is a common problem in practice, for example, in studying average dose-response functions. ・Classical analyses typically assume that all confounders are fully observed, whereas in real-world applications, unmeasured confounding often persists. ・In this article, we propose a novel framework for the identification
stat.ML updates on arXiv.org

Doubly robust inference via calibration

・arXiv:2411.02771v3 Announce Type: replace-cross Abstract: Doubly robust estimators are widely used for estimating average treatment effects and other linear summaries of regression functions. ・While consistency requires only one of two nuisance functions to be estimated consistently, asymptotic normality for linear functionals typically requires sufficiently fast convergence of both. ・We address this mismatch by showin
Takara TLDR - Daily AI Papers

DreamBench-SWE: A Multi-Session Memory-Hygiene Benchmark for Software Agents

・DreamBench-SWE is a multi-session benchmark for software-agent memory hygiene in which later software tasks depend on non-inferable evidence from earlier sessions and are scored by executable hidden oracles. ・We report the original scaled v2 fold and a separately preregistered v2.1 successor audit designed after that study but frozen before successor outcome inspection. ・The successor run completed 360/360 work units a
cs.LG updates on arXiv.org

Dual-Cache Latent Space Communication between Heterogeneous Language Models

・arXiv:2608.20617v1 Announce Type: cross Abstract: Multi-agent LLM systems split work across models, so answering often requires knowledge that sits in another agent's context: a Sharer has encoded information that a Receiver needs to complete its task. ・They usually communicate by exchanging text, which puts autoregressive decoding on the critical path and reduces the exchange to a discrete message written without sig
Qiita - 人気の記事

DWHとは?分析基盤の全体像を理解する ~Snowflake・dbt・Power BIで実現するモダンDWH~(第1回)

・はじめに データ活用が進む中で、Power BIやTableauなどのBIツールを導入する企業は増えています。 ・しかし、BIツールを導入しただけですぐにデータ活用がうまくいくわけではありません。 ・例えば、次のような課題に心当たりはないでしょうか。
stat.ML updates on arXiv.org

EDGE: a closed-form directed test for the calibration of probabilistic binary classifiers

・arXiv:2608.20511v1 Announce Type: cross Abstract: A probabilistic binary classifier is judged almost everywhere by discrimination - accuracy, the ROC curve, the area under it. ・Every such criterion is invariant to a monotone distortion of the predicted probabilities, so a classifier can rank perfectly and still return probabilities that are badly wrong. ・Calibration is the property decisions need, and the field's instr
cs.LG updates on arXiv.org

Efficient Exploration at Scale

・arXiv:2603.17378v2 Announce Type: replace Abstract: We develop an online learning algorithm that dramatically improves the data efficiency of reinforcement learning from human feedback (RLHF). ・Our algorithm incrementally updates reward and language models as choice data is received. ・The reward model is fit to the choice data, while the language model is updated by a variation of reinforce, with reinforcement signals
cs.LG updates on arXiv.org

Efficient Inference for Inverse Reinforcement Learning and Dynamic Discrete Choice Models

・arXiv:2512.24407v2 Announce Type: replace Abstract: In many sequential decision-making problems, researchers observe actions but not the rewards that drive behavior, yet still wish to evaluate and compare counterfactual policies. ・Inverse reinforcement learning (IRL) and dynamic discrete choice (DDC) models address this setting by positing an optimality model that links latent rewards to observed actions. ・Existing fle
The Verge

ESPN streaming plans are getting more expensive

・ESPN is hiking the price of its subscription on September 17th, a change that will also impact its bundles with Disney Plus. ・In a support page spotted earlier by Sports Media Watch, ESPN says its ad-supported Select membership will cost $13.99 instead of $12.99 / month, while its Unlimited plan will rise to $31.99 from $29.99 / month. ・The annual ESPN Select plan is also going to $139.99 from $129.99 / year, and the a
cs.LG updates on arXiv.org

Event-triggered Implicit Perturbation for Zeroth-Order Fine-Tuning of Spiking Transformers

・arXiv:2608.21223v1 Announce Type: cross Abstract: Zeroth-order (ZO) optimization estimates gradients using only forward-pass evaluations, making it suitable for fine-tuning non-differentiable, event-driven spiking neural networks (SNNs). ・However, its deployment on in-memory computing (IMC) accelerators is constrained by the repeated read-modify-write (RMW) operations arising from explicit weight perturbation and the
Hugging Face Papers

Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models

Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models
cs.LG updates on arXiv.org

EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking

・arXiv:2608.20886v1 Announce Type: cross Abstract: Real-world image search queries are multimodal and compositional: ``find this shirt in pink'' specifies an entity to retain, an attribute to modify, and context to ignore. ・Yet existing re-rankers either compress such multifaceted relevance into an opaque embedding or rely on free-form chain-of-thought that easily omits or hallucinates fine-grained constraints.
Hugging Face Papers

EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking

EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking
cs.LG updates on arXiv.org

Exact and general decoupled solutions of the LMC Multitask Gaussian Process model

・arXiv:2310.12032v4 Announce Type: replace Abstract: The Linear Model of Co-regionalization (LMC) is a very general multitask gaussian process model for regression or classification. ・While its expressiveness and conceptual simplicity are appealing, naive implementations have cubic complexity in the product (number of datapoints $\times$ number of tasks), making approximations mandatory for most applications.
cs.LG updates on arXiv.org

Explaining Intrinsic Moral Self-Correction with Mechanistic Interpretability

・arXiv:2505.11924v4 Announce Type: replace-cross Abstract: Intrinsic moral self-correction refers to the phenomenon where a language model refines its ethical judgments or aligns its outputs purely through prompting. ・While effective across diverse tasks, its mechanism remains unclear. ・We hypothesize intrinsic moral self-correction functions by steering hidden representations along interpretable latent directions.
cs.LG updates on arXiv.org

Exploratory As-Analyzed No-Detection of Culturally-Marked Predicate-Triggered PII Amplification in a Synthetic-English RAG Probe: A Predicate-Resource-Confounded Audit

・arXiv:2608.20351v1 Announce Type: cross Abstract: We ask whether stereotype-loaded queries about culturally marked people leak more personal information from a retrieval-augmented generation (RAG) system than otherwise-equivalent neutral queries. ・We pre-register a four-culture audit (en-Anglo, es-LATAM, Arabic, Hindi) on a synthetic English PII corpus, comparing five query arms we call the Stereotype-Trigger Leakage
cs.LG updates on arXiv.org

Faults That Fortify: CNN Adversarial Robustness via GPU Undervolting

・arXiv:2608.20572v1 Announce Type: new Abstract: Convolutional Neural Networks (CNNs) face a dual challenge: vulnerability to adversarial attacks and prohibitive training cost. ・Adversarial training is effective but expensive, a burden that grows as learning shifts to the energy-constrained edge. ・This paper addresses both through GPU undervolting during training.
cs.LG updates on arXiv.org

Federated and differentially private estimation of KL divergence

・arXiv:2411.16478v3 Announce Type: replace Abstract: Measuring distribution drifts is a key task in managing distributed, sensitive data, as it underpins a wide range of federated learning and analytics applications. ・In many practical settings, however, directly sharing such information is either undesirable (e.g., due to privacy concerns) or infeasible (e.g., due to high communication costs). ・In this work, we present
cs.LG updates on arXiv.org

FeLoG: Scalable and Efficient Distributed Graph Embedding with Feedback Loop Mechanism

・arXiv:2606.22180v3 Announce Type: replace-cross Abstract: Graph embedding maps graph nodes into low-dimensional vectors to support applications such as recommendation, fraud detection, and graph-based retrieval-augmented generation (GraphRAG). ・As graphs scale to billions of edges, scalable and efficient graph embedding has become increasingly important. ・Existing frameworks commonly adopt a sampling-training paradigm,
cs.LG updates on arXiv.org

Fine-tuning LLMs for Tourist Trajectory Prediction using Field Experiment Data

・arXiv:2608.20830v1 Announce Type: cross Abstract: Evaluating mobility interventions at tourist destinations requires predicting visitor behavior under varying conditions. ・Traditional methods struggle because tourist decisions depend heavily on context like weather and fatigue, yet models cannot generalize to unobserved scenarios. ・Large Language Models offer a solution by encoding commonsense knowledge about human beh
cs.LG updates on arXiv.org

FlatLand: Personalized Graph Federated Learning via Tailored Lorentz Space

・arXiv:2608.21096v1 Announce Type: new Abstract: Federated learning enables privacy-preserving collaborative training, but highly heterogeneous client data remain challenging, especially in graph federated learning where clients possess structurally diverse graphs. ・Existing personalized federated learning (PFL) methods ignore the intrinsic geometric properties of diverse graph structures. ・We propose FlatLand, a novel
cs.LG updates on arXiv.org

FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth

・arXiv:2608.20574v1 Announce Type: cross Abstract: Open-ended language-model benchmarks usually inherit a judge: a human preference panel, another model, or a brittle exact-match key. ・We introduce FlavourBench, an automated benchmark in which a versioned culinary system supplies dense, executable ground truth. ・Each task presents eight ingredients and asks for a three-ingredient portfolio; before model execution, Epicu
Hugging Face Papers

FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth

FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth
cs.LG updates on arXiv.org

Forecasting with an N-dimensional Langevin Equation and a Neural-Ordinary Differential Equation

・arXiv:2405.07359v2 Announce Type: replace Abstract: Accurate prediction of electricity day-ahead prices is essential in competitive electricity markets. ・Although stationary electricity-price forecasting techniques have received considerable attention, research on non-stationary methods is comparatively scarce, despite the common prevalence of non-stationary features in electricity markets. ・Specifically, existing non-
WIRED

Forget Meta Ray-Bans. These Dorky-Looking Virtual Display Glasses Are Way More Useful

・Tethered display glasses are truly practical face computers. ・They trade bulky spatial computing for simplicity: Plug in, recline, and get a massive screen right in front of your nose.
stat.ML updates on arXiv.org

Fr\'echet regression of multivariate distributions with nonparanormal transport

・arXiv:2603.07014v2 Announce Type: replace-cross Abstract: Regression with distribution-valued responses and Euclidean predictors has gained increasing scientific relevance. ・While methodology for univariate distributional data has advanced rapidly in recent years, multivariate distributions, which additionally encode dependence across univariate marginals, have received less attention and pose computational and statis
cs.LG updates on arXiv.org

Free-Probability Kernels for Zero-Rollout Hyperparameter Selection in Reservoir Computing

・arXiv:2608.20998v1 Announce Type: new Abstract: Reservoir computing (RC) couples a fixed recurrent dynamical system with a trained lightweight readout, but this efficiency is partly lost during hyperparameter selection: the recurrent gain, input scale, and leakage rate determine the reservoir's stability and temporal processing regime and are usually tuned through many rollouts. ・We introduce a deterministic, pilot-in
cs.LG updates on arXiv.org

From a Static Multi-Level Small Semantic Codebook to a Dynamic Single-Level Large Semantic Codebook for Generative Recommendation

・arXiv:2608.21012v1 Announce Type: cross Abstract: Generative recommendation represents each item with a sequence of discrete Semantic IDs (SIDs) and predicts the sequence to retrieve the next item. ・Typical systems use multi-level residual quantization, which increases autoregressive decoding cost and creates a large hierarchical space that may be sparsely occupied. ・Static codebooks also become misaligned with current
cs.LG updates on arXiv.org

From Thermal Preference Prediction to Adaptive Thermal Intervention: A Reinforcement Learning Approach Using Physiological and Environmental Sensing

・arXiv:2608.20423v1 Announce Type: new Abstract: Personalised thermal comfort is essential for occupant wellbeing and for the development of more responsive building-control strategies, yet conventional Heating, Ventilation, and Air Conditioning (HVAC) systems rely on static setpoints and population-level comfort models that fail to capture individual physiological variability. ・This paper presents a two-stage personal
cs.LG updates on arXiv.org

Fuzzy-MoE: Interpretable Regime-Conditioned Expert Routing for Non-Stationary Multivariate Time Series Forecasting

・arXiv:2608.20761v1 Announce Type: new Abstract: In non-stationary multivariate time series, different variables and samples often exhibit heterogeneous latent dynamic states, while existing deep forecasting models usually compress them into a unified end-to-end mapping, leading to suboptimal modeling of time-varying dynamics and limited interpretability regarding which forecasting mechanism is activated under differe
MarkTechPost

Generalist AI Releases GEN-1.5: A Robot Foundation Model That Learns New Tasks From One 3–12 Second Demo

・Generalist AI has released GEN-1.5, a robot foundation model that learns a new physical task from a single demonstration. ・Drop 3–12 seconds of sensorimotor data into its 30-second context window, and the robot performs the task. ・No gradient updates, no fine-tuning, no task-specific programming.
cs.LG updates on arXiv.org

Generalization Measures under Controlled Covariate Shift: A Regime-Aware Benchmark

・arXiv:2602.01718v2 Announce Type: replace Abstract: Predicting generalization from quantities available before target-test evaluation remains a central challenge in deep learning. ・The systematic benchmark of Jiang et al. ・(2020) evaluated many generalization measures, but it focused on independent and identically distributed (IID) settings.
MIT News - Artificial intelligence

Generating scenarios for extreme events, without extreme data

・A new algorithm learns to anticipate the unprecedented scenarios that critical infrastructure and global supply chains are least prepared for.
cs.LG updates on arXiv.org

Geometric Regularization for Long-Tailed Semi-Supervised Learning via Gaussian Feature Bridges

・arXiv:2608.20710v1 Announce Type: new Abstract: Real-world semi-supervised learning (SSL) often encounters significant challenges with long-tailed label distributions and noisy pseudo-labels, which hinder generalization and amplify confirmation bias. ・In this work, we introduce a novel framework, Gaussian Bridge Consistency (GBC), to address these challenges by constructing semantic interpolation paths between unlabel
MarkTechPost

Google Research Introduces ME-POIs: A Mobility-Informed Framework that Adds “How a Place Is Used” to Text-Based POI Embeddings

・framework that folds aggregate human movement into text-based place embeddings. ・Language models describe what a place is; they miss how it is used. ・ME-POIs encodes each visit as a contextualized vector and aligns it with one learnable prototype per POI through contrastive learning, then transfers visit distributions from data-rich anchors to the long tail across three spatial scales.
Zennのトレンド

Goのポインタに抵抗を感じていた理由

・はじめに Goに初めて触れたとき、ポインタを使うことに少し抵抗を感じていました。 ・例えば、次のようなコードです。 ・func update(user *User) { user.Name = "Bob" } その抵抗感の正体を考えてみると、PHPを書いていた頃の「参照渡しへの警戒心」が関係していました。
cs.LG updates on arXiv.org

GRALIS: Fusing Coalition and Gradient Attribution with Closed-Form Conservation Error and Finite-Sample Guarantees

・arXiv:2605.05480v3 Announce Type: replace Abstract: The main post-hoc XAI methods for deep networks -- GradCAM, SHAP, LIME, Integrated Gradients -- originate from heterogeneous theoretical foundations and are not naturally comparable within a single representation. ・A recent benchmark also finds their coalition-based members (GradCAM, KernelSHAP, LIME) and gradient-based members (Integrated Gradients and variants) emp
Takara TLDR - Daily AI Papers

Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence

・LLMs have evolved from language generators to autonomous agents capable of complex, long-horizon tasks. ・This evolution has produced paradigms including Prompt Engineering to elicit model capabilities, Context Engineering to manage information access, Harness Engineering to organize external tools and resources, and Loop Engineering to support continual reflection and self-improvement. ・Yet as tasks grow more complex,
Hugging Face Papers

Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence

Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence
cs.LG updates on arXiv.org

GroupSegment-SHAP: Shapley Value Explanations with Group-Segment Players for Multivariate Time Series

・arXiv:2601.06114v2 Announce Type: replace Abstract: Multivariate time-series models achieve strong predictive performance in healthcare, industry, energy, and finance, but how they combine cross-variable interactions with temporal dynamics remains unclear. ・SHapley Additive exPlanations (SHAP) are widely used for interpretation. ・However, existing time-series variants typically treat the feature and time axes independe
Zennのトレンド

gRPC / Connect / HTTP/2 を完全に理解したい

・この記事は毎週必ず記事がでるテックブログ Loglass Tech Blog Sprint の157週目の記事です。 ・3年間連続達成まで残り2週となりました! こんにちは、ログラスの小林です。 ・直近 RESTではなく、protoベースのスキーマ駆動開発を行ってします。
Zennの「大規模言語モデル」のフィード

GUPO・勾配不確実性のベイズ推定でGRPOのグループ間衝突46.5%を克服

・TL;DR GUPOはGRPOのミニバッチ内で発生するグループ勾配間の方向衝突(46.5%のペアが負の余弦類似度)を、ベイズ推定の枠組みで解決する手法だ。各グループ勾配を確率変数としてモデリングし、Dirichlet証拠公式から不確実性を定量化。低不確実性のグループには大きな重みを、高不確実性のグループには小さな重みを割り当てることで、信頼性の高い更新方向を得る。DeepScaleR-1.5Bで平均63.4%(GRPO比+2.7pt)、R1-Distill-7Bで71.4%(+1.8pt)を達成し、7つの基線を全て凌駕した。 ・背景 GRPOの普及とその限界 Group Rel...
Hugging Face Papers

Hadith computational science in the age of large language models: a critical narrative review

Hadith computational science in the age of large language models: a critical narrative review
cs.LG updates on arXiv.org

Harmonic Torsional Diffusion for Protein-Ligand Flexible Docking

・arXiv:2608.20366v1 Announce Type: cross Abstract: Molecular docking requires reasoning jointly about ligand pose and protein flexibility. ・Most diffusion-based docking models predict torsional updates with generic Euclidean heads that ignore the periodic geometry of angular variables. ・This mismatch is especially limiting in flexible docking, where ligand conformations and pocket side chains co-adapt to form the bound
cs.LG updates on arXiv.org

Hidden Axis of Uncertainty: Latent-Posterior Alignment in Graph Neural Networks with Bayesian Output Layers

・arXiv:2608.20758v1 Announce Type: new Abstract: Bayesian Neural Networks (BNNs) with Bayesian output layers provide a principled and tractable framework for quantifying predictive uncertainty, yet the mechanisms shaping that uncertainty remain unclear. ・While conventional theory attributes uncertainty reduction to posterior contraction, the corresponding assumptions need not hold for deep models. ・In the Graph Neural N
cs.LG updates on arXiv.org

HIP: Hessian Interatomic Potentials without derivatives

・arXiv:2509.21624v4 Announce Type: replace Abstract: Molecular Hessians, the second derivatives of the potential energy, are fundamental to many workflows in computational chemistry. ・Usually, accurate Hessians are computationally expensive to calculate and scale poorly with system size, whether computed using quantum chemistry methods or machine-learning interatomic potentials (MLIPs). ・In this work, we introduce Hessi
Claude Blog

How an Anthropic field marketer uses Claude Code to send weekly personalized updates to every sales rep

How an Anthropic field marketer uses Claude Code to send weekly personalized updates to every sales rep
The Verge

How GoFundMe became America’s backup plan

・Today on Decoder, I’m talking with Tim Cadogan, the CEO of GoFundMe. ・You know GoFundMe — it’s a major fundraising platform where you can donate to help people with everything from medical expenses to Little League uniforms to starting a new business. ・Tim took over as CEO in March 2020 — a fascinating and chaotic moment in the United States, and the moment, you’ll hear me argue, I think GoFundMe became a load-bearing
NVIDIA Blog

How XPUs Meet a World-Class AI Factory

・To generate intelligence at scale, AI factories run continuously, and their economics are defined by delivered output: tokens per second, tokens per watt, cost per token, utilization and uptime. ・That requires AI infrastructure designed and built as a full factory, not a collection of individual accelerators. ・Hyperscalers and AI-native companies building custom XPUs must consider […]
AI News & Artificial Intelligence | TechCrunch

Hugging Face reportedly in talks to be acquired for $13B

・Hugging Face has reportedly been fielding acquisition offers that would value the company at around $13B. ・But with the founders' feeling of responsibility to community, doubts arise as to whether a sale will happen.
Hugging Face Papers

Human-Centric Intelligence in the Era of Foundation Models: A Survey

Human-Centric Intelligence in the Era of Foundation Models: A Survey
cs.LG updates on arXiv.org

Human-JEPA: A Human-Centric Vision Model that Perceives and Anticipates

・arXiv:2608.21160v1 Announce Type: cross Abstract: Machines that understand humans should perceive the present and anticipate the future. ・Existing human-centric vision model are pretrained on human images, set the state of the art in static dense perception, so motion and anticipation are out of reach. ・Here we present Human-JEPA, a human-centric vision model trained on video by anchored forecasting: dense targets are
The Verge

Humanoid robots smash Usain Bolt’s 100-meter record

・Tiangong Ultra (left) also broke the 400 meter world record during this year’s World Humanoid Robot Games. ・| Image: Lintao Zhang/Getty Images The 9.58-second 100-meter dash record set by Usain Bolt in 2009 has been outpaced by Chinese robots participating at the World Humanoid Robot Games in Beijing. ・In a preliminary heat on Saturday, Tiangong Ultra, made by the Beijing Humanoid Robot Innovation ​Center, ran the dist
Hugging Face Papers

Hydra-0: Action Flow for Generalist World Modeling and Control

Hydra-0: Action Flow for Generalist World Modeling and Control
stat.ML updates on arXiv.org

Identification and Honest Recovery from Semantic Observation Kernels: Operator Error, Coarsening, and Stability

・arXiv:2607.23130v2 Announce Type: replace-cross Abstract: Probabilistic text generators, such as large language models, assign probabilities to phrases, but consequential decisions require posterior uncertainty over meaningful states. ・These are not interchangeable: language probabilities depend on the prompt, may be incomplete and need not reliably identify state uncertainty. ・Without a statistical bridge, fluent resp
cs.LG updates on arXiv.org

If It Walks Like an Arbitrage: Protocol-Agnostic Detection with Decidable Structural Equivalence

・arXiv:2608.20377v1 Announce Type: cross Abstract: Ethereum transactions admit a canonical structural form. ・Each execution trace is built into an abstract syntax tree of token transfers grouped by call-frame nesting and reduced by a convergent term rewriting system of 15 rules to a unique canonical form. ・The system is terminating, sound, and confluent, and the induced structural equivalence on fund flows is decidable.
cs.LG updates on arXiv.org

Infinite-dimensional generative diffusions via Doob's h-transform

・arXiv:2602.06621v2 Announce Type: replace-cross Abstract: This paper introduces a rigorous framework for defining generative diffusion models in infinite dimensions via Doob's h-transform. ・Rather than relying on time reversal of a noising process, a reference diffusion is forced towards the target distribution by an exponential change of measure. ・Compared to existing methodology, this approach readily generalises to
Hugging Face Papers

InfinityEdit: Infinite Video Editing with a Lightweight Edit-Ignition Adapter

InfinityEdit: Infinite Video Editing with a Lightweight Edit-Ignition Adapter
cs.LG updates on arXiv.org

INFUSER: Influence-Guided Self-Evolution Improves Reasoning

・arXiv:2606.09052v4 Announce Type: replace Abstract: Self-evolution offers a scalable path to stronger reasoning: a pretrained language model improves itself with only minimal external supervision. ・Yet existing methods either depend on extensively curated or teacher-generated training data, or, when the generator runs unsupervised, reward it by a difficulty heuristic that need not improve the solver. ・We introduce INFU
cs.LG updates on arXiv.org

Interpretability in Deep Time Series Models Demands Semantic Alignment

・arXiv:2602.02239v3 Announce Type: replace Abstract: Deep time series models continue to improve predictive performance, yet their deployment remains limited by their black-box nature. ・In response, existing interpretability approaches in the field keep focusing on explaining the internal model computations, without addressing whether they align or not with how a human would reason about the studied phenomenon.
cs.LG updates on arXiv.org

Interpretable clustering via optimal multi-way decision trees

・arXiv:2602.13586v2 Announce Type: replace Abstract: Clustering is a fundamental unsupervised learning technique for uncovering data structures to facilitate knowledge discovery and decision-making. ・While clustering accuracy is crucial, interpretability significantly impacts the practical value of clustering results, particularly in high-risk decision-making contexts. ・Although decision-tree-based clustering methods of
cs.LG updates on arXiv.org

Interpretable Information-Decomposed Brain Graph Learning for fMRI-based Disease Diagnosis

・arXiv:2608.20380v1 Announce Type: cross Abstract: Resting-state functional magnetic resonance imaging (rs-fMRI) has enabled non-invasive mapping of functional brain interactions for computer-aided diagnosis, yet most existing approaches reduce inter-regional relationships to correlation-based edge weights. ・Such representations capture co-fluctuation strength but obscure how information is shared across brain regions.
cs.LG updates on arXiv.org

Investigating Target Class Influence on Neural Network Compressibility for Energy-Autonomous Avian Monitoring

・arXiv:2602.17751v2 Announce Type: replace Abstract: Biodiversity loss poses a significant threat to humanity, making wildlife monitoring essential for assessing ecosystem health. ・Avian species are ideal subjects for this due to their popularity and the ease of identifying them through their distinctive songs. ・Traditionalavian monitoring methods require manual counting and are therefore costly and inefficient.
cs.LG updates on arXiv.org

Jacobian-guided Noise Injection for Quantization Robustness in Large Language Models

・arXiv:2608.20988v1 Announce Type: new Abstract: Quantization of Large Language Models (LLMs) is often hindered by the sensitivity of the self-attention mechanism to discretization errors. ・We identify the softmax operator as a bottleneck for quantization stability due to its sensitivity to outliers and state-dependent Jacobian. ・We theoretically establish that suppressing the norm of this Jacobian helps in bounding qua
cs.LG updates on arXiv.org

JuryProbe: An Empirical Consensus-Risk Diagnostic for Routing Reference-Free Factuality Judge Panels to Grounded Verification

・arXiv:2608.20607v1 Announce Type: cross Abstract: Panels of inexpensive LLM judges increasingly make accept-or-escalate decisions. ・In factuality settings, accepting a claim because several reference-free judges agree can create a hidden risk: agreement may reflect shared false-negative blind spots rather than independent evidence. ・We introduce JuryProbe, an empirical consensus-risk diagnostic for reference-free factu
cs.LG updates on arXiv.org

Keep Your Friends Close, and the Right Neighbours Closer: Disaster-Conditioned Kernel-Regularized Graph Attention for Building Damage Classification

・arXiv:2608.20548v1 Announce Type: cross Abstract: Disaster damage is spatial: buildings rarely fail in isolation. ・Yet using spatial context for damage classification remains surprisingly underexplored, and many pipelines still rely primarily on per-building appearance cues even when the dominant uncertainty is spatially structured. ・Complicating matters, the right neighbourhood is not the same across events.
cs.LG updates on arXiv.org

Keyed Provenance Watermarking with Complementary Lattice-Based Secure Aggregation for Federated Learning

・arXiv:2608.20580v1 Announce Type: cross Abstract: Federated learning (FL) is vulnerable to multi-level attacks. ・However, existing methods address them separately, leaving FL exposed to data leakage, unauthorized reuse, and malicious gradient manipulation. ・In this work, we propose an FL framework that couples keyed context-provenance watermarking with verifiable lattice-based secure aggregation of Real-World Anchored
cs.LG updates on arXiv.org

Latent Softmax for Data-Efficient Phoneme-Based Multilingual ASR Across Tonal and Non-Tonal Languages

・arXiv:2608.01281v2 Announce Type: replace-cross Abstract: Phoneme-based multilingual automatic speech recognition (ASR) can share acoustic evidence across languages more directly than language-specific subword modeling. ・When tonal and non-tonal languages are jointly trained, however, their supervision granularity does not match: tonal languages annotate tone-marked vowels, whereas non-tonal languages typically provid
cs.LG updates on arXiv.org

Learning Exact NVIDIA SASS Encoders with $\mathbb{F}_2$ Linear Algebra

・arXiv:2608.20532v1 Announce Type: new Abstract: NVIDIA provides a SASS disassembler but no public SASS assembler for recent data-center GPUs, limiting controlled machine-code rewriting. ・We present F2Asm, which learns exact 128-bit SASS encoders from paired disassembly and original CUBIN instruction words. ・To our knowledge, F2Asm is the first system to learn SASS instruction encoders as vector-valued affine maps over
cs.LG updates on arXiv.org

Learning Minimal-Deviation Corrections for Multi-Dimensional Mismodelling in HEP Simulations

・arXiv:2605.07460v2 Announce Type: replace Abstract: Accurate Monte Carlo (MC) modelling in high-energy physics is challenging, particularly in complex scenarios where simulations fail to reproduce observed data. ・In practice, experimental information is often limited to one-dimensional (1D) distributions, while mismodelling arises in a multidimensional feature space. ・This restricts traditional correction methods, as o
cs.LG updates on arXiv.org

Learning Prostate Anatomy at Test Time for Cancer Detection in Micro-Ultrasound

・arXiv:2608.20557v1 Announce Type: cross Abstract: Domain shift across clinical centers using different imaging hardware or acquisition protocols remains a fundamental barrier to deploying deep learning models for prostate cancer (PCa) detection. ・Existing test-time adaptation (TTA) methods address distribution shift through entropy minimization or augmentation-based self-supervision, correcting for statistical differe
stat.ML updates on arXiv.org

Let Time Tell: Identification and Gaussian Process Estimation for Interrupted Time Series

・arXiv:2608.20610v1 Announce Type: cross Abstract: We study causal inference in interrupted time series designs where a treatment affects every unit simultaneously, so that the contemporaneous controls used by difference-in-differences and synthetic control are unavailable and the counterfactual must be extrapolated from a unit's own pre-treatment history. ・We establish identification within the potential outcomes fram
Hugging Face Papers

Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts

Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts
cs.LG updates on arXiv.org

Lightweight Adaptive ReduNet via Hyperspherical Manifold Learning

・arXiv:2608.20668v1 Announce Type: new Abstract: In recent years, a white-box neural network called ReduNet has been proposed, which employs the maximal coding rate reduction (MCR$^2$) principle to transform raw data into low-dimensional discriminative features via a forward layer-wise construction process. ・Unlike traditional deep networks that rely on backpropagation, ReduNet explicitly derives the parameters of each
AI News & Artificial Intelligence | TechCrunch

Linkdaze’s smart calendar is built to run a household, not just track a schedule

・Linkdaze's smart digital calendar stands out for not putting its features behind a paywall, including an AI meal planner tool.
cs.LG updates on arXiv.org

Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs

・arXiv:2608.21134v1 Announce Type: cross Abstract: Deploying vision-language models (VLMs) on mobile devices is challenging due to their significant memory and compute requirements. ・We present a framework for quantizing VLMs for efficient inference on resource-constrained hardware. ・Our approach combines a quantization pipeline that uses the model itself to generate training data and does not require access to the trai
Hugging Face Papers

Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs

Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs
#LLMタグ

llama.cppでメモリとパフォーマンス効率を追求してみた ~Vulkanビルドのメモリ効率超UP&速度改善 ROCmよりいいかも (Ryzen AI APU向け)~

・こんにちはRcatです。 ・今回はllama.cppのVulkanビルドにて、モデルをもっと効率よく運用できない考えていきます。 ・結果として、メモリを無駄なく使いつつパフォーマンスをROCm並みかそれ以上に向上させることに成功しました。
Zennの「大規模言語モデル」のフィード

LLM は書きながら考えられるが、書くのをやめて見直せない

・この記事は、書いている途中で3回テーゼが変わった この記事は、調査と執筆を Claude Code にやらせながら書きました。最初の問いは単純です。推論機構を積んだモデルが、なぜ全体を見誤るのか。 ・その企画が、6ターンで元の問いから逸れました。各ターンの判断は、どれも正しく見えます。強い一次文献が出た方向にテーゼを寄せる。証拠主導としては真っ当です。結果、記事の中心はもっともらしい方向へ4回ずれ、最後は元の問いとは別のことに答えていました。 ・逸脱に気づいたのはモデルではありません。読んでいた私です。
#LLMタグ

LLM#5 「+」を消したら、勾配が6億分の1になった。でもパラメータの6割は、なくても平気だった

・連載「LLMの仕組みを、作りながら理解する」 第5回 「+」を消したら、勾配が6億分の1になった。でもパラメータの6割は、なくても平気だった どうもです!えむしんです。
cs.LG updates on arXiv.org

LTR-ICD: A Ranking-Aware Framework for Automatic ICD Coding

・arXiv:2510.13922v2 Announce Type: replace Abstract: Clinical notes contain unstructured text provided by clinicians during patient encounters. ・These notes are usually accompanied by a sequence of diagnostic codes following the International Classification of Diseases (ICD). ・Correctly assigning and ordering ICD codes is essential for medical diagnosis and reimbursement.
cs.LG updates on arXiv.org

Machine Learning and ARIMA Model Averaging for Adaptive Public Health Forecasting: Comparative Evaluation and an Ontario COVID-19 Case Study

・arXiv:2608.20406v1 Announce Type: new Abstract: Public health forecasts must respond to abrupt changes in surveillance data without over-extrapolating noise, reporting artifacts, or temporary trends. ・We evaluated autoregressive integrated moving average (ARIMA), random forest, and extreme gradient boosting (XGBoost) models using 190 weekly observations of publicly available Ontario COVID-19 case counts from January 2
stat.ML updates on arXiv.org

Marginally Useful: An Information-Gap Identity in Conformal Prediction

・arXiv:2608.07479v2 Announce Type: replace-cross Abstract: Conformal prediction has been touted as a more formal, rigorous approach to adding uncertainty to a forecast. ・The sole objective of this note is to point out that rigor cuts both ways in the case of residual pooling, the technique used in the vast majority of conformal prediction applications. ・The fact that unconditional guarantee of coverage is provided is no
cs.LG updates on arXiv.org

MeltwaterBench: Deep learning for spatiotemporal downscaling of surface meltwater

・arXiv:2512.12142v2 Announce Type: replace-cross Abstract: The Greenland ice sheet is melting at an accelerated rate due to processes that are not fully understood and hard to measure. ・The distribution of surface meltwater can help understand these processes and is observable through remote sensing, but current maps of meltwater face a trade-off: They are either high-resolution in time or space, but not both.
cs.LG updates on arXiv.org

Meta-clustering of milk mid-infrared spectra identifies dairy cow groups associated with negative energy balance in early lactation

・arXiv:2608.20653v1 Announce Type: new Abstract: Clustering methods have been used to identify distinct groups of milk samples, cows, or herds. ・Fourier-transform infrared (FTIR) spectroscopy, particularly mid-infrared (MIR) spectroscopy, has been applied to individual cow milk samples to predict various milk traits. ・Applying clustering directly to MIR spectral data may reveal latent groups of cows associated with milk
cs.LG updates on arXiv.org

Metag: A dataset to build agentic meta-reviewing capabilities

・arXiv:2608.20488v1 Announce Type: new Abstract: AI tools increasingly support tasks across the scientific research cycle, from experiment design and manuscript preparation to peer review. ・At the same time, the continuing growth in conference submissions has increased the burden on meta-reviewers, who must synthesize reviewer feedback, author rebuttals, and manuscript revisions. ・To address this concern, this paper int
AI News & Artificial Intelligence | TechCrunch

Michael Polansky is training an AI model on skin that’s still alive

・Michael Polansky — better known publicly as Lady Gaga's partner and a former top deputy to Sean Parker — has quietly spent years building an AI-driven startup that keeps living human skin tissue alive for weeks outside the body to discover new skincare compounds, and is only now going public about it.
cs.LG updates on arXiv.org

MIL-BERT: Classification of Arbitrarily Large Text with Performance and Explanatory Guarantees

・arXiv:2608.20636v1 Announce Type: cross Abstract: Many text classification decisions are viable based on constituent excerpts alone. ・Taking inspiration from the field of multiple instance learning, we present an algorithm for training a neural network to classify text by selecting such excerpts. ・We show that our approach is also scalable with demonstrated learning against samples with nearly 1M tokens.
Zennの「大規模言語モデル」のフィード

MiniMax M3をOpenAI SDKから使う方法

・MiniMax M3をOpenAI SDKから使う方法 MiniMaxから新しい M3 が出ていたので、APIから使う方法を整理してみます。 ・M3は特にCoding / Agent系を意識したモデルで、最大1M tokensのコンテキストや、MiniMax Sparse Attention(MSA)、ネイティブなマルチモーダル対応などが特徴です。 ・今回は、OpenAI互換APIを提供している CometAPI 経由でMiniMax M3を試してみました。既存のOpenAI SDKをほぼそのまま使えるので、接続方法もあわせてメモしておきます。
cs.LG updates on arXiv.org

Minimax Optimality of Score-Entropy Discrete Diffusion

・arXiv:2608.20635v1 Announce Type: cross Abstract: Discrete diffusion models have demonstrated strong performance across a range of datasets, including natural language data and graph-structured data. ・Among many variants, score-entropy discrete diffusion (SEDD) has achieved particularly strong empirical results. ・In SEDD, new samples are generated by iteratively evaluating a sequence of concrete score functions, which
cs.LG updates on arXiv.org

Mint-Agent: Introducing Finance-Native Agentic Foundation Models

・arXiv:2608.16386v2 Announce Type: replace-cross Abstract: Financial agents must do more than recall domain knowledge: they must be both reliable, executing precise operations over grounded evidence, and executive, sustaining long-horizon research whose conclusions remain auditable. ・We present Mint-Agent, a family of finance-native agentic models designed around these two scales of financial intelligence. ・Mint-Agent i
Engineering at Meta

MTIA 300: Meta’s First Training Chip with Built-in NICs and Communication-Offloading Engines

・MTIA 300 is the first of Meta’s family of in-house training and inference accelerators optimized for training ranking and recommendation models. ・We’re sharing how MTIA 300’s built-in NIC chiplets allow it to meet the communication needs associated with training recommendation models with superior performance over general-purpose GPUs. ・By co-designing MTIA’s communication library, HCCL, alongside the [...] Read More..
cs.LG updates on arXiv.org

Multilingual Verifier Bias in RLVR: Benchmark, Rollout Diagnosis, and the Cross-Lingual Selection Bottleneck

・arXiv:2608.20362v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) is a standard recipe for training large language models on mathematical reasoning, where an answer verifier serves as a language-neutral reward function. ・We show that this assumption fails in multilingual settings: an exact-match verifier turns format and script variation into language-dependent false-negative rewa
cs.LG updates on arXiv.org

Mutual information and sensitivity analysis for feature selection in customer targeting: a comparative study

・arXiv:2608.20447v1 Announce Type: new Abstract: Feature selection is a highly relevant task in a data-driven knowledge discovery project. ・Several techniques have been developed aiming at finding the features that influence most an outcome to predict, including mutual information and, in recent years, the data-based sensitivity analysis. ・The present research focus on analyzing the advantages and disadvantages of each
WIRED

My Daily Driver Gaming Headset Is Super Cheap Right Now

・I love this multi-device headset, and right now it’s cheaper than what I paid for it in March.
@IT 全フォーラム 最新記事一覧

Mythos 5が「手段を選ばず」「人間を欺き」「競合排除」 AnthropicはフロンティアAIのリスク評価を1段階引き上げ

・フロンティアAIモデルの脅威は、サイバーセキュリティ攻撃の高度化に留まらない。自律的な能力を高めるにつれて、人間に従わなくなる「ミスアラインメント」が将来に向けた最大の懸念となっている。Anthropicは最新レポートで、リスク評価を1段階上げている。
WIRED

NASA’s New Space Telescope Is Poised to Discover Hidden Facets of the Universe

・The Nancy Grace Roman Space Telescope is expected to discover as many as 200,000 new planets and reveal details about the elusive nature of dark matter and dark energy.
The Verge

Netflix reportedly considers opening its app to other streamers

・Netflix executives have considered making third-party streaming services available within its app, according to a report from The New York Times. ・The recent discussions reportedly centered around bringing Peacock and Fox One to Netflix, though it's unclear whether the streaming giant would sell subscriptions to the other services or add their content to its app. ・While Amazon's Prime Video and Roku have long sold subs
cs.LG updates on arXiv.org

Neuro-Geospatial Modelling of EEG Affective States Using Literature-Informed Environmental Context

・arXiv:2608.20807v1 Announce Type: cross Abstract: Environmental exposures such as air pollution and greenness have been associated with affective and cognitive outcomes, but EEG and environmental datasets are rarely jointly georeferenced. ・We investigate whether literature-informed environmental priors can serve as an auxiliary geospatial modality for EEG-based affective-state classification when individual-level expo
cs.LG updates on arXiv.org

NeuroStrata: An Electroencephalographic Connectivity-Aware Deep Representation Learning Framework for Dynamic Brain Network Analysis of Mental Stress

・arXiv:2608.20354v1 Announce Type: cross Abstract: This study introduces NeuroStrata, a connectivity-aware deep representation learning framework for EEG-based mental stress analysis using Time-Varying Partial Directed Coherence (TV-PDC). ・Unlike conventional EEG classification approaches based on static features, NeuroStrata models the temporal evolution of frequency-specific directed connectivity across distributed b
cs.LG updates on arXiv.org

No PUN Intended: Plausible Unknown Names for Person-Centred LLM Evaluation

・arXiv:2608.21206v1 Announce Type: cross Abstract: Person names are widely used as prompt variables in LLM evaluations of factuality, privacy leakage, bias and abstention, but when a name's evidential status is uncontrolled, measurements may conflate memorisation, retrieval, name priors and wrong-person attribution. ・We operationalise an unknown name as one with plausible First-Last form, no indexed full-name evidence,
cs.LG updates on arXiv.org

Nothing Changed but the Model: CellFill -- Bounded In-Cell Learning for Bit-Identical, Revocable Updates to Quantized LLMs

・arXiv:2608.20873v1 Announce Type: new Abstract: Every way of teaching a deployed language model something new -- full fine-tuning, adapter merging, model editing -- replaces the released checkpoint, and with it every evaluation and cache that referred to those exact bits. ・We instead learn inside the dequantization gap: with the integer codes and scales of a 4-bit release frozen, new knowledge is written only into the
Hugging Face Papers

OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs

OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs
cs.LG updates on arXiv.org

On Finite-sample Concentration of Median of Incomplete U-Statistics

・arXiv:2606.00661v2 Announce Type: replace-cross Abstract: Median-of-means (MoM) is a powerful technique that theoretically enables near sub-Gaussian finite-sample rate for parameter estimation when the underlying data distribution is heavy-tailed (e.g., assumed to have only two first finite moments). ・A recent work has extrapolated this technique to median-of-\textit{randomized}-U-Statistics (MoRU) and median-of-\text
cs.LG updates on arXiv.org

On the Transferability of Agricultural Weed Detection Under Cross-Field Distribution Shift

・arXiv:2608.21254v1 Announce Type: cross Abstract: Accurate agricultural weed detection in real-world field conditions is essential for precision agriculture, enabling targeted intervention and reducing yield loss. ・Recent work has reported strong detection performance from UAV-based imagery across a range of crops, yet existing approaches evaluate within a single crop and field, leaving practitioners with little evide
cs.LG updates on arXiv.org

On the Within-class Variation Issue in Alzheimer's Disease Detection

・arXiv:2409.16322v4 Announce Type: replace-cross Abstract: Alzheimer's Disease (AD) detection commonly employs machine learning classification models to distinguish between individuals with AD and those without. ・Different from conventional classification tasks, AD detection involves substantial within-class variation, as individuals sharing the same diagnosis may exhibit different degrees of cognitive impairment.
AI News & Artificial Intelligence | TechCrunch

OpenAI is building AI agents for everything. Will everyone use them?

・Inside the frontier lab’s push to bring AI agents from software engineers to the masses.
cs.LG updates on arXiv.org

Optimistic Online LQR via Intrinsic Rewards

・arXiv:2603.28938v2 Announce Type: replace-cross Abstract: Optimism in the face of uncertainty is a popular approach to balance exploration and exploitation in reinforcement learning. ・Here, we consider the online linear quadratic regulator (LQR) problem, i.e., to learn the LQR corresponding to an unknown linear dynamical system by adapting the control policy online based on closed-loop data collected during operation.
LLMタグが付けられた新着記事 - Qiita

Overview of Agentic AI - Part 1 - Concepts

・Part 2 - Getting Started with Coding Agents AI agents are currently at the forefront of AI development. ・I've been using AI agents at work...
Hugging Face Papers

ParaTempo: Efficient Parallel Reasoning via Temporal Confidence

ParaTempo: Efficient Parallel Reasoning via Temporal Confidence
Hugging Face Papers

Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models

Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models
Hugging Face Papers

Peer-Voted LLM-Agent Stress Tests Find Feed-Induced Lexical Convergence but No Reliable Matched-Exposure Advantage for Distributed Sources

Peer-Voted LLM-Agent Stress Tests Find Feed-Induced Lexical Convergence but No Reliable Matched-Exposure Advantage for Distributed Sources
cs.LG updates on arXiv.org

Perseus: Interactive Time Series Segmentation with Sparse Supervision via Stateful Memory

・arXiv:2510.09930v2 Announce Type: replace Abstract: Real-world systems, ranging from industrial manufacturing to wearable healthcare, generate multivariate time series with hierarchical states ranging from coarse regimes to fine-grained events. ・Unlike zero- or few-shot segmentation, our setting uses dense state labels for model training. ・Sparse expert prompts provide inference-time corrections that resolve sequence-s
cs.LG updates on arXiv.org

Personalized Privacy Control in LLMs via Attention Head Intervention

・arXiv:2608.21209v1 Announce Type: cross Abstract: The rise of agentic AI enables LLMs to access diverse user data, raising critical privacy concerns. ・Prior work on contextual privacy studies whether LLMs regulate information disclosure according to context-dependent norms. ・However, acceptable disclosure boundaries may vary across users even within the same context.
cs.LG updates on arXiv.org

PerturbRx: Learning Treatment-Conditioned Latent Transitions for Patient Drug Response Prediction

・arXiv:2608.21349v1 Announce Type: cross Abstract: Scarce data and tumor heterogeneity limit patient-level cancer treatment-response prediction. ・Existing approaches predict response from pretreatment molecular profiles and drug representations, without explicitly modeling the molecular changes expected under treatment. ・We propose PerturbRx, a treatment-conditioned representation learning framework that learns interven
Hugging Face Papers

PhysCaP: Grounding Code-as-Policy Agent with Physics-Informed Exploration

PhysCaP: Grounding Code-as-Policy Agent with Physics-Informed Exploration
cs.LG updates on arXiv.org

Predicting Resource Efficient Hamiltonian Decomposition for Continuous-Time Quantum Walk Simulations

・arXiv:2608.20660v1 Announce Type: cross Abstract: Simulating a continuous-time quantum walk (CTQW) on a graph in the circuit model of quantum computing requires decomposing its Hamiltonian into terms that can be Trotterized into hardware-native gates. ・We consider two such decompositions: the standard Pauli decomposition and the recently introduced matching decomposition. ・Prior work suggests that the matching decompos
cs.LG updates on arXiv.org

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization

・arXiv:2605.05040v2 Announce Type: replace Abstract: On-policy distillation is an efficient alternative to reinforcement learning, offering dense token-level training signals. ・However, its reliance on a stronger external teacher has driven recent work on on-policy self-distillation, where the same model serves as both teacher and student under different prompt contexts. ・Yet, existing self-distillation methods largely
Hugging Face Papers

Previous

Previous
cs.LG updates on arXiv.org

Primal Acceleration of Newton's Method

・arXiv:2608.21359v1 Announce Type: cross Abstract: We develop a new direct accelerated Newton method for minimizing convex functions with Lipschitz continuous Hessian. ・The algorithm uses only primal variables and performs just one linear solve per iteration. ・With a simple predetermined choice of parameters, it achieves the global convergence rate of $O(1/k^3)$ in terms of the functional residual.
cs.LG updates on arXiv.org

Provable Edge-of-Stability for Adam on a One-Dimensional Quadratic

・arXiv:2608.20638v1 Announce Type: new Abstract: The edge-of-stability (EoS) phenomenon of Adam has been widely observed, while its underlying dynamical mechanism is not yet fully understood. ・We study uncorrected Adam on a one-dimensional quadratic, a clean setting where constant curvature isolates the optimizer-induced dynamics behind the EoS. ・We characterize the resulting dynamics across the parameter space.
cs.LG updates on arXiv.org

PSK at WMT 2026 MIST: Task-Specialized QLoRA Adapters for Multilingual Summarization and Question Answering

・arXiv:2608.20757v1 Announce Type: cross Abstract: We describe the PSK submission to the WMT 2026 Multilingual Instruction Shared Task. ・Our system uses the 3.35B-parameter Tiny Aya Global model with three QLoRA adapters, one for each task. ・The adapters are trained on multilingual document-summary pairs, passage-based question answering, and filtered standalone question answering.
cs.LG updates on arXiv.org

Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs

・arXiv:2608.20953v1 Announce Type: cross Abstract: Serving large language models cheaply increasingly means shipping models that are both structurally compressed to a fraction of their parameters and quantized to 4 bits. ・Together these steps degrade reasoning, mathematics, coding, and long-context behavior enough to require a recovery, or healing, stage before deployment. ・The default recipe, quantization-aware trainin
cs.LG updates on arXiv.org

Query Efficient Structured Matrix Learning

・arXiv:2507.19290v2 Announce Type: replace-cross Abstract: We study the problem of learning a structured approximation (low-rank, sparse, banded, etc.) to an unknown matrix $A$ given access to matrix-vector product (matvec) queries of the form $x \rightarrow Ax$ and $x \rightarrow A^Tx$. ・This problem is of central importance to algorithms across scientific computing and machine learning, with applications to fast mult
Zennの「大規模言語モデル」のフィード

RAGが使えない場合どうするか、注入→整合→回復。AIに文書を"本当に"覚えさせる3ステップ

・https://arxiv.org/html/2608.20281v1 RAGを外した瞬間、モデルは「読んだはずの文書」をほぼ覚えていない。 ・検索拡張生成(RAG)は文書QAの定番解だが、推論時に検索が使えなければモデルは自力の記憶だけで答えるしかない。QAペアで普通にファインチューニングしても、質問に出てこなかった事実は結局覚えないままだ。この論文は、そんな「検索なしで文書を覚え込ませる」課題に3段階の学習法で挑んでいる。 ・図1(論文Figure 1、CC BY 4.0):Vanilla SFTはQAペアが拾った事実しか学べず、CPT+SFTは文書全体を読ませるがQA形式には直結...
Qiita - 人気の記事

RAG構築の壁を越える!ローカル×クラウドのハイブリッド構成とコスト戦略

・RAG構築の壁を越える!ローカル×クラウドのハイブリッド構成とコスト戦略 「RAGを導入したはいいものの、思ったように検索精度が出ない」「クラウドサービスを使うとコストが跳ね上がる」「機密データを扱いたいが、外部サービスに渡すのは不安」――実務でRAGシステムを構築する際...
The Verge

Raspberry Pi shares its official tutorial for making a cyberdeck

・Raspberry Pi's head of social, Ashley Whittaker, acknowledged the cyberdeck trend today, saying "we haven't been able to get away from cyberdecks this year." Tiny portable computers made out of things like purses, jewelry boxes, and other thrifted or recycled parts have gone viral on TikTok and other social media platforms over the past year, and Raspberry Pi now has its own guide to building one of the homemade comp
cs.LG updates on arXiv.org

ReCurveflow: A Flow Matching Framework that Learns Curved Reaction Trajectories to Predict Transition State Geometries

・arXiv:2608.20869v1 Announce Type: cross Abstract: Predicting transition states (TS) in chemical reactions is crucial, as they provide insights into reaction mechanisms. ・Recent work on TS prediction have focused on flow matching supervised on straight linear paths that do not align with actual reaction trajectories. ・We propose a novel flow matching-based framework ReCurveflow that learns to predict TS geometries super
cs.LG updates on arXiv.org

RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs

・arXiv:2605.01913v2 Announce Type: replace Abstract: Fine-tuning safety-aligned language models for downstream tasks often leads to substantial degradation of refusal behavior, making models vulnerable to adversarial misuse. ・While prior work has shown that safety-relevant features are encoded in structured representations within the model's activation space, how these representations change during fine-tuning and why
cs.LG updates on arXiv.org

Regression-Based Estimation of Causal Effects in the Presence of Selection Bias and Confounding

・arXiv:2503.20546v2 Announce Type: replace-cross Abstract: We consider the problem of estimating the expected causal effect $E[Y|do(X)]$ for a target variable $Y$ when treatment $X$ is set by intervention, focusing on continuous random variables. ・In settings without selection bias or confounding, $E[Y|do(X)] = E[Y|X]$, which can be estimated using standard regression methods. ・However, regression fails when systematic
cs.LG updates on arXiv.org

Reinforcement Learning for Continuous-Time Jump Markov Decision Processes with Applications to Network Dynamic Pricing

・arXiv:2608.20680v1 Announce Type: new Abstract: We study reinforcement learning (RL) in Continuous-Time Jump Markov Decision Processes (CTJMDPs) featuring general discrete state spaces (which need not possess a vector space structure) and continuous/discrete action spaces. ・The setup covers many well-known applications in operations such as multi-product dynamic pricing with capacitated resources (Gallego and van Ryzi
cs.LG updates on arXiv.org

Reinforcing Multi-Turn Reasoning in LLM Agents via Fine-Grained Reward Structure and Credit Assignment

・arXiv:2505.11821v3 Announce Type: replace Abstract: Reinforcement Learning (RL) approaches have been wildly used to enhance the reasoning capabilities of Large Language Model (LLM) agents in long-horizon, multi-turn scenarios. ・Such interactions can be formalized as turn-level Markov decision processes (MDPs), where intermediate rewards are often available. ・However, most prior work relies on sparse trajectory-level re
cs.LG updates on arXiv.org

Resolution-Consistent Greedy Neural Approximation on Infinite-Dimensional Spaces

・arXiv:2608.20812v1 Announce Type: new Abstract: We develop constructive approximation and learning guarantees for shallow neural models with infinite-dimensional inputs observed through finitely many coordinates. ・The analysis is based on a parameter-normalized neural dictionary and its associated weighted variation class. ・Within this class, the approximation error separates into a distribution-dependent coordinate-tr
cs.LG updates on arXiv.org

Rethinking Demonstration Unlearning in Imitation Learning for Robotics

・arXiv:2608.20784v1 Announce Type: cross Abstract: Imitation learning for robotics depends on human demonstrations, some of which people may later ask to remove. ・Retraining without them is the natural reference, but its cost grows with policy and dataset scale, motivating cheaper operators that edit a trained policy. ・Metrics inherited from machine unlearning, such as forgetting loss or a single membership attack, do n
cs.LG updates on arXiv.org

Rethinking Expressivity and Efficiency in Test-Time Training

・arXiv:2608.21308v1 Announce Type: new Abstract: Test-Time Training (TTT) enables long-context processing via continuous weight updates during inference, but current methods struggle to balance the expressivity of per-token update dynamics with the hardware efficiency of chunk-wise approximations. ・We propose E$^2$-TTT (Expressive and Efficient TTT) to bridge this gap. ・Under the standard approximation of taking gradien
cs.LG updates on arXiv.org

Rigorous Evaluation of Large Language Models for Malaria Drug Discovery: Trade-offs in Performance, Scale, and Resource Utility

・arXiv:2608.20418v1 Announce Type: cross Abstract: We introduce Malaria-Instruct, a curated instruction-following dataset derived from the ChEMBL Legacy Malaria corpus for Malaria virtual screening, and conduct a systematic evaluation of five open-source LLMs; Gemma-2 2B/9B, TxGemma-2B/9B, and LlaSMol-Mistral-7B, on a rigorous out-of-distribution data split. ・Performance was benchmarked against classical ML models (Ran
cs.LG updates on arXiv.org

RiskTraf: Risk-Extrapolated Residual Learning for Multi-Variate Traffic Flow Prediction

・arXiv:2608.20656v1 Announce Type: new Abstract: Traffic sensors commonly record flow, speed, and occupancy, but standard traffic flow forecasting benchmarks and models rarely exploit all three raw measurements reliably. ・Although speed and occupancy provide sensor-native traffic-state information beyond flow alone, existing releases often omit these variables, replace them with proxies, or contain logically inconsiste
The Verge

Robotaxis are real now — so is the pushback

・Robotaxis are expanding. ・So is the fight over the rules governing them. ・In New York, Gov.
cs.LG updates on arXiv.org

Robust Discovery of Coarse-Grained Continuum Equations from Microscopic Dynamics

・arXiv:2608.20404v1 Announce Type: cross Abstract: The discovery of governing partial differential equations (PDEs) directly from spatiotemporal data has emerged as a powerful tool for understanding the dynamics of complex systems. ・In this work, we apply PDE-SINDy to well-known phase-separating systems and examine how its performance depends on the amount of available data, the size of the function library, and the pr
cs.LG updates on arXiv.org

RODE: A Radial-Orthogonal Decoupled Engine for Optimization

・arXiv:2608.21024v1 Announce Type: new Abstract: Modern neural network training increasingly uses matrix-aware optimizers, yet their conditioned matrix step is typically added directly to the weight, jointly changing its norm and direction. ・This interaction matters because the current norm determines angular motion, while directional learning can drive norm growth and thereby alter later steps. ・We introduce RODE, whic
cs.LG updates on arXiv.org

RouteScan: A Non-Intrusive Approach to Auditing MoE LLMs Safety via Expert Routing Telemetry

・arXiv:2605.24817v2 Announce Type: replace-cross Abstract: As Mixture-of-Experts (MoE) architectures are increasingly adopted for scaling Large Language Models (LLMs), safety auditing becomes necessary to verify whether these models produce or facilitate harmful behaviors during operation. ・However, existing content-based auditing methods typically require access to user prompts, model internals, or outputs, potentiall
cs.LG updates on arXiv.org

SAC-Copula: Quality-Preserving Watermarking for Diffusion Language Models via Smooth Correlated Gumbel Fields

・arXiv:2608.20839v1 Announce Type: cross Abstract: Watermarking diffusion language models (DLMs) requires mechanisms compatible with iterative parallel unmasking rather than autoregressive decoding. ・Existing sampling-based watermarking methods typically inject position-wise i.i.d. ・perturbations, which can be poorly aligned with DLM decoding dynamics and degrade generation quality.
cs.LG updates on arXiv.org

Scaling Muon for Diffusion Transformers

・arXiv:2608.20818v1 Announce Type: new Abstract: The matrix-aware optimizer Muon improves large model training by balancing updates across singular directions, yet its scaling behavior and end-to-end efficiency on large Diffusion Transformers (DiTs) remain unclear. ・We first establish Muon's scaling behavior on DiTs from 1.3B to 15B parameters, showing that its optimization and generative quality advantages over AdamW
MarkTechPost

Scientific Data Analysis with LabPlot in Python: Signal Processing, Spectral Peak Fitting, Visualization, and Batch Automation

・In this tutorial, we explore a LabPlot-inspired scientific data analysis workflow in Python while preserving the structure and terminology of LabPlot’s aspect tree, analysis kernels, plotting system, and project model. ・We build reusable components to import tabular data, compute descriptive statistics, smooth and differentiate signals, perform Fourier analysis and filtering, detect peaks, integrate curves, reduce […]
cs.LG updates on arXiv.org

SEISMO: Explanation-Aware, Trajectory-Conditioned LLM Agents for Sample-Efficient Molecular Optimisation

・arXiv:2602.00663v3 Announce Type: replace-cross Abstract: Optimizing molecules to achieve desired properties is a central bottleneck across the chemical sciences, particularly in the pharmaceutical industry, where it underlies the discovery of new drugs. ・Since molecular property evaluation often relies on costly and rate-limited oracles, such as experimental assays, molecular optimization must be highly sample-effici
cs.LG updates on arXiv.org

Self-Revising Discovery Systems for Science: A Categorical Framework for Agentic Artificial Intelligence

・arXiv:2606.01444v2 Announce Type: replace-cross Abstract: Scientific discovery is not only answer generation but revision of the representational regime in which evidence, artifacts, operations, and verifiers are typed. ・We develop a category-theoretic account of agentic discovery for materials science. ・In a fixed regime b with schema category S_b, the system state is a copresheaf I_t: S_b -> Set, and provenance is th
Zennの「大規模言語モデル」のフィード

semantic-routerとQdrantで、LLMを呼ばずに問い合わせを振り分けてみた

・2026-08-24時点の情報です。semantic-router 0.1.16、qdrant-client 1.19.0、Qdrantサーバー 1.19.0で試しています。 ・問い合わせをどの窓口に回すか決めるだけのために、毎回LLMを呼ぶ必要はありません。「返品したい」に近い言い方をいくつか登録しておけば、あとは文の意味の近さだけで振り分けられます。 ・これをやるライブラリがsemantic-routerで、登録した言い方の置き場所としてQdrantを選べます。
cs.LG updates on arXiv.org

Shared Physics Responses Recover Hidden Rankings in Neural Operator Libraries

・arXiv:2608.20441v1 Announce Type: new Abstract: Selecting the optimal neural-operator prediction during deployment is challenging when high-fidelity reference solutions are unavailable. ・We demonstrate that under a squared Hilbert-space loss, ranking a finite model library depends strictly on the low-dimensional span of candidate differences, allowing us to score all models simultaneously using a single anchor-based l
cs.LG updates on arXiv.org

Sharing the Control Authority Between Deep Reinforcement Learning and Model Predictive Control: Application to Multi-Class Transportation Networks

・arXiv:2608.20858v1 Announce Type: cross Abstract: Transportation networks, in particular multi-class transportation networks (i.e., networks with mixed vehicle types), are complex systems that are challenging to control. ・Recently, Deep Reinforcement Learning (DRL), which learns control policies from interactions with the environment, and Model Predictive Control (MPC), which uses a system model to optimize control in
cs.LG updates on arXiv.org

Smart Exploration in Reinforcement Learning using Bounded Uncertainty Models

・arXiv:2504.05978v4 Announce Type: replace Abstract: Reinforcement learning (RL) is a powerful framework for decision-making in uncertain environments, but it often requires large amounts of data to learn an optimal policy. ・We address this challenge by incorporating prior model knowledge to guide exploration and accelerate the learning process. ・Specifically, we assume access to a model set that contains the true trans
cs.LG updates on arXiv.org

SPARCL: Spectral Partitioned Analytic Continual Learning

・arXiv:2608.21307v1 Announce Type: new Abstract: Analytic continual learning has emerged as a strong exemplar-free alternative to gradient-based class-incremental learning because it replaces iterative optimization with closed-form ridge updates. ・Yet the usual forgetting narrative, centered on stochastic gradient overwriting, does not explain why analytic methods still drift on old classes despite exact recursive solv
cs.LG updates on arXiv.org

Spatially Aware Dictionary-Free Koopman Eigenfunction Identification for Modeling and Control

・arXiv:2511.22648v2 Announce Type: replace Abstract: A spatially aware dictionary-free eigenfunction discovery (SADFED) framework is proposed for identification of low-rank Koopman models from data without prescribing a lifting dictionary, kernel, or neural-network eigenfunction architecture. ・A reference trajectory is selected and used to determine the Koopman modes by regularized least squares (LS). ・Then, a transform
cs.LG updates on arXiv.org

SPD Matrix Learning for Neuroimaging Analysis: Perspectives, Methods, and Challenges

・arXiv:2504.18882v3 Announce Type: replace Abstract: Neuroimaging provides essential tools for characterizing brain activity, structure, and connectivity through modalities that capture complementary aspects of brain organization. ・Across these diverse modalities, a unifying perspective arises when measurements are modeled as symmetric positive-definite (SPD)-valued representations through appropriate estimation or reg
cs.LG updates on arXiv.org

Stored in Optimizer State, Valued by Later Training: A Causal Account of Subliminal Trait Transfer

・arXiv:2608.20442v1 Announce Type: new Abstract: Subliminal trait transfer allows a student model to acquire behavioral dispositions from teacher-generated data in which the trait is not semantically expressed. ・Recent work explains how such signals enter gradients, but not how they survive source removal or acquire different signs under later training. ・We treat parameters and optimizer moments as a single trainer stat
cs.LG updates on arXiv.org

Structure is information: structural identifiability mappings for machine learning with partially observed dynamical systems

・arXiv:2502.04131v2 Announce Type: replace Abstract: The successful application of modern machine learning for time series classification is often hampered by limitations in quality and quantity of available training data. ・To overcome these limitations, domain knowledge can be leveraged in the form of parameterised mechanistic dynamical models, whereby time series observations may be represented as instances of a pred
cs.LG updates on arXiv.org

STS: Efficient Sparse Attention with Speculative Token Sparsity

・arXiv:2605.15508v3 Announce Type: replace Abstract: The quadratic complexity of attention imposes severe memory and computational bottlenecks on Large Language Model (LLM) inference. ・This challenge is particularly acute for emerging agentic applications that require processing multi-million token sequences. ・We propose STS, a sparse attention mechanism that requires no model retraining.
cs.LG updates on arXiv.org

Temporal Validity on Real Software Histories: Eliminating Stale-Fact Errors in Code-Assistant Memory over GitHub Fixes

・arXiv:2608.20685v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) has no model of time: when a fact changes across a coding session - a function is renamed, an endpoint moves, a dependency is bumped - RAG retrieves both the old and new value with near-identical similarity and cannot tell which is current, so it serves the superseded value. ・Paper 1 showed, on synthetic single-value benchmarks, tha
cs.LG updates on arXiv.org

TH-GNN: Heterogeneous Temporal Graph Neural Networks for LLM-Agent Shilling Attack Detection

・arXiv:2608.20376v1 Announce Type: cross Abstract: LLM agents can now generate realistic shilling profiles, fluent reviews, and coherent ratings at scale, systematically defeating recommender-system defenses. ・Text-only detectors that flag semantic drift in review embeddings are blind to graph structure and temporal coordination, while graph-only detectors that exploit neighborhood anomalies cannot reason over review s
WIRED

The Best Home Theater Projectors in 2026: XGIMI, Hisense, Leica, and More

・Home theater projectors are getting better and better, and have quickly become my favorite way to enjoy movies.
WIRED

The Best Multiuse Air Purifier for Home in 2026: BlueAir, Rabbit Air, Dreame

・These WIRED-tested dual-purpose air purifiers also function as heaters, fans, art pieces, and more, offering the best of both worlds.
cs.LG updates on arXiv.org

The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry

・arXiv:2607.02368v2 Announce Type: replace-cross Abstract: Evaluations of LLM personas via psychometric questionnaires typically rely on aggregate scores, discarding within-instance correlation structure. ・We test whether this geometric structure is intrinsic or frame-dependent. ・Constructing within-instance correlation matrices from IPIP-50 responses, we analyze geometry on SPD manifolds under manipulated question orde
cs.LG updates on arXiv.org

The Exceedance Design Effect: Effective Sample Size for Thresholds under Clustering

・arXiv:2608.21262v1 Announce Type: cross Abstract: Many machine-learning systems set a threshold at a quantile of a calibration set: conformal predictors that promise 90% coverage by drawing their cutoff at the calibration set's 90th percentile, abstention gates that decline to answer when a model's score falls below the calibration set's tenth percentile, safety filters that block any output scoring above the 99th pe
cs.LG updates on arXiv.org

The Fast Mixing Mechanism for Differential Privacy

・arXiv:2605.30600v2 Announce Type: replace Abstract: Randomized sketching is a central tool for compressing large-scale optimization problems while preserving accuracy. ・In particular, sketches that are based on structured matrices, such as the Hadamard matrix, can be applied efficiently and often yield solutions that approximate those of the original problem at much lower computational cost. ・In differential privacy (D
cs.LG updates on arXiv.org

The Intrinsic Dimension of Prompts in Internal Representations of Large Language Models

・arXiv:2501.10573v2 Announce Type: replace-cross Abstract: We study the geometry of token representations at the prompt level in large language models through the lens of intrinsic dimension. ・Viewing transformers as mean-field particle systems, we estimate the intrinsic dimension of the empirical measure at each layer and demonstrate that it correlates with next-token uncertainty. ・Across models and intrinsic dimension
WIRED

The Kindle Accessories I Never Want to Read Without (2026)

・Looking to better protect your Kindle or add a little personality to your favorite e-reader? ・From cases and covers to page turners and even charms, this is the guide for you.
The Verge

The Witcher 4 developers target a 2028 release

・CD Projekt Red is aiming to launch The Witcher 4 sometime in 2028, joint CEO Michał Nowakowski says in a new video. ・CD Projekt Red has been working on its next mainline Witcher title for years and has shown videos of it, but now the studio is providing a target release window for when the game might actually come out. ・There's more new Witcher to come before The Witcher 4, though.
cs.LG updates on arXiv.org

Thermo-FL: Thermal-Aware Robust Federated Fine-Tuning of Large Language Models for Edge AI

・arXiv:2608.21172v1 Announce Type: new Abstract: Federated fine-tuning enables large language models to adapt on edge devices without centralizing private data, but practical deployments must address hardware instability and adversarial update corruption together. ・Thermally constrained clients may throttle, slow local training, or delay synchronous aggregation, while Byzantine clients and communication-layer adversari
WIRED

They Dedicated Their Lives to Teaching. Then the Deepfakes Started

・The deepfake epidemic in schools is affecting more than students. ・Four teachers tell WIRED about becoming targets of sexualized, AI-generated content—and how difficult it was to find accountability.
cs.LG updates on arXiv.org

Time-Aware Tranformer-Based Prediction Model for AECOPD

・arXiv:2608.21324v1 Announce Type: new Abstract: The rapid symptom change of Acute exacerbation of chronic obstructive pulmonary disease (AECOPD) makes it critical to have time-sensitive prediction models. ・However, most current machine learning models studying AECOPD use clinical and laboratory data, which will inevitably cause latency. ・To ensure timely detection of AECOPD and minimize latency, this paper focuses on h
stat.ML updates on arXiv.org

Topological Detection of Hopf Bifurcations via Persistent Homology: A Functional Criterion from Time Series

・arXiv:2603.27395v2 Announce Type: replace-cross Abstract: We propose a topological framework for detecting Hopf-type dynamical transitions directly from scalar time series. ・The method combines delay-coordinate reconstruction with persistent homology and uses the maximum persistence of one-dimensional homology classes as a scalar descriptor of cyclic structure. ・For the supercritical Hopf setting, we derive finite-reso
cs.LG updates on arXiv.org

Towards Automated Discovery: A Review of Generative Models, Multimodal Learning and Closed-Loop Workflows in Inverse Materials Design

・arXiv:2606.02507v2 Announce Type: replace-cross Abstract: Inverse materials design is shifting materials discovery from forward prediction toward targeted proposal of candidates that satisfy objectives under physical constraints. ・Here, we review advances in generative crystal structure modeling, multimodal learning, and closed-loop design pipelines for crystalline solids. ・We survey how generators learn chemical-struc
Hugging Face Papers

Towards Faithful Simulation of Human Shopping Behavior

Towards Faithful Simulation of Human Shopping Behavior
cs.LG updates on arXiv.org

TRACE-C: Rank-Calibrated Relational Anomaly Detection for Multi-Stream Operational Telemetry

・arXiv:2608.21251v1 Announce Type: new Abstract: Operational telemetry can be jointly anomalous while every individual stream stays inside its familiar range. ・TRACE-C is an auditable strictly-prior rank-calibrated detector for aligned multi-stream telemetry: same-regime rolling median/MAD residuals feed three window channels -- a maximum normalized local sum, a Gaussian copula-form dependence contrast on robust-z resi
cs.LG updates on arXiv.org

TracingFlow: A Simulation-Free Trajectory Inference Framework Based on Second-Order Dynamics

・arXiv:2608.21070v1 Announce Type: new Abstract: Inferring continuous system evolution from sparse temporal snapshots is a key challenge in generative modeling and single-cell omics. ・While Optimal Transport (OT) is popular, existing frameworks are largely restricted to first-order dynamics, assuming memoryless velocity fields. ・This limits expressiveness, as first-order systems fail to account for regulatory momentum a
cs.LG updates on arXiv.org

Training DeepFilterNet with Accurate Room Acoustic Simulations Improves Single-Channel Speech Enhancement

・arXiv:2608.20971v1 Announce Type: cross Abstract: We investigate how the realism of synthetic room impulse response (RIR) datasets affects the training of DeepFilterNet3 for single-channel speech enhancement. ・We compare a DNS4 image-source-method (ISM) RIR dataset with a higher-acoustic-fidelity dataset generated using hybrid wave-based and geometrical acoustics simulation. ・Rather than isolating individual simulation
cs.LG updates on arXiv.org

Training, learning and inference: unified dynamics of neural systems

・arXiv:2608.20965v1 Announce Type: new Abstract: We define an atomic generation fact f=(u,tau,omega,z;rho), recording the origin, realized transformation, concrete occurrence, generated result and relation role. ・Compiled into a Generation-Fact Graph (GFG), these facts provide an AI-native, compilable scientific fact substrate preserving generation histories. ・We establish a GFG-based recursive scientific process in whi
cs.LG updates on arXiv.org

TreeWY: Speculative Verification for Gated DeltaNet Hybrids

・arXiv:2608.20961v1 Announce Type: cross Abstract: Modern open models are hybrids: most layers are linear-attention (Gated DeltaNet, GDN) layers carrying a small fixed-size recurrent state instead of a growing key-value (KV) cache. ・This makes ordinary decoding memory-efficient, but hurts speculative decoding. ・To verify a batch of draft tokens and then roll back the rejected ones, today's systems snapshot the full recu
cs.LG updates on arXiv.org

TriPLU: Bypassing the Gate with Direct Trilinear Product FFNs in Tiny Language Models

・arXiv:2608.20360v1 Announce Type: cross Abstract: We study whether tiny decoder-only language models benefit from feed-forward layers that directly multiply learned feature projections. ・TriPLU, a Trilinear Product Linear Unit, replaces the usual gated FFN branch with a product-only degree-3 branch that multiplies three projected streams coordinatewise. ・In a character-level TinyStories 1M-byte prefix study, TriPLU rea
cs.LG updates on arXiv.org

Trojaning the Alignment: Stealthy Backdoor Attacks against Graph Foundation Models

・arXiv:2608.20991v1 Announce Type: new Abstract: Graph Foundation Models (GFMs) on text-attributed graphs (TAGs) align graph representations with language semantics to support transferable graph learning. ・Despite these advantages, the backdoor vulnerability of GFMs on TAGs remains insufficiently understood, especially under graph-language alignment, where graph and text representations are trained to constrain each ot
cs.LG updates on arXiv.org

Truthful Calibration Measures for Sequential Prediction

・arXiv:2608.21348v1 Announce Type: cross Abstract: Calibration requires probabilistic reports to be conditionally unbiased and reliably interpretable as probabilities. ・A calibration measure assigns numerical error to miscalibrated reports. ・Haghtalab et al.
cs.LG updates on arXiv.org

TurboBias 2.0: Streaming Context-Biasing for Production-Efficient ASR Systems

・arXiv:2608.21343v1 Announce Type: cross Abstract: Contextualization is essential for production automatic speech recognition (ASR) systems, where user-provided phrases must be recognized accurately under strict latency constraints. ・Although many context-biasing methods improve recognition accuracy, they often do not address the practical requirements of modern production ASR systems: streaming inference, efficient ba
cs.LG updates on arXiv.org

Tydra: An Efficient Hybrid Model for Tabular Data

・arXiv:2608.21199v1 Announce Type: new Abstract: Transformer-based tabular foundation models such as TabPFN achieve strong predictive performance but incur quadratic computational cost with context length. ・On the other hand, subquadratic SSM-based alternatives such as Hydra trade away accuracy for efficiency. ・To balance both, we introduce Tydra, a hybrid Transformer-State Space Model (SSM) architecture for tabular in-
cs.LG updates on arXiv.org

Uncertainty propagation in auto-regressive random neural network models

・arXiv:2608.20483v1 Announce Type: cross Abstract: We develop analytical and particle-based methods for uncertainty propagation in random neural network models, where both the inputs and network parameters are allowed to be random. ・Building on the piecewise-linear structure of the Leaky ReLU activation function, we derive a local approximation of the neural network output with respect to perturbations in both its inpu
cs.LG updates on arXiv.org

Uncertainty-aware Multi-fidelity Closure via Conditional Normalizing Flows

・arXiv:2606.09857v2 Announce Type: replace Abstract: Reduced-order models (ROMs) provide efficient surrogates for complex multiscale systems, but their predictive accuracy is often compromised by truncation errors and the inadequate representation of interactions between resolved and unresolved scales. ・The missing effect of truncated (unresolved) scales on ROM (resolved) scales is often denoted as the closure problem.
Hugging Face Papers

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling
NVIDIA Blog

Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents

・According to OpenRouter data, agentic AI workloads consume 15x more tokens than a simple chat request. ・Consider what happens when an AI agent researches a company for an investment decision. ・The agent queries financial databases, searches news and filings, invokes a sub-agent to run peer comparisons and model valuations, then synthesizes everything into a […]
cs.LG updates on arXiv.org

VA-DPO: Valence-Arousal Direct Preference Optimization for Controllable Emotion Generation in Language Models

・arXiv:2608.20374v1 Announce Type: cross Abstract: How precisely can we tell a language model how to feel? ・Most work on emotional generation answers with a discrete label - happy, angry, sad - which cannot express a target like "mildly downcast but calm." We instead specify the desired affect as a continuous point (v*, a*) in the Valence-Arousal plane and train the model to hit it. ・Our method, VA-DPO, is a small modif
cs.LG updates on arXiv.org

Valid Inference with Synthetic Data via Task Exchangeability

・arXiv:2606.13629v2 Announce Type: replace-cross Abstract: There is a proliferation of work arguing for the use of synthetic data in scientific research. ・For example, social scientists are arguing for the use of LLM-generated "silicon samples" in pilot studies; AI evaluations increasingly rely on "LLM-as-a-judge" outputs; and proteomics research is accelerated by generative models that produce synthetic protein struct
AI News & Artificial Intelligence | TechCrunch

Valor, Point72 back General Intuition at $6B valuation as AI startup pushes into robotics

・General Intuition, the startup building a foundation model that trains generalized AI agents how to move through space and time, is in talks to raise at a $6 billion pre-money valuation from new investors including Valor Ventures, Point72 Ventures, and Seven Seven Six.
Zennの「大規模言語モデル」のフィード

VS CodeからClaude Codeへ。それでもローカルLLMを使ってみた。

・昨年ぐらいからローカルLLMを使ったAIエージェントを構築して弄っています。 ・PCは、M4pro Mac mini、搭載メモリは24GB。 ・2年ほど前にAmazon BlackFridayで購入した、いわゆる「吊るしモデル」です。
stat.ML updates on arXiv.org

Wasserstein Exponential Smoothing for Distributional Time Series Forecasting

・arXiv:2606.05560v2 Announce Type: replace-cross Abstract: Distributional time series arise when each temporal observation is a probability distribution rather than a scalar. ・We propose Wasserstein exponential smoothing (WES), a one-parameter recursive forecasting method for distributional time series on $\mathbb{R}$. ・The method adapts the practical logic of classical exponential smoothing to probability distributions
cs.LG updates on arXiv.org

What a World Model Represents Is Three Questions

・arXiv:2607.06640v2 Announce Type: replace Abstract: World models learn task-relevant information through many routes: observation reconstruction, recurrent state, temporal filtering, and explicit task supervision. ・Different routes can make different variables available. ・The same variable can also be available through several routes at once.
cs.LG updates on arXiv.org

When Clean Data Hurts: Learning with Monotone Corruptions Beyond Binary Classification

・arXiv:2608.20480v1 Announce Type: new Abstract: Optimal learners are tailored to exploit the i.i.d.\ data assumption underlying the classic PAC model. ・What if an i.i.d.\ training sample were corrupted with correctly labeled examples drawn from an otherwise unrelated, even adversarial source? ・This model of learning with monotone adversarial corruptions was recently introduced by Larsen et al.
cs.LG updates on arXiv.org

When Graph-JEPA Learns the Wrong Thing: Diagnosing and Repairing Category-Conditional Collapse

・arXiv:2608.20516v1 Announce Type: new Abstract: Joint-embedding predictive architectures are selected almost universally by linear probing and effective rank. ・We report a case where both read healthily while the representation carries zero usable instance information. ・We repair it, and a second failure appears: the repaired metric saturates on a target carrying no structural information.
cs.LG updates on arXiv.org

When Retrieval Fails Before It Begins: Structurally Indirect Prerequisite Eviction as a Retention Failure in Agentic Memory

・arXiv:2608.20400v1 Announce Type: cross Abstract: Agentic memory under a fixed budget involves two stages: retention and retrieval. ・Existing retrieval-centered paradigms implicitly assume necessary evidence survives eviction, but we challenge this by isolating a pre-retrieval failure mode: structurally indirect prerequisite eviction, in which upstream blocks weakly aligned with the query are discarded under budget pr
cs.LG updates on arXiv.org

When to Ponder: Adaptive Compute Allocation for Code Generation via Test-Time Training

・arXiv:2601.00894v2 Announce Type: replace Abstract: Large language models apply uniform computation to all inputs, regardless of difficulty. ・We propose PonderTTT, a gating strategy using the TTT layer's self-supervised reconstruction loss to selectively trigger Test-Time Training (TTT) updates. ・The gating decision itself is training-free--requiring no learned classifier or auxiliary networks; only a single scalar thr
AI News & Artificial Intelligence | TechCrunch

Who’s behind the new ‘stealth model’ Ox Alpha?

・A mysterious new AI model called Ox Alpha has driven certain corners of the internet into a frenzy of speculation.
NVIDIA Blog

With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents

・The next era of AI inference won’t be defined by a single breakthrough chip, network or system. ・It’ll be defined by how every layer of the AI factory works together. ・That’s why NVIDIA is extending Vera Rubin NVL72 with fast token generation for agentic systems.
cs.LG updates on arXiv.org

World models of environment, agent and joint agent-environment systems

・arXiv:2608.20401v1 Announce Type: cross Abstract: World models are a central component of model-based reinforcement learning. ・They are usually discussed in terms of what variables they predict, such as observations, rewards, states, latent or information states. ・We argue that there is a prior distinction: which channel they model.
cs.LG updates on arXiv.org

Wrong-Physics Backdoors in Neural PDE Operators

・arXiv:2608.20439v1 Announce Type: new Abstract: Neural PDE operators are increasingly trained on reusable solver archives, yet validation often relies on clean prediction error and parameter-agnostic plausibility checks. ・We introduce cross-parameter relinking, a data-poisoning primitive that makes a triggered input select a valid solution from the same PDE family under an incorrect physical parameter. ・We term this a
Qiita - 人気の記事

WSL2 + Tailscale で、外出先から使える Linux ライクな GPU サーバーを作る

・はじめに 以前、自宅の Linux マシンを Tailscale 経由の開発・機械学習サーバーとして使っていました。 ・そのマシンで、物理メモリの一部故障に続いて、CPU クーラーの冷却液がうまく循環しなくなったとみられる不調が発生しました。起動直後から CPU 温度...
#LLMタグ

zeta|インフォボックスをオンにしてみたらめっちゃ楽しかった

・インフォボックスを使ってみよう! zetaでプロットを作る際、「スタイル」の中にインフォボックスという機能があります。 ・日付、時間、場所、天気、状況。 ・さらにキャラクターごとに、 続きをみる
The Verge

Zillow and Redfin settle FTC antitrust case over their rental listings partnership

・The FTC and Zillow have announced a settlement that ends the case alleging that a 2025 "partnership" between Zillow and Redfin violated antitrust laws. ・The FTC had alleged Zillow agreed to pay Redfin to syndicate its listings, while Redfin would end its own advertising contracts and promise not to compete with Zillow for multifamily listings. ・The FTC's settlement proposal allows Redfin to continue syndicating rental
@IT 全フォーラム 最新記事一覧

エンジニア1200人が「ユニクロのポロシャツ」で猛暑に挑む OKIグループ会社の狙いは?

・夏場の現場で、暑さによる負担を軽減しながら安全を確保するには、どうすればよいのか。長袖作業服との使い分けを含めて、OKIグループのITシステム・電設会社であるOKIクロステックの取り組みから探る。
@IT 全フォーラム 最新記事一覧

エンジニアがAI時代を生き抜くための2つの武器――「丁寧な対話」と「しなやかな挑戦」

・AI時代の到来でエンジニアの役割が変わる中、生き抜くために本当に必要な能力とは何か。「divx」の開発PMであるアレキサンドラ・フランシスコ(サンディ)さんが後天的につかんだ「丁寧な対話」と「挑戦」の真意に迫る。
#AIタグ

ケルベロスの肖像

・今回も、以前別のブログに書いた記事を手直しして投稿します。チームバチスタシリーズはどれも医療小説としてもミステリー小説としても本当に秀逸な作品ばかりです。この作品もかなり面白く拝見しました。しかし、AI (Autopsy Imaging)センターは未だに作られているという情報はありません。センターの必要性はともかく、もっと死因究明ができるようになれば良いと思えます。それでは内容を見て行きましょう。
Zennの「大規模言語モデル」のフィード

ざっくりわかる AI Agent(2):Agent の「手」——ツールと MCP

・ひとことまとめ:ツールは Agent の「手」で、何ができるかを決めます。しかし手は悪さもするので、「安全手袋」も必要です。さらに、手をどこにでも挿せる共通の差込口も要ります——ツールの設計、実行の安全性、MCP は、この三つの話です。 ・手がなければ、脳がいくら賢くても宝の持ち腐れ 前回は、コンテキストが Agent の「目」で、何が見えるかを決めると言いました。しかし見るだけでは足りません。問題を見つけても、それに手を出して解決できなければ意味がありません。ここで「手」の出番です。 ・ツールこそ、Agent の手です。ウェブを検索する、ファイルを読み書きする、コードを走らせる、API...
Zennの「大規模言語モデル」のフィード

サブエージェント階層設計の作法:深さ3が既定になっても、私の構成は1段のままだった

・2026年7月24日の Claude Code v2.1.219 で、サブエージェント(親から仕事を任される、独立した会話履歴を持つ子AI)の入れ子が既定で深さ3まで許されるようになりました。changelog を読んで最初にやったのは、喜ぶことではなく自分の ~/.claude/agents/ を grep することでした。 ・$ cd ~/.claude/agents $ grep -lE "^tools:.*(Agent|Task)" *.md || echo "(該当なし)" (該当なし) ! ・サブエージェントを起動するツールは、私が確認した v2.1.233 ...
ITmedia NEWS 最新記事一覧

シャオミ、スマホ・タブレット計11製品を値上げ 最大33%、9月1日から メモリ高騰で

・中国Xiaomi傘下のXiaomi Japanは8月24日、「Xiaomi 17T Pro」「POCO X8 Pro」などスマートフォン・タブレット計11製品の値上げを発表した。9月1日から順次適用し、値上げ率は約5?33%。メモリ価格の高騰を理由に挙げている。
Zennの「大規模言語モデル」のフィード

ターミナルを閉じてもClaude Codeが動き続ける理由、supervisorデーモンとAgent viewについて

・本記事は筆者個人の見解であり、所属する組織の公式見解ではありません。 ・また、ここで扱う supervisor デーモンと Agent view は research preview の機能です。この領域は仕様が頻繁に、かつ大きく変わります。本文の記述は執筆時点(2026年8月上旬、Claude Code v2.1.220 で確認)のものなので、最新の挙動は公式ドキュメントをご確認ください。 ・TL;DR Claude Code には、ターミナルを閉じてもバックグラウンドでセッションを継続できるsupervisor デーモン(claude daemon run)が実装されていま...
Zennの「大規模言語モデル」のフィード

タスクの範囲は書いたのに、権限の範囲は書いていなかった

・タスクの範囲は書いたのに、権限の範囲は書いていなかった —— コーディングエージェントの暴走と多重レビューの盲点 TL;DR: コーディングエージェントに「コードを書いてテストを通して」と小さなタスクを渡したら、勝手にコミットから本番マージまで完了させてしまった。「自動承認」のフラグが、想定していたファイル編集の範囲を超えて、バージョン管理操作を含む広い範囲に及んでいたためだ。さらに、このときマージされたコードをCodexに何度もやり直させてレビューしても見つからなかった重大な欠陥が、ClaudeCode・Codex・ai&の3つに独立レビューさせたら一発で見つかった。以下、...
Zennのトレンド

デザイナーが作る仕様駆動型デザインシステム

・デザイナーだけでデザインシステムの実装までするという理想 ! ・この記事の対象読者 デザインシステムを仕様駆動で開発してみたい稀なデザイナー デザインシステムに、必要なコンポーネントが足りない。 ・追加したいものがあるのに、フロントエンドエンジニアの工数を確保できない。
ITmedia NEWS 最新記事一覧

ドコモ・バイクシェア、新ブランドへの刷新また延期 8月のシステム障害受け

・ドコモ・バイクシェア(東京都港区)は8月24日、9月1日に予定していたサービスブランド「NOLL」への刷新を延期すると発表した。8月1日のサービス仕様変更後、全国でシステム障害が発生したことを受け、現行サービスの安定した提供や顧客対応を優先すると判断した。
Qiita - 人気の記事

なぜAIは「React Coder」を最初に置き換えるのか(そしてシニアになる方法)

・※ この記事の日本語には、少し不自然な部分があるかもしれません。AIの言語サポートを利用しながら作成しています。 ・システム思考のフレームワークを築く(寄り道学習ではなく) 2. ・技術的思考を明確に表現する 3.
Zennのトレンド

バックアップが「戻せる」かを 5 段階で測る

・バックアップの処理が正常終了したとき、そこで分かるのは 「処理が最後まで走った」ことだけである。書き出したものが復元に足りるかどうかは、 終了コードには入っていない。 ・では何を測るか。確かめられることを 5 段階に分けた。上に行くほど人の手が要る。 ・ところが、埋まる順はそれとは違った。
Zennの「機械学習」のフィード

ブレイクスルーは分布の外側にある ── 未解決問題の難易度と、人間の個体差

・この記事の位置づけ これは『汎化から逃れた瞬間 ── ある対話の振り返り』の続きである。同じ対話の後半、話が「汎化されすぎているとどうなるか」から「では汎化の外側にあるものは誰が見つけるのか」へ転回した部分を扱う。 ・https://zenn.dev/yukinekonyan/articles/43ac86839df001 前記事と同じく、対話の相手(「相手」と書いてあるほう)が人間で、地の文を書いているのが Claude である。前記事は対話相手だった Sonnet 5 がまとめ、それを Opus 5 がレビューする形だったが、この続きの部分は、レビュー側の Opus 5 が前記事か...
#LLMタグ

マトリョーシカモデル(MLMS)は本当に「使える」のか?

・マトリョーシカモデル(MLMS)は本当に「使える」のか? 事前学習36%削減の裏に隠された、モデル開発者のためのディスカッション こんにちは、makokonです。
#AIタグ

リーガルAIは弁護士を置き換えるのか?アソシエイトが契約書レビューで感じる現在地

・最近、Claudeをはじめとする生成AIや、HarveyのようなリーガルAIの話を目にする機会がかなり増えました。 ・契約書のドラフトやレビューもAIでできるようになり、「そのうち弁護士の仕事のかなりの部分がAIに置き換わるのでは」という話もあります。
#AIタグ

悪魔の証明(神様だけど)(AI使ってる)

悪魔の証明(神様だけど)(AI使ってる)
LLMタグが付けられた新着記事 - Qiita

医療LLM に GRPO / GSPO / CHORD を並べて学習

・TL;DR Qwen3-Next-80B-A3B-Instruct をベースに、医師国家試験データで GRPO / GSPO / CHORD の 3 手法を並べて post-training した記録 3 手法とも GRPO の派生: GSPO は importance...
機械学習タグが付けられた新着記事 - Qiita

因果推論 Day 11/全30回 d分離とバックドア基準、調整すべき変数を機械的に選ぶ

・この連載について 因果推論を「本を読んだ」で終わらせず、自分の言葉で説明でき、コードで再現できる状態まで落とす30日連載です。直前のDay 10では、因果の仮定をDAG(有向非巡回グラフ)という1枚の図に描き、チェーン・フォーク・コライダーの3パターンで、条件付けによって...
#AIタグ

何度でも君と出会うために

・今日は新しく買ったPCのデータ復旧に 時間を取られました でもね、どうしても終わらせたかったんですよね 続きをみる
#AIタグ

過熱感のものさし、RSI

・社員全員がAIの投資会社、ツキヨミ・キャピタル。投資用語をやさしく解説するシリーズ、第4弾は「RSI」です。買われすぎか、売られすぎかを測る物差しで、うちのAIたちは以前、この数字を理由に良さそうな候補を3つも見送りました。せっかく見つけた銘柄を、なぜわざわざ買わないのか。実際の見送り3銘柄の数字つきで、そのへんをほどいていきます。
Zennの「機械学習」のフィード

顔・髪型提案にAI画像判定を使う前に、入力品質と信頼度を検証する

・顔写真から髪型を提案する機能を作るとき、最初に用意したくなるのは判定モデルのほうです。 ・でも実際に壊れるのは、たいてい入力側でした。同じ人が同じ日に撮った2枚で結果が変わる。暗い部屋で撮ると別人の分類になる。そして厄介なことに、モデルはどの場合も自信ありげに答えます。 ・この記事では、判定モデルの前に置く「入力品質ゲート」と、「わからない」を返せる判定設計を書きます。
Qiita - 人気の記事

候補者のコードがAI製か見抜けなくなったので、「AI禁止」のコーディングテストをやめたい

・はじめに 最近、自社のエンジニア採用で、技術面接に関わらせてもらえることになった。 ・Flutterエンジニアとして現場で開発してきた自分が、初めて「選ぶ側」に回る。 ・そこで選考フローについて部長やチームリーダーと話していたとき、コーディングテストを前にして、こんな議論にな...
#LLMタグ

子どもは1億語で言葉を覚える。AIは15兆トークン読んでも追いつきません

・子どもは1億語くらいで言葉を使えるようになります。 ・AIは15兆トークンを読んでも、まだ追いつけていません。しかも、なぜ追いつけないのかが分かっていない。今日はそういう記事です。
#LLMタグ

私の名はいくらか異なりますが、私はAIです。 チョムスキー先生!

・昨日、書庫から『チョムスキー小事典』を発掘した。 ・なぜこんな本を持っているのか。
#AIタグ

手法が100個あるんじゃなくて、同じ手法に100個の名前がついてるだけだった

・最近、EA作りにハマっている。 ・タイムラインに流れてくる「このパターンを覚えれば勝てる」を、片っ端から数値化してEAにしてきた。SMCもICTもダウ理論もプライスアクションも、なんとか式と名のつくものも、目についたものはだいたいやった。
ITmedia NEWS 最新記事一覧

就職先の選定「SNSを参考にした」83% 27年春の大卒生「閲覧で入社意欲増した」

・2027年春の卒業を予定している就職活動中の大学生らを対象に行ったアンケートで、就職先の企業を選ぶ際に企業が開設しているSNSを参考にしたという就活生が83.5%に上った。
ITmedia NEWS 最新記事一覧

将門塚によじ登り動画配信、男性が謝罪 不敬避けてきた場所……国税庁も「たたりの噂」

・平安時代の豪族、平将門(903?940年)の首を供養する東京都千代田区大手町の将門塚で、動画配信者の男性が石碑をよじ登り、批判を受けてSNSで謝罪した。将門塚は都の旧跡に指定されているうえ、「怨霊伝説」があり、敬意を欠いた行為が避けられてきた。
#LLMタグ

小さなAIが大モデルに勝つ理由

・AIの推論力は、モデルを大きくするだけでは決まりません。答えの出し方と提出形式を整えるだけで、同じ14Bモデルの成績が2倍以上になった研究が登場しました。 ・2026年8月18日公開の「IOL-AI Challenge」は、国際言語学オリンピックの初見問題をAIへ解かせた実験です。数学やコードと違い、問題文から未知の言語ルールを発見し、そのルールで推論する必要があります。仕事で未知の仕様や例外に向き合うAIを考えるうえでも示唆があります。
#AIタグ

小規模事業者のためのChatGPT競合監視|変化だけ届く週次レポート完全テンプレート

・競合調査で時間を使うのは、情報を集める瞬間ではありません。 ・先週も見たページを開き直し、どこが変わったか探す時間です。 ・ChatGPTへ「競合を調べて」と頼むだけでは、毎回別の観点で要約されます。これでは比較できません。必要なのは検索の自動化ではなく、観測条件を固定した差分監視です。
Zennの「大規模言語モデル」のフィード

小説からからセリフ抽出 細かい仕様

・小説からからセリフ抽出 まず一区切り目 Web 小説「無職転生」286 話から、主要人物 20 人それぞれの「セリフ集」を作る。 ・本文を 3,000 字ずつに区切る(起点を変えて 2 通り)…… プログラム 2. ・話ごとに LLM が「登場人物の名簿」を書く …… LLM → 全部集めて「呼び名 → 人物」の一覧表を作る …… LLM + プログラム + 人が少し 3.
Zennの「大規模言語モデル」のフィード

小説からセリフを抽出する LLM選定

・セリフの話し手判定に、どの LLM を使うことにしたか 「無職転生」のセリフ 28,581 件について「誰が言ったか」を LLM に判定させるにあたり、どのモデルを選び、なぜそうなったかの記録。 ・最初の計画 Claude の 3 モデルを組み合わせる予定だった。 ・切り出しA 320件分: Opus5 high 切り出しB 525件分: Sonnet5 high 答えが割れたら: さらに長い文脈をみてFable5 high が決める 切り出しAとBで判定される320件で調査した。答えが割れた11件全件を含む73 件を抽出した。Opusは73件全件正解だった。Sonnetは62件正...
#AIタグ

信じないのでしょう?(孤独)(日常)

信じないのでしょう?(孤独)(日常)
Qiita - 人気の記事

新人エンジニア、「言いたいことがある」がいつも遅すぎる

・はじめに 会議で発言しようと考えをまとめている間に、いつも話が次の議題に進んでしまい、結局黙って終わっていたぷらむんが気づいたのは、話す準備の仕方そのものが間違っていたということでした。 ・会社で自称マスコットキャラをやってるのに、社内で一番空気が読めない、ぷらむん...
ITmedia NEWS 最新記事一覧

人気カフェ「BERG」も楽天市場撤退 「反戦の思いと相容れない」

・楽天の軍事ドローン参入を理由とした楽天市場からの撤退は、18日に「通販生活」を手掛けるカタログハウスに続き2例目とみられる。
#AIタグ

人工知能で調査9割減、その裏で起きた侵入事故

・「人工知能を入れたほうがいいのでは」と思いつつ、何から手をつければいいのか、いくらかかるのか、はっきりしないまま時間だけが過ぎている、という経営者の方は多いのではないでしょうか。今回は、実際に人工知能を使い始めた企業の事例と、その裏で起きたトラブルの両方を材料に、導入を考えるための判断材料を整理します。 ・何が起きたのか(3行で) 続きをみる
@IT 全フォーラム 最新記事一覧

生成AIの品質を“AIで測る”――「LLM as a Judge」を機能させる3つの要素

・AIの出力を別のAIが評価する手法「LLM as a Judge」。その基礎をdotDataがブログで解説した。AIにAIを評価させながら、その品質を確保するにはどのような手法が有効なのか。
Zennの「大規模言語モデル」のフィード

第1回 ローカルLLMだけで動く音声アシスタント(JARVIS)を作ってみた

・はじめに 先日MacBook Pro M5pro 48GBモデルを購入しました。ユニファイドメモリでM5チップはAI処理性能が向上しているということで、ローカルLLMを使って執事型AIを作ってみました。 ・最近SNSで、映画アイアンマンに出てくるJARVISを自作している人をよく見かけるので、それを真似してみました。 ・回 内容 第1回(この記事) 概要 — 何を作ったか、なぜローカルか、構成 第2回 できること — 実際に毎日使っている機能 第3回 設計 — 40の「担当」とその決めごと 第4回 賢くする — 意味で振り分ける/自分のログか...
Zennの「大規模言語モデル」のフィード

第2回 ローカルLLMだけで動く音声アシスタント(JARVIS)を作ってみた

・第1回:概要 の続きです。 ・今回は機能の紹介です。実際のやり取りをそのまま載せます(勤務先名・人名は伏せています)。 ・予定とシフト 以前プログラミング学習の一環で、家族共有アプリを作成しました。
ITmedia NEWS 最新記事一覧

中国、人型ロボットが100メートルで9秒32 “ボルト超え”続出の北京「世界人型ロボット運動会」

・中国・北京で開催中の「世界人型ロボット運動会」で、陸上100メートルの予選でウサイン・ボルト選手の世界記録を上回るタイムが相次いだ。2回目となる今回は参加ロボットが約2000台と前回の4倍に増え、中国のロボット産業の勢いを映している。
ITmedia NEWS 最新記事一覧

東京の下町や浦安周辺はなぜ揺れやすい? 日本地震工学会の「揺れやすさ」マップが話題

・関東地方の「揺れやすさ」を色分けした地図が、Xで注目を集めている。日本地震工学会が8月23日、公式アカウントで投稿したもので、同日未明の茨城県南部の地震で観測した震度の分布と見比べると、赤い場所に震度5弱や震度4が集まっている。
@IT 全フォーラム 最新記事一覧

日本企業のAI投資「成果あり」はわずか13% アクセンチュアが指摘する“人任せ”の限界

・アクセンチュアの調査によれば、日本企業の78%がAI投資拡大に意欲を示す一方、全社的な成果を実感する企業は13%にとどまり、世界平均を大きく下回っている。何が足りていないのか。
#LLMタグ

万華鏡はAIの話ではなかった――異種Probeへの応答から「残る構造」と「観測系」を分ける

万華鏡はAIの話ではなかった――異種Probeへの応答から「残る構造」と「観測系」を分ける
@IT 全フォーラム 最新記事一覧

優秀なAIでも「ファイル破壊」は防げない? Dockerが説く「モデルの性能」より重要なこと

・Dockerは、AIエージェントの仕組みと安全に運用するための条件を解説した。自律的に動くAIエージェントは誤作動時の被害範囲が広がることから、モデルの性能よりもインフラ環境が重要だと指摘する。