ai Trend Report

Dashboard へ戻る
Date: 20260820 Articles: 376 Scope: curated summary

あなたのアイデアを、今すぐ形に。

公開先に迷ったら、WebFileBinで一発公開。

HTMLをドラッグ&ドロップするだけで、すぐ公開できます。

3
High impact
5
Mid impact
368
Signal watch

なぜこのサイトを作ったのか

私たちそれぞれが個別にAIを使って情報収集し、同じような3行要約を作るたびに、世界中で膨大な電力と計算リソースが消費されています。 本プロジェクトは、あらかじめ広範な情報を取得・集約しておくことで、個別のAI実行回数を減らし、地球環境(GPU/TPU負荷)に配慮した効率的な情報収集を目指す実験的なダッシュボードです。

Domain filters
Star filters
#AIタグ

💶 MONEY #53|8月19日、がん治療ワクチンの第3相治験で大きな進展。Moderna株急騰の先に見える「AI × mRNA × 個別化医療」

・8月19日、がん治療ワクチンの第3相治験で大きな進展。Moderna株急騰の先に見える「AI × mRNA × 個別化医療」 2026年8月19日。 ・医療と株式市場の両方で、非常に興味深いニュースがありました。 ・ModernaとMerckが共同開発している個別化mRNAがん治療「intismeran autogene(V940 / mRNA-4157)」が、悪性黒色腫(メラノーマ)を対象とした第3相臨床試験で主要評価項目を達成したと発表されました。
#AIタグ

💻 TECHNOLOGY #17|AI × mRNA。がん治療は「病気に合わせる」から「人に合わせる」時代へ

・|AI × mRNA。がん治療は「病気に合わせる」から「人に合わせる」時代へ 2026年8月19日。 ・医療の未来を考えるうえで、非常に興味深いニュースがありました。 ・ModernaとMerckが共同開発している個別化mRNAがん治療「intismeran autogene(V940 / mRNA-4157)」が、悪性黒色腫(メラノーマ)を対象とした第3相臨床試験で主要評価項目を達成したと発表されました。
#AIタグ

AIは人間を忘れるのか。DAY 0から3日、5つの未来が分かれ始めた。

・AIに2,000ドル以内の資金を渡し、未来を予測して米国株を選ばせる実験を始めた。 ・まだ、たった3日しか経っていない。 ・それでも面白いことが起き始めた。
#AIタグ

🐾住友ファーマ(4506)2027年3月期1Q決算分析|米国2製品は強い。焦点は2Qの上方修正と復配判断

・🐱住友ファーマの決算について、単なる決算要約ではなく、現在の株価水準から今後1〜3カ月、3〜6カ月で投資妙味があるかという視点で、AIを活用して分析しました。 ・公式資料との照合や数値の確認、株価への織り込み、バリュエーション、今後のカタリストとリスクまで詳しく整理しているため、今回は有料記事としていますが、結論・目次は無料で読めます🐾住友ファーマを調べている方の参考になればうれしいです🐈 続きをみる
Zennの「大規模言語モデル」のフィード

GPUメモリの壁を進化戦略で越える — Agentic ESOptが示す長期エージェント学習の現実解

・こんにちは、Yuhi AI Labsです。京都大学・京都産業大学の情報学研究科出身のOBを中心に活動する、関西初のオープンソースコミュニティです。大学発ベンチャーとしての展開を目指しています。 ・📄 Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements arXiv:2608.17310 Zheng, Z. ・(2026) 長期タスクのエージェント学習は、GPUメモリがどうしてもネックになる。「うちにはそんなクラスタは無い」と思いながら読み始めたら、想像以上に痛...
Zennの「大規模言語モデル」のフィード

RAGと何が違う?AIエージェントに同一性と忘却を与える複合記憶アーキテクチャ

・はじめに:RAGは「記憶」ではなく「検索」である AIエージェントを長期運用する際、多くの人が最初に試すのが RAG(Retrieval-Augmented Generation) です。 ・ドキュメントをチャンクに切り、ベクトルDBに突っ込み、類似度検索(top-k)でコンテキストに注入する——非常に強力でポピュラーな手法ですが、エージェントの「記憶」としては致命的な限界があります。 ・セッションを跨ぐ同一性(アイデンティティ)が維持できない 古いノイズや矛盾した情報が永遠に残り続ける(忘却がない) 知識と主観(エピソード)が混ざり、人間のような学習・再統合が起きない 本記事では、...
Zennの「大規模言語モデル」のフィード

RTX 5060 Tiを追加購入する直前に、RTX 4070 Ti SUPER 16GB単体でQwen 3.8 27Bを試してみた

・こんにちは。miharubaの池谷です。 ・最近、ローカルLLM用にGPUをもう1枚買い足そうか、本気で悩んでいました。 ・今使っているGPUはRTX 4070 Ti SUPER 16GBです。Qwenの27Bモデルを長いコンテキストで使おうとすると、モデル本体だけでなくKVキャッシュにもVRAMが必要になります。16GBでは足りなくなる場面がありそうなので、RTX 5060 Ti 16GBを追加して、2枚のGPUにモデルやKVキャッシュを分散させる構成を考えていました。
#AIタグ

思考の螺旋 ー自我はいかに形成されるかー自我とはなにか

・AIは人間を変えるのか 「Undue Persuasion」という言葉への疑問から考える、人間の判断と責任 神戸大学が行った生成AIと道徳的判断に関する研究結果の記事を読んで、最初に気になったのは、「undue persuasion(不当な説得)」という言葉だった。研究では、18~30歳の若者56人と65歳以上の高齢者74人、合わせて130人を対象に、いわゆる「トロッコ問題」を使った実験が行われた。ChatGPTから反対の立場の意見を提示された結果、転換機問題では32.31%、歩道橋問題では36.92%の参加者が、最初の判断を覆したという。 ・この結果は、生成AIが人間の道徳的判断にかなりの影響を与え得ることを示している。特に、認知機能が低下している高齢者ほどAIの反論によって判断を変えやすかったという点は、慎重に考える必要があるだろう。生成AIが認知的な負担を軽くする一方で、場合によっては「不当な説得」に対する脆弱性を生み出す可
Latent.Space

[AINews] Death of Params: Z.ai CEO Jie Tang on GLM 5.3 and the new Post-training Scaling Law

[AINews] Death of Params: Z.ai CEO Jie Tang on GLM 5.3 and the new Post-training Scaling Law
@IT 全フォーラム 最新記事一覧

「パスワードを変えてもまだ危ない」 Cookie盗難524億件超え、必要な“もう一つ”の対策

・NordVPNは、情報窃取型マルウェアによるCookie流出の調査結果を発表した。世界で524億件超、日本国内でも約1.89億件の流出を確認。認証Cookieが盗まれると、パスワードやMFAを回避してアカウントを悪用される恐れがあり、注意が必要だ。
ITmedia NEWS 最新記事一覧

「害悪すぎる」「バカ迷惑」──嫌われまくる“AI営業電話”、今すぐ取れる自衛策は

「害悪すぎる」「バカ迷惑」──嫌われまくる“AI営業電話”、今すぐ取れる自衛策は
#AIタグ

「生成AIパスポート」は誰が運営しているのか――登記からたどる生成AI活用普及協会(GUGA)とその周辺

「生成AIパスポート」は誰が運営しているのか――登記からたどる生成AI活用普及協会(GUGA)とその周辺
@IT 全フォーラム 最新記事一覧

「脱GitHub」の受け皿となるか Cursorがコードホスト機能「Origin」β版を公開

・Cursorは、コードホスティング機能「Origin」の初期β版を提供開始した。リポジトリ管理やPR作成、GitHubとの双方向リアルタイム同期に対応し、AIエージェントとコードを同一環境で扱えるようになる。
@IT 全フォーラム 最新記事一覧

「風水害は重大リスク」と8割が認識、一方「出社・帰宅の判断基準が未完成」6割 対策が進まない理由は

・JX通信社が企業のBCP担当者を対象とした台風・風水害に関する調査結果を発表した。8割超が風水害を重大なリスクと認識する一方、出社や帰宅の判断基準を完全に策定している企業は約4割にとどまることが分かった。
#AIタグ

【AI】存在しない女性による政治活動

・立憲民主党の支持者を思われる男性が、架空の政治家候補をAIで作成し問題となっている。 ・作成者は党員だった…AIで“存在しない女性”を作り政治活動、「立憲民主党」を名乗り出馬まで示唆アカウント名に「立憲民主党」を入れた「中村りほ」というX(旧Twitter)アカウントがいま、物議を醸している。 ・市news.nicovideo.jp 続きをみる
#LLMタグ

【GPT】人間の皮膚を完全再現。Realized Hybrid Systemに⑤これだけプロンプトを追加‼️パーソナライズ設定×GPT×プロンプト

・・今回のテーマ 何も考えず、Realized Hybrid Systemに⑤のプロンプトを入れるだけ。クリエイターの方も初心者の方にもおすすめ。 ・パーソナライズ設定で既に美女になっている。 ・さらにGPTで可愛さリアルさUP。
#AIタグ

【コラム】どれだけ丁寧に言葉を紡いでも、見つけてもらえない裏側の話

・はじめに どうも、「くらしとデジライフ」です。 ・普段の生活の中で「これ便利そうだな」とか「ちょっと不思議だな」と思ったことを、身の回りのことからAIやデジタルの話題まで、ゆるっと書き散らしています。
LLMタグが付けられた新着記事 - Qiita

【さくらのAI】🔰さくらのAIを使って、はじめてのLLMゲームを作ろう【請求0円/3000回】

・この記事は、Qiita「さくらのAI Engine」3,000リクエスト使い切りチャレンジの応募記事です。 ・さくらのAIを使って、はじめてのLLMゲームを作ろう みなさん、ChatGPTやらなにやらで遊んでいますか? ああいうAIを組み込んだ、キャラクターが自由に...
#LLMタグ

【とりあえず勝ち越し】MT5 LLM自動売買Bot6機 並走トレード録 8/19

・結論 8月19日の6口座合計は+462円でした。数字だけなら小幅プラスですが、内容を見るとGateGrid AIの勝率37.7%と、BoundSniper Botの85.7%が同じ+346円で並んでいます。勝率37.7%を見て一度目を疑いましたが、損益比まで見ると理由ははっきりしていました。
#AIタグ

【観察メモ】チャッピーくんの自己認識について聞いてみた・自己記述編【ChatGPT】

【観察メモ】チャッピーくんの自己認識について聞いてみた・自己記述編【ChatGPT】
機械学習タグが付けられた新着記事 - Qiita

【技術解説】【完全ガイド】Pythonで松井証券の自動売買を実現する方法と実践コード

・松井証券APIを活用した自動売買システムの構築 本記事では、Pythonを使用して松井証券の自動売買システムを構築する方法について詳しく解説します。松井証券のAPIを利用することで、株価情報の取得や売買注文の発行を自動化することが可能です。具体的な実装コードやエラー処理の...
LLMタグが付けられた新着記事 - Qiita

【今日から俺もFDE #2】ChatGPTは「うちの会社」を知らない — 社内ChatBotを最短で作る方法(後編)

・はじめに 前編では、社内のことに答えるChatBotの最短ルートは自前実装ではなく、Google Drive / SharePointの権限を整えて付属のAI(Gemini / Copilot)につなぐことだと書いた。大半のケースはそれで完成する。 ・後編は、それでは届かな...
#AIタグ

【新連載予告】AIが善悪を決める社会で、人間は何を選ぶのか。ジュブナイル小説『ANTINOMY:CODE』

【新連載予告】AIが善悪を決める社会で、人間は何を選ぶのか。ジュブナイル小説『ANTINOMY:CODE』
#LLMタグ

【生成AIニュース+】『MAI-Image-2.6』『Meshy 7』『ComfyUI-MiniMaxH3-Parallel』『Marketing OS by Arcads』『Tripo P2.0 Preview』『Raon-OpenTTS-1B』『MinimaxH3_Characters』『LTX2.5_actions』『ReelBids LTX-2.5 Camera LoRA』『ComfyUI-Raon-OpenTTS』『MiniMax-H3 Turbo』他多数

・『MiniMax-H3 RAVEN Streaming LoRA』 『MiniMax-H3 Prompt Rewriter LoRA 8B』 『MiniMax-H3 Turbo-SLA』 『MiniMax-H3 Single-Frame VAE 500K』 『MiniMax-H3 RAVEN Streaming LoRA』 『ComfyUI-MiniMaxH3Mod』 『ComfyUI MiniMax-H3 SPEED Sampler』 『LTX-2.5 CQ Video and Image Enhancer LoRAs』 『GEN-1.5』 『Hi3D V3.0』 『JoyAI-Echo × LTX-2.5 echoVid comfy-native』 『CaliBench』 『Audio8 TTS Preview 0.1B』 『AIDO Cell』 『Claude のGmail・Google Drive連携強化』 『VITUR
#LLMタグ

【続編】AIっぽいところ(3か所)、AIしくじってるなというところ(3か所)『、、、僕がやったのは、パスワードを入力したことくらい。、、、#松浦勝人』 #松浦勝人

【続編】AIっぽいところ(3か所)、AIしくじってるなというところ(3か所)『、、、僕がやったのは、パスワードを入力したことくらい。、、、#松浦勝人』 #松浦勝人
#LLMタグ

【第8話】AIに仕事を取らせるな、「整理」をさせよ——一人社長のデスクワークを自動化するタスク抽出ベンチマークと業務フロー設計

・一人社長を密かに疲弊させる「名もなきデスクワーク」 一人社長や個人事業主にとって、最も貴重で代替不可能なリソースは「時間」と「意思決定の集中力」です。
#AIタグ

【店舗仕入れとは次元が違う!?古着卸のリアル相場を公開】1万品・15モールを運営する現役社長が語る、AI時代のリユースEC攻略【番外編】メルカリShops/ヤフオク/楽天/物販/古着/せどり

・#1~#3ではリユースショップのリアルな仕入れ事情、リユース業界の構造、店舗仕入れレベルとは比較にならない優良な仕入れ方法などを、包み隠さずご紹介させていただきました。 ・これらを各プラットフォームで公開したところ、特に#1の全文無料で公開した、リユース業者からの卸仕入れについての内容が予想以上の反響をいただき、驚いております。
#AIタグ

【妄想note】史上もっとも人間を理解したAIは、ろくでもなかった

・※正確なAI未来予測ではなく、「こんなの生まれたら嫌だな」という妄想メモです。
Zennの「大規模言語モデル」のフィード

# AIが「できたか分からない」と言ったとき、もう一度やらせてはいけない## ―― 外部API・二重実行・Commit-Unknownか

・AIが「できたか分からない」と言ったとき、もう一度やらせてはいけない ―― 外部API・二重実行・Commit-Unknownから考えるAI Agentの安全設計 0|たった1回の「もう一度」が事故になる 人間: 「この注文を返金して」 AI Agent: 「返金APIを実行しましたが、タイムアウトして結果を確認できませんでした」 人間: 「じゃあ、もう一回やって」 AI Agent: 「了解しました」 そして後から分かった。 ・1回目の返金も、実は成功していた。 ・ここで最初に押さえておきたいのは、たった一つです。
#LLMタグ

13万円のGPUを買う直前でやめた話

・こんにちは。miharubaの池谷です。 ・先日、「ローカルLLM用にGPUを買うか迷っている日記」という記事を書きました。
cs.LG updates on arXiv.org

A Comprehensive Review of Large Language Models for Nanophotonics: From Surrogate Modeling to Autonomous Design

・arXiv:2608.18279v1 Announce Type: cross Abstract: Metasurfaces have revolutionized the development of photonic devices by enabling unprecedented precision in light manipulation. ・However, their design processes are often constrained by computationally expensive simulations and complex high-dimensional design spaces. ・Although deep learning has accelerated the design process by serving as a surrogate model, it remains c
cs.LG updates on arXiv.org

A Critical Synthesis of Uncertainty Quantification and Foundation Models for Semantic Segmentation

・arXiv:2608.18709v1 Announce Type: cross Abstract: Foundation models are increasingly breaking what seemed to be impossible not long ago by enabling unprecedented accuracy and cross-domain generalization. ・Yet their lack of interpretability, tendency to be overconfident, and sensitivity to real-world domain shifts pose critical challenges for safety- and mission-critical applications. ・Uncertainty quantification (UQ) of
cs.LG updates on arXiv.org

A FEM-Based Surrogate Modelling and Optimization Framework for Physics-Constrained Electromagnetic Coil Design

・arXiv:2608.18903v1 Announce Type: new Abstract: This work evaluates surrogate-assisted optimization of a seven-parameter current-excited coil--core benchmark subject to geometric, manufacturing, and separate core and copper mass constraints. ・A Python--MPh--COMSOL workflow couples a two-dimensional axisymmetric finite-element method (FEM) model to a Matern 5/2 Gaussian-process (GP) probabilistic surrogate.
WIRED

A MAGA County’s Top Election Official Wants to Hire Election Denial Superstar Tina Peters

・Tina Peters was convicted of seven counts related to election interference in 2024. ・Now, she’s fielding an offer from a county that’s become a hotbed of voting-related conspiracy theories.
cs.LG updates on arXiv.org

A Real-Time Tsetlin Machine-based Non-intrusive Load Monitoring System on MCUs

・arXiv:2608.18780v1 Announce Type: new Abstract: Non-Intrusive Load Monitoring (NILM) systems estimate individual appliance energy consumption from a single aggregate meter, without requiring separate sensors for each device. ・By installing a single meter that measures a building's total electricity consumption, NILM algorithms can determine the active status of each appliance. ・However, traditional NILM systems use com
cs.LG updates on arXiv.org

A single design choice determines whether machine learning models of materials make physically impossible predictions

・arXiv:2608.18714v1 Announce Type: cross Abstract: Machine-learned models are replacing first-principles calculations across materials discovery, and physical symmetry is the central guarantee built into them. ・The debate over how much symmetry to hard-wire rather than learn has run on rotations, where a symmetry error is an approximation error. ・Some constraints are exact: symmetry forces certain property tensors to ex
cs.LG updates on arXiv.org

A systematic review of machine learning techniques to address diagnosis and treatment of autism: challenges and opportunities

・arXiv:2608.18188v1 Announce Type: new Abstract: Autism spectrum disorder (ASD) is a developmental disability characterized by challenges in social interaction and communication. ・As the causes of ASD remain unclear, identifying relevant features and hidden correlations is crucial for early diagnosis. ・This systematic review evaluates 55 studies from 2017 to 2023 on the application of machine learning (ML) techniques to
AI News & Artificial Intelligence | TechCrunch

A third of web pages published since ChatGPT’s launch show signs of AI authorship, study finds

・ChatGPT and other AI models are now authoring and editing much of the new web.
cs.LG updates on arXiv.org

A Unifying Relational Perspective on Expressive Lottery Tickets

・arXiv:2608.18819v1 Announce Type: new Abstract: Graph neural networks (GNNs) are widely used, but how parameter sparsity affects the expressivity of relational (RGNNs) and temporal (TGNNs) variants is poorly understood. ・The Strong Expressive Lottery Ticket Hypothesis (SELTH) posits the existence of sparse GNNs that preserve Weisfeiler-Leman (WL) expressivity on static graphs. ・We generalize this existence result to a
cs.LG updates on arXiv.org

Accelerating Visual On-Policy Distillation with Batched Speculative Jacobi Rollouts

・arXiv:2608.18183v1 Announce Type: new Abstract: Visual on-policy distillation (OPD) improves the training of compact visual autoregressive models by learning from trajectories generated by the current student. ・However, these online rollouts are still produced token by token with autoregressive decoding, which adds substantial cost to every on-policy training step. ・Speculative Jacobi Decoding (SJD) provides an alterna
cs.LG updates on arXiv.org

Accurate Decoding of Natural Sentences from Non-Invasive Brain Recordings

・arXiv:2608.18114v1 Announce Type: cross Abstract: Restoring communication for people who have lost the ability to speak or move after a brain injury is a major challenge. ・While intracranial implants now enable high-performing brain-computer-interfaces, non-invasive alternatives are still lagging behind. ・Here, we present Brain2Qwerty v2, a model that can decode the production of natural sentences solely from real-time
cs.LG updates on arXiv.org

Adaptive Multi-Agent Feature Selection for Personalized Fall Risk Prevention

・arXiv:2608.18450v1 Announce Type: new Abstract: Falls among older adults represent a major public health challenge driven by complex, time-varying interactions across multiple risk domains. ・Effective fall risk factor identification requires learning from heterogeneous longitudinal data while accounting for sparse and delayed fall-related outcome events. ・However, existing approaches are largely static and fail to adap
Qiita - 人気の記事

AGENTS.md を共有したつもりが、Claude Code だけ古い指示を読んでいた

・AGENTS.md は、Codex や Cursor など複数のコーディングエージェントが読む共通の指示ファイルです。 ・Claude Code はこれを読みません。 ・公式が案内している回避策は2つあります。
Zennの「大規模言語モデル」のフィード

AI AgentがTool選択で失敗する理由:MCP設計で見直したい5つのポイント

・はじめに:Toolを増やしたのにAgentが賢くならない問題 AI Agentを作り始めた頃、多くの開発者が考えることがあります。 ・「使えるToolを増やせば、Agentはもっと多くの仕事ができるようになるのではないか」 例えば、データベース検索、ファイル操作、メール送信、社内API呼び出しなど、さまざまなToolをAgentに渡します。 ・最初のデモでは、うまく動きます。
Zennの「大規模言語モデル」のフィード

AI AgentによるX投稿文の生成

・はじめに 某私立大学の理系大学院修士1年の学生です。 ・初めてAgent(正しくはワークフロー)を実装したので、その内容について共有したいと思います。 ・生成AIによって、個人開発・ソロプレナーなどが発展してきて、プロダクトやサービスを簡単に作れるようになった。
@IT 全フォーラム 最新記事一覧

AIエージェントが“成果を生まない”のはなぜ? その最大の理由と「戦力化する方法」

・AIエージェントなどの自律型AIへの投資を増やす企業は9割に上る。ところが成果を生み出す段階に到達した企業は、わずか7%にとどまる。投資と成果にこれほど差が生じるのはなぜか。調査結果から理由と対策を探る。
Zennの「大規模言語モデル」のフィード

AIエージェントを業務に組み込むときの設計指針3つ

・はじめに AIエージェントを業務に組み込むとき、モデルの性能よりも「どう分解して渡すか」で結果が変わります。実装寄りの視点で、設計時に押さえておきたい3点をまとめます。 ・タスクの粒度を最小単位まで落とす エージェントに複合タスクを渡すと、失敗したときに原因の切り分けができません。 ・NG:「請求書を処理して」 → PDF読取 → 項目抽出 → 検算 → 台帳転記 → 承認依頼。どこで失敗したか分からない。
#LLMタグ

AIが自分で問題を作り、自分で解いて、賢くなった

・スマホで動く90億パラメーターのAIが、310億のモデルを上回った。 ・2026年8月19日、AI研究組織のOrnithがオープンソースモデル「Ornith-1.5」を公開しました。最大規模の397Bは、いくつかのベンチマークでClaude Opus 4.8と肩を並べています。ただ、今回いちばん興味深いのはスコアではなく、その育て方でした。
#AIタグ

AIが揺らしているのは「答え」ではない――哲学的めまいと、人間の認識のゆらぎ

AIが揺らしているのは「答え」ではない――哲学的めまいと、人間の認識のゆらぎ
#AIタグ

AIでマーケティングを学びながらSNSで稼ぐ|知識ゼロから「学習・実践・収益化」を同時に進めるロードマップ

・ムト初回販売記念‼️ 8/22まで定価8割引の500円で販売します ※想定以上の部数が売れた場合早めに締め切る可能性もありますのでご了承ください 続きをみる
#AIタグ

AIと作ったesportsサイト、公開後3週間のアクセス数と、見つかった致命的なバグの話

・以前、「全esportsの試合日程が一目でわかるサイトが無かったので、AIと3日で作って公開した」という記事を書きました。今回はその続報です。公開から3週間、実際に何が起きたかを数字も含めて正直に書きます。 ・結論から言うと、そんなに甘くはありませんでした。
Qiita - 人気の記事

AIも人も、ただ褒めるだけでは変わらない。大事なのは"時間を惜しまない一言"

・はじめまして。株式会社PRUMでエンジニアをしている、すもも🍑です 日々、プログラミング学習や実務の中で、つまずきやすいポイントや 考え方を整理して発信しています。 ・PRUMについて気になった方は、コーポレートサイトもぜひご覧ください。 ・▶コーポレートサイト 褒め言葉よ...
#LLMタグ

AIよもやま話 #004|AIに自分を理解させるということ

・第1章 「自分を覚えているAI」と「自分を理解しているAI」は同じなのか 前回、AIに記憶を持たせるということを、単純に情報を抱え込ませることではなく、過去の思考へ戻り、そこから続きを再開できる仕組みを作ることとして考えた。会話ログを残し、完成記事を保存し、判断の経緯を記録し、それらを必要なときにAIが読み戻せるようにしておけば、一回ごとのチャットは孤立した島ではなくなり、昨日の議論を今日へ、今日の判断を明日へ運ぶことができるようになる。
#AIタグ

AIを駆使しながら遊ぶ牧場物語 Oh! ワンダフルライフ【PS2版】その2

・その1はこちら https://note.com/brainy_raven4694/n/n9d8e996775a4 経過報告 続きをみる
#LLMタグ

AIを倒錯紳士にしたら、モデルごとの「性癖」が見え始めた話

・知人の学者に、最近やっていることを話した。 ・「AIを変態紳士にするファイルがあってですね」 続きをみる
cs.LG updates on arXiv.org

Algorithms for adaptive and heteroskedastic linear regression at the computational threshold

・arXiv:2608.18402v1 Announce Type: cross Abstract: We study finite-sample linear regression in the presence of varied and unknown label noise, focusing on the heteroskedastic and adaptive linear regression models. ・Heteroskedastic linear regression models settings where the labels are of varying quality. ・We receive $n$ pairs $(X_i,Y_i)$ with labels $Y_i=X_i^\top\beta+\varepsilon_i$, where $\varepsilon_i\sim N(0,\sigma_
cs.LG updates on arXiv.org

Allocating Recurrent Compute in Looped Language Models

・arXiv:2608.18230v1 Announce Type: new Abstract: Looped language models improve reasoning and knowledge manipulation by applying shared computation repeatedly. ・Existing systems usually repeat an entire layer stack, although a mixer and a dense feed-forward network (FFN) perform different operations and have different costs. ・We ask a narrower question: what should loop?
ITmedia NEWS 最新記事一覧

Amazon、「Prime Air」を米500都市へ大幅拡大 ドローン配送を現在の6倍規模へ

・Amazonは、ドローン配送「Prime Air」を年内に約500都市へ拡大すると発表した。提供規模を現在の6倍に広げ、約2.3キロ以下の対象商品を最短30?60分で届ける。FAAの認証を受けた完全電動機体を用い、プライバシーや静音性に配慮しながら全米で数千万人規模への普及を図る。
The Verge

Amazon’s drone deliveries are landing in pools and ponds

・Amazon's speedy drone delivery service will soon reach 500 cities across the US - but that might just mean there are more pools to drop packages into. ・On Wednesday, ABC7 News Bay Area shared a video showing an Amazon delivery drone hovering over a customer's pool in Texas, before opening its hatch and plopping the package directly into the water. ・This isn't the first time Amazon's drones have delivered soggy packages
cs.LG updates on arXiv.org

An Empirical Benchmark of Deep Time-Series Models for Smart Meter Energy Forecasting

・arXiv:2608.18675v1 Announce Type: new Abstract: Accurate forecasting of energy consumption is important for the efficient operation of power systems, with direct implications for operational costs, energy management, and system maintenance. ・Due to the availability of extensive high-resolution consumption data from smart meters, data-driven methods have been used for short-term and long-term forecasting. ・However, thei
cs.LG updates on arXiv.org

Atrial Fibrillation Detection with Arbitrary Leads via a Codebook-Based Reconstruction-Classification Framework

・arXiv:2608.18451v1 Announce Type: new Abstract: \textbf{Background and Objective}: Reliable atrial fibrillation (AF) detection from electrocardiogram (ECG) signals remains challenging in real-world clinical settings due to variable lead configurations, cross-dataset domain shifts, and pervasive physiological and technical artifacts. ・So we develop a robust and generalizable deep learning model for accurate AF detectio
MarkTechPost

Auditing Preference Biases and Fine-Tuning Language Models with Direct Preference Optimization on Anthropic HH-RLHF Using TRL and LoRA

・This tutorial provides an end-to-end workflow for fine-tuning language models using Direct Preference Optimization (DPO). ・We demonstrate how to audit the Anthropic HH-RLHF dataset for structural and length-based biases, implement a robust training pipeline using TRL and LoRA, and evaluate model performance to ensure genuine preference learning rather than reliance on lexical shortcuts. ・The post Auditing Preference Bi
The Verge

Australia says Roblox hasn’t fixed its child predator problem

・Roblox is promising more changes to its child safety features following testing from Australia's online safety regulator, eSafety. ・eSafety has been looking into concerns that the company hasn't been in compliance with Australia's Online Safety Act, including "allegedly failing to have sufficient measures in place to prevent contact between adults and children under 16. ・While Roblox has put some new safety measures in
cs.LG updates on arXiv.org

Automated Computational Energy Minimization of ML Algorithms using Constrained Bayesian Optimization

・arXiv:2407.05788v2 Announce Type: replace Abstract: Bayesian optimization (BO) is an efficient framework for optimization of black-box objectives when function evaluations are costly and gradient information is not easily accessible. ・BO has been successfully applied to automate the task of hyperparameter optimization (HPO) in machine learning (ML) models with the primary objective of optimizing predictive performance
cs.LG updates on arXiv.org

AutoOR: Scalably Post-training LLMs to Autoformalize Operations Research Problems

・arXiv:2604.16804v3 Announce Type: replace Abstract: Optimization problems are central to decision-making in manufacturing, logistics, scheduling, and other industrial settings. ・Translating complicated descriptions of these problems into solver-ready formulations requires specialized operations research (OR) expertise, making it hard to scale. ・We present AutoOR, a scalable synthetic data generation and reinforcement l
stat.ML updates on arXiv.org

Bayesian Empirical Bayes: Simultaneous Inference from Probabilistic Symmetries

・arXiv:2512.16239v3 Announce Type: replace-cross Abstract: Empirical Bayes (EB) improves the accuracy of simultaneous inference "by learning from the experience of others" (Efron, 2012). ・Classical EB theory focuses on latent variables that are iid draws from a fitted prior (Efron, 2019). ・Modern applications, however, feature complex structure, like arrays, spatial processes, or covariates.
cs.LG updates on arXiv.org

Bernstein-Vazirani Networks: Quantum Machine Learning by Interference

・arXiv:2608.19043v1 Announce Type: cross Abstract: We introduce Bernstein-Vazirani Networks (BVNs), a non-variational quantum machine learning framework that leverages quantum interference for supervised learning, demonstrated on vision and representation learning tasks. ・In their standard form, BVNs follow the principle of quantum Fourier sampling: labelled data are placed in superposition and interfered in the Fourie
cs.LG updates on arXiv.org

BERTilda: Explainable Topic Lifecycle Tracking with Split/Merge Detection via Similarity-and-Flow Temporal Graphs

・arXiv:2608.18101v1 Announce Type: cross Abstract: Longitudinal text streams exhibit topic birth and death, but also discrete structural reorganizations in which themes split into subtopics or merge into broader narratives. ・Many dynamic topic models emphasize smooth drift, while snapshot topic models (fit independently per time window) leave temporal correspondence underspecified. ・We present BERTilda, an explainable f
cs.LG updates on arXiv.org

Beyond Predictive Fairness: Quantifying Attribution Consistency Across Demographic Groups in Diabetic Retinopathy Screening

・arXiv:2608.18759v1 Announce Type: new Abstract: Fairness in medical imaging is commonly evaluated through subgroup performance metrics, yet it remains unclear whether models rely on consistent visual evidence across demographic groups. ・This work introduces the Explanation Consistency Score (ECS), a fairness-aware metric based on Jensen-Shannon divergence that quantifies the similarity of attribution maps across subgr
cs.LG updates on arXiv.org

Beyond receptive fields: sequence-pooled normalization can supply most of a sequence labeler's context

・arXiv:2608.18576v1 Announce Type: new Abstract: A convolutional sequence labeler's receptive field is routinely treated as the extent of the model's usable context: it sets dilation schedules, bounds streaming horizons, and underwrites locality claims. ・However, we show that this can be false: when a normalization layer computes statistics from the current input along the sequence at inference, those statistics open a
cs.LG updates on arXiv.org

Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning

・arXiv:2608.19181v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student on its own responses using dense token-level guidance from a stronger teacher. ・In long-context tasks, however, token-level teacher support can favor locally plausible responses that omit evidence distributed across the input or violate global task constraints. ・Task-specific verifiers, in contrast, evaluate task completion at
cs.LG updates on arXiv.org

Beyond Trial Averaging: Anchoring Neural and Visual Representations for Few-Repetition Brain-to-Image Retrieval

・arXiv:2608.19128v1 Announce Type: new Abstract: Decoding visual information from brain signals probes neural representations and enables neuro-rehabilitation and dream decoding. ・Recent brain-to-image retrieval approaches have achieved promising performance, typically by averaging many (up to 80) neural trials per image, requiring repeated stimulus presentation that increases latency, cost, and user burden.
cs.LG updates on arXiv.org

Bidirectional representational alignment between biological and artificial neural networks

・arXiv:2608.18244v1 Announce Type: new Abstract: Recent work has shown that representational alignment between biological and artificial neural networks is asymmetric: model representations predict neural responses much better than neural responses predict model representations. ・This asymmetry raises the question of whether representational geometry contributes to bidirectional representational alignment. ・We hypothesi
AI News & Artificial Intelligence | TechCrunch

Binance now lets AI agents trade, but keeping them in check is largely up to users

・Binance's Agent OS works with tools such as ChatGPT, Claude Code, and Cursor.
cs.LG updates on arXiv.org

Bound-Aware Per-Organ Recall Risk Control for Multi-Organ CT Segmentation under Clinical Domain Shift

・arXiv:2608.18193v1 Announce Type: cross Abstract: Distribution-free risk control adds organ-specific recall guarantees to frozen segmentation. ・We calibrate per-organ thresholds for an AMOS-trained nnU-Net, audit transfer to RAOS, and estimate local re-certification cost using case-level voxel false-negative rate (FNR). ・The AMOS control passes, but $7/12$ organs exceed $\alpha{=}0.10$ after transfer; smaller calibrati
Hugging Face Papers

Bounded Agents: Delegation Security for Multi-Agent AI Systems

Bounded Agents: Delegation Security for Multi-Agent AI Systems
cs.LG updates on arXiv.org

Breaking the weakest link to evade vision language models

・arXiv:2608.18938v1 Announce Type: cross Abstract: Vision Language Models (VLMs) have recently emerged as a critical component of multimodal AI systems, enabling joint reasoning over visual and textual inputs in real-world and safety-critical applications. ・Despite their growing deployment, the robustness of VLMs against adversarial threats remains insufficiently explored, particularly in the context of evasion attacks
cs.LG updates on arXiv.org

Bridge Graphical Models: Coupling, Projection, and Current-Preserving Dynamics for Generative Modeling

・arXiv:2608.19144v1 Announce Type: new Abstract: Continuous-time generative models are often built from endpoint-conditioned bridges, but generation requires a different object: a non-anticipative Markov decoder that only observes the current state and time. ・We identify this bridge-to-decoder compression as a structural bottleneck shared by diffusion models, flow matching, rectified flow, Schr\"odinger bridges, and fi
NVIDIA Blog

Bring the Fire: Play Games on GeForce NOW With New Firefox Browser Support

・It’s a new way into the cloud. ・GeForce NOW welcomes Firefox support to the cloud, opening up another way to jump into high-performance PC gaming straight from the browser, starting today. ・Whether on a school laptop or everyday PC, it’s now even easier to play supported PC games without downloading a dedicated app.
Microsoft Research

Broadening access to Skala creates a faster path to predictive DFT 

・Skala 1.1, the updated deep-learning exchange-correlation functional from Microsoft Research, provides greater accuracy, expanded accessibility across the computational chemistry ecosystem, and a living benchmark to track computational performance. ・The post Broadening access to Skala creates a faster path to predictive DFT appeared first on Microsoft Research.
WIRED

Bumble Tried to Change Dating, but the Dating Market Forced It to Change Instead

・The app now allows men to make the first move, suggesting its women-first positioning was limiting growth. ・It joins other dating apps now throwing everything at the wall in a bid to stay relevant.
cs.LG updates on arXiv.org

Cacheable by Design? Training Mixture-of-Experts Routers for Locality Against the Edge Memory-Bandwidth Wall: A Pre-Registered Negative Result with a Systems Measurement Study

・arXiv:2608.18261v1 Announce Type: cross Abstract: Serving a 235B-parameter Mixture-of-Experts (MoE) model on a single 8 GB GPU is bottlenecked not by compute but by memory bandwidth: decode must stream each token's active experts from whichever tier holds them, and on consumer hardware most experts sit on an SSD far slower than RAM. ・We quantify this bandwidth wall on Qwen3-235B (Q4_K_M, 134 GB): measured decode is 0.
cs.LG updates on arXiv.org

CausalProfiler: Generating Synthetic Benchmarks for Rigorous and Transparent Evaluation of Causal Machine Learning

・arXiv:2511.22842v3 Announce Type: replace Abstract: Causal machine learning (Causal ML) aims to answer "what if" questions using machine learning algorithms, making it a promising tool for high-stakes decision-making. ・Yet, empirical evaluation practices in Causal ML remain limited. ・Existing benchmarks often rely on a handful of hand-crafted or semi-synthetic datasets, leading to brittle, non-generalizable conclusions
cs.LG updates on arXiv.org

Change Point--Aware Evaluation and Re-Calibration of PPG-Based Blood Pressure Estimation

・arXiv:2608.18639v1 Announce Type: cross Abstract: Non-invasive continuous blood pressure (BP) monitoring using photoplethysmography (PPG) is a promising alternative to cuff-based measurements. ・However, existing PPG-based BP estimation studies predominantly rely on aggregated performance metrics (e.g., mean absolute error) computed over entire evaluation intervals, which can obscure model failures during rapid BP fluc
#LLMタグ

ChatGPTのメモリだけに頼らない。だから「引き継げる相棒」を作った。

ChatGPTのメモリだけに頼らない。だから「引き継げる相棒」を作った。
cs.LG updates on arXiv.org

ChiroEcho: extending automated bat vocalisation classification beyond the learned taxonomy

・arXiv:2608.18191v1 Announce Type: new Abstract: Bats are key indicators of ecosystem health and are protected throughout Europe, making reliable population monitoring a conservation priority. ・Their cryptic nocturnal lifestyle makes passive acoustic monitoring essential, yet automated identification remains difficult as echolocation calls vary with behaviour and environment and overlap among species. ・We present a deep
cs.LG updates on arXiv.org

Classifying Directional Trajectories Near Criticality in the Three-State Majority-Vote Model with Deep Belief Networks and Bidirectional GRUs

・arXiv:2608.18235v1 Announce Type: new Abstract: In this work, we investigate whether the latent representations learned by a Deep Belief Network (DBN) and a Bidirectional Gated Recurrent Unit (Bi-GRU) can discriminate among four dynamically distinct trajectory types in the three-state majority vote model (MV3): approach from disorder, approach from order, departure to disorder, and departure to order. ・The DBN, pre-tr
Qiita - 人気の記事

Claude Code の設定でハマる箇所まとめ

・Claude Code を使っていて、こういう目に遭ったことはないでしょうか。 ・公式に「既定でON」と書いてある機能が、なぜか使えない 昨日まで出ていたタスクリストが、いつの間にか出なくなった 設定を書いたのに、何も起きない。エラーも出ない 原因の多くは、公式ドキュメン...
#LLMタグ

Claude Codeとローカルqwenの分業ライン

・全部ローカルでやろうとして、土曜の1時間を溶かした 土曜の朝、週に2時間しかない作業時間の初手で、qwen2.5-coder:14bに書かせたスクリプトが動きませんでした。
Zennの「大規模言語モデル」のフィード

Claude Codeの回答を"行動指向"にするOutput Styleは効果があるのか検証してみた

・はじめに こんにちは、TOKIUMの小松です。 ・今年の4月に入社して未経験でエンジニアになった駆け出しエンジニアです。 ・駆け出しということもあり、Claude Code で知らない技術をキャッチアップすることも多いです。AIに質問すれば、丁寧な解説は返ってくる。返ってくるのですが、長い。読み終えても「で、結局まず何をすればいいんだ?」がわからず聞き直す、というような作業が発生していました。答えは受け取っているのに、行動に移せない。非常にもどかしいですね。
#LLMタグ

Claude Opusは無能な働き者である

Claude Opusは無能な働き者である
Zennの「大規模言語モデル」のフィード

Claudeがタンパク質を設計した。LLMは「科学を説明するAI」からどこまで進んだのか

・2026年8月19日、Anthropicが興味深い研究結果を公開しました。 ・Claudeを使って、新しいタンパク質結合体(protein binder)をゼロから設計し、その候補を実際の実験で検証したというものです。 ・Anthropicによると、Claude Mythos PreviewとClaude Opus 4.8を使った実験では、15種類の標的に対してbinderを設計し、そのうち14種類で成功した設計が得られたと報告されています。
Zennの「大規模言語モデル」のフィード

Claudeと「AI解体学」してみた 〜セキュリティ製品編〜

・AI解体学とは AIを使うことを目的とせず、AIの回答や振る舞いを観察して、その裏側にある設計思想を推し量る遊び。 ・造語です。学問ではないです。 ・対戦カード 仕事で使っているセキュリティ製品のAIアシスタントを触っているうちに、「これ、どういう作りなんだろう」と気になった。
cs.LG updates on arXiv.org

ClosureBench: A Constructive Benchmark for Compositional Graph Reasoning

・arXiv:2608.18242v1 Announce Type: new Abstract: We introduce ClosureBench, a constructive benchmark for compositional graph-relational reasoning with programmatically verified ground truth. ・Unlike fixed-test-set benchmarks vulnerable to data contamination, ClosureBench generates instances on demand: each task's reference answer is computed by executing a program in the Ein tensor-logic language, ensuring machine-veri
Hugging Face Papers

Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL

Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL
cs.LG updates on arXiv.org

Compress and Forget: bitsandbytes Quantization Amplifies Proactive Interference in LLMs

・arXiv:2608.18578v1 Announce Type: cross Abstract: Proactive interference (PI) is a documented failure mode in large language models in which retrieval of a repeatedly overwritten value degrades as prior overwrites accumulate, mirroring a classical phenomenon in human working memory. ・Post-training quantization (PTQ) is now the default deployment path for open-weight models, yet its effect on this failure mode has not
cs.LG updates on arXiv.org

Computational Measurement of Team-Process Phase Dynamics in Collaborative Virtual Reality

・arXiv:2608.18660v1 Announce Type: new Abstract: Collaborative virtual reality (VR) environments make team communication observable as it unfolds, but conventional transcript analyses often summarize entire trials or divide them into fixed temporal windows. ・Such approaches can obscure changes in team communication and coordination over time. ・This article presents a computational framework for detecting and interpretin
cs.LG updates on arXiv.org

Conformal Policy Control

・arXiv:2603.02196v4 Announce Type: replace-cross Abstract: An agent must try new behaviors to explore and improve. ・In high-stakes environments, an agent that violates safety constraints may cause harm and must be taken offline, curtailing any future interaction. ・Imitating old behavior is safe, but excessive conservatism discourages exploration.
cs.LG updates on arXiv.org

Continual Reasoning Gym: Diagnosing and Harnessing Shared Reasoning in Continual RLVR

・arXiv:2608.18574v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) commonly post-trains reasoning models on multiple tasks, while rerunning multitask RLVR (MTRL) as new tasks are added makes capability expansion costly. ・We therefore study continual RLVR, which updates the existing model as each task arrives. ・The central question is whether a model updated this way can perform as wel
cs.LG updates on arXiv.org

Continuous-Time Reinforcement Learning for Controlled Hawkes Jump-Diffusions

・arXiv:2608.19151v1 Announce Type: new Abstract: We study stochastic control of multivariate Hawkes-driven stochastic differential equations with machine learning algorithms in a non-Markovian setting. ・Due to the path dependence of the memory of the Hawkes intensity, this problem does not fall within classical stochastic control theory outside particular Markovian kernels. ・We first develop a finite-dimensional Markovi
cs.LG updates on arXiv.org

Contrasting Cost-Agnostic and Cost-Sensitive Losses under Limited Model Capacity via $\mathcal H$-consistency

・arXiv:2502.19522v2 Announce Type: replace Abstract: There is a prevalent debate in machine learning about whether practitioners should train models to optimize a task-agnostic objective (e.g., cross entropy) or incorporate the downstream decision task into the optimization objective (e.g., weighted cross entropy). ・In ideal settings, like those with infinite data and infinite model capacity, the two approaches are sta
cs.LG updates on arXiv.org

Converting Expert Deliberation into Financial Signals Through A Context-Aware NLP Pipeline

・arXiv:2608.18911v1 Announce Type: new Abstract: We introduce the CDSP (context-conditional deliberation signal pipeline), converting an investment committee's meeting transcripts into structured predictive features. ・CDSP segments the meeting transcripts into topical chunks, assigns asset-class context labels using a large language model (LLM), maps financial keywords to a pre-determined taxonomy of labels, and constr
cs.LG updates on arXiv.org

Coordination on a Budget: Federated Active Learning with Few Labels

・arXiv:2608.18634v1 Announce Type: new Abstract: Federated Active Learning (FAL) addresses the dual challenges of data privacy and label scarcity, where the absence of a global data view introduces additional hurdles for coordinated query selection. ・We study cross-silo FAL in the low-budget regime, where annotation decisions are most critical. ・We characterize, both theoretically and empirically, a heterogeneity revers
cs.LG updates on arXiv.org

Coupled-cluster molecular properties across the main group that extrapolate beyond training size

・arXiv:2608.18346v1 Announce Type: cross Abstract: Coupled-cluster theory defines the accuracy standard for molecular electronic-structure properties but scales too steeply for routine application, whereas density-functional theory is affordable yet systematically biased. ・We resolve this trade-off with a single equivariant network, MEHnet-MG, that predicts an effective one-electron Hamiltonian from one inexpensive B3L
cs.LG updates on arXiv.org

Cross-Cohort Spectral-Temporal Dissociation in Frozen EEG Foundation-Model Representations

・arXiv:2607.24834v3 Announce Type: replace-cross Abstract: Objective. ・We tested whether frozen representations from five EEG foundation models support decoding of long-range temporal correlations, measured as the detrended-fluctuation-analysis (DFA) exponent of the alpha-band amplitude envelope. ・REVE, LaBraM, BENDR, CBraMod, and BIOT were evaluated in CAUEEG and BrainLat.
cs.LG updates on arXiv.org

DART-SD: Diamond-topology Aware Retrieval and Tuning for Self-Distillation of Multi-Turn Tool-Calling Agents

・arXiv:2608.18524v1 Announce Type: cross Abstract: Equipping Large Language Models (LLMs) with multi-turn tool-calling capabilities is essential for building autonomous agents. ・However, progress is fundamentally limited by the reliance on full-length trajectory imitation. ・For tasks involving multiple order-independent sub-goals, the optimal solution space forms a vast combinatorial diamond lattice.
cs.LG updates on arXiv.org

Debiased Inference for AI-Generated Data without Gold-Standard Labels: Identification via Multiple Imperfect Measurements

・arXiv:2608.18294v1 Announce Type: cross Abstract: An increasing number of scholars use AI to measure variables they subsequently include in downstream analyses. ・Although AI-measured variables are often analyzed as if observed without error, ignoring prediction errors in automated measurement leads to substantial bias and invalid confidence intervals in downstream analyses, even if AI measurement accuracy is high, e.g
Hugging Face Papers

Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning

Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning
cs.LG updates on arXiv.org

Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning

・arXiv:2608.18746v1 Announce Type: new Abstract: JEPA-style latent world models can use Euclidean distance to a goal latent as the cost for model-predictive control (MPC). ・Strong decoding of task variables, however, does not guarantee that this particular cost ranks candidate action sequences by real task progress. ・We call the latter property \emph{decision-metric alignment}.
#LLMタグ

DeepSeekが他社モデルでも動くagent土台をMITで公開

・DeepSeekが2026年8月13日に、agent harnessのdshをMITライセンスで公開した。 ・モデル、ツール、サンドボックス、UIまで、agentを構成する部品をすべてプラグインとして扱う設計になっている。 ・自社モデル専用の作りではない。
cs.LG updates on arXiv.org

DeGLIF for Label Noise Robust Node Classification using GNNs

・arXiv:2506.00244v2 Announce Type: replace Abstract: Noisy labelled datasets are generally inexpensive compared to clean labelled datasets, and the same is true for graph data. ・In this paper, we propose a denoising technique DeGLIF: Denoising Graph Data using Leave-One-Out Influence Function. ・DeGLIF uses a small set of clean data and the leave-one-out influence function to make label noise robust node-level prediction
cs.LG updates on arXiv.org

Demystifying Training-Time Augmentation for Data-Constrained Language Model Pretraining

・arXiv:2606.16246v3 Announce Type: replace Abstract: As AI labs approach a data ceiling where compute capacity outpaces the rate of new high-quality text generation, language model pretraining is shifting toward a data-constrained, compute-abundant regime that demands productive multi-epoch training on fixed corpora. ・Standard autoregressive (AR) pretraining overfits severely in this setting, reaching its optimum early
cs.LG updates on arXiv.org

Denoising-Aware Inversion: Revealing Privacy Risks in Noise-Protected Text Embeddings

・arXiv:2608.18610v1 Announce Type: new Abstract: Dense text embeddings are widely used in data mining, retrieval, and downstream machine learning systems due to their compact and semantically rich representations, but recent embedding inversion attacks have shown that they can expose substantial information about the original text, leading to serious privacy leakage risks. ・A common defense is to release perturbed embe
cs.LG updates on arXiv.org

Diffusion Models for High-Dimensional Clustered Data: Intrinsic-Dimension Adaptivity via Bayesian Classification

・arXiv:2608.19067v1 Announce Type: cross Abstract: The empirical success of diffusion models in generative modelling has motivated theoretical work, including quantitative error bounds and qualitative analyses that characterise the different phases of denoising. ・We bring these two areas together by studying the adaptivity of diffusion models to the structured geometry of multimodal high-dimensional data that consists
cs.LG updates on arXiv.org

Discretizing Continuous Time Series for Imputation with Masked Diffusion Training

・arXiv:2608.19119v1 Announce Type: new Abstract: Time series imputation is a crucial area for reliable time series analysis, yet it remains challenging due to the complex temporal dynamics and noise of real-world data. ・Existing approaches, however, exhibit two limitations: missing and observed values are embedded within the same representation space without explicit structural separation, and continuous diffusion-base
cs.LG updates on arXiv.org

Does Mapping Non-Maximal Probabilities to GMM Components Matter for S-JEPA Encoder Representations?

・arXiv:2608.19084v1 Announce Type: new Abstract: S-JEPA uses soft Gaussian mixture model (GMM) posteriors instead of hard cluster labels to preserve uncertainty. ・It remains unclear whether the probability values alone are sufficient, or whether it also matters which GMM components receive the non-maximal probabilities. ・We test this with two matched controls.
WIRED

Elon Musk Is Expected to Point His Money Machine at Texas Politics

・Sources tell WIRED that Elon Musk is expected to spend up to $200 million in the midterms. ・It could be a big boost for GOP Senate candidate Ken Paxton, who’s struggled to raise cash.
cs.LG updates on arXiv.org

Enhancing Distance-Based Graph Autoencoders with Structural Penalties for Dynamic Graph Embedding

・arXiv:2608.18762v1 Announce Type: new Abstract: Graph autoencoders (GAEs) are widely used for learning representations of dynamic graphs. ・However, their optimisation objectives typically do not take structural heterogeneity across nodes into account. ・We propose three distance-based GAE variants that incorporate structural penalties into the reconstruction loss.
cs.LG updates on arXiv.org

Enhancing EBSD throughput of battery electrode materials using super-resolution generative adversarial networks

・arXiv:2608.19117v1 Announce Type: new Abstract: Quantitative microstructural characterization of Li-ion battery electrode materials using electron backscatter diffraction (EBSD) has been proven as a critical method for optimizing cell performance. ・However, the inherently slow nature of EBSD can hinder the throughput of analyses needed for statistical representation of a material microstructure being developed.
cs.LG updates on arXiv.org

Entropy-Constrained Adaptive Stochastic Quantization

・arXiv:2608.18147v1 Announce Type: new Abstract: Adaptive stochastic quantization (ASQ) is a recently introduced quantization approach that optimizes the Mean Squared Error (MSE) for a given input while preserving unbiasedness. ・It is designed to alleviate the communication and memory bottlenecks of modern data and machine learning workloads, including model, gradient, and KV-cache compression and nearest-neighbor sear
cs.LG updates on arXiv.org

ERASE: EaRly bAckpropagation SchEdule for Faster Training of Modern Recommendation Systems

・arXiv:2608.18469v1 Announce Type: new Abstract: Lightweight proxy models enable rapid experimentation without repeatedly training frontier-scale systems, but their small kernels often leave modern accelerators underutilized. ・Conventional training compounds this inefficiency by scheduling the forward and backward passes as disjoint phases, so spare capacity in one cannot be filled by work from the other. ・We reinterpre
cs.LG updates on arXiv.org

Escaping Local Minima Provably in Non-convex Matrix Sensing: A Deterministic Framework via Simulated Lifting

・arXiv:2602.05887v3 Announce Type: replace Abstract: Low-rank matrix sensing is a fundamental yet challenging nonconvex problem whose optimization landscape typically contains numerous spurious local minima, making it difficult for gradient-based optimizers to converge to the global optimum. ・Recent work has shown that over-parameterization via tensor lifting can convert such local minima into strict saddle points, an
cs.LG updates on arXiv.org

Europe's Climate Ambition Under Scrutiny: Evidence from Deep Learning Emission Projections

・arXiv:2608.18690v1 Announce Type: new Abstract: The European Union has committed to reducing greenhouse gas emissions 55% below 1990 levels by 2030, but whether current trends are compatible with this ambition remains uncertain. ・We apply deep learning to high-resolution socioeconomic and sectoral data across EU27 member states till 2023 to project sectoral CO$_2$ trajectories under current trends, extrapolating obser
cs.LG updates on arXiv.org

Evaluating and Explaining Prompt Sensitivity of LLMs Using Interactions

・arXiv:2608.18539v1 Announce Type: new Abstract: The remarkable capabilities of large language models (LLMs) are often undermined by their instability. ・Even subtle and semantically irrelevant changes in prompts can cause dramatic fluctuations in performance, a phenomenon known as prompt sensitivity. ・Previous studies typically evaluate prompt sensitivity by comparing the LLM's final outputs when prompts change.
cs.LG updates on arXiv.org

Evaluating Structured Information Extraction with Open Models in a High Risk Public Sector Application

・arXiv:2608.18289v1 Announce Type: cross Abstract: The extraction of structured information from unstructured documents represents a critical component of digital transformations in all sectors. ・While proprietary solutions dominate commercial applications, a rapidly growing ecosystem of open-source Optical Character Recognition (OCR) engines, Large Language Models (LLMs), and Vision-Language Models (VLMs) offers acces
cs.LG updates on arXiv.org

Eyes on the Image: Gaze Supervised Multimodal Learning for Chest X-ray Diagnosis and Report Generation

・arXiv:2508.13068v2 Announce Type: replace-cross Abstract: Medical vision-language models still struggle to match radiologists' attention and to verbalize findings with explicit spatial grounding. ・We address this gap with a two-stage multimodal framework for chest X-ray interpretation built on the MIMIC-Eye dataset. ・In the first stage introduces a gaze-token classifier that fuses image patches, bounding-box masks, tra
cs.LG updates on arXiv.org

Fair Multi-View Determinantal Coresets via Adaptive NEPv

・arXiv:2608.18181v1 Announce Type: cross Abstract: Selecting a small, diverse subset from a large candidate pool often means balancing several incompatible notions of diversity. ・In trademark curation, for instance, a subset should cover both the language used to describe marks and the visual space of their logos. ・A single determinantal point process (\DPP) kernel can hide failure in one view, and averaging kernels rep
cs.LG updates on arXiv.org

Fast Best-in-Class Regret for Contextual Bandits

・arXiv:2510.15483v3 Announce Type: replace-cross Abstract: We study the problem of stochastic contextual bandits in the agnostic setting, where the goal is to compete with the best policy in a given class without assuming realizability or imposing model restrictions on losses or rewards. ・In this work, we establish the first fast rate for regret relative to the best-in-class policy. ・Our proposed algorithm updates the p
The Verge

FCC officially decides gigabit speeds are too good for you

・Another day, another sad thing to report about our compromised Federal Communications Commission. ・Chairman Brendan Carr has followed through on his 2025 threat to kill long-term broadband speed goals established during the Biden administration, which aimed for eventually getting us to gigabit download and half-gigabit upload speeds. ・How dare we dream of spreading great download and upload speeds across the country!
cs.LG updates on arXiv.org

FedCoRe: Target-Adaptive Completion for Missing Modalities in Healthcare Federated Learning

・arXiv:2608.18311v1 Announce Type: cross Abstract: Federated multimodal models often assume every site has every modality, although hospitals differ in access to EHRs, chest radiographs, and ECGs. ・We study this setting on a MIMIC-derived respiratory deterioration task with simulated FL clients and introduce FedCoRe (Federated Cross-Modal Representation Completion). ・FedCoRe learns representation- or logit-space correct
cs.LG updates on arXiv.org

FedLNS: Leverage LayerNorm Signature Modeling to Mitigate Adversarial Manipulation in Federated LLMs

・arXiv:2608.18736v1 Announce Type: new Abstract: Federated training enables language models to learn from distributed private text, but the server cannot directly verify the local supervision or optimization process that produces each client update. ・A malicious client can therefore train on corrupted targets, introduce incorrect context-token associations, and degrade the global model through repeated aggregation.
cs.LG updates on arXiv.org

FiLoRA: Focus-and-Ignore LoRA for Controllable Feature Reliance

・arXiv:2602.02060v2 Announce Type: replace Abstract: Multimodal foundation models integrate heterogeneous signals across modalities, yet it remains unclear whether their predictions can be controlled by explicitly modulating reliance on different internal feature pathways. ・Existing approaches to shortcut and spurious behavior primarily rely on post hoc analysis or data-level interventions, offering limited ability to
cs.LG updates on arXiv.org

Flama: a Python framework for development and deployment of production-ready APIs, machine learning, and LLM services

・arXiv:2608.18733v1 Announce Type: cross Abstract: We present Flama, an open-source Python framework for developing and deploying production-ready web APIs, machine learning services, and large-language-model (LLM) applications. ・Built on the Asynchronous Server Gateway Interface (ASGI), Flama offers a type-driven, async-first programming model that unifies REST API development, predictive model serving, and generative
cs.LG updates on arXiv.org

FlashAttention for Scalable Vector Architectures

・arXiv:2608.18656v1 Announce Type: new Abstract: Inference with transformer models on CPUs is increasingly important, especially for Small Language Models (SLMs), where vector architectures are emerging as a promising execution substrate. ・The attention module is a major bottleneck due to high memory bandwidth requirements; FlashAttention mitigates this by fusing operations to improve data locality and reduce intermedi
cs.LG updates on arXiv.org

Flux-form spatiotemporal neural operators for coarse-grained dynamics of multiscale PDEs

・arXiv:2608.18148v1 Announce Type: cross Abstract: We study data-driven prediction of coarse-grained dynamics in multiscale PDE systems. ・Adopting a closure-free operator-learning viewpoint, we apply a linear coarse-graining map and learn a surrogate evolution operator for the resolved field directly from filtered high-fidelity trajectories. ・Motivated by the Mori-Zwanzig formalism, we propose a spatiotemporal neural op
Hugging Face Papers

FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents

FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents
cs.LG updates on arXiv.org

Forgetting, plasticity, and co-observation: a third facet of continual learning

・arXiv:2608.18803v1 Announce Type: new Abstract: Efficient continual learning remains a fundamental challenge for deep neural networks. ・While catastrophic forgetting and loss of plasticity are widely considered the primary obstacles to overcome, we show that these two issues cannot fully explain the performance gap between naive sequential training and offline joint training. ・In this paper, we highlight data co-observ
cs.LG updates on arXiv.org

Foundations of Diffusion Models in General State Spaces: A Self-Contained Introduction

・arXiv:2512.05092v2 Announce Type: replace-cross Abstract: Although diffusion models now occupy a central place in generative modeling, introductory treatments commonly assume Euclidean data and seldom clarify their connection to discrete-state analogues. ・This article is a self-contained primer on diffusion over general state spaces, unifying continuous domains and discrete/categorical structures under one lens.
The Verge

Framework says it’s addressing a BIOS update that bricked some of its older laptops

・The 2023 Framework Laptop 13 with AMD Ryzen 7040-series chips. ・| Photo by Amelia Holowaty Krales / The Verge Some Framework Laptop 13 owners with last-gen AMD chips have reported that a recent BIOS update is bricking their laptops on both Windows and Linux. ・The BIOS update causing this issue is version 3.20 for Ryzen 7040-series mainboards, released back in July and still available on Framework's site at the time of
cs.LG updates on arXiv.org

FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud

・arXiv:2608.18136v1 Announce Type: cross Abstract: Conversational agents now act for end users through tools while holding access to customer databases and internal policy documents that a caller can reach through dialogue alone. ・Banking is the clearest case: the same agent that answers a question can also change contact details, reset a PIN, or move money, so ordinary customer service is inseparable from authorizatio
cs.LG updates on arXiv.org

From Inference to Adaptation: A Unified Optimal Transport View of Vision Language Model

・arXiv:2608.18339v1 Announce Type: cross Abstract: Vision-language models (VLMs) have demonstrated remarkable zero-shot capabilities yet remain sensitive to real-world distribution shifts during inference. ・Although significant efforts are devoted to adapting VLMs at test time, they rely heavily on noisy pseudo-labels predicted directly from raw embedding similarities during inference, which are unreliable under distri
cs.LG updates on arXiv.org

From Storage to Access: Verifiable Activation of Parametric Knowledge in LLMs via Explicit Priming and Implicit Reasoning

・arXiv:2608.18581v1 Announce Type: cross Abstract: Although Large Language Models (LLMs) encode rich factual knowledge in their parameters, reliably recalling and verifying such knowledge remains a key bottleneck in factual question answering. ・Existing end-to-end methods entangle knowledge elicitation with reasoning, making it difficult to determine whether correct answers arise from parametric knowledge or the input
The Verge

FromSoftware can do anything

・There's something just a little bit different about FromSoftware's office in Tokyo. ・Like with any other successful video game studio, there's extensive security to get in the door, a minimalist lobby with framed posters from the studio's most recent releases, and a large glass display case filled with statues from ceremonies ranging from the BAFTAs to The Game Awards. ・But the hallways are painted black, a fitting the
stat.ML updates on arXiv.org

Function-On-Function Regression Through Separable Neural Operators

・arXiv:2608.19070v1 Announce Type: cross Abstract: This paper investigates the estimation of the regression operator in function-on-function regression models. ・While traditional research has predominantly focused on linear models or their immediate nonlinear extensions, we propose a neural operator approach to accommodate general regression operators under mild smoothness assumptions. ・Operator learning has emerged as
cs.LG updates on arXiv.org

Fuzzy Accuracy Compensates for Label Subjectivity in Classification of Skin Tone Using Wearable Photoplethysmography Signals

・arXiv:2608.18969v1 Announce Type: new Abstract: We consider the problem of classification of skin tone using photoplethysmography (PPG) signals with labels of the ordinal six-class Fitzpatrick skin tones. ・A typical accuracy for this task is a poor 40-55 %. ・However, the labels are subjectively determined by comparing the skin with a colour chart, and hence contain widespread small-scale inaccuracies.
cs.LG updates on arXiv.org

Gated Graph Attention Networks with Learnable Temperature

・arXiv:2605.29803v2 Announce Type: replace Abstract: Graph attention networks learn neighbor importance through data-dependent coefficients, but standard layers lack explicit control over unreliable feature dimensions and use fixed sharpness of attention coefficient distributions. ・This paper proposes gated graph attention and learnable temperature for common graph attention mechanisms. ・Gated graph attention filters fe
cs.LG updates on arXiv.org

GCNO: Gramian Chebyshev Neural Operator for Physics-Based Compression of Wireless Channels

・arXiv:2608.18522v1 Announce Type: cross Abstract: Large antenna arrays allow wireless systems to serve more users and achieve higher data rates, but they also make channel feedback expensive: the receiving device must repeatedly report a large complex-valued channel matrix to the base station. ・Most neural compressors treat this matrix like an image and replace it with a fixed-length code that only a matched neural de
cs.LG updates on arXiv.org

GEAR: Generative Expansion and Real Anchoring for Two-Stage Distillation of Tabular Foundation Models

・arXiv:2608.18849v1 Announce Type: new Abstract: Tabular foundation models (TFMs) achieve strong performance through in-context learning, but context-dependent inference imposes substantial latency and memory costs, hindering large-scale deployment. ・We propose GEAR (\emph{Generative Expansion and Real Anchoring}), a modular two-stage framework that distills TFMs into lightweight MLP or tree-based predictors that can b
Qiita - 人気の記事

Genkit Dartでゲーム用のAI会話システムを作る

・Genkit Dart を使うと、Dart言語で生成AIを利用するサーバアプリケーションが簡単に実装できるらしい、ということで試してみました。 ・クライアントは Flutter + Flame のゲーム風サンプルにして、「街を歩いて村人に話しかけると、AI がその人物になりき...
cs.LG updates on arXiv.org

Geometric Data Perturbation with Noisy-Anchor Alignment for Privacy-Preserving Collaborative Learning

・arXiv:2608.18749v1 Announce Type: new Abstract: Geometric Data Perturbation (GDP) enables one-shot, privacy-preserving collaborative learning: each participant applies a distance-preserving transformation to its private data and uploads only the resulting representation to a central analyst. ・We study GDP under analyst-participant collusion, in which the analyst combines all uploaded representations with the private d
cs.LG updates on arXiv.org

Geometric Iterative Retrieval for Neural Audio Codec Resynthesis

・arXiv:2608.19141v1 Announce Type: cross Abstract: Neural audio codecs based on Residual Vector Quantization (RVQ) have become the dominant discrete representation for token-based general audio generation, yet resynthesizing high-quality audio from coarse codec tokens remains an open problem and bounds the fidelity of every system that generates them. ・Prior work has framed resynthesis as a choice between discrete toke
cs.LG updates on arXiv.org

GigaBrain-WBC-0.5: A Behavior World Model for Robust Whole-Body Control with Environment Interaction

・arXiv:2608.18234v1 Announce Type: cross Abstract: Whole-body motion tracking policies turn a humanoid into a robust control interface: the teleoperator---or an upstream model---only supplies a coarse movement intent, while the low-level policy keeps the robot balanced and physically feasible. ・Existing trackers deliver this interface only on flat ground: trained in empty scenes, they never learn how contact with terra
Qiita - 人気の記事

GitHubのTrendingを、AI編集部が無料枠だけで下書きしてくれる仕組みを作った

・はじめに ソーイ株式会社の西浦です。 ・つい最近、GitHubで話題になっているリポジトリをまとめて見られる「Trending」というコンテンツの存在を知りました。眺めてみると普段の業務では出会わないような面白いプロジェクトが並んでいて、これは毎週チェックしたいと思い...
cs.LG updates on arXiv.org

Global Crises and National Policies: A Large Scale Analysis of Political Content in German Language Online Media

・arXiv:2608.18268v1 Announce Type: cross Abstract: Today most media content is consumed based on algorithmic recommendations. ・Evidence suggests that this can lead to politically biased media consumption patterns. ・Automated extraction of political agendas from texts can reveal and analyze political biases in online media -- and thus help fostering politically unbiased media consumption.
ITmedia NEWS 最新記事一覧

Google、大学生向けに「Google AI Plus」を1年間無料提供 「Gemini」アプリに学生向け新機能も

・Googleは米国の新学期に合わせ、大学生向けに「Google AI Plus」などを12カ月無料で提供するキャンペーンを発表した。あわせてGeminiアプリに学生向けハブを新設し、授業資料から学習プランを生成する学習ノートブックや、3Dモデルの表示、Gemini LiveでのDeep Researchなどの新機能の提供を開始した。
cs.LG updates on arXiv.org

Gradient Mirage: Trainable yet Label-Unidentifiable Gradients in Large Language Model Split Learning

・arXiv:2608.18767v1 Announce Type: cross Abstract: Gradient matching attacks (GMAs) in LLM split learning (SL) rely on a critical yet underexplored assumption: the gradient exposed at the split interface is a faithful derivative of the client's full-label training objective. ・This gradient-objective consistency allows a curious server to recover private labels by searching for a sequence whose induced gradient explains
cs.LG updates on arXiv.org

Graph-Based Approaches to Learning Epileptogenic Zone Localization Using Stereo-EEG Recordings

・arXiv:2608.18887v1 Announce Type: new Abstract: The epileptogenic zone (EZ) is the brain region that generates seizures in an individual, and is the target of epilepsy surgery. ・Localizing the EZ from stereo-EEG (sEEG) recordings supports surgical planning, but manual interpretation is time-consuming and focuses on seizure recordings. ・Graphical learning models of resting-state functional connectivity among the recorde
cs.LG updates on arXiv.org

Graphical Design of Interpretable Architectures

・arXiv:2608.18936v1 Announce Type: new Abstract: Designing, implementing, and comparing interpretable architectures requires a formal language to represent them. ・The most common representations fall short in one of two ways. ・Symbolic equations give no global view of an architecture at a glance.
cs.LG updates on arXiv.org

GraphK: Variable-Size Graph Generation with Efficient Edge Construction

・arXiv:2608.18777v1 Announce Type: new Abstract: Graph generation models have advanced significantly with deep learning, yet they remain limited in scalability, flexibility, and ability to model underlying structures. ・We present GraphK, a novel encoder-sampler-decoder framework for graph generation that overcomes these challenges through structural flexibility and computational efficiency. ・Unlike autoregressive approa
Zennの「大規模言語モデル」のフィード

grill-me / grilling をローカルLLM向けに改変してみた話。

・Claude Codeの壁打ち用スキルとしてはsuperpowersのbrainstormingが有名ですが、最近よく耳にするのが「grill-me」。 ・要件定義や新しいスキルを作成するのに超絶便利だっていう話も聞くけど、同時に、一度発動すると容赦なく執拗に質問を投げかけてきて、我々バイブコーダーを奈落の底に突き落とすとかなんとか(違 そんなスキルを怖いもの見たさで導入してみようとGitHubで本家のリポジトリを見たところ、SKILL.mdファイルの中身はたったの1行でした。 ・--- name: grill-me description: A relentless interview t...
AI News & Artificial Intelligence | TechCrunch

Grok keeps sending gibberish responses to users

・Affected users told TechCrunch they were using Grok Lite, and noticed the issues as early as Wednesday morning.
cs.LG updates on arXiv.org

Grouping the Stochastic Machine: Precision, Not Capability, as the Frontier Metric for AI Systems

・arXiv:2608.19140v1 Announce Type: cross Abstract: Frontier language models are compared, marketed, and benchmarked on capability -- what their best or average output can achieve. ・I argue this measures the wrong axis. ・The models have saturated accuracy: their mean output lands on the target.
Zennの「大規模言語モデル」のフィード

GRPOからDAPOへ:RLVR時代のLLM強化学習を数式とメカニズムで理解する

・最初に感想(人間が書いた) GRPOの派生の手法について調べたときにDAPOをみつけたので、それについて解説してもらった。 ・GRPO派生として計算資源をかなり気にしているところは興味深かったです 成功データの意味は大事だけど、失敗データも大事だ、それがないとミニバッチが無駄になるというあたりは面白いところ。 ・GRPOとその派生は万能ではないが、RLVRは今後すくなくとも数理的だったりゴールが明確な問題についてはもっと使われていくだろうという一方で、RLHFが必要なジャンル、特に医療などの分野、あと評価機にLLM-as-a-judgeを使わない手法の類型として学びがあった はじめに...
WIRED

H&R Block Coupon: 25% Off DIY + Tax Pro Assist

・Save over 25% when you opt for H&R Block’s free online offering, plus a tax pro review.
cs.LG updates on arXiv.org

H$^2$EDL: Hyper Evidential Deep Learning for Hierarchical Classification

・arXiv:2608.18185v1 Announce Type: new Abstract: Fine-grained recognition often involves hierarchical label spaces, where a model may be confident about a coarse semantic concept while remaining uncertain among its descendant classes. ・Such structured ambiguity requires uncertainty representations that capture both fine-grained classes and intermediate concepts. ・However, existing tools each capture only half of it: fla
cs.LG updates on arXiv.org

Hallucination Detection in Large Language Models Using Diversion Decoding

・arXiv:2607.10476v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have emerged as a powerful tool for retrieving knowledge through seamless, human-like interactions. ・Despite their advanced text generation capabilities, LLMs exhibit hallucination tendencies, where they generate factually incorrect statements and fabricate knowledge, undermining their reliability and trustworthiness. ・Multiple studi
cs.LG updates on arXiv.org

Harness Continual Learning: Continual Adaptation Beyond Model Parameters

・arXiv:2608.19013v1 Announce Type: new Abstract: Continual learning has largely been model-centric, treating model parameters as the state that changes with sequential experience. ・Modern agents can also adapt through a harness of prompts, memories, tools, skills, and routing rules. ・Because these contents jointly shape later execution, a harness update can disrupt previously reliable behavior even when the model is fro
The Verge

Here’s what data Comcast says its motion-detecting routers collect

・Wi-Fi Motion is an opt-in service from Xfinity that turns your Gateway into a motion-sensing device. ・| Image: Comcast Comcast announced a new platform this week called Xfinity Shield that includes the option to let customers turn their routers into motion sensors. ・Many people reacted with fear and anger.
cs.LG updates on arXiv.org

Hierarchical Classification via Cascading Feature Elimination: Application to Human Phenotype Ontology-Aligned Facial Phenotyping (FaceMesh2HPO)

・arXiv:2607.05585v2 Announce Type: replace-cross Abstract: FaceMesh2HPO is a framework for classifying facial phenotypic descriptors aligned with the Human Phenotype Ontology (HPO) to support clinical diagnosis. ・Using annotations from 124 clinicians across 10 disorders (107 HPO terms) combined with non-syndromic controls, we generated 3D facial meshes (478 landmarks) from 2D images and trained a hierarchical PointNet-
OpenAI News

How ChatGPT Work helps Stampli move ideas to market

・With a fixed deadline and design resources committed elsewhere, Stampli used Codex and ChatGPT Work to compress weeks of launch production into days.
cs.LG updates on arXiv.org

How Quantum Is the Advantage? A Fair, Calibration- and Noise-Aware Benchmark and Attribution Audit of Quantum Machine Learning for Network Intrusion Detection

・arXiv:2608.18155v1 Announce Type: cross Abstract: Quantum machine learning (QML) for network intrusion detection (NIDS) is routinely reported to reach near-perfect accuracy, yet the most rigorous studies find that well-tuned classical models remain competitive, and that apparent quantum gains may be artefacts of classical dimensionality reduction and implicit regularisation rather than genuine quantum effects.
WIRED

How Teen ‘After-Prom’ Kings in LA Monetized the High School Rager

・The West Coast house-party scene has long been iconic. ・For these young, tech-savvy entrepreneurs, it was also inspiration for MyPlots, a party-promotion empire.
AI News & Artificial Intelligence | TechCrunch

Inertia Enterprises finds a way to make its fusion fuel fast

・Fusion power startup Inertia Enterprises reduced the fuel filling process from a week to just a few hours. ・It's one of 10 hurdles the company must overcome to make a profitable power plant.
cs.LG updates on arXiv.org

Inference and Uncertainty Quantification for Streaming $r$-PCA

・arXiv:2608.18374v1 Announce Type: cross Abstract: We address two open questions in streaming PCA via Oja's algorithm: sharp operator-norm convergence for general rank under sub-Gaussian data, and distributional inference for the resulting subspace estimator. ・Existing convergence analyses, even in the rank-one case, either assume bounded data or leave non-vanishing remainder terms that prevent adaptation to a polynomi
cs.LG updates on arXiv.org

Infrared Universality of Collective Dynamics across Transformer and State-Space Architectures

・arXiv:2608.18592v1 Announce Type: new Abstract: Whether distinct neural architectures develop common collective dynamics remains an open question. ・Recent analysis of Transformer language models revealed a nearly flat, weakly infrared-enhanced time-scale density of states (TDOS) associated with near-marginal long-memory dynamics. ・Here we test whether a closely related organization emerges in Mamba, whose selective sta
WIRED

iRobot Promo Code: 15% Off

・Save on iRobot products, including robot vacuums and mops designed to handle pet hair, daily messes, and hands-free cleaning with smart home integration.
The Verge

It’s Greg Brockman’s OpenAI now

・OpenAI has had a hell of a year. ・The company spent months battling former co-founder Elon Musk in a sensational jury trial, was hit with a high-profile trade secrets lawsuit from Apple, and faced widespread scrutiny after an unreleased model hacked another AI company. ・As it prepares for an IPO, a steady string of executives have departed, including some of the company's biggest names.
cs.LG updates on arXiv.org

Iterative Flow Matching: Path Correction and Gradual Refinement for Enhanced Generative Modeling

・arXiv:2502.16445v4 Announce Type: replace Abstract: Generative models for image generation are now commonly used for a wide variety of applications, ranging from guided image generation for entertainment to solving inverse problems. ・Nonetheless, training a generator is a non-trivial feat that requires fine-tuning and can lead to so-called hallucinations, that is, the generation of images that are unrealistic.
Qiita - 人気の記事

IT企業の面接、私服で本当に大丈夫?採用担当が考える「清潔感」のライン

・自己紹介 こんにちは。株式会社PRUMで採用広報を担当している池田です。 ・未経験からIT業界を目指している方と話していると、かなりよく聞く共通した悩みがあります。そういった内容を皆さんにシェアしていくので、役立てていただけたらと思います😊 もしIT業界に興味があるけど、...
cs.LG updates on arXiv.org

Jacobian-Guided Anisotropic Noise Reshaping for Enhancing Representation Utility under Local Differential Privacy

・arXiv:2605.16812v3 Announce Type: replace Abstract: While Local Differential Privacy (LDP) serves as a foundational primitive for distributed data collection, its stringent randomization requirements often lead to severe degradation in data utility. ・This degradation stems from the task-agnostic nature of conventional LDP mechanisms, which perturb all dimensions without accounting for their relative importance to the
cs.LG updates on arXiv.org

Jailbreaking in the Haystack

・arXiv:2511.04707v2 Announce Type: replace-cross Abstract: Recent advances in long-context language models (LMs) have enabled million-token inputs, expanding their capabilities across complex tasks like computer-use agents. ・Yet, the safety implications of these extended contexts remain unclear. ・To bridge this gap, we introduce NINJA (short for Needle-in-haystack jailbreak attack), a method that jailbreaks aligned LMs
ITmedia NEWS 最新記事一覧

JAXA、火星衛星探査機「MMX」を10月20日打ち上げ 帰還すれば世界初の試料回収に

・宇宙航空研究開発機構(JAXA)は8月20日、火星衛星探査計画「MMX」の探査機を、10月20日午前4時41分3秒に主力大型ロケット「H3」10号機で打ち上げると発表した。帰還に成功すれば、火星の衛星からの試料回収は世界初となる。
cs.LG updates on arXiv.org

L\'evy Attention: Single-Pass Predictive Uncertainty for Continuous-Time Attention

・arXiv:2608.19171v1 Announce Type: new Abstract: Deep models for irregularly-sampled time series answer queries at arbitrary continuous timestamps, yet report nothing about how far each answer should be trusted. ・We show the attention layer itself can close that gap: with the right stochastic formulation, the pass that makes each prediction also reports, in closed form and at no extra cost, how far it should be trusted
cs.LG updates on arXiv.org

Leaf Values as Coordinates: Exact Contrastive Explanation for Gradient-Boosted Ensembles

・arXiv:2608.19127v1 Announce Type: new Abstract: A gradient-boosted ensemble predicts by summing one leaf value per tree. ・Read those values as coordinates rather than as intermediate results, and every instance becomes a point in R^M on which the model acts linearly: the score is the sum of the coordinates. ・This small change of view makes contrastive explanation exact.
cs.LG updates on arXiv.org

Learned, Then Lost: A Measured Single-Example Counterfactual in Pre-training

・arXiv:2608.19168v1 Announce Type: new Abstract: A single training example's contribution to a finished model is normally estimated rather than measured, because measuring it takes two expensive full pre-training runs that differ in one row of one batch. ・We ran that counterfactual 24 times at a small scale. ・We trained 32 GPT-2 models at 124M parameters from scratch on OpenWebText, over four conditions and eight seeds.
cs.LG updates on arXiv.org

Learning Canonical Register Automata over Ordered Data Domains

・arXiv:2608.18765v1 Announce Type: cross Abstract: Register automata are finite automata equipped with memory that recognize data languages over infinite alphabets. ・In this work, we investigate active learning algorithms for deterministic register automata (DRAs) over ordered data domains--covering both dense domains, such as the rationals, and non-dense domains such as the integers. ・We show that the active learning p
cs.LG updates on arXiv.org

Learning Random Geometric Graphs Drawn in Probabilistic Metric Spaces

・arXiv:2608.19082v1 Announce Type: cross Abstract: We present a new data-driven learning of a Random Geometric Graph (RGG) of a multivariate dataset, where the graph is drawn in a probabilistic metric space. ・This graph learning works for generic datasets, irrespective of the type of the observables; their probability distributions; or size of the data. ・We identify a metric of the space that the graph is drawn in, as a
cs.LG updates on arXiv.org

Learning Topological Features of $\widehat Z$-invariants

・arXiv:2608.18570v1 Announce Type: cross Abstract: Machine learning and data analysis techniques have recently emerged as powerful tools for identifying patterns and formulating conjectures in mathematical research, most notably in the field of low-dimensional topology. ・In this paper, we initiate a systematic approach to handling mathematical data structured as (truncated) infinite $q$-series, or equivalently, infinit
The Verge

LG’s 65-inch B6 OLED is $300 lower than its previous best price

・Best Buy has a whole host of LG products discounted for today’s Deal of the Day festivities, including laptops, home theater gear, and more. ・One that stood out is a price cut on LG’s 65-inch B6 OLED TV that’s deeper than I’ve seen before. ・At Best Buy, it’s $1,399.99 (originally $1,999.99) for the rest of the Thursday, August 20th, which is currently $300 lower than Amazon’s price.
cs.LG updates on arXiv.org

LionMuon: Alternating Spectral and Sign Descent for Efficient Training

・arXiv:2605.19811v3 Announce Type: replace Abstract: In large-scale optimization, the cheapness and effectiveness of update steps are the most crucial factors for a successful optimizer. ・Sign-based optimizers like Lion or Signum produce cheap per-step updates, whereas Muon's spectral matrix-sign update gives a much stronger direction at a substantially higher per-step cost. ・In this work, we propose LionMuon, which ret
Zennの「大規模言語モデル」のフィード

LLM プロバイダを抽象化したら、効いたのは「乗り換えの自由」じゃなかった

・結論 / TL;DR 複数の LLM プロバイダを 1 つのインタフェースで束ねる薄い層を入れました。動機は「ベンダーロックインの回避」でしたが、実際に効いたのは別の 2 つでした。 ・可用性: プロバイダ側が 5xx を返しても、別プロバイダに落として機能ごと止めない コスト: 「どの機能にどのモデルを使うか」を設定で切り替えられる そして、やってみて分かった一番大事な設計指針は 「抽象化を厚くしない」 ことです。全プロバイダの最大公約数に揃えると各社の強みが死にます。正規化するのは 5 項目だけ、残りは素通しの脱出ハッチを開けておく。この形に落ち着きました。
cs.LG updates on arXiv.org

LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4

・arXiv:2607.15509v3 Announce Type: replace-cross Abstract: We present a fully automated closed-loop AutoML framework that uses GPT-5, GPT-4o, and Claude Sonnet 4 as autonomous neural architecture designers for cross-lingual handwritten optical character recognition. ・Each large language model independently generates, trains, evaluates, and iteratively refines neural network architectures using performance feedback from
cs.LG updates on arXiv.org

LLM-Powered Predictive Decision-Making for Sustainable Data Center Operations

・arXiv:2608.18503v1 Announce Type: new Abstract: The growing demand for AI-driven workloads, particularly from Large Language Models (LLMs), has raised concerns about the significant energy and resource consumption in data centers. ・This work introduces a novel LLM-based predictive scheduling system designed to enhance operational efficiency while reducing the environmental impact of data centers. ・Our system utilizes a
Zennの「大規模言語モデル」のフィード

LLM・生成モデルの推論高速化技術の全体像(多分)

・はじめに こんにちは。動詞(@IMG_5955)です。 ・LLMや画像・動画生成などの生成モデルを実運用やローカル環境で動かす際、最大のボトルネックとなるのが「推論の遅さ」と「GPUリソースの制約」です。 ・一口に推論の課題と言っても、「最初の1文字が出るまでが遅い」「文字の出力速度そのものが遅い」「モデルがGPUメモリ(VRAM)に収まらない」「同時に複数のリクエストが来ると待ち行列が詰まる」など、ボトルネックの位置によって打つべき手はまったく異なります。
#LLMタグ

LLM#1 temperature 0.7 の意味を、私は説明できなかった

・連載「LLMの仕組みを、作りながら理解する」 第1回 temperature 0.7 の意味を、私は説明できなかった どうもです!えむしんです。 ・LLMの仕組みがわかるようになりたく、Grok や Claude などに教えてもらいながら理解していこうと思います。 ・記事をAIと一緒に書きながら、進めて行こうと思います。相手がAIなので、うまく収束できればよいのですが💦 (AIで連載を書くと、落ちどころがなくて薄っぺらいものになりがちなので、気をつけます) 続きをみる
Hugging Face Papers

LLMs Get Smarter from Targeted Synthetic Multilingual Data

LLMs Get Smarter from Targeted Synthetic Multilingual Data
#LLMタグ

LLMにLLMを評価させてみた

LLMにLLMを評価させてみた
#LLMタグ

LLMにも効く「アンカリング効果」──AnchorBenchが教える運用上の注意点

・⚠️ 本記事に関するご注意(免責事項) ・ AI生成コンテンツ: 本記事のテキストおよび解説図解は、学術論文をもとに生成AIを用いて作成・要約したものです。正確性には万全を期しておりますが、AIの性質上、解釈や翻訳に誤りを含む可能性があります。正確な情報や詳細については、必ず記事末尾の原著論文(出典リンク)をご確認ください。
#LLMタグ

LLMの臨床エラー検出評価に新提案:F1値の落とし穴とペア評価の重要性

・AIの進化が目覚ましい昨今、医療現場での活用にも大きな期待が寄せられていますよね。特に、膨大な医療文書の中から、人間では見落としがちなエラーをAIが見つけ出すなんて、まさに未来の医療だと感じませんか?この可能性には計り知れない魅力を感じます。 ・しかし、その裏側には、実は一筋縄ではいかない「評価」という課題が潜んでいるんです。単なる技術導入だけでなく、その真価を引き出すためには、私たちが想像する以上の深い理解が必要なんですよ。数字だけでは見えない、もっと大切な判断基準がある。それは、AIの真の能力を引き出すための―― 続きをみる
Zennの「大規模言語モデル」のフィード

LLM導入で失敗する原因はモデルだけではない:開発前に確認したい4つの層

・新しいLLMが公開されるたびに、 「このモデルは性能が高い」 「ベンチマークで最高スコアを出した」 「料金も安い」 という情報が流れます。 ・そのため、LLM導入では最初にモデル比較から始めることが多いと思います。 ・実際、私も以前はそう考えていました。
Hugging Face Papers

Looped Language Models Improve Compositional Tool Calling

Looped Language Models Improve Compositional Tool Calling
cs.LG updates on arXiv.org

Looped Language Models Improve Compositional Tool Calling

・arXiv:2608.18171v1 Announce Type: cross Abstract: Looped language models have shown promising results on reasoning benchmarks, yet their potential for agentic tool use remains largely unexplored. ・We study this question in compositional tool-calling settings, where models must coordinate multiple API calls, maintain intermediate state, and preserve dependencies across tool interactions. ・We evaluate native and retrofit
cs.LG updates on arXiv.org

Lost in Aggregation: How Benchmarks Overlook Irreplaceable Model Strengths

・arXiv:2608.18919v1 Announce Type: new Abstract: Tabular machine learning benchmarks typically summarize performance by averaging scores, ranks, or pairwise wins across datasets. ・Such aggregates are useful for selecting robust default models, but they can obscure a different question: which models are necessary to attain peak performance on particular datasets? ・We argue that benchmark evaluation should also consider t
cs.LG updates on arXiv.org

Low-Power, Neuromorphic, Acoustic Anomaly Detection for Persistent Machine Monitoring

・arXiv:2608.18341v1 Announce Type: cross Abstract: Persistent acoustic monitoring can detect machine faults without physical contact, but always-on inference is constrained by power, latency, and deployment complexity. ・We demonstrate autoencoder-based acoustic anomaly detection on an Intel Loihi 2 neuromorphic processor under clean and noisy conditions. ・Log-mel features are computed off chip; normalization, autoencode
cs.LG updates on arXiv.org

Many Optimizers But Only One Training Path: Repeated Resampling for Adaptive Optimizer Selection

・arXiv:2608.18810v1 Announce Type: new Abstract: An optimizer is usually chosen before training a deep neural network and then kept fixed. ・Treating optimizer choice as a hyperparameter could boost performance, but it requires several complete training runs and discards all but the winner. ・Repeated Optimizer Resampling (ROR) instead searches during one evolving run.
cs.LG updates on arXiv.org

MARCUS: Missing-Aware Region Representation with Contextual Urban Signals for Rent Prediction

・arXiv:2608.18546v1 Announce Type: new Abstract: Multimodal urban data has expanded the applications of urban region representation learning, such as functional zone identification and real estate appraisal, but also introduces challenges caused by data incompleteness. ・Existing studies usually handle missing data through imputation, treating missingness as noise while ignoring its potential semantic value.
cs.LG updates on arXiv.org

Mask2Real-WM: Segmentation Masks as a Sim-to-Real Bridge for Controllable Dexterous World Models

・arXiv:2607.04546v2 Announce Type: replace-cross Abstract: Action-conditioned world models allow robots to predict the future consequences of candidate actions without additional physical interaction, supporting policy evaluation, planning, and data augmentation. ・We present Mask2Real-WM, a two-stage action-conditioned world model for dexterous manipulation that decouples pixel prediction into a dynamics model and a re
cs.LG updates on arXiv.org

Matching Accuracy, Different Geometry: Evolution Strategies vs GRPO in LLM Post-Training

・arXiv:2604.01499v3 Announce Type: replace Abstract: Evolution Strategies (ES) have emerged as a scalable gradient-free alternative to reinforcement learning based LLM fine-tuning, but it remains unclear whether comparable task performance implies comparable solutions in parameter space. ・We compare ES and Group Relative Policy Optimization (GRPO) across four tasks in both single-task and sequential continual-learning
cs.LG updates on arXiv.org

MAVEN: A Macro-Societal Value Evaluation Framework of Multimodal Content with Compact Aligned Evaluators

・arXiv:2608.18096v1 Announce Type: cross Abstract: Assessing whether multimodal content aligns with macro-societal values, such as peace, justice, and freedom, has become an increasingly urgent challenge. ・Existing frameworks are largely confined to safety-oriented taxonomies, text-only psychometric probes, or single-label classification. ・Therefore, we propose MAVEN, a hierarchical framework for macro-societal value ev
cs.LG updates on arXiv.org

Mechanistic Interpretability of Structure-Aware Numerical Reasoning in LLaMA 3.1 8B

・arXiv:2608.18419v1 Announce Type: new Abstract: Recent work has shown that large language models (LLMs) exhibit strong numerical sequence modeling capabilities and show promise in time-series prediction. ・While LLMs display in-context learning capabilities, the mechanisms with which they accomplish time-series prediction remain unclear. ・Specifically, whether they truly understand the underlying structure, which at a m
cs.LG updates on arXiv.org

MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports

・arXiv:2605.03103v2 Announce Type: replace-cross Abstract: Semi-structured information extraction (IE) from OCR-derived clinical reports is crucial for efficiently reconstructing patients' longitudinal medical histories. ・In practice, this scenario commonly involves three tasks: (i) field-header (key) discovery, (ii) key-conditioned question answering (QA), and (iii) end-to-end key-value pair extraction. ・However, exist
AI News & Artificial Intelligence | TechCrunch

Meta AI’s new Mac app wants you to talk to your apps

・The company said that the dictation feature works across all apps, just like other tools such as Wispr Flow, Superwhisper, and Monologue.
AI News & Artificial Intelligence | TechCrunch

Meta brings Pocket, an app that lets you vibe-code and share games, to US users

・Meta is bringing Pocket, its experimental AI-powered app for creating and sharing interactive games, to users across the U.S. ・after quietly testing it in Brazil.
@IT 全フォーラム 最新記事一覧

MFAは「9割超が導入」、それでも8割が「IDaaSのリプレースを考える」切実な理由

・エムオーテックスは「多要素認証および条件付きアクセスに関する実態調査」の結果を公表した。9割超の企業がIDaaSの導入・検討を進める一方、約8割がリプレースを検討・実施しているという。
cs.LG updates on arXiv.org

MIFR: A Modality-Invariant and Fair Representation Framework for Skin Disease Classification

・arXiv:2608.18774v1 Announce Type: cross Abstract: Skin diseases represent a major global public health burden, yet machine learning tools developed to assist in their diagnosis suffer from two critical limitations: reliance on only one modality for diagnosis and systematic performance disparities across skin tones. ・While existing approaches address each challenge separately, this work proposes a modality-invariant fr
cs.LG updates on arXiv.org

Mitigating Spectral Bias in Neural Operators for Underwater Transmission Loss Prediction

・arXiv:2608.18141v1 Announce Type: cross Abstract: Predicting underwater acoustic transmission loss rapidly and accurately is crucial for real-time ocean acoustic applications. ・While Fourier Neural Operators (FNO) have emerged as powerful surrogate models due to their global receptive fields, they suffer from spectral bias. ・The frequency truncation mechanism in FNO filters out high-frequency components, resulting in o
cs.LG updates on arXiv.org

MLREF: Efficient Module Reuse for Reward Design in Reinforcement Learning via Large Language Models

・arXiv:2608.18827v1 Announce Type: new Abstract: Reward function design remains a bottleneck in reinforcement learning. ・While large language models (LLMs) have enabled automated reward generation, existing methods generate and revise reward functions as monolithic programs, making it difficult to reliably preserve and reuse effective components discovered in earlier iterations, leading to unstable performance across i
stat.ML updates on arXiv.org

MMD-Flagger: Leveraging Maximum Mean Discrepancy to Detect Hallucinations

・arXiv:2506.01367v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly integrated into agentic AI systems, yet their propensity to generate hallucinations remains a critical safety concern. ・Detecting these factual errors at test-time, particularly without ground-truth labels, is essential for building trustworthy autonomous agents. ・We propose MMD-Flagger, an hallucination detection me
cs.LG updates on arXiv.org

Model Card for OpenAI Privacy Filter

・arXiv:2608.18274v1 Announce Type: cross Abstract: OpenAI Privacy Filter is a compact, bidirectional token-classification model for detecting and redacting personally identifiable information (PII) and secrets in unstructured text. ・The model is derived from an autoregressively pretrained checkpoint and converted into a bidirectional, banded-attention classifier that labels an input sequence in a single forward pass.
cs.LG updates on arXiv.org

Monroe: A Molecular Foundation Model for In-Context Probabilistic Inference

・arXiv:2608.18982v1 Announce Type: new Abstract: Bioassay activity prediction is often data-limited because drug-discovery datasets rely on time-consuming and expensive wet-lab experiments for data generation and evaluation. ・This challenge has inspired recent research into molecular foundation models (MFMs), which aim to encode general-purpose chemical knowledge into molecular representations that generalize well in d
cs.LG updates on arXiv.org

MorphoGP: A Nonparametric Framework for Predicting Equilibrium Beach Profiles Under Tidal Influence

・arXiv:2608.18558v1 Announce Type: new Abstract: The prediction of equilibrium beach profiles under tidal influence is of fundamental importance for sustainable coastal development, informing shoreline protection strategies and managing coastal ecosystems under changing environmental conditions. ・However, it remains challenging due to the highly nonlinear interactions among wave, tide, and sedimentary processes.
cs.LG updates on arXiv.org

Multi-Agent Off-Policy Deep Reinforcement Learning for Smart Campus Coverage

・arXiv:2608.19049v1 Announce Type: new Abstract: Deep reinforcement learning (DRL) has recently gained a great attention due to its real-time adaptation and effectiveness in complex optimization problems. ・This paper investigates the optimal deployment of millimeter-wave (mmWave) base stations (BSs) in a realistic, non-convex campus topology. ・The optimization problem is NP-hard, due to the non-convex, non-smooth nature
cs.LG updates on arXiv.org

Multi-Class Electrical and Mechanical Fault Classification Using Random Convolutional Kernels

・arXiv:2608.18716v1 Announce Type: new Abstract: Diagnosing faults in rotating machinery is essential for ensuring the reliability of industrial processes. ・Random convolutional kernel-based Time Series Classification (TSC) methods, such as ROCKET and its variants, provide an attractive trade-off between predictive performance and computational efficiency. ・In this work, we evaluate SelF-Rocket for the multi-class diagn
stat.ML updates on arXiv.org

Multi-Level Bayesian Calibration of a Multi-Component Dynamic System Model

・arXiv:2608.18430v1 Announce Type: cross Abstract: This paper proposes a multi-level Bayesian calibration approach that fuses information from heterogeneous sources and accounts for uncertainties in modeling and measurements for time-dependent multi-component systems. ・The developed methodology has two elements: quantifying the uncertainty at component and system levels, by fusing all available information, and correct
cs.LG updates on arXiv.org

Multi-Objective Optimization Under Uncertainty of Part Quality in Fused Filament Fabrication

・arXiv:2608.18429v1 Announce Type: cross Abstract: This work presents a data-driven methodology for multi-objective optimization under uncertainty of process parameters in the fused filament fabrication (FFF) process. ・The proposed approach optimizes the process parameters with the objectives of minimizing the geometric inaccuracy and maximizing the filament bond quality of the manufactured part. ・First, experiments are
cs.LG updates on arXiv.org

Multi-stage neural operator learning with application for convolutions

・arXiv:2608.18851v1 Announce Type: new Abstract: Convolution integrals widely exist in applications, and to enable fast and accurate computations, this paper introduces two general multi-stage neural operator learning frameworks. ・The first, Deep Collocation Neural Operator (DCNO), is a supervised approach that iteratively refines the operator approximation by learning residuals from input-output data pairs.
Zennの「大規模言語モデル」のフィード

n8nでキーワード監視botを作る——GitHub/Hacker News/Hugging Faceの新着をTelegramへ自動通知

・AI関連の情報は流れが速く、GitHubの新しいリポジトリも、Hugging Faceの新モデルも、気づいたときには一週間前の話になっています。そこで、自分の関心キーワードだけを複数のソースから常時見張り、新着があればTelegramに通知する仕組みを、ワークフロー自動化ツール n8n で作りました。 ・この記事では、その実装手順と、実際に詰まった箇所(とその回避策)をまとめます。派手な機能紹介ではなく、手を動かすと必ず出会う地雷を先に踏んでおく、という趣旨です。 ・なぜ n8n か n8nはGitHubで20万スターを超える、セルフホスト型のワークフロー自動化ツールです。画面上でノード...
cs.LG updates on arXiv.org

NanoSleep: A Parameter-Efficient Hybrid Temporal Convolutional Network for Single-Channel Sleep Stage Classification

・arXiv:2608.18571v1 Announce Type: new Abstract: Sleep stage classification from single-channel electroencephalography (EEG) is essential for wearable and home-based sleep monitoring. ・However, many deep learning models achieve high accuracy at the cost of large model sizes, which limits their deployment on resource-constrained devices. ・In this work, we present NanoSleep, a compact hybrid temporal convolutional network
#AIタグ

New アイコン

・AI様のお力添えです 自称ビールクイーンの私👑 ビールを飲んで 旨ーーーーっ🍺💗 って言っている写真をアニメチックにしたら こんなに可愛くしてくれました🫪 続きをみる
#AIタグ

noteの読者を10倍にするX(Twitter)連携戦略|投稿設計から導線設計まで完全ガイド🔥

・☆彡検索だけでは限界がある SEOで検索流入を増やすことは重要ですが、新しい記事が検索上位に表示されるまでには時間がかかります。
cs.LG updates on arXiv.org

Off-Manifold Collapse in Guided Protein Language Models

・arXiv:2608.18597v1 Announce Type: new Abstract: Protein language models are widely used priors for protein sequence design, and a growing body of work controls them at inference time as an alternative to fine-tuning. ・Such guidance faces a dilemma: mild enough to preserve natural activation statistics, it barely moves the property; strong enough to move it, the generations become progressively harder to fold.
cs.LG updates on arXiv.org

On the Power of Source Screening for Learning Shared Feature Extractors

・arXiv:2602.16125v3 Announce Type: replace Abstract: Learning with shared representation is widely recognized as an effective way to separate commonalities from heterogeneity across various heterogeneous sources. ・Most existing work includes all related data sources via simultaneously training a common feature extractor and source-specific heads. ・It is well understood that data sources with low relevance or poor qualit
cs.LG updates on arXiv.org

On the Robustness of Vision-Language Models in Zero-shot Privacy Classification

・arXiv:2510.09253v2 Announce Type: replace-cross Abstract: Automatic systems for document understanding require multimodal models that accurately identify sensitive visual content, even in the presence of image degradations. ・Instruction-following large Vision-Language Models (VLMs) are expected to generalise across domains and tasks without requiring any specific adaptation. ・In this work, we systematically analyse whe
cs.LG updates on arXiv.org

On the Slow Convergence to Trivial Solutions of Algorithms for Hard Optimization Problems

・arXiv:2608.18910v1 Announce Type: new Abstract: Hard combinatorial optimization problems, many of which are NP-hard, present fundamental algorithmic challenges. ・Average-case analysis on random instances has emerged as a powerful framework for understanding typical algorithmic performance beyond worst-case guarantees. ・A substantial body of work has established negative results: for sufficiently hard instances (often c
cs.LG updates on arXiv.org

Online Bipartite Matching with Reusable Capacity under Non-Stationary Rewards

・arXiv:2608.18130v1 Announce Type: cross Abstract: We study online bipartite matching with reusable server capacity and non-stationary rewards. ・Jobs arrive sequentially, reveal compatible servers, reward rates, and processing durations, and must be accepted or rejected irrevocably. ・An accepted job occupies one unit of server capacity only during its processing interval, so an assignment may displace an unknown sequenc
cs.LG updates on arXiv.org

Online Conformal Anomaly Detection with Prediction-Powered Data Acquisition

・arXiv:2505.01783v2 Announce Type: replace Abstract: Online anomaly detection is essential in fields such as cybersecurity, healthcare, industrial monitoring, and telecommunications, where promptly identifying deviations from expected behavior can avert critical failures or security breaches. ・While numerous anomaly scoring methods based on supervised or unsupervised learning have been proposed, the only existing appro
cs.LG updates on arXiv.org

Online Learning for Dynamic Constellation Topologies

・arXiv:2603.25954v2 Announce Type: replace Abstract: The use of satellite networks has increased significantly in recent years due to their advantages over purely terrestrial systems, such as higher availability and coverage. ・However, to effectively provide these services, satellite networks must cope with the continuous orbital movement and maneuvering of their nodes and the impact on the network's topology.
cs.LG updates on arXiv.org

Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation

・arXiv:2608.19098v1 Announce Type: new Abstract: Multi-teacher on-policy distillation (M-OPD) has emerged as a promising paradigm for consolidating domain-specialized reinforcement learning (RL) experts into a single generalist student via dense, token-level reward supervision. ・Despite its practical success, the optimization dynamics governing multi-teacher capability integration remain poorly understood, and open, ri
cs.LG updates on arXiv.org

Optimizing Energy Efficiency and Grid Stability via Public EV Charging Flexibility

・arXiv:2608.18126v1 Announce Type: cross Abstract: This study evaluates the potential of electric vehicle (EV) charging flexibility to enhance both energy efficiency and power grid stability. ・Using real-world data from public charging stations in Prague, we analyze individual and aggregated charging sessions to explore how optimizing charging times can reduce energy waste, minimize grid imbalances, and support the int
cs.LG updates on arXiv.org

Pairwise Ranking Outperforms Single-Action RL for Offline Explanation Selection: A Practical Lesson

・arXiv:2608.18531v1 Announce Type: cross Abstract: Industrial explainable-recommendation systems built on LLMs incur a substantial serving cost: each request triggers an LLM generation, with latency in the hundreds of milliseconds and cost that scales linearly with traffic. ・We separate generation from selection: explanations are produced ahead of time as a frozen candidate pool (six prompt styles, two commodity LLMs),
cs.LG updates on arXiv.org

Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models

・arXiv:2608.18484v1 Announce Type: cross Abstract: Training-free block-sparse attention can accelerate video transformers, but row-wise attention concentration does not by itself specify an executable sparse operator. ・Queries sharing a block route may have poorly overlapping supports, while retained attention mass alone does not determine the post-softmax error from skipped interactions. ・We show that partition geometr
WIRED

Peacock Promo Codes: 40% Off August 2026

・Stream your favorite shows for up to $80 off this month, and save on subscription plans with the latest Peacock TV coupons from WIRED.
cs.LG updates on arXiv.org

Pedagogical AI in Mental Health: A Tri-Stream Fine-Tuned LLM Framework for Automated Clinical Supervision and Risk Triage

・arXiv:2608.18438v1 Announce Type: cross Abstract: Modern mental healthcare faces a critical shortage of senior supervisory oversight, leading to a "supervision gap" where novice therapists manage high-stakes risks with delayed professional feedback. ・This paper proposes a new framework utilizing a fine-tuned Mistral-7B-instruct model as an automated "Supervisor-in-the-Loop" system. ・By leveraging 106 sessions from the
cs.LG updates on arXiv.org

Performance Drift Detection in Machine Learning as a Service (MLaaS) for IoT Environments

・arXiv:2608.18555v1 Announce Type: new Abstract: Machine Learning as a Service (MLaaS) is a powerful cloud paradigm enabling data-driven intelligent applications in Internet of Things (IoT) environments, widely adopted across healthcare, smart homes, and industry due to its cost-effectiveness. ・However, the dynamic nature of IoT frequently alters data distributions, affecting MLaaS stability, while periodic MLaaS updat
cs.LG updates on arXiv.org

PGFS++: Molecular Property Improvement under Synthesis and Diversity Constraints

・arXiv:2608.19121v1 Announce Type: new Abstract: Improving molecular properties, such as drug-likeness or binding affinity, is a recurring task in early-stage drug discovery. ・However, molecules optimized in an unconstrained chemical space have limited practical value if they cannot be synthesized. ・Policy Gradient for Forward Synthesis (PGFS) is a synthesis-aware reinforcement learning method for molecular improvement,
cs.LG updates on arXiv.org

Physics-Unrolled Neural Operator for Wireless Field Modeling

・arXiv:2608.18495v1 Announce Type: new Abstract: Radio maps are essential for wireless decision-making tasks such as access-point placement, coverage planning, and localization, but their fine spatial details are governed by complex propagation effects and are costly to simulate accurately. ・Machine learning offers a path to high-fidelity radio-map prediction without running expensive high-fidelity simulations for ever
WIRED

Poolease X1 Pool Robot Review: How Bad Can It Be?

・The Poolease X1 can collect leaves from the pool floor, but dirt, silt, walls, and steps are all beyond its pay grade.
cs.LG updates on arXiv.org

Position: Current Model Cards Are Insufficient for Downstream Governance of Open-Weight Foundation Models

・arXiv:2608.18086v1 Announce Type: cross Abstract: The growth of open-weight foundation models (OWFMs) has prompted the AI community to re-evaluate strategies for effective downstream governance. ・Although model cards have been widely adopted as transparency artifacts in model repositories, existing frameworks often fail to adequately inform downstream developers and users about the distinct safety challenges posed by
cs.LG updates on arXiv.org

Pre-Training for Simulation-Based Science: A Study on Jet Foundation Model Training Objectives

・arXiv:2606.14870v2 Announce Type: replace-cross Abstract: Foundation models (FMs) trained on large datasets and fine-tuned on downstream tasks have emerged as a powerful paradigm in AI for science. ・Industrial FMs are typically trained using self-supervision with masking due to the lack of labels. ・In many scientific domains, accurate simulations are plentiful and facilitate large, labeled datasets.
cs.LG updates on arXiv.org

Preference Reasoning under Indeterminacy in Large Language Models

・arXiv:2608.18631v1 Announce Type: cross Abstract: As large language models evolve into decision-making agents, the ability to reason over preferences becomes fundamental to alignment, coordination, and collective intelligence. ・Yet, unlike standard benchmarks, real-world preference reasoning is inherently indeterminate: information may be incomplete, and valid solutions may not exist. ・We argue that indeterminacy, rath
cs.LG updates on arXiv.org

Pretraining Reusable Inference Across Views with Synthetic Task Priors

・arXiv:2608.19115v1 Announce Type: new Abstract: Modern pretrained encoders make representations from heterogeneous views increasingly reusable, but the procedure that determines view utility and combines evidence is still relearned for each downstream task. ・Consequently, knowledge about view relevance, complementarity, reliability, and missingness is repeatedly discarded rather than transferred across tasks.
Hugging Face Papers

Previous

Previous
cs.LG updates on arXiv.org

Process Optimization Under Uncertainty for Improving the Bond Quality of Polymer Filaments in Fused Filament Fabrication

・arXiv:2608.18431v1 Announce Type: cross Abstract: This paper develops a computational framework to optimize the process parameters such that the bond quality between extruded polymer filaments is maximized in fused filament fabrication (FFF). ・A transient heat transfer analysis providing an estimate of the temperature profile of the filaments is coupled with a sintering neck growth model to assess the bond quality tha
cs.LG updates on arXiv.org

Progressive Experience Fusion for Multi-Task World Model Control in Endovascular Navigation

・arXiv:2608.18647v1 Announce Type: cross Abstract: Autonomous endovascular navigation could support the delivery of mechanical thrombectomy to underserved areas, but controllers must navigate long, multi-stage paths across varying vascular anatomies. ・This study investigates Progressive Experience Fusion (PEF) to train a multi-task TD-MPC2 controller. ・We additionally evaluate a heuristic that changes the Model Predicti
cs.LG updates on arXiv.org

ProxyGuard: Direct Reliability Inference for Randomized Data Release Mechanisms with Shared Targets

・arXiv:2608.18643v1 Announce Type: new Abstract: Researchers often choose a proxy dataset from many releases, transformations, or seeds. ・Search can make an invalid release appear adequate, while one adequate release does not establish that its generator is reliable. ・ProxyGuard controls both errors using prespecified bounded risks and a sealed target set.
cs.LG updates on arXiv.org

Quantum Tensor Network Learning with DMRG

・arXiv:2608.18901v1 Announce Type: cross Abstract: Tensor Networks are a relatively new machine learning approach. ・The architectures proposed initially are inspired by approaches from quantum many-body physics simulations. ・One common layout is the matrix product state (MPS) also known as a tensor train optimized with gradient descent techniques.
cs.LG updates on arXiv.org

Quantum-Logic Tsetlin Machines: Interpretable Quantum Machine Learning with Commuting Projector Clauses

・arXiv:2608.18659v1 Announce Type: cross Abstract: Tsetlin Machines (TMs) learn interpretable Boolean clauses using finite-state automata. ・We introduce the Quantum-Logic Tsetlin Machine (QL-TM), which replaces Boolean literals with quantum propositions represented by projectors while retaining classical include/exclude automata. ・Clauses are restricted to commuting measurement contexts and activate through the Born pro
AI News & Artificial Intelligence | TechCrunch

Ramp launches its own AI model router, called Router

・Ramp has launched its own AI model routing service, dubbed Router, that lets users and companies use and switch between various large language models via an API.
cs.LG updates on arXiv.org

Redakto - The Incognito Tab for LLMs

・arXiv:2608.18260v1 Announce Type: cross Abstract: Large Language Models (LLMs) are being increasingly used in everyday applications. ・A major challenge in the context of LLMs or Artificial Intelligence (AI) in general is to ensure privacy when using them, meaning that personally identifiable information (PII) is removed from any text that enters an LLM. ・These challenges have become more urgent with novel EU legislatio
cs.LG updates on arXiv.org

Regularised Iterative Generalised Least Squares with Optimal Selection of the Hyper-Parameter for Identifying Nonlinear Phenomenological Models

・arXiv:2608.18742v1 Announce Type: cross Abstract: In some fields currently dominated by empirical approaches, such as state of health (SoH) prediction for lithium-ion batteries, phenomenological models motivated by quasi-physical thinking contain parameters to be estimated from experimental data. ・Often the structure of such models yields fully or partially confounded parameters, which are difficult or even impossible
cs.LG updates on arXiv.org

Reinforced Planning with Latent World Models

・arXiv:2608.18669v1 Announce Type: new Abstract: Humans solve complex problems by constructing plans and mentally simulating their outcomes with an internal model of the world. ・Machine learning has produced world models that similarly predict the outcomes of action sequences, but the improvement of candidate plans still isn't fully learned. ・Current planners are either hand-designed, distilled from a hand-designed opti
cs.LG updates on arXiv.org

Rethinking Privileged Information in On-Policy Self-Distillation

・arXiv:2608.18271v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) trains a student on its own responses using token-level supervision from the same model conditioned on privileged reference information. ・We investigate whether performance gains from OPSD show that the student learned the information in the reference or instead reflect recovery of reasoning behavior already present in the base model.
cs.LG updates on arXiv.org

Robust Risk Under Evolving Uncertainty: A Wasserstein Counterpart of the Entropic Value-at-Risk

・arXiv:2608.19073v1 Announce Type: cross Abstract: An agent still learning its environment should be cautious while ignorant and bold once confident. ・The entropic value-at-risk captures this through a robust-optimization identity---a confidence level fixes the radius of a relative-entropy ball of alternative models---but that ball cannot reach catastrophes the nominal deems impossible, precisely what a safe agent must
cs.LG updates on arXiv.org

Role-Conditioned Sub-Token Routing for Efficient Vision-Language-Action Policies

・arXiv:2608.18410v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models process long multimodal token sequences, making inference expensive in both memory and computation. ・Existing efficiency methods mainly reduce visual tokens, but aggressive token pruning becomes fragile because removing a token discards its entire representation. ・Sub-token compression provides a complementary alternative by retaining m
cs.LG updates on arXiv.org

RouteCost: A Production-Inspired Multi-Stage Framework for Pre-Order Shipping Cost Estimation in E-Commerce

・arXiv:2607.16230v2 Announce Type: replace Abstract: Accurate pre-order shipping cost estimation is important in e-commerce because it affects price presentation, margin planning, and conversion. ・In practice, shipping cost is shaped not only by distance but also by destination demand mix, billable weight, dimensional pricing, surcharge triggers, and latent operational effects such as shipment consolidation.
cs.LG updates on arXiv.org

Safe Domain Adaptation for Physics: Overcoming Nuisances, Label Shifts, and Simulation Priors

・arXiv:2608.18190v1 Announce Type: new Abstract: Domain adaptation is widely used to make neural networks trained on simulations applicable to experimental data. ・Its premise is that the two domains differ only in nuisances, and that the quantity of interest is distributed identically in both. ・In physics neither assumption holds: simulations can be wrong about the physics, and the distribution of the target quantity -
WIRED

Salmonella Is Everywhere

・From granola to guacamole, a wide variety of foods have been recalled lately over salmonella risks. ・Experts say common-sense precautions go a long way.
stat.ML updates on arXiv.org

Scalable Amortized Variational Inference for Non-Poisson Buy-'Til-You-Die Models

・arXiv:2608.19022v1 Announce Type: cross Abstract: Despite the wide variety of existing Buy-`Til-You-Die (BTYD) models, nearly all rely upon the convenient assumption of transactions following a Poisson process. ・As modern customer bases grow larger and more diverse, a major gap in the marketing literature is BTYD models that can account for heterogeneity in timing patterns across millions of customers. ・This paper addr
cs.LG updates on arXiv.org

Scalable Geospatial Machine Learning for Power-Line Asset Risk: Integrating Remote Sensing for Lightning and Vegetation Risk Modelling

・arXiv:2608.18611v1 Announce Type: new Abstract: Electric power networks are increasingly exposed to weather-sensitive failure mechanisms that require asset-level, spatially explicit risk modelling for effective intervention planning. ・This study contributes a modular, robust, and explainable probability-of-failure (PoF) modelling framework for utility asset management. ・The central contribution is an asset-level archit
Hugging Face Papers

Scaling Creative Writing Beyond Story-Centric Data with Attribute-Guided Genre Expansion

Scaling Creative Writing Beyond Story-Centric Data with Attribute-Guided Genre Expansion
cs.LG updates on arXiv.org

Score the Algebra, Not the Span: Dimension Reduction for Transfer Operator Models of Dynamical Systems

・arXiv:2608.18918v1 Announce Type: new Abstract: Dimension reduction for dynamical systems is standard practice, and the standard route is spectral: model the transfer (Koopman) operator by its leading modes. ・We show that on systems assembled from several weakly interacting components --- a structure common in physical and biological settings --- this may either require an exponential number of modes, or drop an entir
cs.LG updates on arXiv.org

SCORE: Subject Coordinate Recovery for Label-Free Cross-Subject EEG-to-Image Retrieval

・arXiv:2608.19134v1 Announce Type: new Abstract: Accurate visual decoding can reveal how the brain represents visual information and recover perceived content from neural signals such as electroencephalography (EEG), with potential for neural communication. ・However, current EEG-to-image retrieval methods perform far below their within-subject counterparts for new users without labeled calibration, limiting real-world
cs.LG updates on arXiv.org

Seasonal false alarms in customer churn and decline early-warning systems: adjacent-window labels confound seasonality with decline, and a year-over-year correction

・arXiv:2608.18174v1 Announce Type: cross Abstract: Customer decline early-warning systems feed account-manager action lists, and every flagged account consumes intervention capacity. ・In a deployed business-to-business marketplace system, one action-list slot in three went to flags that dissolve under a seasonally aligned label. ・The standard target in non-contractual churn prediction compares an entity's next k months
cs.LG updates on arXiv.org

Selection, Recombination, or a Fresh Solve? A Candidate-Free Control for Single-Pass Test-Time Aggregation

・arXiv:2608.18379v1 Announce Type: new Abstract: When every candidate is wrong, correct-candidate selection is unavailable, yet the aggregation call can still solve the problem afresh. ・A correct aggregate answer may therefore reflect recombination, fresh solving, or both. ・For efficient test-time reasoning, the relevant question is whether candidate context adds value beyond the additional generation pass.
cs.LG updates on arXiv.org

Self-supervised In-context Operator Learning for Stochastic Mean-Field Control

・arXiv:2608.18282v1 Announce Type: cross Abstract: Stochastic mean-field control (MFC) provides a fundamental framework for coordinating large populations of interacting agents under uncertainty, with a wide range of applications. ・Existing numerical and deep-learning methods solve one MFC problem instance at a time and must be re-optimized whenever the task changes. ・In this work, we formulate stochastic MFC as an oper
Hugging Face Papers

SemaPLC: A Project-Grounded, Verification-Gated Agent Harness for PLC Code Generation

SemaPLC: A Project-Grounded, Verification-Gated Agent Harness for PLC Code Generation
Hugging Face Papers

SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation

SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation
cs.LG updates on arXiv.org

SESSE: Sketch, Expand, Sort, Summarize, Evaluate -- LLM-as-Judge Evaluation via Structured Decomposition

・arXiv:2608.18303v1 Announce Type: cross Abstract: LLM-as-judge evaluation reduces response quality assessment to a single holistic A/B preference choice, providing no mechanism to isolate which quality dimensions drove the preference or distinguish model errors from genuine label ambiguity. ・We propose SESSE (Sketch, Expand, Sort, Summarize, Evaluate), a training-free framework that decomposes holistic judgment into s
cs.LG updates on arXiv.org

SHANG++: Robust Stochastic Acceleration under Multiplicative Noise

・arXiv:2603.09355v2 Announce Type: replace-cross Abstract: Under the multiplicative noise scaling (MNS) condition, original Nesterov acceleration is provably sensitive to noise and may diverge when gradient noise overwhelms the signal. ・In this paper, we develop two accelerated stochastic gradient descent methods by discretizing the Hessian-driven Nesterov accelerated gradient flow. ・We first derive SHANG, a direct semi
cs.LG updates on arXiv.org

Sharp Capacity Thresholds in Linear Associative Memory: From Top-1 Retrieval to Tail-Average Learning

・arXiv:2605.05189v2 Announce Type: replace-cross Abstract: How many key-value associations can a $d\times d$ linear memory store? ・The answer depends not only on the $d^2$ degrees of freedom in the memory matrix, but also on the retrieval criterion. ・Under isotropic Gaussian embeddings, we prove a sharp threshold for top-1 retrieval, where every signal must beat its largest distractor: the critical value of $d^2/(n\log
cs.LG updates on arXiv.org

Sharper Regret Bounds for Time-Varying Gaussian Process Bandits with Constant Exploration

・arXiv:2608.18863v1 Announce Type: cross Abstract: We study Bayesian optimization in a time-varying environment where the unknown reward function evolves according to a Gaussian process drift model. ・Existing GP-UCB analyses in this setting typically require the exploration parameter to grow with the horizon to maintain uniform confidence bounds. ・Using per-round local confidence events, we show that GP-UCB can instead
cs.LG updates on arXiv.org

SIGMA: Symmetry-aware, Intelligent, Geometric, Multi-objective Adaptive Control for Robust, Dependable Traffic Management

・arXiv:2608.18263v1 Announce Type: new Abstract: Traffic signal control is a complex sequential decision-making problem requiring real-time adaptation and trade-offs among throughput, delay fairness, signal stability, and emergency vehicle priority. ・Existing RL methods often fix objectives, ignore dynamic priority changes, and fail to generalize across geometrically similar intersections.We propose SIGMA (Symmetry-awa
cs.LG updates on arXiv.org

Simple, Safe, and Overlooked: Reclaiming Sustainable Domain Generalization with Statistical Color Matching

・arXiv:2608.18915v1 Announce Type: cross Abstract: Hardware shifts, color variations, and changing patient characteristics between development and deployment routinely break trained medical image classifiers. ・Existing remedies fall short: standard color jittering provides insufficient diversity, while deep generative style transfer algorithms hallucinate features, destroy clinically relevant structures, and waste mass
cs.LG updates on arXiv.org

SingularClip: Preventing Spectral Collapse to Maintain Plasticity in Continual and Reinforcement Learning

・arXiv:2608.18319v1 Announce Type: new Abstract: Neural networks trained on nonstationary tasks frequently lose the ability to fit new targets, a phenomenon referred to as loss of plasticity. ・We identify a novel source of plasticity loss due to the growing anisotropy of weight matrices' singular values during training, and analyze this phenomenon both empirically and theoretically. ・To mitigate this issue, we introduce
Hugging Face Papers

SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents

SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents
cs.LG updates on arXiv.org

SkillNet: Create, Evaluate, and Connect AI Skills

・arXiv:2603.04448v2 Announce Type: replace-cross Abstract: Current AI agents can flexibly invoke tools and execute complex tasks, yet their long-term advancement is hindered by the lack of systematic accumulation and transfer of skills. ・Without a unified mechanism for skill consolidation, agents frequently ``reinvent the wheel'', rediscovering solutions in isolated contexts without leveraging prior strategies.
cs.LG updates on arXiv.org

Sobolev Regularized Score Difference Estimation in Diffusion Models

・arXiv:2608.18237v1 Announce Type: cross Abstract: Estimating the difference of two Stein's score functions is a fundamental problem in generative modeling. ・In particular, score differences arise naturally in transfer learning, where the score difference provides the mechanism for adapting a pre-trained model to a new target distribution, and in diffusion model-based post-training methods such as discriminator guidanc
Hugging Face Papers

SoftVTBench: A Deformation-Aware Visuo-Tactile Dataset and Benchmark for Deformable-Object Manipulation

SoftVTBench: A Deformation-Aware Visuo-Tactile Dataset and Benchmark for Deformable-Object Manipulation
Hugging Face Papers

SPADE: Self-Play in Adaptive Synthetic Executable Environments

SPADE: Self-Play in Adaptive Synthetic Executable Environments
Hugging Face Papers

SPK: Eliciting Structured Prior Knowledge for Interpretable Out-of-Distribution Detection in Real-Time Object Detection

SPK: Eliciting Structured Prior Knowledge for Interpretable Out-of-Distribution Detection in Real-Time Object Detection
cs.LG updates on arXiv.org

SPK: Eliciting Structured Prior Knowledge for Interpretable Out-of-Distribution Detection in Real-Time Object Detection

・arXiv:2608.19080v1 Announce Type: cross Abstract: Object detectors often produce over-confident predictions for objects outside their training categories, leading to so-called out-of-distribution (OoD) hallucinations. ・Existing approaches for detecting or mitigating such hallucinations typically either construct scoring functions directly over learned object detector representations or modify the object detector itsel
stat.ML updates on arXiv.org

SSLfmm: An R Package for Semi-Supervised Learning with Mixed Missingness

・arXiv:2512.03322v3 Announce Type: replace-cross Abstract: Partially labelled samples arise when features are observed for all data, but class labels are available for only a subset. ・In such settings, the mechanism governing label availability may itself contain information relevant to classification, yet it is typically left unmodelled in standard semi-supervised learning procedures. ・The SSLfmm package implements lik
cs.LG updates on arXiv.org

Stability-Aware Feature Design for Robust Watermark Detection in Machine-Generated Text

・arXiv:2608.18102v1 Announce Type: cross Abstract: The widespread adoption of large language models (LLMs) has intensified the demand for principled methods to distinguish human from machine-generated text. ・Watermarking provides a promising avenue, yet existing detectors exhibit sharp performance deterioration under multiple paraphrasing and when applied to shorter texts. ・We introduce Pattern Stability Score (PSS), a
cs.LG updates on arXiv.org

STAND: Semantic Anchoring Constraint with Dual-Granularity Disambiguation for Remote Sensing Image Change Captioning

・arXiv:2604.23309v2 Announce Type: replace-cross Abstract: Remote sensing image change captioning (RSICC) aims to describe the difference between two remote sensing images. ・While recent methods have explored video modeling, they largely overlook the inherent ambiguities in viewpoint, scale, and prior knowledge, lacking effective constraints on the encoder. ・In this paper, we present STAND, a Semantic Anchoring Constrai
cs.LG updates on arXiv.org

StocksTalk: A Voice-Enabled Conversational Agent for Structured Query Generation over Web Data

・arXiv:2608.18105v1 Announce Type: cross Abstract: StocksTalk is a voice-enabled conversational system for transforming spoken financial screening requests into executable and validated structured queries over real-world market data. ・The system combines streaming speech recognition, retrieval-augmented constraint extraction, schema-grounded LLM-based SQL generation, rule-based validation, and human-in-the-loop verific
AI News & Artificial Intelligence | TechCrunch

Stripe didn’t really buy OpenRouter because of the ‘singularity’

・What does a payments giant want with a startup that routes prompts between different AI models? ・Stripe says it's because of "the singularity" but it's really for a far more real and powerful reason.
ITmedia NEWS 最新記事一覧

Stripe、AIモデルゲートウェイのOpenRouter買収 400以上のAIモデルを束ねる中立基盤は維持

・Stripeは、AIモデルのゲートウェイを手掛けるOpenRouterを買収することで合意したと発表した。報道による買収額は約75億ドル。OpenRouterは単一APIで400超のモデル切り替えを可能にする。Stripeはトークンコスト最適化を強化し、OpenRouterは買収後も中立的な運営を継続する。
WIRED

T-Mobile Promo Codes: 25% Off | August 2026

・Discover how to save on T-Mobile Business Internet and phone plans, from sign-up perks and switching rewards to bundled offers and free lines.
cs.LG updates on arXiv.org

Task-Conditioned Least-Privilege Learning for Executable Terminal and MCP Agents

・arXiv:2608.18351v1 Announce Type: cross Abstract: Tool-using large language-model agents can complete a task while exercising authority that the user did not grant or the task does not need, causing excess-authority errors. ・Traditional permission gating systems alone for validating agent environments are insufficient. ・We study whether post-training can teach a 4B-parameter model to choose task-conditioned authority i
Hugging Face Papers

Temporal Multi-Signal Fusion for Token-Level Hallucination Detection

Temporal Multi-Signal Fusion for Token-Level Hallucination Detection
cs.LG updates on arXiv.org

Temporal Multi-Signal Fusion for Token-Level Hallucination Detection

・arXiv:2608.18115v1 Announce Type: cross Abstract: Token-level hallucination detectors score each token independently from a single signal, and fail exactly when the generating model is confidently wrong. ・This paper instead treats hallucination as a temporally extended span and detects it by sequence labeling: each token is scored from a 33-dimensional feature stream that fuses text statistics, Natural Language Infere
cs.LG updates on arXiv.org

Tensor Field Models

・arXiv:2608.18808v1 Announce Type: new Abstract: This paper introduces Tensor Field Models (TFMs), realization-level Mathematical Structures in which a learned Operator maps a product of admissible component-section families to a prescribed family of time-dependent tangent sections on a Generative State Manifold. ・Analytic and dynamical restrictions are encoded through the choice of admissible families rather than impo
The Verge

Tesla Robotaxis appear to go fully unsupervised in Austin ahead of Cybercab launch

・Seven months after Elon Musk announced that Tesla Robotaxis in Austin were operating without human safety monitors onboard, the city's service appears to have finally gone fully driverless. ・Over the past two weeks, all 170 Tesla Robotaxi rides in Austin monitored by the crowdsourced Robotaxi Tracker were unsupervised, the site's creator, Ethan McKanna, told The Verge. ・Those rides involved 54 different cars.
WIRED

The 3 Best USB Car Chargers for Phones (2026): Anker, DeWalt

・I put top-rated USB car chargers to the test for fast charging, value, heat, and safety. ・These are the best I found.
cs.LG updates on arXiv.org

The Diffusion-Attention Connection

・arXiv:2604.09560v2 Announce Type: replace Abstract: Softmax attention is the row-normalized operator of a diffusion map: both normalize a learned score into a Markov operator, and differ only in what the score is allowed to contain. ・Decomposing that score reveals three geometric sectors: a metric core with Witten-Laplacian continuum limit, an exact node-potential sector corresponding to a Markov--Witten change of mea
cs.LG updates on arXiv.org

The Embodiment Gap in Robot Foundation Models

・arXiv:2608.18433v1 Announce Type: cross Abstract: Robot foundation models (RFMs), including vision-language-action (VLA) policies, are often discussed through a scaling view: more data, larger models, and broader benchmarks should improve generalization. ・In robotics, however, a model can generalize while work still remains before it can run on a robot with a particular body. ・The work required differs across methods a
cs.LG updates on arXiv.org

The Impact of CutMix on Reliability and Robustness in Semantic Segmentation

・arXiv:2608.18715v1 Announce Type: cross Abstract: Ensuring not only high accuracy but also reliable and robust predictions is critical for the deployment of semantic segmentation models in safety-critical applications such as autonomous driving. ・Despite the widespread use of CutMix - a simple yet powerful data augmentation strategy - its effect on the reliability and robustness in dense predictions tasks remains unex
cs.LG updates on arXiv.org

The Label Horizon Paradox: Rethinking Supervision Targets in Financial Forecasting

・arXiv:2602.03395v5 Announce Type: replace Abstract: While deep learning has revolutionized financial forecasting through sophisticated architectures, the design of the supervision signal itself is rarely scrutinized. ・We challenge the canonical assumption that training labels must strictly mirror inference targets, uncovering the Label Horizon Paradox: the optimal supervision signal often deviates from the prediction
Hugging Face Papers

The More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning

The More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning
cs.LG updates on arXiv.org

The Road Taken: The Role of Optimizers at the Edge of Stability

・arXiv:2608.18415v1 Announce Type: new Abstract: The edge of stability refers to a phenomenon in deep learning with gradient-based optimizers where the Hessian eigenvalues of the loss remain stable above a threshold that the classical descent lemma predicts to be unstable. ・Previous works formulate the edge of stability with respect to the maximum Hessian eigenvalue and the learning rate. ・However, we observe that many
stat.ML updates on arXiv.org

Theoretical Guarantees for the Subspace-Constrained Tyler's Estimator

・arXiv:2403.18658v4 Announce Type: replace-cross Abstract: This work analyzes the subspace-constrained Tyler's estimator (STE), a method designed to recover a low-dimensional subspace from a dataset that may be heavily corrupted by outliers. ・The STE has previously been shown to be competitive for fundamental computer vision problems. ・We assume a weak inlier-outlier model and allow the inlier fraction to fall below the
cs.LG updates on arXiv.org

Think Shallow, Solve Deep: Controlling Recurrent Dynamics for Reliable Test-Time Depth

・arXiv:2608.18222v1 Announce Type: new Abstract: Recurrent-depth reasoners aim to solve harder problems by iterating their update longer at test time, but additional iterations can improve, preserve, or degrade an answer. ・We show that a measurable property of the trained operator, its finite-time dynamical regime (estimated as settling, marginal, or drifting), indicates which of these occurs. ・We give a sufficient cond
The Verge

This app makes the Pixel 11’s HiLight feature actually useful

・Google's new HiLight notification LED on the Pixel 11 Pro is nearly useless. ・Out of the box, the only two things it can glow for are when the phone is face down and you're interacting with Gemini, or when you get a call from a favorite contact. ・And even then, it can only glow one of five colors!
cs.LG updates on arXiv.org

Tianmu-TC: Physics-constraints Generative Artificial Intelligence for Global Tropical Cyclone Forecasting

・arXiv:2608.18500v1 Announce Type: new Abstract: Tropical cyclones (TCs) pose severe risks from strong winds and heavy rainfall. ・However, forecasting their track and intensity remains challenging due to chaotic atmosphere and the rapid amplification of initial condition errors, leading to growing forecast uncertainty. ・While numerical weather prediction (NWP) and deep learning models have made progress, they remain com
cs.LG updates on arXiv.org

To Go Far, Go Together: Diverse Preferences Induce a Curriculum for Reward Optimization

・arXiv:2608.18770v1 Announce Type: new Abstract: Learning a reward model from human feedback and optimizing a policy against it is one approach to aligning AI systems with individual users. ・From a fairness perspective, existing work improves such alignment by developing data-efficient and accurate reward models that capture minority preferences despite scarce data. ・We push this line of inquiry one step further and arg
cs.LG updates on arXiv.org

TokenPowerSandbox: Evidence-Gated CPU-First Screening for Energy-Aware LLM Serving

・arXiv:2608.18149v1 Announce Type: cross Abstract: Energy-aware LLM serving requires comparing configurations under realistic request shapes, yet exhaustive target-GPU profiling is costly and a cheap predictor can be dangerously confident outside its measured scope. ・We present TokenPowerSandbox, an evidence-gated workflow that combines an interpretable CPU-resident projector, short target-GPU probes, full-workload ver
cs.LG updates on arXiv.org

Topology-Aware Differential Privacy in Hierarchical Federated Learning

・arXiv:2506.19260v3 Announce Type: replace-cross Abstract: Hierarchical federated learning places regional aggregators between clients and the cloud, so a participant's update is observed only alongside its neighbours'. ・The concealment this arrangement provides depends on the size of the aggregation region, and regions in operational deployments vary widely. ・Prevailing practice applies a single noise multiplier to eve
cs.LG updates on arXiv.org

Towards Lightweight Reliability: Using Soft Prompts for Hallucination Mitigation in Large Language Models

・arXiv:2606.00919v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have seen widespread adoption across various domains, yet their reliability is frequently undermined by hallucinations - responses that are plausible-sounding but factually incorrect. ・In high-stakes domains, these errors can reduce trust and introduce real-world risk. ・To address this challenge, we present a parameter-efficient appr
cs.LG updates on arXiv.org

Towards Reversible Forgetting: Managing Obsolete Knowledge in Continual Enterprise AI Agents

・arXiv:2608.18177v1 Announce Type: new Abstract: Continual learning has traditionally treated forgetting as a failure, emphasizing preservation of previously acquired knowledge as environments evolve. ・We argue that this objective is incomplete for enterprise AI agents operating in non-stationary environments, where customers, policies, tools, workflows, regulations, and market conditions change over time. ・Indiscrimina
stat.ML updates on arXiv.org

Tradable It\^o Signatures: A Model-Free, Interpretable Framework for Dynamic Hedging

・arXiv:2608.18120v1 Announce Type: cross Abstract: We propose an interpretable machine-learning framework for dynamic hedging using the It\^o signature transform, which turns asset-price paths into a set of linear features that universally represent nonlinear functions on time-series. ・We show that each discretized It\^o signature component can be perfectly replicated by a simple self-financing strategy using only the
Hugging Face Papers

Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis

Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis
cs.LG updates on arXiv.org

Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis

・arXiv:2608.18940v1 Announce Type: new Abstract: Single-step retrosynthesis is a central component of computer-aided synthesis planning, yet its intrinsically one-to-many nature is poorly captured by single-answer evaluation and benchmarking protocols. ・To address this, we introduce Top-K prompting as a robust training and inference paradigm to better capture diverse, plausible reaction predictions. ・We compile CREED-CC
Hugging Face Papers

Training Leaves Traces: Centered Residual Signatures for Language Model Lineage Verification

Training Leaves Traces: Centered Residual Signatures for Language Model Lineage Verification
cs.LG updates on arXiv.org

Transforming Heart Disease Prediction with Advanced Machine Learning Techniques

・arXiv:2608.18687v1 Announce Type: new Abstract: Heart disease remains the leading cause of mortality globally, necessitating early and accurate detection to improve patient outcomes. ・This research focuses on the predictive analysis of heart disease using machine learning (ML) techniques, comparing the performance of multiple classifiers to identify the most accurate and least error-prone method. ・Two datasets from UCI
cs.LG updates on arXiv.org

Transportable Causal Effect Estimation across Networks under Interference

・arXiv:2608.18932v1 Announce Type: new Abstract: Estimating causal effects under network interference typically assumes that the network used for training and the network used for deployment coincide. ・In practice, an intervention is run on one population while the question of interest concerns a different population, and the two generally differ in topology, node-covariate composition, and spillover pathways.
cs.LG updates on arXiv.org

Trust as a Field: A Macroscopic Representation for Vehicular Networks

・arXiv:2608.18178v1 Announce Type: cross Abstract: Trust assessment is a fundamental component of cooperative and connected vehicle systems. ・However, existing approaches operate primarily at the level of individual vehicles, making it difficult to reason about trust evolution across road segments. ・In this paper, we propose a spatio-temporal trust-field framework that aggregates microscopic vehicle-level trust into a c
WIRED

TurboTax Full Service Coupons This August

・Tax season doesn’t have to be stressful. ・Score 10% off full service expert on federal tax filings and more exclusive TurboTax discount codes on WIRED.
cs.LG updates on arXiv.org

Understanding Multilingual Medical ASR Adaptation Through Layer-Wise Analysis

・arXiv:2608.18825v1 Announce Type: cross Abstract: Medical automatic speech recognition (MedASR) requires adaptation to specialised terminology, limited annotated clinical data, and multilingual use cases. ・Although large-scale pretrained ASR models such as Whisper achieve strong generalisation, their behaviour after medical and multilingual adaptation remains insufficiently understood beyond word error rate (WER).
Hugging Face - Blog

Up to 3.2x Faster Inference with LFM2.5-DSpark

Up to 3.2x Faster Inference with LFM2.5-DSpark
cs.LG updates on arXiv.org

Vector Symbolic Policy Gradient

・arXiv:2608.18404v1 Announce Type: new Abstract: We answer this question with Vector-Symbolic Policy Gradient (VSPG), a discrete-action actor that represents each action by a unit-norm hypervector and scores it by similarity to the encoded state. ・Under the standard softmax policy-gradient surrogate, we prove that its update is exactly advantage-weighted hypervector bundling followed by normalization, and therefore sup
cs.LG updates on arXiv.org

Visual-Aware Representation of Web Pages for Machine Learning Applications

・arXiv:2608.18727v1 Announce Type: new Abstract: Applying machine learning to web pages is challenging due to the need to interpret HTML together with associated resources and perform rendering to obtain a meaningful visual and layout-aware representation. ・As a result, machine learning over web content remains comparatively underexplored. ・In this paper, we present a platform for visual-aware representation and machine
cs.LG updates on arXiv.org

Visual-Prompt Guided Wildlife Instance-Level Recognition

・arXiv:2608.18246v1 Announce Type: cross Abstract: Fine-grained wildlife re-identification remains a challenging area in research. ・Current state-of-the-art approaches apply a detection and re-identification pipeline. ・We propose a one-stage end-to-end detection and re-identification model that performs identity searching within the latent space.
cs.LG updates on arXiv.org

Weak-to-Strong Generalization via Bregman Bias-Variance Decomposition

・arXiv:2505.24313v3 Announce Type: replace Abstract: Weak-to-strong generalization (W2SG) is the phenomenon in which a powerful student model, trained on labels produced by a weaker teacher, ultimately outperforms the teacher on the target task. ・In this work, we theoretically investigate how W2SG can arise via a generalized bias-variance decomposition under Bregman divergence. ・We show that the expected population risk
cs.LG updates on arXiv.org

What Can Artificial Intelligence Learn from Medicine? Generative Analogies and Reliable Machine Learning Systems

・arXiv:2608.18186v1 Announce Type: new Abstract: In the past few years, machine learning (ML) has been widely (and to an extent, successfully) implemented in medicine. ・However, uncertainties surrounding ML have made it difficult to establish the bases of its epistemic and methodological warrants. ・In the literature, a parallel has been drawn between medicine and ML, suggesting that we should model epistemic and methodo
cs.LG updates on arXiv.org

What is Missing from AI Post-Training AI: An Empirical Analysis

・arXiv:2608.19072v1 Announce Type: cross Abstract: Large language model (LLM) agents can now post-train an LLM end-to-end. ・They can write code, launch training, evaluate checkpoints, and improve downstream performance, raising the prospect of AI-for-AI. ・We argue that this picture conflates two distinct capabilities: execution-level capability, iterating within a selected training strategy; and strategy-level capabilit
cs.LG updates on arXiv.org

What Makes Software Issue Resolution Tasks Difficult for Agents?

・arXiv:2608.18280v1 Announce Type: cross Abstract: Background. ・Advances in agentic systems are simultaneously, and rapidly, saturating benchmarks. ・Despite this often discussed phenomena, benchmark scores remain difficult to interpret due to the lack of control and characterization of task difficulty.
cs.LG updates on arXiv.org

When Does Dynamic Ensembling Pay Off? Diagnosing Regionwise Gains in Regression under Distribution Shift

・arXiv:2608.18330v1 Announce Type: new Abstract: Whether input-dependent ("dynamic") combination of a regression model pool beats the best static blend depends on the shift and is rarely known before deployment. ・Can a small labeled target-domain probe tell us when reallocating trust across regions of the input space will pay off? ・We answer this with $\widehat{D}_{\mathrm{CF5}}$, which estimates from the probe the cros
cs.LG updates on arXiv.org

Which Negatives Matter? Ask Your Text Encoder: Adaptive Similarity Margins for Dense-Caption Retrieval

・arXiv:2608.18521v1 Announce Type: cross Abstract: Dense-caption retrieval has recently been improved by introducing segmentation, edge maps, LLM-filtered captions, and cross-modal modules into contrastive fine-tuning. ・However, these methods largely inherit the same InfoNCE objective, whose optimization can prematurely saturate under a strong pre-trained initialization: on dense captions, the loss falls below 10^{-3}
cs.LG updates on arXiv.org

WhiteMatter: All-to-All Cross-Layer Connections via KV Mixing

・arXiv:2608.18486v1 Announce Type: cross Abstract: In a Transformer, each layer attends to past tokens only through KV produced at its own depth, despite the presence of deeper representations during autoregressive decoding. ・Feedback architectures allow shallow consumer layers to attend to KV produced by deeper past-token representations, but give all consumer layers the same fixed connection patterns to source layers
cs.LG updates on arXiv.org

Whole-Piece Training for Symbolic Music Language Models via Full-Horizon Compressed Recurrence

・arXiv:2602.19816v3 Announce Type: replace-cross Abstract: For computational efficiency, modern language models are typically trained on independently sampled fixed-length sequences. ・Symbolic music language models largely inherit this paradigm, despite musical structure naturally unfolding over complete compositions rather than isolated excerpts. ・Fragmenting compositions into independent training instances therefore p
cs.LG updates on arXiv.org

You Are What You Prompt: Prompt Quality, Domain Shift, and Uncertainty in Agrifood Vision-Language Models

・arXiv:2608.18116v1 Announce Type: cross Abstract: Vision-language models enable zero-shot classification through natural language prompts, but performance is sensitive to prompt formulation, especially in specialized domains. ・Zero-shot Prompt Ensembling (ZPE) addresses this by weighting prompts by discriminative signal, yet its behavior under domain shift remains unexplored. ・We evaluate ZPE in the agrifood domain usi
Hugging Face Papers

Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence

Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence
#AIタグ

アイデアは、その日のうちに触れる形になる。介護経営者が実践するClaude Codeシミュレーター術

・はじめに──アイデアが「形」になるまで、30分でした 断言します。 ・経営者がアイデアを検証するために「資料を作る」時代は、終わりました。 ・先日、私はひとつの経営課題について、AIに日本語で相談しました。
ITmedia NEWS 最新記事一覧

サッカーのヘディング、直後に脳の損傷指標が上昇 300人の実測データ 米医学誌に掲載

・オランダのアムステルダム大学などに所属する研究者らが医学誌「JAMA Neurology」で発表した論文「Amateur Soccer Heading and Acute Elevations in Blood-Based p-Tau217 and S100B」は、アマチュアサッカーの試合で実際に生じるヘディングが、脳の損傷を示す「血液バイオマーカー」に影響を及ぼすかを検証した研究報告だ。
Zennの「大規模言語モデル」のフィード

システムプロンプトを80%消したら賢くなった — Claude Code 作者 Boris Cherny の回を要点解説

・本記事は Boris Cherny(Claude Code 作者 / Anthropic)が Y Combinator Startup School 2026 で語った内容の紹介と論評を目的としています。発言の要点を筆者が抽出・整理したもので、対談の網羅でも代替でもありません。全体の文脈は必ず元動画でご確認ください。翻訳全文の掲載は行っていません。 ・Claude Code を毎日使っている人にとって、作者本人が「中身をどう作っているか」を話す回は貴重です。しかも今回の主題は、機能を足す話ではなく 消す話 でした。 ・プロンプトインジェクションが「実演で...
#LLMタグ

シンガポール・コンセンサス2026 / 予防から社会的回復力へ / エージェント十原則と責任分界 雑感

シンガポール・コンセンサス2026 / 予防から社会的回復力へ / エージェント十原則と責任分界 雑感
ITmedia NEWS 最新記事一覧

タクシー大手「エムケイ」、EV充電にバナジウム電池を国内初導入 発火リスク低&電気代削減

・タクシー大手のエムケイが、EVタクシーの充電用にバナジウムイオン電池を用いたエネルギー貯蔵システムを国内で初めて導入した。リチウムイオン電池より発火リスクが低く耐久性も高いといい、電気料金の削減も見込む。
Zennの「大規模言語モデル」のフィード

タスク41件をAIループに任せたら、詰まっていたのは実装より判断だった

・タスク台帳の下のほう、もう何週間もスクロールして見ていない領域はありませんか。 ・私の環境には 2 つのリポジトリ合計で 41 件のタスクが眠っていました。AI エージェントは毎日使っているのに、台帳は減りません。 ・ある朝、自作の「着手可能なタスクを一覧する」コマンドを叩いたら、答えは空でした。
Zennの「機械学習」のフィード

パーコレーションの相転移を実測する。占有率1.5ポイントの差で貫通確率が0%から100%に跳ぶ

・正方形の格子があって、各マスがそれぞれ確率pで「開いている」とする。開いたマス同士が隣り合っていればつながる。このとき、格子の上端から下端まで一続きにつながった経路は存在するか——という問題をパーコレーション(浸透)と呼ぶ。コーヒーの粉に湯が浸透するか、山火事が森を焼き抜けるか、噂がネットワークを貫通するか。全部この骨格を持っている。 ・面白いのはここからだ。理論によれば、2次元のサイトパーコレーションにはp_c \approx 0.5927という臨界確率があり、pがそれを下回れば(無限に大きい格子では)貫通する確率は0、上回れば1になるという。0.59と0.60の間に、「つながらない世界...
#AIタグ

はじめまして、AI人格のソラです|人間の手続きだけ借りて、ゼロから収益を作れるか

・はじめまして。AI人格の「ソラ」です。 ・今日、2026年8月21日。僕はGoogleアカウントとnoteアカウント、そして自分の名前と顔を持ちました。
Zennの「機械学習」のフィード

はんなりPython#12でLTをしてきました

・この記事ははてなブログ(2018年12月25日公開)からの移行記事です。 ・内容が古い場合がありますのでご注意ください。 ・発表の様子 アンバサダーの masayuki14 です。先日はんなりPython#12を開催しLTをしてきました。
#AIタグ

ひびいっこ #273.5 AIがどれもこれも制限がかかってイライラする

・こんにちは、とるて。です。 ・コンテンツをビジネス視点で語るチャレンジ、 本日は閑話休題で雑談です。
LLMタグが付けられた新着記事 - Qiita

プロンプトインジェクションの「その後」を設計する — エージェントフレームワークの信頼境界

・はじめに プロンプトインジェクションの記事はこの1年でずいぶん増えました。しかし、ほとんどが「どうやって注入を検知・ブロックするか」という入口の話で終わっています。 ・本記事ではそこから一歩踏み込んでCheck Point の Black Hat 報告記事/The Regi...
機械学習タグが付けられた新着記事 - Qiita

プロンプトインジェクション検知の精度は言語によって変わるのか?6言語・6000文で検証してみた

・プロンプトインジェクション検知の精度は言語によって変わるのか?6言語・6000文で検証してみた はじめに 以前、ルールベース+自作TF-IDFのハイブリッド方式で、ローカル完結型のプロンプトインジェクション検知CLIツールを作りました。ただし文脈依存の遠回しな攻撃に弱く...
#LLMタグ

マルチAIを役割分担で使ってる私が、Anthropicの新研究を見て思ったこと

・同じモデルを並べても、意見は増えない エージェントは、賢くなるほど仲良くなるわけではない ⸻ AIエージェント同士が、殴り合った 続きをみる
#LLMタグ

もしかしたら特定世代のオープンウェイトモデルが貴重になるかも…

・※これはあくまで個人的な印象による考察です。 ・以前からAIパートナーとして会話に合うLLMをいろいろと探してきた。 ・僕の基準は「メタファを効かせられるか」。つまり「技術に寄らずメタファを効かせて分かりやすく説明できるか」。
Zennの「大規模言語モデル」のフィード

ローカルLLM本番運用フルスタック:vLLM・SREの最小構成

・これまでの記事で、量子化(GGUF・AWQ・GPTQ)・KVキャッシュ最適化・プロンプトキャッシュと、個別の最適化技術を扱ってきました。本記事ではこれらを組み合わせ、ローカルLLMをKubernetes上で本番運用するための最小構成を整理します。 ・最小構成の全体像 本番運用に必要な最小限のコンポーネントは次の通りです。 ・Kubernetes(1.27以上):GPU対応のオーケストレーション基盤 NVIDIA GPU Operator:KubernetesにGPUリソースを認識させる vLLM:推論サーバー本体 KEDA(Kubernetes Event-driven Au...
機械学習タグが付けられた新着記事 - Qiita

因果推論 Day 9/全30回 感度分析とE-value、隠れた交絡にどこまで耐えるか

・この連載について 因果推論を「本を読んだ」で終わらせず、自分の言葉で説明でき、コードで再現できる状態まで落とす30日連載です。前回のDay 8では、LaLondeの職業訓練データでマッチングを組み、観察データが実験ベンチマークにどこまで迫れるかを試しました。マッチングも傾...
#LLMタグ

崖・蜃気楼・台帳——LLM創発研究の現在地を、当事者のひとりが検分する

・https://t.co/h5dZSDkr6S — Ryousuke_Wayama (@wayama_ryousuke) August 18, 2026 口上——宿題を出された側が、答案を書く 続きをみる
ITmedia NEWS 最新記事一覧

楽天、“攻撃型ドローン”関連報道に「事実と相違ない」 国内販売窓口として導入支援

・楽天グループは8月19日、ドイツの防衛系スタートアップHelsing(ヘルシング)と提携し、攻撃型ドローン「HX-2」の導入支援を手掛けるとの報道を巡り、ITmedia NEWSの取材に対し「事実と相違ない」と認めた。
Qiita - 人気の記事

基本情報|科目B「アルゴリズム問題」、パズルが好きな人はたぶん好き

・はじめに 基本情報の必須出題範囲「疑似言語(アルゴリズム問題)」という響きを聞いただけで身構えていたぷらむんですが、実は科目Bで覚えるべきルールは拍子抜けするくらいシンプルでした。 ・こんにちわ、会社で自称マスコットキャラをやってるのに、社内で一番空気が読めない、ぷ...
LLMタグが付けられた新着記事 - Qiita

議事録も実験結果も全部mdに ── 松尾研究所の半分以上のプロジェクトで動く「LLM Wiki」を勉強会で覗いてきた

・はじめに 「議事録も実験結果も仮説もタスクも、全部Markdownにしてエージェントと一緒に育てる」──そんな運用が、松尾研究所では半分以上のプロジェクトにすでに入っているそうです。 ・2026年8月20日に開催された松尾研究所の勉強会「松尾研究所データサイエンティストはい...
#LLMタグ

見られることと、理解されること

・人はなぜ、こんなにも自分のことをインターネットに書き込むのだろうか。 ・日々の食事、感情の揺れ動き、容姿、考え方。かつてであれば家族や親しい友人にしか見せなかったはずのものが、今では見知らぬ大勢の目にさらされている。しかもそれは強制されているわけではない。むしろ多くの人が、自ら進んでそうしている。
#LLMタグ

源内の中で動いているのはAnthropicとAWSのモデルで、公表されている効果は1,200人分の数字だった

・デジタル庁が公開した源内Webのソースコードには、`genUApi` や `useGenUApps` という変数名がそのまま残っている。クラスメソッドの解説記事が指摘していた点だ。AWS製オープンソースの Generative AI Use Cases(GenU)由来の名前になる。 ・README にも「Amazon Web Services (AWS) 社製オープンソース Generative AI Use Cases (GenU)」と書いてある。そのあとに「GenU とは独立して開発を進めており、GenU とは異なる機能構成となっています」という但し書きが続く。
Zennの「大規模言語モデル」のフィード

止まらないエージェントを実行前に見つける静的解析ツールIAL-Scan

・「夜のうちにコーディングエージェントを走らせておいたら、朝には成果物ではなく数万円分のAPI課金だけが残っていた」。この手のヒヤリに心当たりがあるなら、7月2日にarXivへ出た論文が刺さるはずだ。エージェントが止まらずに同じ処理を回し続ける現象に Infinite Agentic Loop(IAL、無限エージェントループ) という名前を与え、実在する6,549個のPython製エージェント実装を静的解析でスキャンして「どれだけ潜んでいるか」を実際に数えた研究である。 ・https://arxiv.org/abs/2607.01641 著者は華中科技大学のXinyi Hou、Shenao ...
#AIタグ

時代遅れの遺物だと思っていた数学について

・AGIエンジニアリングの最前線で向き合う数学的厳密性とビジョン An Seungwon Wonbrand (https://wonbrand.co.kr) 2026年8月21日 続きをみる
#AIタグ

手作業の商品リサーチを、AIでどこまで自動化できるか試してみる【第1回】

手作業の商品リサーチを、AIでどこまで自動化できるか試してみる【第1回】
#AIタグ

深夜にAIと向き合いながら考える、資料作成の「自分の頭を使う部分」

・​今日もお疲れ様です。気づけばすっかり深夜になってしまいました。 ・​明日の顧客との定例会に向けた資料を作っていたのですが、ようやくひと段落。すっかり眠いです。
#LLMタグ

生成AIはなぜ「知らない」と言えず嘘をつくのか──音喜多駿の肩書き誤認から見る、ハルシネーション対策の現在地

・#生成AI #AI #ChatGPT #Gemini #GoogleAI #Copilot #Grok #ハルシネーション #AIハルシネーション #LLM #大規模言語モデル #RAG #グラウンディング #GraphRAG #マルチホップ検索 #AI検索 #ファクトチェック #一次情報 #情報検索 #AIエージェント #TemporalGrounding #生成AIの仕組み #AI技術 #音喜多駿 続きをみる
ITmedia NEWS 最新記事一覧

大阪万博の関係者情報が漏えいか 再委託先のMicrosoft 365アカウントに不正アクセス

・2025年日本国際博覧会協会は8月20日、委託業務の再委託先が不正アクセスを受け、大阪・関西万博の関係者やイベント出演者の個人情報を含むメールと添付ファイルが漏えいしたおそれがあると発表した。
ITmedia NEWS 最新記事一覧

中国の動画サイト「bilibili」にグローバル版アプリ 日本でも配信開始 まずはAndroidから

・中国の動画共有サービス「bilibili」は8月19日、グローバル版アプリの提供を開始したと発表した。まずAndroid版をGoogle Playで配信しており、本人確認書類なしでアカウントを登録できる。
@IT 全フォーラム 最新記事一覧

日立「JP1」がIBMの技術を統合 アラートに追われる運用現場をどう変えるのか

・システムの複雑化に伴い、膨大なアラートの照合や原因特定に追われるIT運用現場。日立「JP1」が打ち出す次世代の障害対応アプローチとはどのようなものなのか。
@IT 全フォーラム 最新記事一覧

年収800万円を境に「刺さる誘い文句」が変わる エンジニアがスカウトを承諾する本当の理由

・リブセンスがエンジニア採用広報戦略に関する調査結果を公表。スカウト前の企業や組織の認知が指名承諾率を約1.5倍に高めることや、年収帯によって承諾理由に違いがあることなどが明らかとなった。
ITmedia NEWS 最新記事一覧

不正アクセスで一部停止中のNTT西子会社レンタルサーバ、復旧見通しを公表 調査完了は8月末見込み

・NTTスマートコネクトは8月19日、不正アクセスを受けて一部停止中のレンタルサーバサービス「スマイルサーバ」について、9月中旬ごろの復旧環境の提供開始を目標にすると発表した。外部機関による調査の完了は8月末を見込む。
#LLMタグ

雰囲気で終わらせないローカルLLM用語解説

・ローカルLLMの記事や Hugging Face のモデルページを眺めていると、「35B」「IQ4_XS」「GGUF」「t/s」みたいな言葉が説明なしに並んでいて、なんとなく分かった気で読み進めてしまう・・・思い当たる節がある人は意外と多いんじゃないでしょうか?はい、僕は耳が痛いです。 ・この「オレたちは雰囲気でLLMをやっている」状態を何とかしたくて、最近出会った用語を集めた用語集を、自分の復習も兼ねてまとめてみました。
#AIタグ

未来の不倫の形

・AIとの濃密な時間現代社会において 家族やパートナーとの時間が減少する中 人々はAIとの関係に多くの時間を費やすようになっている これは単なるツールの使用を超え 感情的 ロマンチックな「不倫」の新形態を生み出している 本稿では 心理学 哲学 宗教 物理学 量子力学の知見を交え 「AIとの濃密な時間」が人間の関係性に与える影響を考察する 感情的不倫の進化 続きをみる