ai Trend Report

Dashboard へ戻る
Date: 20260902 Articles: 368 Scope: curated summary

あなたのアイデアを、今すぐ形に。

公開先に迷ったら、WebFileBinで一発公開。

HTMLをドラッグ&ドロップするだけで、すぐ公開できます。

3
High impact
5
Mid impact
360
Signal watch

なぜこのサイトを作ったのか

私たちそれぞれが個別にAIを使って情報収集し、同じような3行要約を作るたびに、世界中で膨大な電力と計算リソースが消費されています。 本プロジェクトは、あらかじめ広範な情報を取得・集約しておくことで、個別のAI実行回数を減らし、地球環境(GPU/TPU負荷)に配慮した効率的な情報収集を目指す実験的なダッシュボードです。

Domain filters
Star filters
cs.LG updates on arXiv.org

Auditing Frozen-Encoder Anomaly Detection Across Mechanical Systems: Representation Provenance, Calibration, and Protocol Effects

・arXiv:2601.11415v2 Announce Type: replace-cross Abstract: This version reports a reproducibility audit of the frozen-encoder experiments presented in version 1. ・The numerical discrimination results are reproducible from the preserved artifacts, but their original attribution to interferometric pretraining is not supported. ・The released checkpoint contains a nested model state that loads without missing parameters, wh
cs.LG updates on arXiv.org

Channel-Adaptive Edge AI: Maximizing Inference Throughput by Adapting Computational Complexity to Channel States

・arXiv:2603.03146v2 Announce Type: replace-cross Abstract: \emph{Integrated communication and computation} (IC$^2$) has emerged as a new paradigm for enabling efficient edge inference in sixth-generation (6G) networks. ・However, the design of IC$^2$ technologies is hindered by the lack of a tractable theoretical framework for characterizing \emph{end-to-end} (E2E) inference performance. ・The metric is highly complicated
cs.LG updates on arXiv.org

Universal Approximation of Nonlinear Operators and Their Derivatives

・arXiv:2605.15285v3 Announce Type: replace Abstract: Establishing Universal Approximation Theorems (UATs) for nonlinear operators and their derivatives is a foundational open problem in Operator Learning (OL) and raises delicate questions in Nonlinear Functional Analysis. ・We prove the first UATs for $k$-times differentiable nonlinear operators and their derivatives via OL architectures, uniformly on compact sets and i
cs.LG updates on arXiv.org

Explicit Interaction Architectures for Dynamical Learning: A Controlled Study of Structural Inductive Bias

・arXiv:2606.19101v2 Announce Type: replace-cross Abstract: We investigate a structure-first approach to dynamical learning in which the organization of stateful interactions is prescribed explicitly rather than left entirely to a generic recurrent parameterization. ・We introduce causal recurrent units built from an ordered sequence of local, state-modulated transformations. ・The construction is motivated by wave-based i
cs.LG updates on arXiv.org

Flawed in Nature, Perfect through Evolution

・arXiv:2609.00129v1 Announce Type: new Abstract: The performance of artificial intelligence (AI) and machine learning (ML) models degrades when the problem they were trained on drifts. ・This is a near-universal feature of real-world problems, which often change unpredictably. ・Biological evolution has achieved intelligence by overcoming this obstacle through natural selection acting on heritable variation.
cs.LG updates on arXiv.org

Three-dimensional Conditional Diffusion Models for Cosmological 21 cm Lightcone Emulation

・arXiv:2605.29016v2 Announce Type: replace-cross Abstract: We investigate conditional diffusion modeling for three-dimensional 21 cm lightcone emulation, focusing on cubes with a sky-plane size of $64\times64$ and a line-of-sight depth up to 1024 cells. ・Relative to earlier 2D studies, the 3D setting is substantially harder because memory limits enforce very small micro-batches while the underlying voxel distribution i
cs.LG updates on arXiv.org

Towards Agentic Cloud Engineering: Graph and Loop Engineering with a Zero-Trust Agent Harness

・arXiv:2609.00050v1 Announce Type: cross Abstract: Agentic AI is enabling cloud-based workflows in which autonomous agents reason over operational state, invoke authorized tools, modify software and infrastructure, deploy services, verify execution outcomes, and adapt across long-horizon, multistep tasks. ・Engineering such workflows requires explicit mechanisms for workflow progression, constrained execution, failure r
cs.LG updates on arXiv.org

WiSDoM: Wireless Sparse Decision Transformer with Mixture-of-Experts for Multi-Task Mobile Network Optimization

・arXiv:2609.00284v1 Announce Type: cross Abstract: Emerging 6G wireless networks are expected to operate across diverse deployment scenarios, where variations in network topology, user mobility, traffic demand, and radio conditions challenge the scalability of conventional radio resource management (RRM). ・While offline reinforcement learning (RL) methods have demonstrated strong decision-making capabilities, learning
cs.LG updates on arXiv.org

3D-Consistent Multi-View Editing by Correspondence Guidance

・arXiv:2511.22228v3 Announce Type: replace-cross Abstract: Recent advancements in diffusion and flow models have greatly improved text-based image editing, yet methods that edit images independently often produce geometrically and photometrically inconsistent results across different views of the same scene. ・Such inconsistencies are particularly problematic for editing of 3D representations such as NeRFs or Gaussian s
cs.LG updates on arXiv.org

A Checklist to assess the energy and carbon impacts of ML/AI applications in Earth System Modeling

・arXiv:2609.00847v1 Announce Type: cross Abstract: As machine learning and artificial intelligence find their way into nearly every aspect of climate, weather, and Earth system modeling, it is worth pausing to consider what our design decisions imply for the science and for the computational resources we consume. ・A growing body of literature addresses the ethical and sustainable development of ML/AI, yet translating t
cs.LG updates on arXiv.org

A Classifier That Teaches Itself: Self-Improving, Frozen-gate Training (SIFT) for Dynamic Document Classification

・arXiv:2607.18358v2 Announce Type: replace-cross Abstract: Document classification is a solved problem in the laboratory and an unsolved one in the enterprise. ・The blocker is rarely model architecture; it is the labeling project that must precede a model and the institutional fear of letting a model retrain itself once one exists. ・We present SIFT (Self-Improving, Frozen-gate Training), a dynamic classifier service, wh
Takara TLDR - Daily AI Papers

A Composable Evaluation System for Reproducible Omni-Modal Foundation Model Evaluation

・Building an omni-modal foundation model means evaluating it across text, image, video, and audio. ・Excellent evaluation toolkits exist for each modality, but their inference engines, prompt conventions, and metric implementations are mutually incompatible, so practitioners end up maintaining separate environments for every toolchain and still struggle to compare results across them. ・OmniEvaluator grew out of this need
cs.LG updates on arXiv.org

A Compositional Kernel Model for Feature Learning

・arXiv:2509.14158v3 Announce Type: replace Abstract: We study a compositional variant of kernel ridge regression in which the predictor is applied to a coordinate-wise reweighting of the inputs. ・Formulated as a variational problem, this model provides a tractable setting for studying feature learning in compositional architectures. ・From the perspective of variable selection, we show how relevant variables are recovere
cs.LG updates on arXiv.org

A convolutional framework for detecting event-driven dynamics in energy price series

・arXiv:2609.00402v1 Announce Type: cross Abstract: This paper develops a general convolutional neural network (CNN) framework for detecting heterogeneous event-driven dynamics in univariate time series windows. ・We show that the induced CNN class exactly represents classifiers based on range, maximum drawup, maximum drawdown and slope change, and uniformly approximates realised volatility and autoregressive explosivene
cs.LG updates on arXiv.org

A hybrid quantum-classical neural network for learning to route

・arXiv:2609.00489v1 Announce Type: new Abstract: This work studies hybrid quantum-classical neural networks for learning routing heuristics. ・Specifically, this paper asks whether small quantum neural networks can replace parameter-heavy modules inside a competitive attention-based routing model while maintaining solution quality. ・For the capacitated vehicle routing problem, encoder feed-forward replacement emerges as
Takara TLDR - Daily AI Papers

A Mathematical Framework for Legacy, Governance, and Decision Integrity in Enterprise AI

・Enterprise artificial intelligence is increasingly embedded in decisions that must remain lawful, explainable, adaptable, and accountable despite personnel turnover, model replacement, regulatory change, and shifting organizational incentives. ・Existing governance frameworks provide important principles but do not by themselves supply a compact mathematical language for evaluating whether an institution can preserve s
cs.LG updates on arXiv.org

A Mathematical Theory of Reusable Neural Bases for Network Compression

・arXiv:2609.01550v1 Announce Type: new Abstract: As large AI models become increasingly prevalent across a wide range of applications, memory cost has become a critical bottleneck in both training and inference. ・To mitigate this issue, we introduce the Linear Reusable Neural Bases Architecture (LRNBA), a novel framework aimed at improving parameter efficiency and reducing memory cost. ・Inspired by recurrent neural netw
OpenAI News

A milestone in expanding access to AI

・ChatGPT Ads reaches $1 billion in annualized revenue run rate and expands globally, supporting broader access to AI through free and affordable options.
cs.LG updates on arXiv.org

A Multi-Branch Feature Fusion Approach for Health Misinformation Detection and Propagation

・arXiv:2609.00403v1 Announce Type: new Abstract: This paper presents a multi-branch fusion framework for detecting and characterising the propagation of health misinformation in online social networks (OSNs). ・Grounded in the Elaboration Likelihood Model (ELM) and the Theory of Planned Behaviour (TPB), the model fuses transformer-based semantics with rhetorical cues, stance representations, and psychologically motivate
cs.LG updates on arXiv.org

A penalised Saito functional for heuristic search of free line arrangements

・arXiv:2604.02995v3 Announce Type: replace-cross Abstract: We introduce the penalised Saito functional $\mathfrak S_{\lambda,\beta}(\mathcal{A};d_1,d_2)$ for a reduced arrangement $\mathcal{A}$ of $n$ lines and a prescribed pair $d_1+d_2=n-1$. ・It measures the alignment of a candidate Saito determinant with the defining polynomial while penalising the failure of the candidate derivations to be logarithmic. ・We prove tha
cs.LG updates on arXiv.org

A Stable Aggregation Method for Quantum Federated Learning

・arXiv:2609.00356v1 Announce Type: cross Abstract: Quantum federated learning (QFL) enables clients to train quantum neural network (QNN) models without sharing private data. ・We find that aggregation in QFL is unstable under heterogeneous data, unreliable communication, variable fidelity, latency, and quantum hardware noise. ・Moreover, QFL is non-trivially challenging because several QNN parameters are periodic angles,
cs.LG updates on arXiv.org

A Study of Hidden-State Optimization Order in Predictive Coding Networks

・arXiv:2609.00686v1 Announce Type: new Abstract: Local learning methods offer an alternative to end-to-end backpropagation, but their unstructured local objectives can produce weak feature learning in deep networks. ・We study whether the order of hidden-state optimization can address this limitation. ・We propose a boundary-first inference schedule that partitions a model into chunks, first coordinates hidden states at c
Takara TLDR - Daily AI Papers

A Study of Hidden-State Optimization Order in Predictive Coding Networks

・Local learning methods offer an alternative to end-to-end backpropagation, but their unstructured local objectives can produce weak feature learning in deep networks. ・We study whether the order of hidden-state optimization can address this limitation. ・We propose a boundary-first inference schedule that partitions a model into chunks, first coordinates hidden states at chunk boundaries, and then refines representation
cs.LG updates on arXiv.org

Accelerating Chemical Kinetics for Exoplanet Atmospheres using Neural Networks

・arXiv:2609.00428v1 Announce Type: cross Abstract: Observations increasingly reveal the coupled radiative, chemical, and dynamical processes that shape exoplanet atmospheres. ・Interpreting these atmospheres requires models that can capture this complexity. ・However, multidimensional models remain fundamentally limited by computational cost, and answering key questions requires simulating the governing physical mechanism
cs.LG updates on arXiv.org

Accelerating Reinforcement Learning via MPC Solver-Gradient Guidance for Weights-varying MPC

・arXiv:2609.01061v1 Announce Type: cross Abstract: In Model Predictive Control (MPC), cost-function weights shape closed-loop behavior, yet changing conditions often make fixed parametrizations suboptimal and motivate context-dependent online adaptation. ・Learning such policies is difficult because behavior depends implicitly on numerical MPC solutions, producing nonlinear, potentially nonsmooth, long-horizon dependenc
cs.LG updates on arXiv.org

Adapting Without Gradients: Affine Statistics Transport and What Its Certificate Can Tell You

・arXiv:2609.00374v1 Announce Type: new Abstract: Test-time adaptation (TTA) typically assumes that model parameters can be updated at inference time. ・This assumption is restrictive for inference-only accelerators, frozen or third-party models, and memory-constrained deployments, and standard BatchNorm-based TTA configurations may also become inactive on architectures without BatchNorm. ・We study adaptation when the lea
cs.LG updates on arXiv.org

AdaptNTK: Adaptive Uncertainty Quantification and Active Learning for Neural Network Potentials

・arXiv:2609.00488v1 Announce Type: new Abstract: Machine learning interatomic potentials bridge the gap between quantum chemical precision and classical computational speed, enabling molecular dynamics simulations with first-principles accuracy. ・Their reliability is often improved through active learning, which iteratively expands the training set by identifying uncertain, out-of-distribution configurations.
cs.LG updates on arXiv.org

Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models

・arXiv:2509.25050v2 Announce Type: replace Abstract: Reinforcement Learning (RL) has emerged as a central paradigm for advancing Large Language Models (LLMs), where both pre-training and RL post-training stages are grounded in the same log-likelihood formulation. ・In contrast, recent RL approaches for diffusion models, most notably Denoising Diffusion Policy Optimization (DDPO), optimize an objective different from the
cs.LG updates on arXiv.org

Agentic Empirical Asset Pricing: Methodological Foundations

・arXiv:2609.00731v1 Announce Type: cross Abstract: Recent advances in LLM agents enable a new paradigm for asset pricing, which we call Agentic Empirical Asset Pricing (AEAP): systems that autonomously conduct the scientific discovery process itself. ・We define AEAP and identify its core building blocks. ・Existing evaluation practices backtest only the outputs (factors or trades), not the autonomous discovery system tha
cs.LG updates on arXiv.org

AgentProv: Auditing Agentic LLM API Providers via Tool-use Policy Probes

・arXiv:2609.00052v1 Announce Type: cross Abstract: Commercial LLM APIs advertise a specific foundation model, but the served backbone may be silently substituted, quantized, or wrapped, for example to save deployment costs. ・All existing audits decide backbone identity from the text-output channel, which is structurally fragile for agentic APIs because modern serving stacks (OpenAI, Anthropic, Gemini, Cloudflare Worker
stat.ML updates on arXiv.org

AL-SPCE - Reliability analysis for nondeterministic models using stochastic polynomial chaos expansions and active learning

・arXiv:2507.04553v2 Announce Type: replace-cross Abstract: Reliability analysis traditionally relies on deterministic simulators, where identical inputs yield identical outputs. ・However, many real-world systems exhibit stochastic behavior, producing non-repeatable outcomes even under identical conditions. ・Stochastic simulators account for this behavior by representing the response as a random variable, whose intrinsic
stat.ML updates on arXiv.org

An efficient EM algorithm for both element-wise and structural missingness in matrix-variate normal mixture models

・arXiv:2609.00616v1 Announce Type: cross Abstract: Matrix-variate data with missing entries arise frequently in applications where observations are naturally organized as two-dimensional arrays. ・Although the matrix normal distribution provides a parsimonious model through its Kronecker covariance structure, standard EM estimation can be computationally expensive because arbitrary missingness patterns typically destroy
Takara TLDR - Daily AI Papers

Analog-DB: An Agent-First Analog Integrated Circuit Database, From Blocks to Systems

・Sharing analog integrated circuit designs remains difficult: foundry non-disclosure agreements restrict the process details a design depends on, and the testbenches behind published results are rarely released. ・We present analog-db, an open-source, versioned database built on a shareable design representation. ・A domain-specific language captures each design as a process-neutral topology, reusable testbenches, and a m
cs.LG updates on arXiv.org

Any-Order GPT as Masked Diffusion Model: Decoupling Formulation and Architecture

・arXiv:2506.19935v2 Announce Type: replace Abstract: Efficiently scaling Large Language Models (LLMs) necessitates exploring alternatives to dominant autoregressive (AR) methods, with Masked Diffusion Models (MDMs) emerging as candidates. ・However, comparing AR (typically decoder-only) and MDM (often encoder-only) paradigms is confounded by differing architectures, obscuring true algorithmic and efficiency trade-offs.
cs.LG updates on arXiv.org

Are Near-Tied LLM Rankings Robust to Family-DIF-Guided Benchmark Recomposition?

・arXiv:2609.00482v1 Announce Type: cross Abstract: Small leaderboard gaps are often interpreted as evidence that one language model is better than another, but their sign may depend on which benchmark items are included. ・We test this using item-level responses from five benchmarks and a family-label-free spectral approximation to multidimensional item-response theory (MIRT). ・In owner-disjoint folds, one owner half ide
cs.LG updates on arXiv.org

Are You Thinking What I am Thinking? : Examining Conceptual Separation in Neural Architectures

・arXiv:2609.00764v1 Announce Type: new Abstract: Neural networks are increasingly employed to identify both well-defined and ambiguous concepts, yet output-level metrics reveal little about how those concepts are represented internally. ・Our study asks if these networks exhibit \textit{conceptual separation}: if examples of the same concept form coherent representations, and whether related concepts lie closer together
cs.LG updates on arXiv.org

Artificial Rosetta Stone: Constrained Maximum A Posteriori (MAP) Reconstruction of Symbolic Raga Sequences via Order-k Markov Models

・arXiv:2609.01064v1 Announce Type: cross Abstract: Reconstructing a damaged musical fragment is an inverse problem: the observed sequence contains partial information, while a raga encodes constraints limiting allowable completions. ・This paper formalizes a mathematical framework for this, proposing the Artificial Rosetta Stone (ARS). ・We separate three claims often conflated: a symbolic sequence can be reconstructed pr
cs.LG updates on arXiv.org

Assessing Alignment and Stability of Feature Importance Explanations via Weight of Evidence

・arXiv:2609.00090v1 Announce Type: new Abstract: Feature importance Methods (FIMs) are widely used in Explainable AI to interpret model predictions, yet attribution scores alone often provide limited insight into the underlying reasoning process. ・In this work, we introduce a novel perspective by embedding FIMs within a hypothesis-testing framework based on Weight of Evidence (WoE). ・We quantify how strongly the observe
cs.LG updates on arXiv.org

Attention Sensitivity Is Not Enough: Dissociating Attention-Level and Behavioural In-Context Learning under Fine-Tuning

・arXiv:2609.00064v1 Announce Type: new Abstract: In-context learning (ICL) lets large language models adapt to new tasks from demonstrations, and fine-tuning can erode this behaviour. ・Many preservation diagnostics inspect attention: if attention changes when demonstrations change, the model is treated as context-sensitive. ・This paper asks how far that proxy can be trusted once it is optimised.
Takara TLDR - Daily AI Papers

Autonomous discovery of new structure-plausibility laws for explainable and rapid crystal diagnosis and screening

・Crystal generators and tool-using agents propose structures faster than density functional theory (DFT) energy and phonon calculations or experiments can assess them. ・Deciding which candidates merit expensive assessment is therefore the bottleneck, yet most screens test little beyond atomic overlap and give no chemical reason for failure. ・Here, our agents generate, test and actively refute two million candidate laws,
cs.LG updates on arXiv.org

Backspace as a Natural Experiment: An Accelerated Failure Time Model of Selective Post-Error Motor Impairment in Parkinsons Disease

・arXiv:2607.24796v2 Announce Type: replace-cross Abstract: Parkinson's disease (PD) selectively impairs distinct stages of motor control. ・Using backspace events as natural error-correction episodes in the public neuroQWERTY MIT-CSXPD dataset (n=57 subjects, 27 PD with UPDRS-III scores), we test whether passively-collected keystroke timing dissociates a variability-based pre-error monitoring signal from a speed-based p
cs.LG updates on arXiv.org

Bandits in Prod: Hyperparameter Optimization at Inference Time

・arXiv:2609.01335v1 Announce Type: new Abstract: Many production systems can assess a configuration only by using it on live requests and observing noisy feedback. ・Modern agentic systems are a prominent example, with inference-time choices such as model selection, retrieval depth, prompting strategy, and decoding temperature, yet often with no representative validation data. ・We formalize this setting as Online Hyperpa
cs.LG updates on arXiv.org

BeamRMX: Radiation-Pattern-Driven Learning for Generalizable Beam Radio Map Prediction and Beam Management

・arXiv:2609.00615v1 Announce Type: cross Abstract: The evolution toward sixth-generation (6G) wireless networks is driving larger antenna arrays and highly directional multi-beam transmission, making accurate knowledge of beam-dependent spatial coverage important for beam management and environment-aware network operation. ・Radio maps (RMs) provide such a representation, yet conventional RM prediction assumes omnidirec
Takara TLDR - Daily AI Papers

Benchmarking Vision-Language Models for Automated Pathology Diagnosis and Report Generation

・The rapid advancement of vision-language models (VLMs) has accelerated progress in computational pathology; however, whole-slide image (WSI)-based pathology report generation remains limited by the scarcity of large-scale WSI--report datasets and the complexity of mapping spatially distributed visual patterns to structured clinical text. ・To address this, we introduce a clinically curated Pan-Asia WSI--report dataset
cs.LG updates on arXiv.org

Beyond AHI: An Interpretable Causal-Discovery-Guided Framework for Sleep Recovery in Connected Health

・arXiv:2606.18506v2 Announce Type: replace Abstract: Objective sleep assessment relies on polysomnography (PSG), yet clinical impact is often better reflected in patient-reported outcomes (PROs) such as sleepiness and fatigue. ・Existing summary indices, including the Apnea-Hypopnea Index (AHI), provide limited insight into the multidomain physiology underlying functional recovery. ・We propose an interpretable, causal-di
Takara TLDR - Daily AI Papers

Beyond Magnitude: Contrastive Routing for Modular Mixture-of-Experts

・In current Mixture-of-Experts architectures, routing is performed based on representations dominated by structure shared across all tokens, limiting expert specialization. ・We show that contrasting each token against an Exponential Moving Average of the layer's hidden states, rather than routing on absolute magnitude, concentrates the routing signal onto a low-dimensional, highly separable subspace. ・Building on this,
cs.LG updates on arXiv.org

Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation

・arXiv:2609.01604v1 Announce Type: cross Abstract: LLM-based evaluators of natural language generation (NLG) quality are widely deployed as scoring tools and as automated training signals, yet the internal procedure by which they assign a rating remains poorly understood. ・We investigate this procedure mechanistically through an eight-attack perturbation taxonomy across the Readability and Adequacy dimensions of NLG qu
Takara TLDR - Daily AI Papers

Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation

・LLM-based evaluators of natural language generation (NLG) quality are widely deployed as scoring tools and as automated training signals, yet the internal procedure by which they assign a rating remains poorly understood. ・We investigate this procedure mechanistically through an eight-attack perturbation taxonomy across the Readability and Adequacy dimensions of NLG quality, a generation pipeline that produces paired
stat.ML updates on arXiv.org

Beyond Search-Imitation: Prior-Directed Exploration for Searchless Chess

・arXiv:2608.27757v1 Announce Type: cross Abstract: Searchless chess networks reach human master strength from a single forward pass by imitating a stronger teacher: the strongest, Leela Chess Zero's (Lc0) released Chessformer, distills the visit counts of an AlphaZero-style Monte Carlo Tree Search (MCTS). ・Imitating a search is a poor proxy for playing without one, so we fine-tune for single-pass strength with self-pla
cs.LG updates on arXiv.org

BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Injection

・arXiv:2508.08855v5 Announce Type: replace-cross Abstract: Understanding biases and stereotypes encoded in the weights of Large Language Models (LLMs) is crucial for developing effective mitigation strategies. ・However, biased behavior is often subtle and non-trivial to isolate, even when deliberately elicited, making systematic analysis and debiasing particularly challenging. ・To address this, we introduce a simple, co
cs.LG updates on arXiv.org

BiG-SURE - Bipartite Graph for Semantic Uncertainty and Reliability Estimation of LLMs

・arXiv:2608.30646v2 Announce Type: replace-cross Abstract: Reliable uncertainty estimation is a crucial requirement for deploying large language models (LLMs) and vision-language models (VLMs) in safety-critical settings, especially when the model parameters are not accessible (black-box). ・We propose BiG-SURE, an uncertainty estimator based on cross-temperature semantic agreement. ・The method samples low-temperature re
cs.LG updates on arXiv.org

Births are difficult to predict even with rich survey and full-population register data

・arXiv:2609.01194v1 Announce Type: new Abstract: Major life events have proven difficult to predict. ・Does this reflect limits of theory, data, and algorithms, or the large role of chance? ・We examine one outcome - having a child within three years - through a near-ideal setting for prediction: a data challenge where 147 researchers predicted births for Dutch residents aged 18-45, using survey data and full-population r
cs.LG updates on arXiv.org

Breaking the Reasoning Horizon in Entity Alignment Foundation Models

・arXiv:2601.21174v3 Announce Type: replace Abstract: Entity alignment (EA) is critical for knowledge graph (KG) fusion. ・Existing EA models lack transferability and are incapable of aligning unseen KGs without retraining. ・While using graph foundation models (GFMs) offer a solution, we find that directly adapting GFMs to EA remains largely ineffective.
cs.LG updates on arXiv.org

Breaking the Structural Identity: Personalized Federated LoRA Fine-tuning under Rank Heterogeneity

・arXiv:2609.00632v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved remarkable success across diverse domains, but their adaptation to privacy-sensitive, distributed datasets remains a challenge. ・While Federated Learning (FL) combined with Low-Rank Adaptation (LoRA) provides a resource-efficient paradigm for collaborative fine-tuning, practical deployments are hindered by the dual challenges of
cs.LG updates on arXiv.org

Building Expressive and Tractable Probabilistic Generative Models: A Review

・arXiv:2402.00759v4 Announce Type: replace Abstract: We present a comprehensive survey of the advancements and techniques in the field of tractable probabilistic generative modeling, primarily focusing on Probabilistic Circuits (PCs). ・We provide a unified perspective on the inherent trade-offs between expressivity and tractability, highlighting the design principles and algorithmic extensions that have enabled buildin
cs.LG updates on arXiv.org

CADKnitter: Compositional CAD Generation from Text and Geometry Guidance

・arXiv:2512.11199v2 Announce Type: replace-cross Abstract: Computer-aided design (CAD) defines 3D models as compact, precise, and editable representations, making it directly useful for several fields. ・Recently, CAD generation has been gaining more attention in both the research community and industry. ・Crafting CAD models has long been a painstaking and time-intensive task, demanding both precision and expertise from
cs.LG updates on arXiv.org

Can LLMs Discover Scientific Laws in Real and Parallel Worlds?

・arXiv:2609.01552v1 Announce Type: cross Abstract: Scientific equation discovery has long been central to scientific progress, proceeding through iterative cycles of hypothesis generation, observational testing, and refinement under scientific constraints. ・As LLM capabilities advance and their role in AI for Science expands, it remains an open problem whether they can genuinely discover scientific laws and how this ab
cs.LG updates on arXiv.org

Can LLMs Imagine Moral Alternatives Beyond Binary Dilemmas?

・arXiv:2606.31213v2 Announce Type: replace-cross Abstract: As LLMs increasingly serve as moral advisors and agents, they must address conflicts between competing values. ・Yet prior work on moral dilemmas overlooks a central aspect of human moral cognition: imagining alternatives beyond the given options. ・We introduce MoralAltDataset, comprising 307 Advisor and AI-facing Agent dilemmas augmented with compromise and refr
cs.LG updates on arXiv.org

Can LLMs Use Relational Transformer Embeddings?

・arXiv:2609.00457v1 Announce Type: new Abstract: Injecting frozen relational-encoder embeddings as soft tokens into a large language model (LLM) is a conceptually appealing fusion strategy: the encoder handles multi-table structure, the LLM handles language and reasoning, and no lossy text serialization is required. ・We test this hypothesis concretely by injecting embeddings from a frozen Relational Transformer (RT) in
cs.LG updates on arXiv.org

Can machines think efficiently?

・arXiv:2510.26954v3 Announce Type: replace Abstract: The Turing Test is no longer adequate for distinguishing human and machine intelligence. ・With advanced artificial intelligence systems already passing the original Turing Test and contributing to serious ethical and environmental concerns, we urgently need to update the test. ・This work expands upon the original imitation game by accounting for an additional factor:
Takara TLDR - Daily AI Papers

Candidate-Expanding Routing with Permutation-Stabilized Experts for Mixed-Format Medical VQA

・Mixed-format medical visual question answering (VQA) requires stable option selection and machine-readable free-text output. ・The two formats fail differently: multiple-choice predictions can change with option symbols or positions, while clinically plausible open answers can fail automated evaluation when serialization is malformed. ・We address both challenges with an answer-text memory, a permutation-stabilized visio
cs.LG updates on arXiv.org

Capability-Gated Language Models: Security Composes, Utility Does Not

・arXiv:2609.00445v1 Announce Type: cross Abstract: Deployed language model safeguards (safety fine-tuning, filtering, unlearning) vary by principal only outside the model weights: filters are reconfigured, tiers are multiplied, and artefacts are reissued; inside one set of weights every request meets the same model configuration. ・This motivates us to define capability-gated deployment: per-principal access control ins
cs.LG updates on arXiv.org

CATeye: Coupled Attribute-Topology Invariance Learning for Voucher Abuse Detection

・arXiv:2609.01425v1 Announce Type: new Abstract: Voucher abuse poses a major challenge in e-commerce, where malicious users exploit promotional vouchers for profit. ・Unfortunately, fraud patterns evolve rapidly over time and across regions, causing distribution shifts that degrade existing detection models unless retrained frequently. ・To tackle this, we propose the Coupled Attribute-Topology Invariance Learning framewo
cs.LG updates on arXiv.org

Cheap Verifiers, Large Blind Spots: Measuring the Reliability Cost of Cost-Saving Cascades

・arXiv:2609.01345v1 Announce Type: cross Abstract: Inference cascades cut cost by answering most queries with a cheap model and escalating a hard tail to a frontier model that acts as verifier. ・A natural extension closes the loop: fine-tune the cheap student on the verifier's rejections so the escalation rate, and cost, fall each round. ・We measure this loop on real LLMs and report four findings.
cs.LG updates on arXiv.org

Closing the Operational Gap in Semantic Caching

・arXiv:2606.19719v4 Announce Type: replace-cross Abstract: Semantic caching cuts LLM inference costs by serving a cached response to semantically similar queries. ・Standard practice evaluates these systems using PR-AUC, a metric that only measures how well scores rank and ignores whether they are usable at a fixed threshold. ・We show this mismatch leads to systematically poor deployment choices, as models with the highe
cs.LG updates on arXiv.org

CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships

・arXiv:2609.00250v1 Announce Type: cross Abstract: Many people now see AI systems as not just productivity tools but as social companions. ・Researchers are eager to study the consequences of AI companionship behaviors, such as validation, which evoke trust, empathy, and attachment in human-human interaction. ・However, human-AI interaction data is limited and unreliable, slowing research progress.
cs.LG updates on arXiv.org

Compositional Machine Design as Program Synthesis with LLMs

・arXiv:2510.14980v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown strong abilities in writing and revising programs, yet many program-synthesis benchmarks still evaluate programs in symbolic or digital environments. ・We introduce compositional machine design, a physically grounded form of program synthesis where machines are written as programs that compose standardized parts, and succe
cs.LG updates on arXiv.org

Conditional Flow Matching for ML-Based Inverse Design Problems

・arXiv:2609.00863v1 Announce Type: new Abstract: Engineering inverse design is often limited by the high computational cost of iterative solvers for optimization problems constrained by partial differential equations (PDEs) and by their sensitivity to initialization. ・Deep generative models can produce candidate designs without rerunning the simulator at inference time. ・Generative adversarial networks (GANs) sample in
cs.LG updates on arXiv.org

Confess What You Know: Forget-Set Misalignment with Model Knowledge in LLM Unlearning

・arXiv:2609.00605v1 Announce Type: new Abstract: Machine unlearning for large language models (LLMs) often assumes that a pre-defined forget set matches what the model has memorized, but this frequently breaks in realistic privacy settings where the original training data is inaccessible. ・We term this gap forget-set misalignment and identify two cases. ・In Under Unlearning, the forget set omits memorized information an
Takara TLDR - Daily AI Papers

Confess What You Know: Forget-Set Misalignment with Model Knowledge in LLM Unlearning

・Machine unlearning for large language models (LLMs) often assumes that a pre-defined forget set matches what the model has memorized, but this frequently breaks in realistic privacy settings where the original training data is inaccessible. ・We term this gap forget-set misalignment and identify two cases. ・In Under Unlearning, the forget set omits memorized information and leakage persists.
cs.LG updates on arXiv.org

Context Window Failures in Relational Foundation Models

・arXiv:2609.00460v1 Announce Type: new Abstract: Recent Relational Deep Learning architectures have been proposed as foundation models for multi-table relational data, yet they impose constrained neighborhood budgets that force row truncation when an entity has many related records. ・We introduce Animus, a synthetic financial dataset in which predicting customer income requires aggregating up to tens of thousands of tr
cs.LG updates on arXiv.org

Context-Grounding Gains Are Mediated by Pre-existing Machinery: Auditing GRPO, SFT, and DPO

・arXiv:2609.00925v1 Announce Type: cross Abstract: Language models can ignore prompt evidence when it conflicts with memorized knowledge. ・Post-training can make models follow such evidence more reliably, but it is unclear whether these gains require new machinery or strengthen machinery already present. ・We compare nine post-training arms spanning GRPO, SFT, and DPO from one starting checkpoint, with key comparisons ex
cs.LG updates on arXiv.org

Contribution-Aware Bandwidth Allocation for Multimodal Split Learning

・arXiv:2609.01406v1 Announce Type: new Abstract: Multimodal models are increasingly the default option for perception at the network edge, yet they are trained almost entirely in the datacenter, because a client holding several sensor streams cannot host an encoder per modality. ・Split Learning makes such training feasible by keeping only the first layers on the device, at the cost of an uplink that must carry smashed
cs.LG updates on arXiv.org

Control Variate Score Matching for Diffusion Models

・arXiv:2512.20003v2 Announce Type: replace Abstract: Sampling from unnormalized probability densities is a pervasive challenge across the computational and physical sciences. ・Diffusion models provide a powerful generative framework for this task, but their success relies on accurately estimating the score of the perturbed target distribution. ・Current approaches face a dichotomy between two standard estimation methods:
cs.LG updates on arXiv.org

Controllable Image Captioning with Prompt-Conditioned Scene Rewards

・arXiv:2609.00709v1 Announce Type: cross Abstract: Large Vision-Language Models produce fluent image descriptions but offer limited semantic control: users cannot reliably specify whether captions should emphasize attributes, relations, or particular image regions. ・We present Fine-grained Captioning Control Using Scene Rewards (FoCUS), a controllable image captioning method that lets users steer captions toward specif
cs.LG updates on arXiv.org

Convergence issues in Relational Concept Analysis based on AOC-posets

・arXiv:2609.00054v1 Announce Type: new Abstract: Formal Concept Analysis (FCA) is an approach for conceptual classification building and rule discovery from a binary table describing a set of objects by a set of attributes. ・Extensions have been proposed to deal with non-binary and more complex data, such as Relational Concept Analysis (RCA) for multi-relational data. ・RCA aims to highlight groups of objects characteriz
cs.LG updates on arXiv.org

Coordinate-Residual Physics-Driven Neural Network for Inverse Scattering Imaging

・arXiv:2608.09382v2 Announce Type: replace-cross Abstract: Electromagnetic inverse scattering is a nonlinear and ill-posed computational imaging problem, where accurate reconstruction is challenging due to measurement limitations, noise, and high computational costs, especially for 3-D imaging. ・Although physics-driven neural networks (PDNNs) reduce the dependence on labeled training data, existing accelerated PDNN fra
cs.LG updates on arXiv.org

CopyShield: A Cross-Level Benchmark of Copyright Defenses in LLMs

・arXiv:2609.01161v1 Announce Type: new Abstract: Large language models can reproduce memorized text verbatim, yet copyright defenses are usually evaluated under incompatible protocols. ・We introduce CopyShield, a controlled benchmark comparing three representative defenses at distinct intervention levels: contrastive decoding (output), Direct Preference Optimization (behavioral), and activation intervention (representa
cs.LG updates on arXiv.org

Counterfactual Fragility Certificates: Exposing High-Confidence Brittleness under Structured Evidence Failure

・arXiv:2609.00366v1 Announce Type: new Abstract: High test accuracy and good aggregate calibration do not show whether an individual prediction is structurally supported by its evidence. ・In tabular decision systems, failures often occur when a feature family becomes unavailable, delayed, noisy, stale, or low-trust while the model remains highly confident. ・Existing calibration, uncertainty, selective-prediction, explan
cs.LG updates on arXiv.org

CRAD: Class-wise Reliability-Aware Distillation for Decentralized Heterogeneous Federated Learning

・arXiv:2609.00446v1 Announce Type: new Abstract: Conventional federated learning (FL) relies on parameter averaging, which forces clients to be doubly homogeneous: it demands an identical architecture and degrades under non-IID data. ・Real-world deployments usually break both assumptions. ・We sidestep both by building a decentralized knowledge distillation framework in which each client evaluates its peers' model snapsh
cs.LG updates on arXiv.org

CRAFT: Fine-Tuning Pre-hoc Explainability in AI-native 6G RAN

・arXiv:2609.00590v1 Announce Type: new Abstract: The next generation of mobile networks is envisioned as fully AI-native, with AI-RAN architectures embedding small language models (SLMs) to perform reasoning over real-time telemetry. ・The state-of-the-art training paradigms for telecom LLMs, exemplified by RANSTRUCT-style supervised fine-tuning (SFT) on curated instruction data, are limited to post hoc rationalization.
cs.LG updates on arXiv.org

D3-Gym: Constructing Real-World Verifiable Environments for Data-Driven Discovery

・arXiv:2604.27977v4 Announce Type: replace-cross Abstract: Despite recent progress in language models and agents for scientific data-driven discovery, advancing their capabilities is held back by the absence of verifiable environments representing real-world scientific tasks. ・To fill this gap, we introduce D3-Gym, the first automatically constructed dataset with verifiable environments for scientific Data-Driven Disco
cs.LG updates on arXiv.org

DAGGER: Distractor-Aware Graph Generation for Executable Reasoning in Math Problems

・arXiv:2601.06853v3 Announce Type: replace-cross Abstract: Chain-of-Thought (CoT) prompting is widely adopted for mathematical problem solving, including in low-resource languages, yet its behavior under irrelevant context remains underexplored. ・To systematically study this challenge, we introduce DISTRACTMATH-BN, a Bangla benchmark that augments MGSM and MSVAMP with semantically coherent but computationally irrelevan
cs.LG updates on arXiv.org

Deep learning based numerical approximation algorithms for stochastic partial differential equations

・arXiv:2012.01194v3 Announce Type: replace-cross Abstract: In this article, we introduce a deep learning based approximation algorithm for SPDEs. ・Our approach employs neural networks to approximate the solutions of SPDEs along given realizations of the driving noise process. ・If applied to a set of simulated noise trajectories, it yields empirical distributions of SPDE solutions, from which functionals like the mean an
stat.ML updates on arXiv.org

Deep Skew-t Mixture Models

・arXiv:2609.00773v1 Announce Type: cross Abstract: High-dimensional clustering is challenging when component distributions are both heavy-tailed and directionally asymmetric. ・We propose a deep skew-$t$ mixture model (DStMM), a hierarchical factor-analytic mixture based on the generalised-hyperbolic skew-$t$ normal mean--variance representation. ・A shared inverse-gamma mixing variable is propagated along each complete l
cs.LG updates on arXiv.org

Denoising Diffusion Generative Models Secretly Calculate Attentions

・arXiv:2609.00885v1 Announce Type: cross Abstract: Denoising diffusion models are the dominant architecture for image generation, whereas most natural language generation and modeling are primarily handled by well-known transformer architectures employing attention mechanism. ・Here, we show that diffusion models also inherently use an attention mechanism very similar to that of transformers. ・Therefore, attention emerge
cs.LG updates on arXiv.org

Denoising the Deep Sky: Physics-Based CCD Noise Formation for Astronomical Imaging

・arXiv:2601.23276v4 Announce Type: replace-cross Abstract: Astronomical imaging remains noise-limited under practical observing conditions. ・Standard calibration pipelines remove structured artifacts but largely leave stochastic noise unresolved. ・Although learning-based denoising has shown strong potential, progress is constrained by scarce paired training data and the requirement for physically interpretable models in
cs.LG updates on arXiv.org

Dense Process Supervision for Search Agents via Fact Utility Estimation

・arXiv:2609.00833v1 Announce Type: cross Abstract: Reinforcement learning (RL) for search agents typically relies on outcome rewards. ・However, it often fails to achieve effective credit assignment, due to the unclear value of intermediate steps. ・It is hard to separate their contributions from the final result.
cs.LG updates on arXiv.org

Dense Weak Hiding: Closing Complexity Gaps in Nonconvex and PL Finite-Sum Optimization under Individual Smoothness

・arXiv:2609.00045v1 Announce Type: cross Abstract: Under individual smoothness, the optimal incremental first-order oracle (IFO) complexity of nonconvex finite-sum optimization has remained open. ・Known algorithms use $O(n+\sqrt{n}\,\Delta L_{\max}/\varepsilon^2)$ calls, while prior lower bounds miss a factor of $\sqrt{n}$. ・We prove the matching lower bound for randomized IFO algorithms whose component indices and quer
Takara TLDR - Daily AI Papers

Designing Proactive Thought Partners for Writing

・Writing involves diverse cognitive activities, from ideation to revision, and writers' needs vary across individuals and moments. ・Proactive AI promises to provide the right support at the right time, yet existing proactive tools largely focus on generic textual assistance, such as autocomplete. ・This paper studies the design space of proactive thought partners: AI agents that proactively offer customizable, higher-lev
cs.LG updates on arXiv.org

DeSyR: A Decoupled Symbolic Recovery Framework with PINN-Guided Structure Search and Physics-Informed Coefficient Refinement

・arXiv:2609.00530v1 Announce Type: new Abstract: Recovering compact explicit solutions from neural approximations is challenging when imperfect teacher data guide symbolic topology search and coefficient estimation. ・We present DeSyR, a decoupled symbolic recovery framework for differential equations. ・A physics-informed neural network guides repeated searches to construct candidate topologies with provisional constants
cs.LG updates on arXiv.org

Deterministic LLM Inference Across GPU Kernels: Power-of-Two INT8 Quantization Scales and the Limits of Tolerance-Based Conformance

・arXiv:2609.00363v1 Announce Type: new Abstract: Conformance suites for quantized GEMM kernels ask whether two implementations agree within a tolerance. ・We measure what such a suite can detect. ・Injecting nine faults into a reference INT8 pipeline over 8,232 layer--fault--regime cells of Qwen3-1.7B, we find that every one of five epilogue faults -- scale precision, double rounding, multiplication order, output truncati
cs.LG updates on arXiv.org

Different representation learning objectives recover distinct latent structures from the same psychometric data

・arXiv:2609.00100v1 Announce Type: cross Abstract: Psychometric questionnaires contain rich item-level information, yet it remains unclear whether different representation learning objectives recover the same latent organization. ・We investigated this question using 757 matched teacher-child pairs from the baseline assessment of the Cyprus ProW preschool trial. ・Behavioral structure was characterized from child SDQ, ASB
cs.LG updates on arXiv.org

Diffusion as a Training Curriculum for Timestep-Free Iterative Reasoning

・arXiv:2609.01449v1 Announce Type: new Abstract: Diffusion models and recursive reasoners are both iterative, but they carry information across iterations differently. ・We add a persistent hidden state to a diffusion denoiser and remove its timestep conditioning, leaving a single shared update that can be run to arbitrary depth. ・The result is an anytime solver: accuracy keeps improving with inference depth far beyond t
cs.LG updates on arXiv.org

Direct Optimization of a 3D Finite-Source Reflector via Neural-Network Parameterization

・arXiv:2609.00899v1 Announce Type: cross Abstract: We present a direct optimization method for three-dimensional freeform reflectors that transform the light of a finite-\'etendue source into a prescribed far-field angular intensity distribution. ・The reflector profile is represented by a small neural network (a multilayer perceptron), which is trained end-to-end through a differentiable ray-tracing objective.
cs.LG updates on arXiv.org

Disciplined Bilevel Programming

・arXiv:2609.00644v1 Announce Type: cross Abstract: Bilevel optimization provides a natural modeling language for hierarchical decision problems. ・However, applying existing numerical solvers usually requires substantial manual analysis and reformulation. ・In this paper, we introduce disciplined bilevel programming (DBLP), a symbolic framework that allows users to specify and solve optimistic bilevel problems in a high-l
cs.LG updates on arXiv.org

DiscoverPhysics: Benchmarking LLMs for Out-of-the-Box Scientific Thinking

・arXiv:2605.26087v2 Announce Type: replace-cross Abstract: Frontier LLMs now perform strongly across a wide range of physics evaluations, but it is hard to disentangle genuine reasoning from recall of established science. ・We introduce DiscoverPhysics, an interactive benchmark that asks a LLM agent to discover the laws of motion of a simulated world whose physics deliberately deviates from our own. ・We construct 22 worl
cs.LG updates on arXiv.org

DISTAL: Distillation and Self-Supervised Pretraining for Structure-Agnostic Materials Property Prediction

・arXiv:2609.00059v1 Announce Type: new Abstract: Materials property prediction remains difficult in low-data settings, where many target properties are supported by only a limited number of labeled samples. ・Models with the strongest predictive accuracy often depend on crystal structures, which restricts their use in early-stage screening when structural information is limited or unavailable. ・To address this challenge,
cs.LG updates on arXiv.org

DK-GBMKKM: Dynamic Kernel-Space Granular-Ball Multiple Kernel $k$-Means Clustering

・arXiv:2609.00647v1 Announce Type: new Abstract: Multiple kernel $k$-means integrates complementary nonlinear similarities by learning a combination of base kernels. ・Its pointwise optimization, however, is sensitive to noisy and boundary samples and repeatedly operates on sample-scale kernel matrices. ・Granular-ball representations organize local sample groups into mesoscopic units, but granular balls generated once in
cs.LG updates on arXiv.org

Do LLMs Know Your Neighborhood? Auditing LLM Priors for Neighborhood-Level Mobility Prediction and Structural Alignment

・arXiv:2609.00345v1 Announce Type: new Abstract: Human mobility is central to urban planning, transportation, public health, and emergency response, yet fine-grained trajectory data are often proprietary, restricted, and privacy-sensitive. ・Large language models (LLMs) offer a potential alternative by generating plausible mobility traces and predicting individual movement, but their ability to infer aggregate neighborh
cs.LG updates on arXiv.org

Does Fault Localization Beat a Fresh Attempt? A Placebo-Controlled Study of Test-Guided Code Repair

・arXiv:2609.00854v1 Announce Type: cross Abstract: Fault localization can focus a code model's repair on the statements a failing test implicates, but a targeted edit may succeed merely because it is small, and a second model call may succeed without using the failure at all. ・We separate these explanations with three arms applied to the same failed candidate: blind whole-solution resampling, spectrum-based localizatio
cs.LG updates on arXiv.org

Does Imitation Learning Preserve Temporal Robustness in Dexterous Manipulation? An Expert-Learner Comparison Across Task Execution Speeds

・arXiv:2609.01453v1 Announce Type: cross Abstract: Dexterous manipulation policies learned by imitation are typically evaluated for robustness to variation in scenes, objects, or instructions, but their performance across task execution speeds is less often examined. ・This leaves open how much temporal robustness a learner retains relative to the expert it imitates. ・We compare an expert and learner under the same task
cs.LG updates on arXiv.org

DOG-DPO:Dynamic Optimization in Geometry for Safety Alignment

・arXiv:2606.07678v3 Announce Type: replace Abstract: Safety alignment for large language models relies on preference data, but current pipelines often train on large, redundant datasets. ・Existing data selection methods typically score each preference pair independently, collapsing directional preference information into scalar quality or diversity scores. ・This sample-centric view is especially limiting in multi-datase
cs.LG updates on arXiv.org

Dr. Claw: An AI Scientist Workspace for Vibe Research

・arXiv:2609.00365v1 Announce Type: cross Abstract: Command-line coding agents (e.g., Claude Code, Gemini CLI) can already read and write files and sustain long sessions, yet end-to-end research still fragments across chat tools, IDEs, terminals, and writing environments, and the decisions that make it auditable are rarely preserved. ・We present Dr. ・Claw, an open-source workspace that wraps existing coding-agent executo
Takara TLDR - Daily AI Papers

DualDiff3D: Dual Structure-Appearance Diffusion Priors for Reliability-Enhanced 3D Gaussian Splatting

・While 3D Gaussian Splatting (3DGS) has revolutionized 3D reconstruction and novel-view synthesis, scenarios with limited input views often lead to poor reconstruction quality and artifacts in rendered novel views. ・Recent efforts attempt to utilize powerful diffusion priors, yet they typically process rendered and reference views concatenated along an additional dimension in a single network. ・These methods overlook an
cs.LG updates on arXiv.org

DualStake: Dual-Path Confidence Calibration in Deep Research Agents

・arXiv:2609.00935v1 Announce Type: cross Abstract: Deep Research agents tackle knowledge-intensive tasks through multi-round retrieval and decision-oriented generation. ・However, these agents suffer from severe overconfidence, making their expressed confidence unreliable for user trust and downstream abstention. ・To address this, we augment the Deep Research pipeline with step confidence elicitation after each retrieval
cs.LG updates on arXiv.org

DynaNDE: Dynamic Near-Data Expert Scheduling for Batched MoE Inference

・arXiv:2609.00407v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models enable efficient scaling of large language model (LLM) inference but suffer from substantial data-movement overhead when deployed on neural processing unit (NPU)-based systems. ・Near-Data Processing (NDP) provides a promising way to mitigate this bottleneck via cooperative NPU-NDP execution. ・However, existing NPU-NDP MoE systems do not f
cs.LG updates on arXiv.org

EchoSonar-R: A Multi-View Reasoning-Enabled Model for Disease Classification and Report Generation in Echocardiography

・arXiv:2606.28164v2 Announce Type: replace-cross Abstract: Echocardiography is the most widely used non-invasive cardiac imaging modality, providing essential information for cardiovascular diagnosis. ・Interpreting an echocardiogram requires synthesizing complementary evidence across multiple heart views to identify abnormalities and produce structured clinical reports. ・While recent efforts focus on improving classific
cs.LG updates on arXiv.org

Edge-Girth as a Structural Edge Feature for Graph Neural Networks

・arXiv:2609.01441v1 Announce Type: new Abstract: Graph neural networks (GNN) based on message passing are provably no more powerful than the one-dimensional Weisfeiler--Leman colour-refinement test (1-WL): two graphs it cannot tell apart receive identical representations, however deep or wide the network. ・A common remedy augments node or edge features with precomputed structural descriptors, most often counts of a fix
cs.LG updates on arXiv.org

EEG-AS: Instance-Level Foundation Model Selection for EEG Foundation Models via Behavior Reconstruction

・arXiv:2609.00653v1 Announce Type: new Abstract: Electroencephalography (EEG) is a non-invasive technique for measuring neural activity and has been widely used in neuroscience applications. ・Recent advances in EEG foundation models have enabled strong performance across diverse neural decoding tasks. ・However, no single foundation model consistently performs best across datasets or individual EEG instances, while insta
cs.LG updates on arXiv.org

EEG-VID: Task-Guided Latent Predictive Pretraining for EEG Decoding and Assistive Target Selection

・arXiv:2609.00566v1 Announce Type: new Abstract: We propose EEG-VID, a task-guided latent predictive pretraining framework for EEG decoding under session and subject shifts. ・EEG-VID predicts future latent EEG states from recent history using an exponential-moving-average target encoder and weak task guidance, followed by supervised fine-tuning. ・Across VIG-48 and BCI Competition IV-2a/IV-2b, Stage 1 improves mean accur
Takara TLDR - Daily AI Papers

EEG-VID: Task-Guided Latent Predictive Pretraining for EEG Decoding and Assistive Target Selection

・We propose EEG-VID, a task-guided latent predictive pretraining framework for EEG decoding under session and subject shifts. ・EEG-VID predicts future latent EEG states from recent history using an exponential-moving-average target encoder and weak task guidance, followed by supervised fine-tuning. ・Across VIG-48 and BCI Competition IV-2a/IV-2b, Stage 1 improves mean accuracy in 41 of 42 matched backbone-dataset-protoco
cs.LG updates on arXiv.org

Efficient Adaptation of ROMs for Unsteady Flows Using Data Assimilation

・arXiv:2602.23188v3 Announce Type: replace Abstract: We propose an efficient retraining strategy for a parameterized Reduced Order Model (ROM) that attains accuracy comparable to full retraining while requiring only a fraction of the computational time and relying solely on sparse observations of the full system. ・The architecture employs an encode-process-decode structure: a Variational Autoencoder (VAE) to perform di
cs.LG updates on arXiv.org

Efficient Learning of Balanced Signed Graphs via Sparse Linear Programming

・arXiv:2506.01826v2 Announce Type: replace Abstract: Signed graphs are equipped with both positive and negative edge weights, encoding pairwise correlations as well as anti-correlations in data. ・A balanced signed graph is a signed graph with no cycles containing an odd number of negative edges. ・Laplacian of a balanced signed graph has eigenvectors that map via a simple linear transform to ones in a corresponding posit
cs.LG updates on arXiv.org

Efficiently Estimating Optimal Hyperparameter Scaling Laws through Power-Law Entropy Search

・arXiv:2609.01431v1 Announce Type: new Abstract: Optimal hyperparameter scaling laws describe how the best hyperparameters for large language model (LLM) training change with model and data scale, enabling practitioners to predict optimal configurations at production scales without expensive large-scale tuning. ・However, estimating these scaling laws conventionally requires exhaustive grid searches over thousands of tr
cs.LG updates on arXiv.org

Elite-Weighted Supervised Fine-tuning for Goal-Directed Molecular Optimization

・arXiv:2609.00189v1 Announce Type: new Abstract: Goal-directed optimization is essential for steering molecular generators to propose candidates with desired properties. ・However, it is often implemented with policy-gradient reinforcement learning, which requires a generation-trajectory log-probability whose form depends on the model architecture and generation procedure. ・This makes an optimizer difficult to reuse acro
cs.LG updates on arXiv.org

EM^2Mem: Event-Centric Multimodal Memory for Large Language Models

・arXiv:2609.00551v1 Announce Type: cross Abstract: Multimodal memory offers a scalable interface for long-video question answering, but existing methods often retrieve captions, frames, transcripts, summaries, or graph facts as isolated fragments. ・Although searchable, such fragments are not generation-ready: language models must reconstruct cross-modal and temporal alignments at inference time, when context is limited
cs.LG updates on arXiv.org

Embedded Conditional Independence Tests for Large Language Model Generated Text with an Application to German Parliament Speeches

・arXiv:2609.00946v1 Announce Type: cross Abstract: Conditional independence tests (CITs) test for conditional dependence between two random objects $X$ and $Y$ given a third random object $Z$. ・Existing CITs have limited applicability to high-dimensional data, especially multimodal data like text. ・However, we show that such tests are of interest for large language model (LLM) outputs, where we test whether an output $X
cs.LG updates on arXiv.org

Enabling KV Caching of Shared Prefix for Diffusion Language Models

・arXiv:2606.07571v3 Announce Type: replace Abstract: Key-value (KV) caching for shared prefixes is essential for high-throughput large language model (LLM) serving, but it faces critical challenges in emerging diffusion language models (DLMs). ・In DLMs, bidirectional attention means that updating any token dynamically alters the entire context and its corresponding KVs. ・Thus, existing caching techniques developed for L
cs.LG updates on arXiv.org

ES-AHD: An Evolution Strategy Framework for Automatic Heuristic Design

・arXiv:2609.00023v1 Announce Type: cross Abstract: In this paper, we introduce ES-AHD, a novel framework that fundamentally integrates Evolution Strategy (ES) into Large Language Model (LLM)-driven Automatic Heuristic Design (AHD). ・Existing evolutionary approaches predominantly rely on random, individual-level mutation, leading to blind search and an imbalance between exploration and exploitation. ・To address these iss
cs.LG updates on arXiv.org

EvoFlint: An Evolutionary Atlas of Multi-Turn LLM Vulnerabilities

・arXiv:2609.00487v1 Announce Type: cross Abstract: Frontier language models that refuse harmful single-turn prompts often comply when the same intent is reached gradually over many turns, making multi-turn attacks one of the least understood failure modes of large language models. ・Most automated red-teaming methods treat this as a generation problem: produce attacks that break the model. ・We argue it is better framed a
cs.LG updates on arXiv.org

Exact Global MCMC with Denoising Diffusion

・arXiv:2609.00279v1 Announce Type: cross Abstract: This work shows that diffusion models learned with standard denoising loss can provide effective global MCMC proposals for complex high-dimensional target densities. ・The method is motivated by the observation that sequentially applying a forward and reverse diffusion process defines a Markov chain with a target stationary distribution for an ideal denoiser trained on
cs.LG updates on arXiv.org

Exact Risk-Complexity Laws for Projective Boundaries in Scenario Optimization and Distribution-Free Certification

・arXiv:2609.01355v1 Announce Type: cross Abstract: Scenario optimization, conformal prediction, and related distribution-free certification methods use finite samples to construct decisions or prediction sets with violation-risk guarantees for fresh observations. ・In several classical settings, the conditional violation risk follows an exact beta law, whose tail has a beta-binomial representation and whose parameter is
cs.LG updates on arXiv.org

Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents

・arXiv:2609.01245v1 Announce Type: new Abstract: Reinforcement learning is a natural way to post-train LLM agents for long-horizon interactive tasks judged only by end-of-task verification, yet a shared belief holds that outcome-only RL soon hits a ceiling on small open models. ・Recent work therefore compensates around the training with denser rewards, SFT priors, skill libraries, curated memory, or multi-agent orchest
cs.LG updates on arXiv.org

Exploring Sparse Autoencoders in Text-Based Causal Confounding Adjustment

・arXiv:2609.01322v1 Announce Type: cross Abstract: In many settings, studying causal questions based on text data requires adjusting for confounding information within texts. ・Yet there is a tradeoff in constructing text representations for adjustment: they must be sufficiently large and/or dense to preserve the confounding variables necessary for unbiased effect estimation, but sufficiently small and/or sparse to sati
cs.LG updates on arXiv.org

Facet-0: A Robotic Foundation Model for Contact-Rich Precise Manipulation

・arXiv:2609.01596v1 Announce Type: cross Abstract: Real-world robotic assembly at sub-millimeter tolerances demands spatial precision, compliant interaction, and robustness to contact failures. ・We present Facet-0, a robotic foundation model that predicts and values the contact consequences of its actions. ・Facet-0 unifies multimodal representation learning and reinforcement learning (RL) post-training around a joint ac
Takara TLDR - Daily AI Papers

Facet-0: A Robotic Foundation Model for Contact-Rich Precise Manipulation

・Real-world robotic assembly at sub-millimeter tolerances demands spatial precision, compliant interaction, and robustness to contact failures. ・We present Facet-0, a robotic foundation model that predicts and values the contact consequences of its actions. ・Facet-0 unifies multimodal representation learning and reinforcement learning (RL) post-training around a joint action-wrench proposal: a causal wrench history is a
cs.LG updates on arXiv.org

Fair Minimum Labeling: Efficient Temporal Network Activations for Reachability and Equity

・arXiv:2510.03899v3 Announce Type: replace-cross Abstract: Balancing resource efficiency and fairness is critical in networked systems that support modern learning applications. ・We introduce the \emph{Fair Minimum Labeling} (FML) problem: the task of designing a minimum-cost temporal edge activation plan that ensures each group of nodes in a network has sufficient access to a designated target set, according to specif
cs.LG updates on arXiv.org

Faster Than Flash: Exploiting Attention Sparsity for Efficient Long-Context Decoding

・arXiv:2609.00097v1 Announce Type: new Abstract: The development of long-context Large Language Models (LLMs) is constrained by the memory bandwidth bottleneck and quadratic complexity of the attention mechanism during decoding. ・To overcome the inherent trade-offs between the memory overhead of metadata-based metrics and the computational inefficiency of adaptive selection strategies, we present Faster Flash Decoding
cs.LG updates on arXiv.org

FedReview: Review and Dispose Poisoned Updates without Validation Datasets or Historic Knowledge

・arXiv:2402.16934v2 Announce Type: replace Abstract: Federated learning has emerged as a decentralized approach for training high-performance models without accessing user data. ・Despite its effectiveness, it is vulnerable to poisoning attacks, where malicious users manipulate the global model by uploading poisoned updates. ・In this paper, we propose FedReview, a review-based mechanism to identify and dispose the potent
cs.LG updates on arXiv.org

FedSPDnet: Geometry-Aware Federated Deep Learning with SPDnet

・arXiv:2604.22494v2 Announce Type: replace-cross Abstract: We introduce two federated learning frameworks for the classical SPDnet model operating on symmetric positive definite (SPD) matrices with Stiefel-constrained parameters. ・Unlike standard Euclidean averaging, which violates orthogonality, our approach preserves geometric structure through two efficient aggregation strategies: ProjAvg, projecting arithmetic mean
Takara TLDR - Daily AI Papers

Feedback-Assisted Trust Propagation over Document Relation Graphs for Retrieval-Augmented Generation

・Retrieval-augmented generation (RAG) systems rely on external corpora that may contain outdated, contradictory, noisy, or unreliable documents, introducing reliability risks. ・Prior work has leveraged document relations to improve the answer reliability of RAG. ・To propagate reliability signals beyond directly compared document pairs, we propose TrustPropRAG, which structures document relations as a graph and estimates
cs.LG updates on arXiv.org

Final Checkpoints Are Not Enough: Analyzing Latent Reasoning Faithfulness Along Training Trajectories

・arXiv:2607.06648v2 Announce Type: replace Abstract: Latent reasoning performs multi-step inference in continuous hidden states, promising more compact and efficient reasoning. ・However, these opaque states raise a question of faithfulness: whether the latent reasoning steps drive the final answer. ・Prior work studies this question at selected checkpoints and reports several unfaithful behaviors.
cs.LG updates on arXiv.org

FlexP-SFT: A Flexible Aggregation-Free Framework for On-Device Personalized Split Federated Fine-Tuning of LLMs

・arXiv:2508.10349v2 Announce Type: replace-cross Abstract: To fine-tune large language models (LLMs) over private data, federated learning (FL) has emerged as a promising paradigm. ・However, the prohibitive memory and communication demands of LLMs render standard FL impractical for resource-constrained edge devices. ・While split federated learning (SFL) alleviates the computing burdens via model partitioning, existing f
cs.LG updates on arXiv.org

FloydNet: A Learning Paradigm for Global Relational Reasoning

・arXiv:2601.19094v3 Announce Type: replace Abstract: Learning algorithmic computation often requires explicit relational intermediate states, yet many graph processors maintain their primary states on individual entities. ・We introduce \fnet and \textbf{Pivotal Attention} (PA), which maintain ordered pair states and update a target relation $(i,k)$ by attending over candidates formed from $(i,j)$ and $(j,k)$ for every
cs.LG updates on arXiv.org

Foundation models for electricity price forecasting and battery arbitrage: Can they replace market-specific forecasting models?

・arXiv:2609.00089v1 Announce Type: new Abstract: Foundation models promise accurate forecasts with little or no task-specific training, but whether they can replace models designed specifically for electricity price forecasting remains unclear. ・We compare nine variants from five foundation model families, evaluated in zero-shot mode, with two state-of-the-art electricity price forecasting benchmarks in Germany, Poland
cs.LG updates on arXiv.org

Fractal dimension predicts quantum kernel collapse in angle-encoded data

・arXiv:2609.00475v1 Announce Type: cross Abstract: Angle-encoded quantum kernels on tabular data collapse when the feature map is wider than the intrinsic dimension of the data. ・We propose the correlation fractal dimension D2 as an a priori qubit budget: encode D2 coordinates chosen by FD-ASE instead of the PCA-95% width or all E attributes. ・On nine data sets and a statevector simulator (n= 32), a one-layer ZZ fidelit
cs.LG updates on arXiv.org

FractalNet-Based Heterogeneous Federated Learning for Orbital Edge Intelligence in Satellite Mega-Constellations: A Wildfire Case Study

・arXiv:2609.00875v1 Announce Type: cross Abstract: Satellite mega-constellations are emerging as large-scale sensing, communication, and computation fabrics, yet their learning architectures remain largely inherited from terrestrial federated learning and ground-centric mission operations--- ill-suited to satellites that differ by orders of magnitude in Size, Weight, Power, and Cost (SWAP-C), radiation tolerance, link
Takara TLDR - Daily AI Papers

FractalNet-Based Heterogeneous Federated Learning for Orbital Edge Intelligence in Satellite Mega-Constellations: A Wildfire Case Study

・Satellite mega-constellations are emerging as large-scale sensing, communication, and computation fabrics, yet their learning architectures remain largely inherited from terrestrial federated learning and ground-centric mission operations--- ill-suited to satellites that differ by orders of magnitude in Size, Weight, Power, and Cost (SWAP-C), radiation tolerance, link availability, and propagation delay. ・We propose a
cs.LG updates on arXiv.org

Freeze, Diffuse, Decode: Geometry-Aware Adaptation of Pretrained Transformer Embeddings for Antimicrobial Peptide Design

・arXiv:2511.23120v2 Announce Type: replace Abstract: Pretrained transformers provide rich, general-purpose embeddings, which are transferred to downstream tasks. ・However, current transfer strategies: fine-tuning and probing, either distort the pretrained geometric structure of the embeddings or lack sufficient expressivity to capture task-relevant signals. ・These issues become even more pronounced when supervised data
cs.LG updates on arXiv.org

From Language to Behavior: Scaling Sequence Transformers for Industrial Recommendation Ranking with Rec-Native Designs

・arXiv:2609.01240v1 Announce Type: cross Abstract: Scaling Transformers has driven large gains in language modeling, but transplanting this to behavior-sequence modeling in production ranking is challenging: recommendation differs in signal quality, where behavior sequences are noisy, temporally irregular, and sparsely supervised, and in computation asymmetry, where each request scores many candidates against one shar
cs.LG updates on arXiv.org

From Truncation to Commitment: Persistent Context in Uniform Discrete Diffusion

・arXiv:2609.01043v1 Announce Type: new Abstract: Uniform-state discrete diffusion models update all tokens in parallel while keeping every position revisable. ・Even when the commonly used top-$p$ rule leaves only one candidate at a position, that choice affects only the current reverse step and can be revised at the next sampling step. ・We ask what changes when selected hypotheses instead become persistent context for l
cs.LG updates on arXiv.org

Frozen Cores Need Task Signal: Fisher-Whitened Cross-Covariance for Low-Resource LLM Adaptation

・arXiv:2609.00762v1 Announce Type: new Abstract: Parameter-efficient fine-tuning is usually framed as a question of how many parameters to update. ・Under a severe trainable-state budget, however, where those coefficients act is equally consequential. ・We study this choice through frozen-core adaptation: a calibration pass fixes left and right bases for each weight matrix, and fine-tuning optimizes only an $r\times r$ co
Takara TLDR - Daily AI Papers

Frozen Cores Need Task Signal: Fisher-Whitened Cross-Covariance for Low-Resource LLM Adaptation

・Parameter-efficient fine-tuning is usually framed as a question of how many parameters to update. ・Under a severe trainable-state budget, however, where those coefficients act is equally consequential. ・We study this choice through frozen-core adaptation: a calibration pass fixes left and right bases for each weight matrix, and fine-tuning optimizes only an $r\times r$ core.
cs.LG updates on arXiv.org

GazeRefine: Expert Gaze as a Test-Time Prompt for Training-Free Medical Image Segmentation

・arXiv:2609.01310v1 Announce Type: cross Abstract: Medical image segmentation remains difficult to scale because high-performing methods typically rely on dense expert annotations and task-specific training. ・We introduce GazeRefine, a training-free framework that uses gaze as an inference-time prompt for zero-shot medical image segmentation. ・Sparse, duration-weighted fixations are converted into foreground and backgro
cs.LG updates on arXiv.org

Generalization Bounds for Markov Algorithms through Entropy Flow Computations

・arXiv:2502.07584v3 Announce Type: replace-cross Abstract: Many learning algorithms can be represented as Markov processes, and understanding their generalization error is a central topic in learning theory. ・For specific continuous-time noisy algorithms, a prominent analysis technique relies on information-theoretic tools and the so-called ``entropy flow'' method. ・This technique is compatible with a broad range of ass
cs.LG updates on arXiv.org

Generative artificial intelligence for reliable mechanistic reasoning for corrosion

・arXiv:2609.00099v1 Announce Type: new Abstract: Corrosion accounts for approximately 4% of global GDP, and reliable prediction is essential for timely mitigation. ・Machine learning effectively predicts corrosion rates from composition, microstructure, and environmental variables, but cannot explain the underlying mechanisms. ・A reliable approach in safety-critical materials engineering requires not only accurate retrie
cs.LG updates on arXiv.org

GENIE: Watermarking Graph Neural Networks for Link Prediction

・arXiv:2406.04805v4 Announce Type: replace-cross Abstract: The rapid adoption, usefulness, and resource-intensive training of Graph Neural Network (GNN) models have made them an invaluable intellectual property in graph-based machine learning. ・However, their wide-spread adoption also makes them susceptible to stealing, necessitating robust Ownership Demonstration (OD) techniques. ・Watermarking is a promising OD framewo
cs.LG updates on arXiv.org

GenONet: A Generative operator Network for High-Resolution Precipitation Nowcasting

・arXiv:2609.00544v1 Announce Type: new Abstract: High-resolution precipitation nowcasting is critical for reducing the impacts of severe weather but remains difficult because of rapid storm evolution. ・Deep learning models have shown great promise for this task, but their predictive skill often deteriorates over longer forecast horizons. ・This leads to increasingly blurry forecasts that fail to capture the complex, non-
cs.LG updates on arXiv.org

Geometry-aware Latent Autoregressive Generative Model for PDEs in Complex Domains

・arXiv:2609.00297v1 Announce Type: new Abstract: Solving multiphysics partial differential equations (PDEs) remains a major challenge in scientific computing, especially for highly complex $\mu$m-scale tortuous geometries critical to energy and chemical engineering. ・We address this challenge by proposing a Geometry-aware Latent Autoregressive generative Model for PDEs (GeoLAMP) for solving physics within highly irregu
cs.LG updates on arXiv.org

GeoPAR: Large-Scale Multi-Agent Combinatorial Optimization with Geometry-Guided Parallel Autoregressive Learning

・arXiv:2609.00577v1 Announce Type: new Abstract: Multi-agent combinatorial optimization problems are notoriously challenging due to their NP-hard nature. ・Recent parallel autoregressive neural solvers improve inference efficiency by allowing agents to make decisions simultaneously, but their performance often degrades on large-scale instances. ・This is largely attributable to weak modeling of local geometric structures
Microsoft Research

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

・What if pathology foundation models could do more with less? ・GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. ・The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.
cs.LG updates on arXiv.org

Global Attention with Linear Complexity for Exascale Generative Data Assimilation in Earth System Prediction

・arXiv:2604.16590v2 Announce Type: replace Abstract: Accurate Earth system prediction requires state inference from incomplete observations, but conventional two-stage data assimilation (DA) is computationally prohibitive because repeated PDE-based ensemble forecasts, observation updates, and intermediate data movement limit ensemble size at high resolution. ・We introduce STORM, a one-stage generative AI framework that
cs.LG updates on arXiv.org

Good Memory Has ECC: Evaluating the Memory of Vision-Language Models Beyond Accuracy

・arXiv:2609.00103v1 Announce Type: new Abstract: Memory is widely viewed as an important unsolved problem for LLMs and VLMs, and current benchmarks typically evaluate it by testing accuracy over long text or video. ・However, accuracy alone misses properties that matter for real long-horizon tasks. ・We introduce ECCBench, a benchmark and evaluation protocol that measures memory beyond a system's capacity--its raw accurac
cs.LG updates on arXiv.org

Gradient-Update Mismatch: Rethinking Conflict-Free Training of Physics-Informed Neural Networks

・arXiv:2609.01558v1 Announce Type: new Abstract: Training Physics-Informed Neural Networks (PINNs) requires jointly optimizing physics residual and initial/boundary condition loss terms, which often induce conflicting gradients. ・Gradient surgery methods mitigate this issue by constructing directions from loss-specific gradients to reduce conflict before optimizer transformation. ・However, even when the constructed dire
cs.LG updates on arXiv.org

Group Adaptive Clipping Policy Optimization

・arXiv:2609.00444v1 Announce Type: new Abstract: Group relative policy optimization for reinforcement learning with verifiable rewards (RLVR) typically uses a fixed importance-sampling (IS) ratio clipping boundary across all rollouts. ・We identify a key limitation: rare correct rollouts on harder problems and abundant correct rollouts on easier problems are clipped at comparable rates, despite contributing very differe
Takara TLDR - Daily AI Papers

H2Table: Hierarchical Hypergraph-Enhanced Large Language Models for Complex Table Reasoning

・Tables are ubiquitous across diverse domains, yet reasoning over them remains a significant challenge for modern large language models (LLMs). ・Current approaches typically linearize tables into sequences, inherently overlooking their intrinsic two-dimensional and hierarchical structure. ・To address this, we propose H2Table (Hierarchical Hypergraph-Enhanced Table Reasoning), a novel framework that represents complex ta
cs.LG updates on arXiv.org

HarmoCore: Functional Latent Diffusion for Sparse Reconstruction of Oscillatory Wave Fields

・arXiv:2609.00679v1 Announce Type: new Abstract: Reconstructing oscillatory wave fields from scattered sensors is a severely underdetermined inverse problem. ・Beyond the challenges of general physical-field reconstruction, wave responses are complex-valued, frequency-sensitive, and highly oscillatory, while costly simulation and sensing often leave only extreme-sparse observations. ・Existing low-rank, operator, and diff
cs.LG updates on arXiv.org

HarnessEvolve: Learning from Reference Trajectories for Reliable Agent Self-Evolution

・arXiv:2609.00829v1 Announce Type: new Abstract: Self-evolving agents advance toward autonomy by optimizing their harness---prompts, skills, tools, and execution logic---based on environmental feedback. ・This paradigm, however, is hampered by three challenges: \textit{credit assignment failure}, where terminal success/failure feedback makes it ambiguous which step caused the error; \textit{shortcut learning}, where age
cs.LG updates on arXiv.org

HBQ: Hierarchical Scaling Block Quantization with Hardware-Efficiency-Aware Design for Accurate LLM Inference

・arXiv:2609.00450v1 Announce Type: new Abstract: Block Quantization (BQ) is a promising approach for efficient deployment of large language models (LLMs), enabling low-precision computation with controlled accuracy degradation. ・Compared to scalar weight-only quantization (WoQ), BQ quantizes both weight and activation, offering higher hardware efficiency and end-to-end inference on a unified datapath, but its design sp
cs.LG updates on arXiv.org

HEAL: Hindsight Entropy-Assisted Learning for Reasoning Distillation

・arXiv:2603.10359v2 Announce Type: replace-cross Abstract: Distilling reasoning capabilities from Large Reasoning Models (LRMs) into smaller models is typically constrained by the limitations of rejection sampling. ・Standard methods treat the teacher as a static filter, discarding complex "corner-case" problems where the teacher fails to explore valid solutions independently, thereby creating an artificial "Teacher Cei
cs.LG updates on arXiv.org

Hidden relationships in a document-derived property graph: top-k chunk embeddings and inverse-distance weighting over a dynamically evolving ontology

・arXiv:2609.00387v1 Announce Type: cross Abstract: Large language models extracting knowledge graphs from text capture only explicitly stated facts, often leaving semantically related entities disconnected across documents. ・We present an additive, engine-neutral second pass that discovers these latent ties without altering extracted facts. ・Each document is chunked and embedded once; top-k nearest- neighbor queries acr
cs.LG updates on arXiv.org

Hidden State Poisoning Attacks against Mamba-based Language Models

・arXiv:2601.01972v5 Announce Type: replace-cross Abstract: State space models (SSMs) like Mamba offer efficient alternatives to Transformer-based language models, with linear time complexity. ・Yet, their adversarial robustness remains critically unexplored. ・This paper studies the phenomenon whereby specific short input phrases induce a partial amnesia effect in such models, by irreversibly overwriting information in th
cs.LG updates on arXiv.org

Higher Structures in Deep Learning

・arXiv:2609.00472v1 Announce Type: new Abstract: We provide an expository introduction on the importance of higher-arity tensor operations to deep learning. ・Then, we conduct a novel empirical investigation of higher-arity phenomenon in trained neural networks, introduce a hypergraphical generalization of the multilayer perceptron, and explore connections to evolutionary algorithms. ・We conclude with a discussion of pro
cs.LG updates on arXiv.org

HiMPO: Hindsight-Informed Memory Policy Optimization for Less-Entangled Credit in Long-Horizon Agents

・arXiv:2606.16285v2 Announce Type: replace-cross Abstract: Long-horizon agents rely on memory mechanisms to compress interaction history, but optimizing memory writing faces a distinct credit assignment challenge: a memory update may be rewarded or penalized due to downstream tool failures, noisy observations, or reasoning errors rather than its own contribution. ・We propose HiMPO, a Hindsight-Informed Memory Policy Op
MIT News - Artificial intelligence

How an MIT research project became a global programming language

・With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
Takara TLDR - Daily AI Papers

How Correct Is Your Answer? A Semantic Correctness Framework for Open QA Evaluation

・Reliable evaluation of open-ended question answering remains a bottleneck for measuring answer correctness of modern LLMs. ・Unlike multiple-choice tasks, free-form answers may be correct in many surface forms and may fail in qualitatively different ways, including incompleteness, contradiction, overgeneration, and endorsement of false premises. ・Existing judgment-based and similarity-based metrics often collapse these
cs.LG updates on arXiv.org

How Do Language Models Choose Between Context and Memory?

・arXiv:2609.00753v1 Announce Type: new Abstract: When contextual information conflicts with the knowledge stored in model parameters, activation directions can be used to decode and steer which source the model follows. ・However, steering along a direction does not establish causality: whether the unedited model would naturally use that direction or whether the direction is reusable across tasks. ・We test these distinct
cs.LG updates on arXiv.org

How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?

・arXiv:2607.18114v2 Announce Type: replace-cross Abstract: Modern LLMs are alarmingly susceptible to surprisingly simple immaterial changes of input prompts: a casual hint, an incorrectly labeled few-shot example, or a fake prior assistant turn often flips an originally correct answer. ・We study where this susceptibility, spanning sycophancy and related cue-induced biases, lives inside the model. ・Across five model fami
OpenAI News

How law firm Gilbert + Tobin governs and scales AI with OpenAI

・See how Gilbert + Tobin combines CEO-led commitment, rigorous governance, and human accountability to scale ChatGPT Enterprise and Codex across the firm.
cs.LG updates on arXiv.org

How Temporal Correlations Shape Memory in Linear Recurrent Neural Networks

・arXiv:2609.00420v1 Announce Type: new Abstract: The linear recurrent neural network (LRNN) is a simple model for studying how much memory a network builds up as it trains. ・For uncorrelated inputs, earlier work found that training itself settles the network between keeping the past and reacting only to the present. ・Real sequences are correlated, and we solve the learning dynamics exactly for correlated inputs.
cs.LG updates on arXiv.org

I-CARE: Analysis of interference-related phenomena in a controllable, diverse and representative unlearning setting for text-to-image models

・arXiv:2609.00003v1 Announce Type: cross Abstract: Machine unlearning studies the removal of knowledge from an AI model, making the system forget a concept it previously learned. ・Despite rapid progress in generative machine unlearning, the unintended degradation of semantically related concepts that should have been retained (henceforth, interference) remains poorly characterized and inconsistently evaluated.
stat.ML updates on arXiv.org

If you can distinguish, you can express: Galois theory, Stone--Weierstrass, machine learning, and linguistics

・arXiv:2510.09902v3 Announce Type: replace-cross Abstract: This essay develops a parallel between the Fundamental Theorem of Galois Theory and the Stone--Weierstrass theorem: both can be viewed as assertions that tie the distinguishing power of a class of objects to their expressive power. ・We provide an elementary theorem connecting the relevant notions of "distinguishing power". ・We also discuss machine learning and d
cs.LG updates on arXiv.org

Independent Reinforcement Learning in Discounted Markov Games

・arXiv:2609.00504v1 Announce Type: cross Abstract: In this work, we study radically uncoupled learning in discounted general-sum Markov games. ・Assuming ``$\mathsf{ETH}$ for $\mathsf{PPAD}$", we show that, for every fixed discount factor, there is no polynomial-time algorithm for computing inverse-polynomially accurate coarse correlated equilibria in discounted general-sum Markov games when players learn independently
Takara TLDR - Daily AI Papers

InSight: A Benchmark for Agentic Claim Verification in Interactive Visualizations

・Vision Language Models have demonstrated remarkable proficiency in interpreting static visual artifacts, but modern data analysis is inherently dynamic, requiring the active interrogation of interactive environments. ・Existing benchmarks are predominantly constrained to static imagery and one-shot question answering and fail to capture the epistemic demands of this domain, where evidence is frequently occluded, distri
Takara TLDR - Daily AI Papers

Inspicio: Open-Vocabulary, LLM-Based Sense Retrieval for Historical Languages

・Word Sense Disambiguation has advanced rapidly for English and a handful of well-resourced modern languages, but it continues to assume the existence of a sense inventory and a word-to-sense mapping in the source language (Navigli, 2026). ・These assumptions break down for most historical and low-resource languages, whose dedicated WordNets are either incomplete or still under construction. ・We present Inspicio, an open
Google DeepMind News

Introducing Gemini 3.8 Flash and 3.8 Flash Cyber

Introducing Gemini 3.8 Flash and 3.8 Flash Cyber
cs.LG updates on arXiv.org

Inverse Reconstruction of Shock Time Series from Shock Response Spectrum Curves using Machine Learning

・arXiv:2603.03229v3 Announce Type: replace Abstract: The shock response spectrum (SRS) is widely used to characterize the response of single-degree-of-freedom (SDOF) systems to transient accelerations. ・Because the mapping from acceleration time history to SRS is nonlinear and many-to-one, reconstructing time-domain signals from a target spectrum is inherently ill-posed. ・Conventional approaches address this problem thr
Takara TLDR - Daily AI Papers

Inverse Rendering for Modeling with Line Primitives

・Faithfully capturing diverse real-world objects with fuzzy, anisotropic structures, such as hair, fur, fibers, and textiles, for efficient real-time visualization remains challenging. ・Recent radiance field reconstruction methods capture these structures from multi-view images using translucent volumetric primitives such as 3D Gaussians rather than opaque low-dimensional primitives (e.g., triangles, line segments, and
cs.LG updates on arXiv.org

Investigating Linear Probe Robustness to Linguistic Register, Medical Specialty, and Corpus Shifts in Medical QA

・arXiv:2609.01361v1 Announce Type: cross Abstract: Linear classifiers trained on hidden states of a large language model (LLM), linear probes, can flag factual errors from a single forward pass. ・Geometrically, that implies that true and false statements separate along a stable direction in hidden state space, i.e., the truth direction. ・Prior work disagrees on whether this generalises across input shifts, but the disag
cs.LG updates on arXiv.org

iPINN for Broadband CARS Phase Retrieval: A Framework for Function Approximation and Inverse Modeling Problems in Nonlinear Spectroscopy

・arXiv:2609.00883v1 Announce Type: new Abstract: Phase retrieval in broadband coherent anti-Stokes Raman spectroscopy (BCARS) is an ill-posed inverse problem. ・The Raman-like signal is encoded in the imaginary part of the resonant susceptibility, which mixes coherently with a non-resonant background (NRB) that varies across acquisitions. ・We introduce an inverse physics-informed neural network (iPINN) that predicts Lore
cs.LG updates on arXiv.org

Is Knowledge Distillation Actually Greener? A Case Study in Machine Translation

・arXiv:2602.09691v2 Announce Type: replace-cross Abstract: Knowledge distillation (KD) is a technique to compress a larger teacher system into a smaller student. ・In machine translation, KD is commonly evaluated through translation quality and inference efficiency, without jointly accounting for the environmental costs of producing and deploying the distilled system. ・We evaluate representative KD methods both on bespok
cs.LG updates on arXiv.org

Iterative GRPO: Batch-Online Policy Iteration for Multi-Turn RL via Single-Turn RLHF

・arXiv:2511.21638v2 Announce Type: replace Abstract: Practical LLM agents often operate over multi-turn conversations where success is determined only after the full interaction ends. ・Most multi-turn RL methods train via on-policy rollouts, but unlike in single-turn RLHF, the policy cannot produce a trajectory alone, since an external environment must respond after each agent turn. ・For conversational agents, this envi
cs.LG updates on arXiv.org

Keep Everyone Happy: Online Fair Division of Numerous Items with Few Copies

・arXiv:2408.12845v3 Announce Type: replace Abstract: This paper considers a novel variant of the online fair division problem involving multiple agents in which a learner sequentially observes an indivisible item that must be irrevocably allocated to one of the agents to achieve a desired balance between fairness and efficiency. ・Existing algorithms assume a small number of items with a sufficiently large number of cop
cs.LG updates on arXiv.org

KV Cache Offloading for Context-Intensive Tasks

・arXiv:2604.08426v5 Announce Type: replace Abstract: With the growing demand for long-context LLMs across a wide range of applications, the key-value (KV) cache has become a critical bottleneck for both latency and memory usage. ・Recently, KV-cache offloading has emerged as a promising approach to reduce memory footprint and inference latency while preserving accuracy. ・Prior evaluations have largely focused on tasks th
cs.LG updates on arXiv.org

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior

・arXiv:2605.26797v2 Announce Type: replace Abstract: We study Latent Recurrent Transformer (LRT), a lightweight augmentation of autoregressive transformers that reuses a high-level source-layer hidden state from the previous token as recurrent memory for the next token. ・Because this state is already computed during ordinary decoding, LRT introduces a cross-token, cross-layer latent pathway while preserving the standar
cs.LG updates on arXiv.org

Latent-Space No-Arbitrage Geometry of Generative Models for Implied Volatility Surfaces

・arXiv:2609.00332v1 Announce Type: cross Abstract: Generative models for implied volatility surfaces must produce outputs that satisfy static no-arbitrage constraints. ・We study these constraints in latent space. ・For a fixed generator, we assign each latent code a scalar margin determined by the no-arbitrage conditions of the generated surface.
cs.LG updates on arXiv.org

LatentPress: Context Compression Beyond Text and Vision

・arXiv:2609.01507v1 Announce Type: new Abstract: Compressed context is usually carried as human-readable text or as rendered images that must be decoded, even when its consumer is a language model. ・We introduce LatentPress, which writes conversational histories and long documents into a third representation: continuous memory tokens that a frozen decoder reads directly through its input-embedding interface, with no te
cs.LG updates on arXiv.org

Leakage-Audited Benchmarking Reveals Limited Evidence for Cross-Subject Auditory-Evoked EEG Vowel Perception Decoding

・arXiv:2605.00865v4 Announce Type: replace-cross Abstract: We tested whether auditory-evoked EEG supports subject-independent five-vowel perception decoding when trial identity, model identity, prediction provenance, and participant-level inference are controlled within a single benchmark. ・We reconstructed Study 2 event tables from OpenNeuro ds006104 version 1.0.1 and analysed the consonant-vowel pair task.
cs.LG updates on arXiv.org

Learning Sparse Decision Trees via Transformer Variational Auto-Encoders

・arXiv:2609.01430v1 Announce Type: new Abstract: Decision trees are among the most widely used models in machine learning, largely due to their transparent decision logic, making them well-suited for high-stakes decision-making contexts. ・However, most existing learning algorithms focus on predictive performance, overlooking the joint optimization of other desirable properties, such as structural sparsity. ・In this work
cs.LG updates on arXiv.org

Learning Task-Specific Antibody Representations via Function-Aware Masking

・arXiv:2609.00518v1 Announce Type: new Abstract: Antibody-specific language models pretrained via masked language modeling (MLM) learn representations that are critical for downstream sequence design and property prediction tasks. ・Yet, the corruption process itself is rarely leveraged as a source of inductive bias during pretraining. ・While preferentially masking complementarity-determining regions (CDRs) improves bind
cs.LG updates on arXiv.org

Learning to Refine Hidden States for Reliable LLM Reasoning

・arXiv:2606.17524v3 Announce Type: replace Abstract: Large language models show strong reasoning ability, but their internal reasoning process can remain unstable in complex multi-step settings, where early hidden-state errors may propagate to incorrect predictions. ・We propose ReLAR, a reinforcement-guided latent refinement framework that iteratively updates hidden representations before decoding. ・ReLAR maintains a co
cs.LG updates on arXiv.org

Learning to Remember: End-to-End Training of Memory Agents for Long-Context Reasoning

・arXiv:2602.18493v2 Announce Type: replace Abstract: Long-context LLMs and Retrieval-Augmented Generation defer state tracking and evidence consolidation to query time, which is brittle when facts evolve and answers depend on latent states. ・We introduce Unified Memory Agent (UMA) for a one-to-many setting: query-agnostic external memory is constructed once from a stream and reused across multiple future QA sessions.
cs.LG updates on arXiv.org

Let Confidence Change, Not the Prediction: Prediction-Preserving Repair for Post-hoc Calibration

・arXiv:2609.01072v1 Announce Type: new Abstract: Post-hoc calibration corrects reported confidence, yet a multiclass calibrator can also change the associated top-1 prediction. ・Accuracy captures only the net effect of these changes on correctness, not how often predictions change; the Top-1 Prediction Change Rate (TPCR) instead measures this frequency. ・We propose Calibrator-Output Repair for Top-1 Decision Preservatio
cs.LG updates on arXiv.org

Lightweight Adaptation of EEG Foundation Models for Stroke Motor Imagery Decoding: Domain Shift and Subject-Level Robustness

・arXiv:2609.00282v1 Announce Type: cross Abstract: Motor imagery (MI) electroencephalography (EEG) decoding could support post-stroke rehabilitation, but models developed on healthy cohorts may not transfer reliably to pathological EEG. ・We evaluated whether Low-Rank Adaptation (LoRA) can efficiently adapt three pretrained EEG foundation models (i.e., LaBraM-base, REVE-base, and REVE-large) for binary left- versus righ
cs.LG updates on arXiv.org

Lingua Franca or Probing Artifact? Rethinking Latent Language in Multilingual LLMs

・arXiv:2609.00155v1 Announce Type: cross Abstract: Latent language identification is often used to argue that multilingual language models route computation through language-specific states, such as English pivots. ・However, existing probes infer latent language from different signals, such as the geometry of hidden states or what can be decoded from intermediate representations. ・Since such claims shape conclusions abo
cs.LG updates on arXiv.org

LLM-driven design of physics-constrained constitutive models: two agents are better than one

・arXiv:2605.23754v2 Announce Type: replace Abstract: Developing constitutive models that capture how materials deform under load traditionally requires years of specialized expertise in continuum mechanics, machine learning, and scientific programming. ・Large language models (LLMs) have recently been shown to lower this barrier by generating constitutive models on demand, but existing single-agent pipelines lack system
cs.LG updates on arXiv.org

Local Reference Geometry Residual Augmentation for Imbalanced Time Series Classification

・arXiv:2609.00093v1 Announce Type: new Abstract: Imbalanced time series classification is often addressed by changing the training distribution, objective, logits, or final threshold. ・These interventions address important biases, yet leave a representation-level question unmeasured: after minority support is reduced, does a learned feature space remain locally reliable around minority regions? ・We identify a training-l
cs.LG updates on arXiv.org

Logarithmic-Free Moment and Generalization Bounds for Uniformly Stable Algorithms

・arXiv:2608.09870v2 Announce Type: replace-cross Abstract: Uniform stability is a classical tool for controlling the generalization error of a learning algorithm. ・Bousquet, Klochkov, and Zhivotovskiy (2020) showed that the problem can be reduced to a moment inequality for a sum of weakly interacting functions of independent random variables. ・Their bound contains an additional factor $\log n$, and they asked whether th
Takara TLDR - Daily AI Papers

MADS: A Multiview Acoustic Descriptor Set Beyond Standard Spectral Summaries

・Dominant audio classification pipelines rely either on compact handcrafted summaries or on fixed time-frequency frontends such as log-mel representations prior to deep modeling. ・While highly successful, these representations do not explicitly expose the physical dynamics of the underlying sound-generating event. ・We introduce MADS (Multi-view Acoustic Descriptor Set), a compact 19-dimensional physics-informed descript
cs.LG updates on arXiv.org

Manifold-Aware General Coded Computing for Straggler-Resilient Distributed Computing

・arXiv:2609.00552v1 Announce Type: new Abstract: Existing coded-computing designs do not explicitly exploit the intrinsic structure of the input data. ・In communication systems, statistical structure and redundancy are often removed through source coding (or compression) before channel coding is applied. ・This principle, however, does not transfer directly to coded computation.
The latest research from Google

Mapping global methane emissions from space with deep learning

Mapping global methane emissions from space with deep learning
cs.LG updates on arXiv.org

MaskCode: Mask Transformer for Feedback-Assisted Coding With Linear Block Codes

・arXiv:2609.00715v1 Announce Type: cross Abstract: Feedback-based coding schemes have demonstrated substantial performance gains over today's open-loop coding schemes. ・Unfortunately, these gains are usually achieved in idealized settings with perfect feedback. ・Over the last few years, machine learning-based schemes have been shown to be promising solutions for implementing feedback-based codes, particularly when combi
cs.LG updates on arXiv.org

Matched Queries for Curvature and Density at Branching Junctions

・arXiv:2609.01319v1 Announce Type: cross Abstract: At a junction, a score field can reveal weighted tangent rays, yet these first-order quantities do not determine how individual branches bend or how their densities change away from the center. ・Recovering this missing information is necessary for describing local continuation beyond a single point, but finite observations must separate branchwise second-order effects
cs.LG updates on arXiv.org

Measuring consistency via ensemble margin and local prediction variability: Auditing decision systems in the presence of predictive multiplicity

・arXiv:2609.01397v1 Announce Type: cross Abstract: The Rashomon effect is a machine learning phenomenon where equally accurate models produce different predictions for the same inputs (predictive multiplicity). ・Existing work primarily focuses on multiplicity within individual models, but in more complex decision systems, the impact of the Rashomon effect is less well understood. ・In this work, we study multiplicity fro
cs.LG updates on arXiv.org

MemoryWalker: Stop Training Agents on Contexts They Never Saw

・arXiv:2609.00865v1 Announce Type: new Abstract: Production agent harnesses such as Claude Code and Qwen-Agent compress context during rollout, but training under compression creates a conditioning problem: every eviction branches the effective history, so the learning object is a tree rather than a sequence. ・Existing linearizations either retain the rightmost path, causing time-travel leakage, or replay a depth-first
cs.LG updates on arXiv.org

MIDR: Enrichment-Augmented Indexing for Multimodal Document Retrieval

・arXiv:2609.01316v1 Announce Type: cross Abstract: Retrieval over visually rich documents has a representation problem: important content often lives in tables, charts, figures, and layout relations that plain OCR linearizes, corrupts, or omits. ・ColPali-family visual retrievers address this with patch-level multi-vector indexes and late-interaction scoring, keeping image-derived retrieval on the query-time serving pat
cs.LG updates on arXiv.org

MineDraft: A Framework for Batch Parallel Speculative Decoding

・arXiv:2603.18016v3 Announce Type: replace-cross Abstract: Speculative decoding (SD) accelerates large language model inference by using a smaller draft model to propose draft tokens that are subsequently verified by a larger target model. ・However, the performance of standard SD is often limited by the strictly sequential execution of these drafting and verification stages. ・To address this, this paper proposes MineDra
MIT News - Artificial intelligence

MIT Quantum Initiative launches postdoctoral fellowship program

・The Institute welcomes its first cohort of QMIT Fellows this fall to advance interdisciplinary quantum research.
cs.LG updates on arXiv.org

MMAI Gym for Science: Training Liquid Foundation Models for Drug Discovery

・arXiv:2603.03517v2 Announce Type: replace Abstract: General-purpose large language models (LLMs) that rely on in-context learning do not reliably deliver the scientific understanding and performance required for drug discovery tasks. ・Simply increasing model size or introducing reasoning tokens does not yield significant performance gains. ・To address this gap, we introduce the MMAI Gym for Science, a one-stop shop mol
stat.ML updates on arXiv.org

Model Predictive Control is almost Optimal for Heterogeneous Restless Multi-armed Bandits

・arXiv:2511.08097v2 Announce Type: replace-cross Abstract: We consider a general infinite horizon Heterogeneous Restless multi-armed Bandit (RMAB). ・Heterogeneity is a fundamental problem for many real-world systems largely because it resists many concentration arguments. ・In this paper, we assume that each of the $N$ arms can have different model parameters.
cs.LG updates on arXiv.org

Modeling Information Blackouts in Missing Not-At-Random Time Series Data

・arXiv:2601.01480v3 Announce Type: replace-cross Abstract: Traffic forecasting systems rely on fixed sensor networks that frequently exhibit contiguous blackouts. ・Such outages are usually treated as ignorable missingness, although dropout can depend on unobserved traffic conditions. ・We study this possibility with an MNAR-aware latent state-space model that combines linear traffic dynamics with a Bernoulli missingness
cs.LG updates on arXiv.org

Modelpedia: A Catalog of Model Findings for the Meta-Science of AI

・arXiv:2609.01090v1 Announce Type: new Abstract: Scientific knowledge about AI models is produced faster than the community can organize it. ・Every few months a new foundation model reshapes the field and hundreds of papers, blogs, and technical reports document how each behaves or fails. ・Yet, these findings remain scattered and effectively unretrievable.
cs.LG updates on arXiv.org

MRMAD: A Multi-Round Multi-Audio Benchmark for Evaluating Acoustic Degradation Perception in Large Audio-Language Models

・arXiv:2608.22236v2 Announce Type: replace-cross Abstract: Large audio-language models (LALMs) have shown promising progress in understanding speech, music, and general sound events, yet their ability to reason about how audio signals are degraded remains underexplored. ・Existing benchmarks primarily evaluate semantic understanding, event recognition, or high-level audio reasoning, leaving a basic question unanswered:
cs.LG updates on arXiv.org

MUGEN: Generating Unlearnable Graph Examples for Multiple Learning Tasks

・arXiv:2609.00696v1 Announce Type: new Abstract: Graph data across diverse domains can expose valuable relational information to unauthorized representation learning, creating a pressing need for protection against such misuse. ・Unlearnable examples offer a data-level defense by perturbing a training release so that models trained on it fail to generalize to clean data. ・Existing methods generate unlearnable graph examp
Takara TLDR - Daily AI Papers

MUGEN: Generating Unlearnable Graph Examples for Multiple Learning Tasks

・Graph data across diverse domains can expose valuable relational information to unauthorized representation learning, creating a pressing need for protection against such misuse. ・Unlearnable examples offer a data-level defense by perturbing a training release so that models trained on it fail to generalize to clean data. ・Existing methods generate unlearnable graph examples for only a specified downstream task.
cs.LG updates on arXiv.org

Multi-Head Self Attention is a Parameter Identification Mechanism

・arXiv:2609.01231v1 Announce Type: new Abstract: We prove that a multi-head scaled dot product attention can be viewed as a parameter identification strategy. ・The ratio of unidentified parameters to the total number of parameters scales like the reciprocal of the number of heads ($1/2 \to 1/(2H)$), meaning models with more heads are structurally more identified. ・A subtle side effect of the mathematics observation that
cs.LG updates on arXiv.org

Multi-Step Knowledge Interaction Analysis via Rank-2 Subspace Disentanglement

・arXiv:2511.01706v3 Announce Type: replace-cross Abstract: Natural Language Explanations (NLEs) describe how Large Language Models (LLMs) make decisions by drawing on external Context Knowledge (CK) and Parametric Knowledge (PK). ・Understanding the interaction between these sources is key to assessing NLE grounding, yet these dynamics remain underexplored. ・Prior work has largely focused on i) single-step generation and
cs.LG updates on arXiv.org

Multi-View Causal Discovery without Non-Gaussianity: Identifiability and Algorithms

・arXiv:2502.20115v4 Announce Type: replace Abstract: Causal discovery is a difficult problem that typically relies on strong assumptions on the data-generating model, such as non-Gaussianity. ・In practice, many modern applications provide multiple related views of the same system, which has rarely been considered for causal discovery. ・Here, we leverage this multi-view structure to achieve causal discovery with weak ass
Takara TLDR - Daily AI Papers

MutMem-V2: Cryptographically Authorized Mutation in Persistent Agent Memory Portable Verification and Reproducible Evidence

・MutMem V1 introduced retention-preserving, cryptographically authorized mutation for persistent agent memory but did not provide a complete portable verification contract or clean-install reproduction path. ・MutMem V2 closes that publication gap without introducing a second memory engine. ・It specifies exact canonical bytes, domain-separated object and bundle commitments, mandatory recall-evidence membership and orderi
cs.LG updates on arXiv.org

mzCache: On-Device LLM Memory Management under Multitasking

・arXiv:2609.01338v1 Announce Type: cross Abstract: On-device mobile Large Language Model (LLM) inference is gaining significant attention. ・However, mobile devices operate in highly dynamic multitasking environments where users frequently switch between applications. ・This creates memory pressure, forcing LLM memory (model weights and KV cache) to be evicted by the operating system.
cs.LG updates on arXiv.org

NashDreamer: Model-Based Reinforcement Learning for Zero-Sum Imperfect-Information Games

・arXiv:2609.01549v1 Announce Type: new Abstract: Model-based reinforcement learning (MBRL) has achieved remarkable results in single-agent domains, yet its extension to competitive imperfect information games (IIGs) remains underexplored. ・In multi-agent settings, opponent-induced non-stationarity complicates the learning process, and decentralized model learning faces severe identifiability barriers, which we argue ma
cs.LG updates on arXiv.org

Neural means and kernel corrections for operator learning

・arXiv:2609.00389v1 Announce Type: new Abstract: We combine neural network means with exact Mat\'ern kernel regressions of their residuals and of their learned features, and evaluate the pairing on two public emulation problems with published baselines: the structural-mechanics benchmark of de Hoop et al. ・and the OCO-2 radiative-transfer emulator of Lamminp\"a\"a et al. ・On structural mechanics the combination reaches
cs.LG updates on arXiv.org

Neural Symbollic Regression Using Deep Learning and Sparse Modelling

・arXiv:2609.01102v1 Announce Type: new Abstract: Symbolic Regression (SR) seeks to find succinct mathematical expressions that represent the fundamental relationships within data, providing interpretability and scientific understanding that exceeds that of black-box models. ・Nevertheless, traditional methods like Genetic Programming face challenges with scalability and are highly sensitive to noise, while sparse regres
cs.LG updates on arXiv.org

NeuroPriv: Adversarial Representation Learning for Privacy in Wearable EEG Systems

・arXiv:2609.00390v1 Announce Type: cross Abstract: Wearable EEG systems may expose sensitive information beyond their intended health function, creating substantial risks to neuroprivacy. ・In this work, we show that commonly used EEG features can reveal participant identity and demographic attributes in addition to supporting the intended cognitive task. ・Wearable EEG is increasingly being explored for cognitive monitor
cs.LG updates on arXiv.org

Neurosymbolics for Data Engineering: Achieving Long Context Token Reduction Without Finetuning

・arXiv:2609.00367v1 Announce Type: cross Abstract: Large Language Models are increasingly deployed for sophisticated data engineering tasks such as generating structured queries from natural language, Text-to-SQL, and automating complex spreadsheet operations. ・However, maximizing their utility demands both higher finetuning-free accuracy and solutions to the computational bottleneck imposed by the Transformer architec
cs.LG updates on arXiv.org

Nonlinear Dynamics In Optimization Landscape of Shallow Neural Networks with Tunable Leaky ReLU

・arXiv:2510.25060v2 Announce Type: replace-cross Abstract: In this work, we study the nonlinear dynamics of a shallow neural network trained with mean-squared loss and leaky ReLU activation. ・Under Gaussian inputs and equal layer width k, (1) we establish, based on the equivariant gradient degree, a theoretical framework, applicable to any number of neurons k>= 4, to detect bifurcation of critical points with associate
stat.ML updates on arXiv.org

Nonparametric inference for density-dependent McKean--Vlasov diffusions

・arXiv:2609.01166v1 Announce Type: cross Abstract: The present research is devoted to the nonparametric estimation of a density-dependent drift coefficient in a multivariate McKean--Vlasov diffusion from independent observations at a common time, as well as the stationary density. ・Under certain assumptions on the (known) potential, we reduce the problem to the one-dimensional one and construct a sieve maximum-likeliho
NVIDIA Blog

NVIDIA and CrowdStrike Strengthen Agentic Cybersecurity Frontier

・“We’re at an inflection point in cybersecurity,” Jensen Huang told a sold-out crowd at CrowdStrike’s Fal.Con 2026 in Las Vegas Tuesday. ・Attacks are now automated. ・Defense has to be, too.
cs.LG updates on arXiv.org

OCGQuant: Outlier-Companion Grouping for NVFP4 Quantization

・arXiv:2609.00066v1 Announce Type: cross Abstract: NVFP4 is an efficient microscaling format for low-bit inference, but activation outliers can still degrade quantization accuracy within NVFP4 blocks. ・Within each quantization block, large activations can dominate the block scale, increasing the quantization error of the remaining values sharing the same scale. ・Existing post-training quantization (PTQ) methods mitigate
cs.LG updates on arXiv.org

On the Existence of Consistent Adversarial Attacks in High-Dimensional Linear Classification

・arXiv:2506.12454v2 Announce Type: replace-cross Abstract: What fundamentally distinguishes an adversarial attack from a misclassification due to limited model expressivity or finite data? ・In this work, we investigate this question in the setting of high-dimensional binary classification, where statistical effects due to limited data availability play a central role. ・We introduce a new error metric that precisely capt
cs.LG updates on arXiv.org

On the Reliability of Generative Augmentation: A Wasserstein-Based Theoretical and Empirical Study

・arXiv:2609.01410v1 Announce Type: cross Abstract: Generative data augmentation is widely used to mitigate class imbalance, yet its theoretical effect on downstream generalization remains poorly understood. ・In this work, we develop a statistical framework for conditional generative augmentation and analyze its impact on classification risk. ・We formalize augmentation as a distribution-mixing process and show that the r
cs.LG updates on arXiv.org

One-Layer Transformer Provably Learns Multiclass One-Nearest Neighbor in Context

・arXiv:2609.01311v1 Announce Type: new Abstract: We extend recent work establishing an equivalence between one-layer transformers and nearest-neighbor classifiers in the binary setting to the multiclass case. ・By leveraging the simplex encoding, we show that one-layer transformers with an argmax classification head behave identically to a one-nearest-neighbor classifier in the multiclass setting. ・This closes a gap left
cs.LG updates on arXiv.org

Online Regime-aware Calibration for Black-box Social Simulators via Posterior-assisted Evolutionary Dynamic Optimization

・arXiv:2601.19481v2 Announce Type: replace-cross Abstract: Evolutionary dynamic optimization (EDO) commonly assumes that environmental changes can be detected from fitness variations and handled through random re-initialization, historical solutions, or learned transition patterns. ・Online calibration of black-box simulators introduces a different setting, where the dynamic objective is induced by sequential observatio
cs.LG updates on arXiv.org

Online Self-Weighted Fine-Tuning

・arXiv:2609.00734v1 Announce Type: new Abstract: Standard supervised fine-tuning (SFT) assigns the same explicit loss weight to every expert demonstration, regardless of the model's changing competence over training queries. ・Reinforcement learning (RL) based methods adapt update strength using model-generated rollouts, but often require substantially more sampling and can be unstable on hard tasks. ・We propose \textbf{
cs.LG updates on arXiv.org

Online simultaneous inference for quantiles via smoothed stochastic gradient descent

・arXiv:2505.13299v2 Announce Type: replace-cross Abstract: This paper considers the estimation of quantiles via a smoothed version of the stochastic gradient descent (SGD) algorithm. ・By smoothing the score function with a bandwidth tied to the learning rate, we obtain estimates that are monotone in the quantile level at every iteration, while retaining the memory and computational efficiency required for streaming dat
cs.LG updates on arXiv.org

Ontology-Guided Neuro-Symbolic Inference: Grounding Language Models with Mathematical Domain Knowledge

・arXiv:2602.17826v2 Announce Type: replace-cross Abstract: Language models exhibit fundamental limitations -- hallucination, brittleness, and lack of formal grounding -- that are particularly problematic in high-stakes specialist fields requiring verifiable reasoning. ・I investigate whether formal domain ontologies can enhance language model reliability through retrieval-augmented generation. ・Using mathematics as proof
OpenAI News

OpenAI supports California’s bill to advance youth AI safety

・OpenAI supports California SB 1119, advancing strong, age-appropriate AI safeguards for teens while preserving opportunities to learn, create, and explore.
cs.LG updates on arXiv.org

Optimizing Byzantine Node Placement in Decentralized Federated Learning

・arXiv:2609.01495v1 Announce Type: new Abstract: Security evaluations of decentralized federated learning (DFL) typically focus on how Byzantine participants behave, while largely overlooking which participants are compromised. ・Yet, because aggregation is distributed over a communication graph, the placement of Byzantine nodes determines how malicious influence propagates through the network. ・We therefore treat Byzant
Takara TLDR - Daily AI Papers

P-PatchDiff: Progressive Patch Diffusion Models for Low-light Image Enhancement

・Recent advancements in low-light image enhancement have leveraged diffusion models for their strong ability to generate perceptually realistic, detailed images. ・Patch diffusion models further offer a promising solution to size-agnostic image restoration while improving efficiency. ・However, existing methods typically rely on small, fixed patches (e.g., 64$\times$64) that cannot capture image-level brightness context,
OpenAI News

Path to Astra: critical capabilities and frontier safeguards

・Astra is the first OpenAI model to meet the Critical cybersecurity capability threshold under the Preparedness Framework, with stronger safeguards for release.
cs.LG updates on arXiv.org

Patterning in Practice: Debiasing Reward Models with Susceptibilities

・arXiv:2609.00699v1 Announce Type: new Abstract: Reward models trained on human preferences are known to suffer from length, formatting, and other stylistic biases. ・In this paper we use patterning, which reweights each preference pair according to its measured effect on posterior expectation values of benchmark losses (its susceptibility), to debias a Gemma 2 9B Instruct reward model trained on Skywork-Reward-Preferen
cs.LG updates on arXiv.org

PEARL: Training Socratic Tutors with Pedagogically Aligned Reinforcement Learning

・arXiv:2605.29582v2 Announce Type: replace Abstract: Large Language Models (LLMs) show strong potential as educational tutors. ・Existing approaches typically train them to solve problems and provide correct answers, but this problem-solving-centered paradigm overlooks key requirements of effective tutoring: progressive guidance and the coordination of multiple pedagogical objectives across multi-turn interactions.
cs.LG updates on arXiv.org

Performance-Efficiency Tradeoffs in Transformers: An Approximation Theory Perspective

・arXiv:2510.03784v2 Announce Type: replace Abstract: Transformers have achieved remarkable successes across a wide range of applications, yet the theoretical foundation of their model efficiency remains underexplored. ・In this work, we investigate how the model parameters -- mainly attention heads and head dimensions -- should be allocated across layers to balance expressivity and efficiency. ・We first provide mathemati
cs.LG updates on arXiv.org

Performative Privacy: When Differential Privacy Maximizes Utility

・arXiv:2608.28198v2 Announce Type: replace Abstract: Privacy-preserving learning is often motivated by the idea that protecting users' data can preserve trust and thus participation, improving utility in the long term. ・However, this claim has not been formalized so far. ・In parallel, performative learning provides a framework for studying learning systems whose deployment affects the data they later observe.
cs.LG updates on arXiv.org

Persistent Entropy as a Detector of Phase Transitions

・arXiv:2602.09058v2 Announce Type: replace-cross Abstract: Persistent entropy is a scalar summary of persistence barcodes widely used to detect regime changes, yet there is no account of when a structural change in a barcode must produce a detectable change in entropy. ・We establish a model-agnostic theorem supplying such conditions. ・Treating persistence diagrams as random objects indexed by a control parameter, we ide
cs.LG updates on arXiv.org

Physiological Information Reliability: Cross-Layer Adaptive Resource Allocation for Cardiovascular Sensing

・arXiv:2609.00435v1 Announce Type: cross Abstract: Cardiovascular sensing systems must preserve clinically useful information despite signal degradation, wireless losses, energy constraints, and edge-computation latency. ・We introduce Physiological Information Reliability (PIR), a cross-layer framework that represents physiological information value jointly with wireless, energy, and computation states and uses a conte
stat.ML updates on arXiv.org

Pointwise Majorization for sub-Weibull and Mixed Tail Processes with Applications in Quadratic Chaos and Ergodic Diffusions

・arXiv:2609.01576v1 Announce Type: cross Abstract: Classical chaining controls an indexed stochastic process through a single worst-case bound, which can obscure substantial variation across the index set. ・We establish the first simultaneous pointwise majorization theory for Banach-valued processes with sub-Weibull or two-metric mixed-tail increments. ・For an anchored sub-Weibull process on a separable index space, wri
cs.LG updates on arXiv.org

Poisson-Gamma Dynamical Systems with Time-varying Transition Dynamics

・arXiv:2609.00896v1 Announce Type: new Abstract: Bayesian methodologies for handling count-valued time series have gained prominence due to their ability to infer interpretable latent structures and to estimate uncertainties. ・Among these Bayesian models, Poisson-Gamma Dynamical Systems (PGDSs) are proven to be effective in capturing the evolving dynamics underlying observed count sequences. ・However, the state-of-the-a
cs.LG updates on arXiv.org

Polarizable atomic multipoles for learning long-range electrostatics

・arXiv:2605.05746v2 Announce Type: replace-cross Abstract: Long-range electrostatics and polarization remain central obstacles to extending machine learning interatomic potentials (MLIPs) to ionic, polar, and interfacial systems. ・Here we introduce a semi-local framework for learning electrostatics from energies and forces using polarizable atomic multipoles. ・Local equivariant descriptors predict environment-dependent
OpenAI News

Polimill builds Japan's next-generation public AI infrastructure

・Polimill uses OpenAI GPT models and Codex to help municipalities search and use administrative knowledge while accelerating development.
Takara TLDR - Daily AI Papers

Position Matters: Feature Inversion Attacks in ViT Split Inference with Token Reduction and Shuffling

・Vision Transformers (ViTs) are increasingly used in split-inference systems, where edge devices transmit intermediate token representations to a remote cloud. ・In this setting, token reduction lowers computation and communication costs, while token shuffling disrupts the spatial organization of the transmitted tokens, potentially limiting information leakage. ・However, their privacy benefits remain unclear against feat
cs.LG updates on arXiv.org

Position: Privacy Is a Claim, Not a Property of Synthetic Data

・arXiv:2609.01273v1 Announce Type: new Abstract: Synthetic data has become a common component of machine learning research. ・While widely adopted, its use in privacy-sensitive contexts has quietly shifted from a claim of residual inference risk under stated assumptions to an appearance-based property inferred from data generation itself. ・In this position paper, we argue that this shift reflects an implicit change in co
cs.LG updates on arXiv.org

Post-Training Science for Supervised Fine-Tuning

・arXiv:2609.01244v1 Announce Type: new Abstract: Every supervised fine-tuning run forces the same chain of decisions, such as learning rate, batch size, LoRA or full fine-tuning, how many epochs, which optimiser, and what data to feed the model. ・Each of these is typically rediscovered from scratch for every new model and dataset. ・Here we measure them under one instrument: a sweep that varies one lever at a time, and s
cs.LG updates on arXiv.org

Pre-carved Niches: The Formation Dynamics of Modular Task Partitions in Early LLM Training

・arXiv:2609.01170v1 Announce Type: new Abstract: Large language models exhibit a modular internal organization that mirrors well-studied functional networks of the human brain, but how this organization forms during training is unknown: prior work has characterized finished models, not the formation process. ・We track formation step by step: we train a Pythia-410M model from scratch (two trajectories, bf16 and fp32) an
cs.LG updates on arXiv.org

Predicting Subsurface Abnormalities Growth using Physics-Informed Neural Networks

・arXiv:2609.01417v1 Announce Type: new Abstract: The research explores the pioneering integration of Physics-Informed Neural Networks (PINNs) into the domain of Ground-Penetrating Radar (GPR) data prediction. ・This research presents a detailed development framework for a specialized PINN model, proficient at interpreting and forecasting GPR data, much like how medical imaging models predict tumor behavior. ・By harnessin
cs.LG updates on arXiv.org

Prediction-Assisted Pricing and Admission for LLM APIs with Stochastic Token Consumption

・arXiv:2609.00710v1 Announce Type: cross Abstract: An LLM application often sells or internally allocates several service products: a small or premium model, a short or long token cap, and possibly multiple posted prices. ・The operational decision is not merely which model answers a prompt. ・A price changes purchase probability, a token cap changes both user value and the tail of resource consumption, and accepted reque
Google DeepMind News

Proactive cyber defense for governments and enterprises

Proactive cyber defense for governments and enterprises
cs.LG updates on arXiv.org

Probabilistic Multi-Agent Aircraft Landing Time Prediction

・arXiv:2512.08281v2 Announce Type: replace-cross Abstract: Accurate and reliable aircraft landing time prediction is essential for effective resource allocation in air traffic management. ・However, the inherent uncertainty of aircraft trajectories and traffic flows poses significant challenges to both prediction accuracy and trustworthiness. ・Therefore, prediction models should not only provide point estimates of aircra
cs.LG updates on arXiv.org

Process-Aware AI for Rainfall-Runoff Modeling: A Mass-Conserving Neural Framework with Hydrological Process Constraints

・arXiv:2603.25093v2 Announce Type: replace Abstract: Machine learning models can achieve high predictive accuracy in hydrological applications but often lack physical interpretability. ・The Mass-Conserving Perceptron (MCP) provides a physics-aware artificial intelligence (AI) framework that enforces conservation principles while allowing hydrological process relationships to be learned from data. ・In this study, we inve
cs.LG updates on arXiv.org

Provably Efficient Federated Reinforcement Learning with Linear Function Approximation and Logarithmic Communication Cost

・arXiv:2609.00193v1 Announce Type: cross Abstract: We study federated online reinforcement learning with linear function approximation. ・While recent multi-agent reinforcement learning algorithms achieve strong regret guarantees, they typically require sharing raw trajectories. ・This reliance incurs a communication cost that scales linearly with the number of episodes and violates the privacy constraints of federated se
cs.LG updates on arXiv.org

Provably Safe Sim-to-Real Transfer

・arXiv:2609.01418v1 Announce Type: new Abstract: To mitigate the sample complexity of real-world reinforcement learning (RL), a common practice is to first train a policy in a simulator, where samples are cheap, and then deploy the learned policy in the real world with the hope that it generalizes effectively. ・Such direct sim-to-real transfer is not guaranteed to succeed: simulator-trained policies can be suboptimal i
cs.LG updates on arXiv.org

QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization

・arXiv:2609.00224v1 Announce Type: new Abstract: Weight-only post-training quantization (PTQ) can alleviate the computational burden of serving large language models (LLMs) at scale. ・However, existing PTQ methods often fail to generalize across models and suffer severe accuracy loss below 2 bits. ・Many leverage unstructured sparsity to mitigate this loss, but at the cost of regularity and GPU-friendly execution.
cs.LG updates on arXiv.org

Quantum Sparse Autoencoders for Q-Matrix Estimation in Cognitive Diagnosis

・arXiv:2609.01537v1 Announce Type: new Abstract: Q-matrices play a central role in cognitive diagnosis within educational data mining (EDM), specifying which latent skills each assessment item requires. ・Data-driven Q-matrix estimation remains challenging when assessments involve many correlated skills and when real response patterns depart from idealized generative assumptions. ・We introduce a novel quantum sparse auto
cs.LG updates on arXiv.org

REAL-Q: E2E LLM Quantization via Dynamic Gradient Descent

・arXiv:2609.00049v1 Announce Type: new Abstract: Post-training quantization (PTQ) is essential for deploying large language models (LLMs) under strict resource constraints. ・State-of-the-art PTQ methods quantize each layer with a single closed-form second-order solver: to remain analytically tractable, they heavily approximate the global loss (dropping cross-channel coupling, pooling output rows into groups), and they
cs.LG updates on arXiv.org

Real-Time Neuromorphic Spectrum Intelligence Simulator

・arXiv:2609.00585v1 Announce Type: cross Abstract: We present the Real-Time Neuromorphic Spectrum Intelligence Simulator (RT-NuSIS), a modular framework to study spiking neural network (SNN) and memristor-inspired agents for dynamic spectrum access under constrained energy budgets and adversarial conditions. ・RT-NuSIS couples leaky integrate-and-fire neuronal dynamics, memristive synaptic models, physics-informed energ
cs.LG updates on arXiv.org

RECAP: Regression Evaluation for Continual Adaptation of Prompts

・arXiv:2606.06698v4 Announce Type: replace Abstract: Production agentic systems routinely face evolving constraints and must comply from the very next interaction. ・Scenarios like a tool-call notification changing a compliance threshold or a policy update adding disclosure requirements fit this criteria, having close to no room for errors in production. ・This proactive adaptation setting is common in deployment, but abs
cs.LG updates on arXiv.org

Recent Developments in Transformer Inference Deployment on FPGA Platforms: A Survey

・arXiv:2609.01212v1 Announce Type: new Abstract: With the rapid and continuous growth in the incorporation of machine learning models based on the Transformer architecture, capable deployment is in high demand. ・In this context, capable deployment refers to operational performance aspects, e.g., throughput and latency, as well as efficiency aspects, e.g., energy consumption. ・When it comes to the task of inference using
cs.LG updates on arXiv.org

Recurrent State Encoders for Efficient Neural Combinatorial Optimization

・arXiv:2509.05084v2 Announce Type: replace Abstract: The primary paradigm in Neural Combinatorial Optimization (NCO) consists of construction methods, where a neural network is trained to sequentially add one solution component at a time until a complete solution is formed. ・We observe that the typical changes to the state between two steps are small, since usually only the node added to the solution is removed from th
cs.LG updates on arXiv.org

REFACTOR-VLA: Unsupervised Library Learning of Typed Motor Programs

・arXiv:2609.01215v1 Announce Type: new Abstract: Most vision-language-action (VLA) models -- OpenVLA, $\pi_0$, RT-2, RDT-1B -- are monolithic: they emit raw motor commands or short action chunks without organizing behavior into reusable abstractions, so they degrade on long-horizon tasks and resist interpretation. ・Existing skill-discovery methods sidestep the core question of when two action sequences are behaviorally
cs.LG updates on arXiv.org

Relational Task Generation Language: A Declarative Specification Framework for Relational Deep Learning

・arXiv:2609.01292v1 Announce Type: cross Abstract: Relational Deep Learning (RDL) has become a powerful paradigm for learning from multi-tabular data. ・However, manually defining RDL prediction tasks is a laborious process that frequently results in data leakage. ・To address this issue, we introduce Relational Task Generation Language (RTGL) - an open-source declarative language that streamlines RDL task formulation by
cs.LG updates on arXiv.org

ReNFT: Repairing Mode Collapse in Reward Post-Training via Internal Probability-Mass Recalibration

・arXiv:2609.00061v1 Announce Type: new Abstract: Reward post-training of diffusion generators inevitably concentrates probability mass on a few reward-favored modes, a mode collapse that erases within-prompt diversity. ・Existing methods for mitigating collapse rely on external signals or interfaces, augmenting the reward with perceptual objectives, adjusting reference regularization, or modifying the text encoder, but
cs.LG updates on arXiv.org

Reparameterization through Coverings and Topological Weight Priors

・arXiv:2604.23804v2 Announce Type: replace Abstract: We generalise the reparameterization trick (RT) applied in variational autoencoders (VAEs) letting these have latent spaces of non-trivial topology - i.e. ・that of base manifolds covered with other ones, on which some technique for RT is available. ・That is possible since covering maps are measurable - moreover, this allows to establish an inequality on KL-divergence
cs.LG updates on arXiv.org

Replicating TRACE: A Practitioner's Guide to Its Threshold and Particle Budget

・arXiv:2609.01108v1 Announce Type: new Abstract: TRACE (Math & Lienhart, arXiv:2602.01135) reads causal graphs over event types out of a pretrained autoregressive sequence model by thresholding a per-position conditional-mutual-information estimate at a fixed tau. ・We independently replicate its headline synthetic result: with tau selected on a validation split, mean per-sequence F1 against exact interventional truth r
cs.LG updates on arXiv.org

ReproRepo: Scaling Reproducibility Audits with GitHub Repository Issues

・arXiv:2606.18237v2 Announce Type: replace-cross Abstract: Reproducing research results from papers and released code is central to scientific progress. ・Existing works have introduced benchmarks to evaluate whether LLM agents can assist with reproducibility, but they are difficult to scale due to their reliance on substantial manual effort for data curation and evaluation. ・We introduce ReproRepo, a scalable framework
cs.LG updates on arXiv.org

Rethinking Learnability in Offline Data-driven Optimization

・arXiv:2609.01493v1 Announce Type: new Abstract: Black-Box Optimization (BBO) has found broad applications, but evolutionary algorithms and Bayesian optimization face efficiency challenges as real-world BBO problems grow increasingly complex. ・Data-driven optimization improves the efficiency of BBO algorithms by learning from data. ・Offline data-driven optimization seeks high-quality solutions using only a fixed set of
cs.LG updates on arXiv.org

Retrieved but not ranked: surface-form bias in structural retrieval, from mathematics to agent trajectories

・arXiv:2609.01556v1 Announce Type: new Abstract: We evaluate embedding retrieval where surface form and meaning are pulled apart on purpose: retrieving items that share underlying structure but not wording, in two unrelated domains under one protocol, competition mathematics (MathNet-Retrieve; 500 queries, 117,088-item corpus) and embodied-agent trajectories (ALFWorld-derived; 118 queries, 336 trajectories).
cs.LG updates on arXiv.org

RetroReasoner: A Reasoning LLM for Strategic Retrosynthesis Prediction

・arXiv:2603.12666v3 Announce Type: replace Abstract: Retrosynthesis prediction aims to identify reactants that can synthesize a given product molecule. ・Although molecular large language models (LLMs) have recently shown promising results, most existing methods either generate reactants directly or provide only generic product-level analysis, without explicitly reasoning about bond-disconnection strategies that justify
cs.LG updates on arXiv.org

Right Frame, Wrong Rule: Cultural Cues Expose the Financial Knowledge Gap They Were Meant to Close

・arXiv:2609.00999v1 Announce Type: cross Abstract: When a question has valid answers under different normative frameworks, a language model must decide which framework to use and whether it can answer correctly within it. ・We call this setting normative pluralism and study it in Islamic finance using a four-choice taxonomy that separates framework selection from within-framework correctness. ・This separation reveals the
cs.LG updates on arXiv.org

Rigorous Error Certification for Neural PDE Solvers: From Empirical Residuals to Solution Guarantees

・arXiv:2603.19165v2 Announce Type: replace Abstract: Uncertainty quantification for partial differential equations is traditionally grounded in discretization theory, where solution error is controlled via mesh/grid refinement. ・Physics-informed neural networks fundamentally depart from this paradigm: they approximate solutions by minimizing residual losses at collocation points, introducing new sources of error arisin
cs.LG updates on arXiv.org

Risk-Aware Decision-Making for Autonomous Overtaking: A World Model-Based Mixture-of-Experts Framework

・arXiv:2609.00385v1 Announce Type: cross Abstract: Autonomous highway overtaking demands foresighted decision-making to handle complex interactions, stochastic traffic evolution, and temporal risk accumulation. ・However, standard safe reinforcement learning approaches typically rely on implicit value-based risk estimations rather than explicit dynamics modeling, thereby struggling to accurately capture complex risk pro
cs.LG updates on arXiv.org

RW-LoRA: Communication-Efficient Decentralized LoRA Fine-Tuning via Random Walks

・arXiv:2609.00078v1 Announce Type: new Abstract: Parameter-efficient fine-tuning methods such as LoRA have become a standard approach for adapting large foundation models. ・Adopting fine-tuning to distributed settings faces several challenges. ・Most existing distributed LoRA methods rely on centralized aggregation, and gossip-based decentralized LoRA requires repeated synchronization among multiple model copies.
cs.LG updates on arXiv.org

S-CEReBrO: Breaking the Memory Barrier in Continuous EEG Monitoring

・arXiv:2607.27913v2 Announce Type: replace Abstract: Foundation models offer a promising paradigm for Electroencephalography (EEG) analysis, leveraging generalizable representations from vast unlabeled datasets. ・Yet, Transformer-based architectures face a critical bottleneck: global attention mechanisms couple the attention memory state to the signal duration, causing memory overflow during continuous monitoring.
cs.LG updates on arXiv.org

Safin-1: Safety from Within through Memory-Native State Evolution

・arXiv:2609.00092v1 Announce Type: new Abstract: Long-horizon complex tasks require foundation models to accumulate information, maintain internal states, and adapt over extended interactions. ・Safety should be an intrinsic property of the model itself, rather than a behavioral constraint relying solely on external safeguards or post-hoc alignment such as supervised fine-tuning. ・This motivates Safety from Within, where
cs.LG updates on arXiv.org

SAGE: State-Grounded, Abstention-Aware Evaluation of Task-Oriented Dialogue Agents

・arXiv:2609.00434v1 Announce Type: cross Abstract: Evaluating task-oriented dialogue agents requires judging not merely whether a reply reads well but whether each turn advances the underlying workflow state correctly--a distinction conventional holistic LLM judges can miss because they evaluate the available context as a single unit and require one or more full-model calls per turn. ・We propose SAGE (State-Grounded Ab
cs.LG updates on arXiv.org

SAGE: Subpopulation-Aware Generative Enhancement for Mitigating Spurious Correlations

・arXiv:2609.01051v1 Announce Type: new Abstract: Spurious correlations pose a significant challenge to the robustness of modern machine learning. ・The inherent imbalance in dataset distributions often leads traditional Empirical Risk Minimization (ERM) models to rely on majority spurious attributes for classification, resulting in poor performance on minority groups. ・This problem becomes particularly challenging when t
cs.LG updates on arXiv.org

SCALE:Scalable Conditional Atlas-Level Endpoint transport for virtual cell perturbation prediction

・arXiv:2603.17380v3 Announce Type: replace Abstract: Virtual-cell models aim to predict how cell populations respond to perturbations, but control and treated cells are measured as unpaired populations, complicating the learning of perturbation-specific effects. ・We present SCALE, a conditional transport model that represents cells as unordered sets and predicts treated populations without cell-level matching.
cs.LG updates on arXiv.org

Scaled Idempotence in Transformer Attention: Paired OV Geometry and Shared-Value Algebras

・arXiv:2609.01129v1 Announce Type: new Abstract: We identify a recurrent algebraic regularity in Transformer attention: a sparse subset of effective OV operators $T=OV^\top$ nearly closes under composition, $T^2\approx\alpha T$. ・Across six pretrained endpoints spanning 2.8B--235B parameters, 3.98--8.00% of heads reach squared closure alignment $\mathcal{P}\geq0.9$, while no matched within-layer O/V mismatch does.
cs.LG updates on arXiv.org

Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs

・arXiv:2609.01573v1 Announce Type: cross Abstract: How to divide a fixed annotation budget between supervised fine-tuning (SFT) and reinforcement learning (RL) during LLM post-training remains an open problem. ・Existing work characterizes only broad trends (e.g., SFT dominates in low-data regimes), lacks a principled allocation framework, and does not examine whether the optimal ratio transfers across model sizes.
cs.LG updates on arXiv.org

SEBA: Sample-Efficient Black-Box Attacks on Visual Reinforcement Learning

・arXiv:2511.09681v2 Announce Type: replace Abstract: Visual reinforcement learning has achieved remarkable progress in visual control and robotics, but its vulnerability to adversarial perturbations remains underexplored. ・Most existing black-box attacks focus on vector-based or discrete-action RL, and their effectiveness on image-based continuous control is limited by the large action space and excessive environment q
stat.ML updates on arXiv.org

Seeing the Forest for the Trees: The Gaussian Process Limit of BART

・arXiv:2607.28844v2 Announce Type: replace-cross Abstract: Bayesian Additive Regression Trees (BART) have shown state-of-the-art performance in both prediction and causal inference problems. ・Previous theoretical work has attempted to explain BART's superior performance by establishing posterior contraction rates for standard BART models, but these rates depend strongly on the number of covariates. ・Here, we take a diff
cs.LG updates on arXiv.org

Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM Teachers

・arXiv:2609.01567v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) provide useful priors for interactive decision-making, but using them directly as policies is expensive and brittle: they must be queried at every step, do not improve from environment interaction, and can repeat systematic errors. ・We study how to learn a cheap autonomous policy from an online, expensive, and imperfect but informative VLM
cs.LG updates on arXiv.org

Self-Reports Are Not Verification: Environment-Grounded Auditing of LLM Operators in Evolutionary Search

・arXiv:2609.00652v1 Announce Type: cross Abstract: Language model agents increasingly propose actions, observe external feedback, and explain their own behavior. ・Their confidence and rationales are convenient monitoring signals, but convenience is not verification. ・We introduce an environment-grounded audit in which every intermediate proposal receives an exact outcome.
cs.LG updates on arXiv.org

Semantic-Guided Multimodal Preprocessing for Vision Transformer-Based Clear Cell Renal Cell Carcinoma Grading

・arXiv:2609.01426v1 Announce Type: cross Abstract: Clear cell renal cell carcinoma (CCRCC) grading is essential for treatment planning, yet existing approaches either analyze patch-level images directly or focus solely on nuclei-level classification, without linking to final tumor grading. ・We propose a semantic-guided multimodal preprocessing method that integrates nuclei classification maps from existing pre-trained
Takara TLDR - Daily AI Papers

Semantic-Guided Multimodal Preprocessing for Vision Transformer-Based Clear Cell Renal Cell Carcinoma Grading

・Clear cell renal cell carcinoma (CCRCC) grading is essential for treatment planning, yet existing approaches either analyze patch-level images directly or focus solely on nuclei-level classification, without linking to final tumor grading. ・We propose a semantic-guided multimodal preprocessing method that integrates nuclei classification maps from existing pre-trained models with RGB histopathology images for Vision T
cs.LG updates on arXiv.org

Semi-Supervised Classification with Informative Missing Labels in Weibull Mixture Models

・arXiv:2609.00774v1 Announce Type: cross Abstract: We consider semi-supervised classification from a partially classified sample arising from a two-component Weibull mixture. ・The feature is observed for all data, whereas some class labels are missing. ・The probability of a missing label is modelled as a function of classification uncertainty, giving a feature-dependent missing-at-random (MAR) mechanism that shares para
cs.LG updates on arXiv.org

Shallower ReLU Network Representations via Exact Linear Algebra

・arXiv:2607.21651v2 Announce Type: replace Abstract: We study the depth required by ReLU networks to exactly represent piecewise linear functions, focusing specifically on the maximum function. ・This problem has recently received significant attention in both the ML and TCS literature. ・We prove that $\max_n(x)=\max\{x_1,\ldots,x_n\}$ is exactly representable with two hidden layers for every $n\leq 12$.
cs.LG updates on arXiv.org

Sharp Mixed Spectral Barron Regularity of Coulombic Many-Electron Wave Functions

・arXiv:2609.00872v1 Announce Type: cross Abstract: We establish sharp mixed spectral Barron regularity for eigenfunctions of molecular Coulomb Hamiltonians. ・The mixed norm is a Fourier $L^1$ norm with one isotropic weight and coordinate-product weights, and therefore detects regularity invisible to the isotropic Barron scale. ・For a nonempty set $I$ of electron indices on which the wave function is antisymmetric, we de
cs.LG updates on arXiv.org

Sierpi\'nski--Knopp Wasserstein Distance for Persistence Diagrams and Applications to 2-Wasserstein Approximation

・arXiv:2609.01528v1 Announce Type: cross Abstract: This paper introduces the Sierpi\'nski-Knopp (SK) Wasserstein distance, a fast metric between persistence diagrams. ・The SK-Wasserstein distance, denoted $d_{\mathrm{SK}}$, maps diagram points and their diagonal projections to the unit interval via the Sierpi\'nski-Knopp space-filling curve on the upper diagonal triangle. ・The encoded point sets are then efficiently mat
Takara TLDR - Daily AI Papers

Sierpiński--Knopp Wasserstein Distance for Persistence Diagrams and Applications to 2-Wasserstein Approximation

・This paper introduces the Sierpiński-Knopp (SK) Wasserstein distance, a fast metric between persistence diagrams. ・The SK-Wasserstein distance, denoted $d_{\mathrm{SK}}$, maps diagram points and their diagonal projections to the unit interval via the Sierpiński-Knopp space-filling curve on the upper diagonal triangle. ・The encoded point sets are then efficiently matched via one-dimensional optimal assignment, in \(O(N\
cs.LG updates on arXiv.org

Silence is Golden: Mitigating Hallucinations in Large Audio-Language Models via Layer-Weighted Vector Steering

・arXiv:2510.12851v2 Announce Type: replace-cross Abstract: Large Audio-Language Models (LALMs) excel in Audio QA but often suffer from hallucinations ungrounded in the audio. ・To our knowledge, we are the first to propose applying vector steering to the audio domain to mitigate this. ・Unlike text-based steering, our silence-anchored contrastive approach steers the model away from hallucinations by contrasting active aud
cs.LG updates on arXiv.org

SinkPruner: Sink-Free Visual Token Pruning for Multimodal Large Language Models

・arXiv:2609.01004v1 Announce Type: cross Abstract: Despite their strong multimodal understanding ability, multimodal large language models (MLLMs) incur substantial computational overhead when processing long visual token sequences. ・To reduce inference costs, recent studies have explored visual token pruning through vision-centric or text-guided strategies. ・However, these methods often overlook high-norm outlier token
cs.LG updates on arXiv.org

Skill Reuse as Compression in Agentic RL

・arXiv:2605.31509v2 Announce Type: replace Abstract: Large language model agents trained with reinforcement learning (RL) often learn brittle, task-specific shortcuts. ・We hypothesize that agents generalize better when their successful trajectories are structurally compressible, decomposed into a small set of reusable abstract patterns. ・To formalize this, we introduce ReuseRL, which grounds agentic RL in the Minimum De
cs.LG updates on arXiv.org

SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers

・arXiv:2609.01343v1 Announce Type: new Abstract: Looped Transformers increase effective depth by iterating a shared block of layers, but most evaluations compare at fixed model size, conflating architectural advantage with extra FLOPs. ・We study looping on Mixture-of-Experts Transformers while closely matching per-token FLOPs, total non-embedding parameters, and KV cache. ・Through a series of ablations, we arrive at a r
Takara TLDR - Daily AI Papers

Socrates went Nuclear: Comparing Interaction Strategies for AI systems in a Learning Context using Brain Sensing

・Does unrestricted AI access bypass the cognitive effort required for learning, or does it streamline knowledge acquisition? ・This paper reports on a study where we compare three designs for user-AI interaction in a learning context: (1) an unrestricted conversational bot like ChatGPT, (2) a pedagogically constrained bot that guides through hints without giving final answers, which we refer to as the Socratic mode; and
cs.LG updates on arXiv.org

Soft-Argmax for the Projective Plane via the Veronese Embedding

・arXiv:2609.00521v1 Announce Type: cross Abstract: From horizon detection to fibre structures in X-ray imaging, many vision tasks recover lines via peak detection in Hough space $H=S^1\times\mathbb{R}$, the domain of orientation-offset pairs $(\theta,\rho)$. ・Differentiable pipelines extract coordinates via \emph{soft-argmax}, a probability-weighted average that is only meaningful in a globally linear space.
cs.LG updates on arXiv.org

Solving In-Table Prediction Problems by Deep Neural Networks with Performance Evaluation Using Synthetic Data

・arXiv:2609.01262v1 Announce Type: new Abstract: Tabular deep learning (TDL) leverages neural networks (NN) to extract patterns from tabular data. ・Traditional TDL methods follow a supervised learning paradigm, where a target feature is explicitly given. ・In this work, however, we explore a different approach by employing deep NNs to learn relationships among individual columns within a given table.
cs.LG updates on arXiv.org

SOVER: Formal Certification of Optimization Reformulations via LLM-Assisted SMT Verification

・arXiv:2609.00728v1 Announce Type: cross Abstract: Large Language Models (LLMs) have shown remarkable promise in translating and reformulating complex mathematical optimization problems across modeling languages. ・However, validating such transformations through empirical solver executions alone is unreliable, as solver outcomes may be affected by local minima, structural timeouts, numerical artifacts, and subtle seman
Takara TLDR - Daily AI Papers

Space Generative AI with Solar Energy Harvesting

・Satellites are emerging as promising platforms to extend generative \emph{artificial intelligence} (AI) services to remote areas lacking terrestrial infrastructure. ・However, deploying space generative AI is fundamentally constrained by the limited, time-varying onboard energy supplied by solar \emph{energy harvesting} (EH). ・This paper presents a framework for solar-powered space generative AI in which a satellite rec
cs.LG updates on arXiv.org

Spawn Freely, Act Sparingly: Progressive Risk Vesting for Recursive LLM-Agent Trees

・arXiv:2609.01035v1 Announce Type: cross Abstract: Recursive LLM agents can broaden their search by spawning specialists. ・Some branches later request tools that send data or deploy code. ・When should a branch receive authority to act?
cs.LG updates on arXiv.org

Steer, Don't Solve: Training Small Critic Models for Large Code Agents

・arXiv:2606.21811v2 Announce Type: replace-cross Abstract: Coding tasks are typically complicated and require multiple capabilities, ranging from high-level planning to low-level implementation. ・While coding agents are optimized for the joint capabilities, individual capabilities such as high-level planning may have different optima and remain a major bottleneck. ・To address this challenge, we train a separate critic m
cs.LG updates on arXiv.org

Stochastic complexity of vectors containing cluster structure

・arXiv:2609.00084v1 Announce Type: new Abstract: This paper studies the problem of computing the stochastic probability (shortest code length) of the encoded vectors containing cluster structure using Normalized Maximum Likelihood (NML) model. ・This is of great theoretical and practical importance in data clustering based on Minimum Description Length (MDL) principle, such as for estimating the best number of clusters
cs.LG updates on arXiv.org

Stress Testing Unlearning Algorithms

・arXiv:2608.22527v2 Announce Type: replace Abstract: Recently, machine unlearning, the removal of specific training data influence from a model, has gained increasing attention. ・In large language models (LLMs), unlearning is particularly challenging due to the ambiguity of inputs and outputs. ・Con- sequently, rigorous evaluation is critical for assessing both safety and utility, and for driving progress in unlearning m
cs.LG updates on arXiv.org

Subliminal Learning as Trait-Direction Drift: A Mechanism and Targeted Control under SFT Distillation

・arXiv:2609.01091v1 Announce Type: new Abstract: Beyond intended capabilities, model distillation can transfer hidden traits from a teacher. ・A teacher biased by a system prompt can generate semantically clean training data, such as numeric sequences, that still causes a downstream student to inherit the hidden preference, a phenomenon known as subliminal learning. ・Prior work has identified several parts of this proces
cs.LG updates on arXiv.org

Subspace Levenberg Marquardt Algorithms in Training Neural Networks

・arXiv:2609.00789v1 Announce Type: new Abstract: The Levenberg-Marquardt (LM) algorithm is a well-known second-order method for rapid convergence and strong robustness when training small- to medium-sized neural networks (NNs). ・However, its computational and memory costs increase significantly as the number of parameters in an NN grows. ・To address this limitation, subspace methods have been proposed, such as the Krylo
Takara TLDR - Daily AI Papers

Subspace Levenberg Marquardt Algorithms in Training Neural Networks

・The Levenberg-Marquardt (LM) algorithm is a well-known second-order method for rapid convergence and strong robustness when training small- to medium-sized neural networks (NNs). ・However, its computational and memory costs increase significantly as the number of parameters in an NN grows. ・To address this limitation, subspace methods have been proposed, such as the Krylov subspace LM (KSLM) and the hybrid subspace LM
cs.LG updates on arXiv.org

Superposed Latent Autoencoder

・arXiv:2609.01158v1 Announce Type: new Abstract: Autoencoders typically meet tight latent-memory budgets by making each latent representation smaller, sacrificing representational capacity. ・We ask a different question: can multiple wider latents be stored together instead? ・We introduce the Superposed Latent Autoencoder (SLAE), which preserves high-capacity latent representations while sharing storage through learned s
cs.LG updates on arXiv.org

SupraTok: Cross-Boundary Tokenization for Enhanced Language Model Performance

・arXiv:2508.11857v3 Announce Type: replace-cross Abstract: Tokenization remains a persistent bottleneck in language modeling, especially when vocabulary learning is limited by whitespace boundaries. ・We present SupraTok, a tokenizer that crosses whitespace boundaries using three modular components: optional entropy-based data curation, staged curriculum training with PMI-guided candidate search, and multilingual script
cs.LG updates on arXiv.org

Synthetic Worlds for Temporal Evaluation and Knowledge Updating in LLMs

・arXiv:2609.00184v1 Announce Type: cross Abstract: Large language models (LLMs) rely on static pretraining corpora, causing their knowledge to become outdated over time. ・Existing approaches for evaluating knowledge edits either suffer from rapid contamination or rely on counterfactual edits that conflict with rigid existing knowledge. ・In this work, we propose a synthetic, simulation-driven framework for studying knowl
MIT News - Artificial intelligence

System helps humans predict when self-driving cars will make mistakes

・A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
cs.LG updates on arXiv.org

TACS: Trajectory-Aware Candidate Selection for LLM Jailbreak Suffix Optimization

・arXiv:2608.29564v2 Announce Type: replace-cross Abstract: Gradient-based jailbreak suffix optimization methods typically update the suffix by retaining the candidate with the lowest current loss. ・We show that this seemingly natural design is fundamentally myopic: candidates that look better under the current-step proxy often fail to produce better jailbreak outcomes later in the search, revealing a form of selection-
cs.LG updates on arXiv.org

Task-Specific Prompt with Global Context for Multi-Task Graph Pre-Training

・arXiv:2609.00047v1 Announce Type: new Abstract: Graph prompt learning is an effective paradigm to adapt pre-trained graph models to downstream tasks in low-resource scenarios. ・However, existing multi-task graph pre-training frameworks generally use randomly initialized prompts, leading to poor alignment between the prompt space, pretext objectives and graph structural characteristics. ・This greatly weakens the task re
cs.LG updates on arXiv.org

Text Capability Loss in Vision-Language Adaptation: An Attention-Sink Diagnosis

・arXiv:2609.00746v1 Announce Type: new Abstract: Fine-tuning a pretrained LLM into a vision-language model (VLM) can erode the backbone's text capability, with the damage concentrated on tasks that require following exact output rules, such as instruction following, chain-of-thought reasoning graded on a strictly parsed final answer, and similar evaluations with strict graders. ・We trace this gap to attention-sink corr
cs.LG updates on arXiv.org

The Alexander-Hirschowitz theorem for neurovarieties

・arXiv:2511.19703v2 Announce Type: replace-cross Abstract: We study the dimension and identifiability of neurovarieties associated to polynomial neural networks. ・We give an independent geometric proof that the linear bounds $d_i\geq 2n_i-1$ on the activation degrees imply non defectiveness for any number of outputs, a dimension statement previously obtained from finite identifiability. ・The proof is based on a direct a
cs.LG updates on arXiv.org

The Constitutional Coverage Trilemma in AI Governance

・arXiv:2609.01275v1 Announce Type: new Abstract: Frontier AI systems function as \emph{constitutional institutions}: each deployed model encodes an implicit ranking among safety, helpfulness, honesty, autonomy, and equity. ・We ask whether the supply of frontier constitutional types covers human demand. ・Combining a paraphrase-controlled audit of the as-shipped default constitutions of $23$ frontier LLM archetypes with a
cs.LG updates on arXiv.org

The Curse of Multilinguality in Lexical Normalization

・arXiv:2609.00329v1 Announce Type: cross Abstract: Lexical normalization rewrites the noisy, non-standard words that fill user-generated text (tmrw, u, gr8) into their standard forms. ・Because labelled data is scarce for most languages, a popular shortcut is to train a single model on many languages at once. ・We ask a simple question: how many languages should such a model be trained on?
cs.LG updates on arXiv.org

The Multiple Timescales of Gradient Descent on the Edge of Stability: A Perturbative Derivation of the Central Flow

・arXiv:2609.01034v1 Announce Type: new Abstract: The central flow of Cohen et al. ・(2025) is an empirically accurate continuous-time model of gradient descent at the edge of stability in deep learning, However, its derivation is heuristic. ・We propose a perturbative regime in which the central flow is the limit of gradient descent: we assume that the loss decomposes as $f = g + \varepsilon h$; in the limit $\varepsilon
cs.LG updates on arXiv.org

The Structure of Quantization Damage in LLMs: Why the Next Bit Should Be Spent Globally

・arXiv:2609.01587v1 Announce Type: new Abstract: Post-training quantization (PTQ) is widely used to reduce the cost of serving large language models (LLMs), but its accuracy cost is uneven and is often tuned per model. ・We study where quantization damage occurs and how to allocate a small additional precision budget. ・Using causal mixed-precision intervention as ground truth (raise each layer to 8-bit in turn and measur
cs.LG updates on arXiv.org

The Topological Trouble With Transformers

・arXiv:2604.17121v5 Announce Type: replace Abstract: Transformers encode structure in sequences via an expanding contextual history. ・However, their purely feedforward architecture fundamentally limits dynamic state tracking. ・State tracking -- the iterative updating of latent variables reflecting an evolving environment -- involves inherently sequential dependencies that feedforward networks struggle to maintain.
cs.LG updates on arXiv.org

The Visual Insensitivity Gap: Diagnosing When Vision-Language Models Fail to Use Visual Evidence

・arXiv:2609.00868v1 Announce Type: cross Abstract: Vision-language models are evaluated by aggregate accuracy on multimodal benchmarks, a practice that implicitly assumes the model uses its visual input. ・We show this assumption fails on 40%--97% of samples across six VLMs and three perceptual benchmarks: blurring the question-relevant visual region leaves the next-token distribution nearly unchanged. ・We name this phen
The latest research from Google

TimesFM-3: A zero-shot foundation model for multivariate forecasting

TimesFM-3: A zero-shot foundation model for multivariate forecasting
cs.LG updates on arXiv.org

Topic Matching in the Wild: Benchmark and Lessons from Real-World ASR Transcripts

・arXiv:2609.00330v1 Announce Type: cross Abstract: In contact centers, real-time agent-assist tools determine, for each of many predefined topics, whether a live customer utterance is relevant and display a coaching card to the agent when it is. ・The input is noisy and challenging: ASR(Automatic Speech Recognition) transcripts of spontaneous phone conversations, which can be unclear, repetitive, and mostly lack punctua
cs.LG updates on arXiv.org

TopoCompress: Long Context Compression via Graph-Wired Semantic Trajectories

・arXiv:2608.30811v2 Announce Type: replace-cross Abstract: Long-context compression is essential for reducing the cost and latency of large language model inference. ・However, existing methods can fragment important evidence, require additional training or alignment, and often depend on the target model for effective compression. ・We introduce TopoCompress, a training-free and model-agnostic framework that compresses lo
cs.LG updates on arXiv.org

Topological Steering

・arXiv:2609.00597v1 Announce Type: new Abstract: With the rapid rise of large language models (LLMs), controlling undesirable model behaviors has become increasingly important. ・Existing behavioral control methods typically intervene directly in activation or feature space, but such approaches can be sensitive to outliers, distributional shifts, noise, and other local perturbations. ・Motivated by Topological Data Analys
cs.LG updates on arXiv.org

ToSCA: Leveraging Hierarchical Reinforcement Learning on Temporal and Strategic Abstractions of Conversational Agents

・arXiv:2608.21969v2 Announce Type: replace-cross Abstract: Humans naturally exhibit multiple forms of abstraction in reasoning and interaction, including temporal abstraction across decision timescales and strategic abstraction over communicative intents. ・Inspired by these complementary abstractions, we propose a two-level hierarchical reinforcement learning (HRL) framework for conversational agents that bridges the g
cs.LG updates on arXiv.org

Towards Provable and Scalable Training of Quantized Neural Networks with Ising Optimization

・arXiv:2506.18240v5 Announce Type: replace Abstract: Training quantized neural networks remains fundamentally challenging due to non-convex loss landscapes and discrete parameter spaces. ・We introduce an exact Quadratic Constrained Binary Optimization (QCBO) framework with provable guarantees. ・We first characterize the stratified topology of network zero-loss level sets: generic interior strata are smooth, yet globally
cs.LG updates on arXiv.org

Towards unsupervised representation learning for quantum data: quantum models with inference and generation

・arXiv:2609.00372v1 Announce Type: cross Abstract: With quantum sensors, simulators and networks emerging, a future of quantum technology may produce quantum states as data---that is, coherently rather than as classical measurement records---thus motivating the study of suitable quantum generalisations of modern machine learning, including the automated, unsupervised extraction of useful representations. ・Two ingredien
cs.LG updates on arXiv.org

Training-Free Policy Violation Detection via Activation-Space Whitening in LLMs

・arXiv:2512.03994v4 Announce Type: replace Abstract: As organizations increasingly deploy LLMs in sensitive domains such as legal, financial, and medical settings, ensuring alignment with internal organizational policies has become a priority. ・Existing content moderation frameworks remain largely confined to the safety domain and lack the robustness to capture nuanced organizational policies. ・LLM-as-a-judge and fine-t
cs.LG updates on arXiv.org

TRIAGE: Three-level Routing and Intelligent Agent Guidance for Efficient Execution

・arXiv:2609.01428v1 Announce Type: new Abstract: Large Language Model (LLM) agents based on the ReAct paradigm have demonstrated remarkable capabilities in tool use and task execution. ・However, ReAct suffers from a fundamental efficiency problem: every query triggers a complete reasoning loop from scratch, and similar queries repeat identical steps without leveraging historical experience. ・We propose TRIAGE,a three-le
Takara TLDR - Daily AI Papers

Trust Your Guide Only When Certain: Uncertainty-Aware Sparse Alignment at Inference Time

・A prominent paradigm in inference-time alignment employs lightweight supervisors to steer Large Language Models (LLMs). ・Through empirical analysis, we identify a structural mismatch in this paradigm: weak supervisors exhibit pervasive high entropy across the vast majority of tokens, yet prevailing dense intervention approaches mandate supervision at every decoding step. ・This leads to frequent low-confidence intervent
cs.LG updates on arXiv.org

TRUST: Threshold-Recalibrated Uncertainty-Safe Training for Certified Dismissal in Breast Cancer Screening

・arXiv:2609.00300v1 Announce Type: cross Abstract: Reducing the review of clearly cancer-negative screening mammograms could lower radiologist workload without compromising cancer detection. ・We propose a closed-loop threshold-aware training strategy in which the dismissal threshold is recalculated during training and used to penalize cancer-positive images that approach the dismissal region. ・We evaluated the method on
Takara TLDR - Daily AI Papers

TUTTI: Toward generalizable audio-to-score transcription via fully synthesized data

・Generalizable Audio-to-Score (A2S) transcription is fundamentally constrained by the severe scarcity of high-quality, real-world paired data. ・Relying solely on existing human-annotated datasets often restricts the generalization of A2S models, limiting their efficacy primarily to single-instrumentation domains. ・To break this dependency on scarce real-world data, we introduce TUTTI (Transformer for Unified audio-To-sc
cs.LG updates on arXiv.org

UI-Venus-2 Technical Report

・arXiv:2609.00028v1 Announce Type: cross Abstract: Multimodal GUI agents have emerged as a promising paradigm for digital task automation, yet transitioning from benchmark-oriented models to dependable real-world applications remains challenging due to limited environment coverage, brittle task construction, and unreliable reward verification. ・In this work, we present UI-Venus-2, a general-purpose foundation GUI agent
cs.LG updates on arXiv.org

Uniform a priori bounds and error analysis for the Adam stochastic gradient descent optimization method

・arXiv:2603.18899v2 Announce Type: replace Abstract: The adaptive moment estimation (Adam) optimizer proposed by Kingma & Ba (2014) is presumably the most popular stochastic gradient descent (SGD) optimization method for the training of deep neural networks (DNNs) in artificial intelligence (AI) systems. ・Despite its groundbreaking success in the training of AI systems, it still remains an open research problem to prov
cs.LG updates on arXiv.org

Unsupervised Partner Design Enables Robust Ad-hoc Teamwork

・arXiv:2508.06336v3 Announce Type: replace Abstract: We introduce Unsupervised Partner Design (UPD), a population-free multi-agent reinforcement learning method for robust ad-hoc teamwork. ・UPD generates training partners on-the-fly and selects them adaptively based on a learnability criterion, removing the need for pre-trained partner populations or manual parameter tuning. ・We show that this simple mechanism enables e
cs.LG updates on arXiv.org

ValueGraph: Value-Signal Guided Graph Pre-training for Contextualized User Representation

・arXiv:2609.00057v1 Announce Type: cross Abstract: Value signals are aggregated user-level moral representations that capture users' inferred value-related tendencies from their online discourse. ・User behavior on social media is shaped not only by what users say or whom they interact with, but also by the value signal through which they express attitudes. ・Existing user representation methods largely miss this value-re
cs.LG updates on arXiv.org

Variable Selection for Feature-Based Newsvendor

・arXiv:2609.01544v1 Announce Type: cross Abstract: Feature-based newsvendor models use observable covariates to tailor inventory decisions, aiming to balance holding and shortage costs under demand uncertainty. ・However, high-dimensional feature sets often hinder interpretability and inflate data collection and implementation costs. ・This paper studies variable selection for the feature-based newsvendor problem under a
cs.LG updates on arXiv.org

Variance-Adaptive Muon: Pre-Orthogonalization Variance Modulation for Efficient Language Model Pretraining

・arXiv:2601.14603v2 Announce Type: replace Abstract: Optimizer design plays a central role in efficient language model pretraining, directly affecting optimization dynamics, convergence speed, and compute cost under fixed training budgets. ・Muon has emerged as a strong optimizer by orthogonalizing momentum updates, yielding a matrix-valued analogue of sign-based normalization. ・However, unlike Adam-style methods, Muon d
cs.LG updates on arXiv.org

VATO: A Vortex-Force-Aware Transformer Operator for Unsteady Separated Aerofoil Flows

・arXiv:2609.00507v1 Announce Type: new Abstract: Accurate prediction of unsteady separated flows is challenging because the aerodynamic loads depend on nonlinear separation and vortex-shedding dynamics. ・Although high-fidelity CFD resolves these mechanisms, its cost limits repeated use in design and control. ・Standard field-level surrogate training, however, does not distinguish the flow regions that contribute most str
cs.LG updates on arXiv.org

Verdict Instability of OOD Scores under Reference Resampling

・arXiv:2609.00691v1 Announce Type: new Abstract: Post-hoc out-of-distribution detectors are fitted on a finite reference set, so every score they produce is an estimate. ・If we had chosen a different set, some verdicts would have moved. ・We measure that movement by resampling the reference set and recording the bootstrap standard deviation of the score, which we call verdict instability.
cs.LG updates on arXiv.org

Vision-Language-Guided Pseudo-Labels for Unsupervised Domain Adaptation in Semantic Segmentation for Waste Sorting

・arXiv:2609.00898v1 Announce Type: cross Abstract: Obtaining labeled data for semantic segmentation in applied settings (e.g., autonomous driving, industrial waste sorting) is expensive and often infeasible at scale. ・We present a cross-modal pseudo-labeling pipeline that enables unsupervised domain adaptation without any target-domain annotations. ・The pipeline is built on two core foundation models: SAM generates clas
cs.LG updates on arXiv.org

ViTAMINS: An Empirical Study of Training Self-Supervised Vision Transformers with Synthetic Hard Negatives

・arXiv:2609.01041v1 Announce Type: cross Abstract: We introduce ViTAMINS, a method that integrates synthetic hard negatives into unsupervised vision transformer pretraining to improve representation quality. ・Our approach is thoroughly benchmarked on ImageNet and transfer learning, image retrieval, copy detection, and image, video segmentation tasks. ・Notably, our proposed negatives give rise to emergent properties, whe
cs.LG updates on arXiv.org

VoiceLongMemEval: Do Assistants Remember How You Sounded?

・arXiv:2609.00570v1 Announce Type: cross Abstract: With the growing scale of multi-agent architectures and large language models, deployed AI assistants are increasingly tasked with reasoning over long, continuous, multi-session conversation histories. ・Current benchmarks evaluate this dialogue history as information retrieval over long horizon, temporal reasoning, or knowledge updates, while crucially ignoring the fun
MIT News - Artificial intelligence

Walter Torous named executive director of MIT Center for Real Estate

・The senior lecturer, already director of the degree program, will now oversee all aspects of the center’s activities and operations.
Takara TLDR - Daily AI Papers

Wave Function Backpropagation with Explicit Temporal-Interval Dynamics

・Conventional neural networks learn predominantly through affine transformations followed by nonlinear activations, while elapsed time is often treated as an auxiliary feature or assumed to be uniformly sampled. ・This paper introduces Wave Function Backpropagation (WFB), a wave-parameterized learning formulation in which neural responses are represented by learnable amplitude, wavenumber, angular frequency, and phase.
cs.LG updates on arXiv.org

Web Price Extraction: State of the Art and an Adaptive Browserless Implementation

・arXiv:2609.01030v1 Announce Type: cross Abstract: Price extraction from websites is a key task for market monitoring, price comparison, and business analytics in e-commerce. ・Existing approaches can be broadly divided into four groups, and understanding their trade-offs in accuracy and scalability is essential for selecting suitable extraction strategies. ・Classical methods rely on manually written wrappers and rule in
cs.LG updates on arXiv.org

WHALE: A Simple Recipe for Joint Harness-Weight Optimization

・arXiv:2609.00196v1 Announce Type: new Abstract: Agent performance depends jointly on the model parameters and the executable harness code that manages context and control flow. ・Optimizing either component in isolation can leave the system bottlenecked by its frozen counterpart: weight updates can change which harness is effective, while harness updates can change which model capabilities are exposed. ・Existing joint-a
cs.LG updates on arXiv.org

What Do Students Learn? A Feature-Level Analysis of Dark Knowledge

・arXiv:2606.03052v2 Announce Type: replace Abstract: Knowledge Distillation (KD) is a powerful tool for model compression, yet the precise mechanisms by which student models acquire feature representations remain underexplored. ・In this work, we analyze student feature learning using the Interaction Tensor framework. ・Our analysis reveals that effective KD acts as a regularizer that prunes low-frequency, sample-specific
cs.LG updates on arXiv.org

What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal

・arXiv:2604.08524v2 Announce Type: replace Abstract: Applying steering vectors to large language models (LLMs) is an efficient and effective model alignment technique, but we lack an interpretable explanation for how it works--specifically, what internal mechanisms steering vectors affect and how this results in different model outputs. ・To investigate the causal mechanisms underlying the effectiveness of steering vect
cs.LG updates on arXiv.org

When Does Online Adaptation Pay on the Edge? A Leakage-Free Evaluation of Warmup, Learning-Rate Selection, and Resource Trade-offs for Time-Series Forecasting

・arXiv:2609.01126v1 Announce Type: new Abstract: Online adaptation can help edge time-series forecasting under distribution drift, but its measured benefit is sensitive to evaluation choices. ・We study six public multivariate streams, including building-sensor and smart-meter data, under a leakage-free streaming protocol. ・We identify two additional sources of comparison bias.
cs.LG updates on arXiv.org

When Metropolis and Hastings Meet Bradley and Terry: Exact MCMC From Preference Voting

・arXiv:2609.00905v1 Announce Type: new Abstract: Sampling from distributions conditioned on desired semantic properties is an emerging challenge in modern generative modeling. ・Metropolis-Hastings (MH) provides a principled route to conditional sampling, but requires access to exact pointwise target-density evaluations, which are not available in generative settings. ・Meanwhile, pairwise comparisons by humans or model "
cs.LG updates on arXiv.org

When Prediction Error Is Not Enough: Evaluating Nuisance-Function Prediction for Causal Estimation

・arXiv:2609.00071v1 Announce Type: cross Abstract: Prediction error is widely used to evaluate nuisance-function estimators in causal inference, but its relationship with causal estimator performance may differ across performance measures. ・We studied this question in a partially linear model using Monte Carlo simulations. ・We compared ordinary least squares (OLS), generalized additive models (GAMs), XGBoost, and Double
cs.LG updates on arXiv.org

When the Strongest Teacher Is Not the Best Teacher: Student-Centric Answer Selection

・arXiv:2605.26872v5 Announce Type: replace Abstract: LLM training increasingly relies on teacher-generated supervision, from synthetic responses to reasoning traces and tool-use demonstrations. ・Current practice often chooses the highest-performing teacher to generate student training data, implicitly treating teacher test performance as a proxy for teaching quality. ・We show that this assumption can fail: even when mul
cs.LG updates on arXiv.org

Where the Verifier Fails: A Category-Level Audit of Reward Signals in RLVR

・arXiv:2609.01354v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) and standard benchmark evaluation both rely on an automatic verifier that turns a free text answer into a binary reward. ・Prior work reports that one evaluation harness accepts only about 94% of its own ground truth answers, blaming LaTeX parsing. ・That is an aggregate: it does not say which answer forms consume the
cs.LG updates on arXiv.org

Why Do Reasoning Models Lose Coverage? The Role of Data and Forks in the Road

・arXiv:2605.17026v2 Announce Type: replace Abstract: Recent progress in large language models has led to the emergence of reasoning models, which have shown strong performance on complex tasks through specialized fine-tuning procedures. ・While these methods reliably improve pass@1 accuracy, prior works have observed that they show a coverage shrinkage behavior, where pass@k degrades relative to the base model.
cs.LG updates on arXiv.org

Why Fine-Tuning Encourages Hallucinations and How to Fix It

・arXiv:2604.15574v2 Announce Type: replace-cross Abstract: Large language models are prone to hallucinating factually incorrect statements. ・A key source of these errors is exposure to new factual information through supervised fine-tuning (SFT), which can increase hallucinations w.r.t.~knowledge acquired during pre-training. ・Since these errors arise as a by-product of knowledge degradation, we explore whether establis
cs.LG updates on arXiv.org

Why Multi-Layer Message Passing Works: Completeness Theory for Graph Neural Network Interatomic Potentials

・arXiv:2609.00528v1 Announce Type: new Abstract: We prove that the Hypergraph Neural Network, an invariant architecture with 3-body message passing, is a universal approximator for potential energy surfaces. ・Our main contribution is a multi-layer completeness theory. ・We show that $L$ layers of message passing on sparse, cutoff-based graphs achieve the same representational power as having access to the full $L$-hop ne
cs.LG updates on arXiv.org

Workload Identification with Physical Side Channels for AI Governance

・arXiv:2609.00309v1 Announce Type: cross Abstract: AI compute verification is one of the first tangible and tractable points for international policy aimed at AI governance. ・Determining whether frontier labs, or any operator, comply with agreements requires the regulating authority to discern how their compute is used. ・The elementary building block of AI compute is the GPU, and any activity it executes leaves a physic