New AI & machine learning papers
434 new AI & machine learning papers appeared in the last 3 days. Pipette read every one, and these are the 60 with the broadest interest, a real step forward and careful claims. Each shows the sentence of its abstract that states the main result, exactly as its authors wrote it.
Language models 56Agents & reasoning 41Computer vision 49Image, video & audio generation 17Reinforcement learning 23Safety, alignment & fairness 26Learning theory & optimization 20Efficiency & hardware 19AI for science & medicine 48Speech & audio 48Search & recommendation 19Benchmarks & evaluation 36Other machine learning 32
The best of the last 3 days
Double descent without digital computation
Here we demonstrate double descent in a decentralized analog network of self-adjusting resistive elements.
Peer-reviewed journalClaims a big stepPUBG Ally: A Conversational Embodied Agent as an AI Teammate
We introduce PUBG Ally, an embodied agent for PUBG: BATTLEGROUNDS that can reason, act autonomously, and play alongside players as a voice-enabled teammate.
PreprintReal-world useSpot, Separate, and Enhance: Fully Generative Approach for Audio Mixing
We introduce Spot, Separate, and Enhance (SSE), the first multimodal, user-guided generative model for audio remixing and enhancement.
PreprintClaims a big stepReal-world useVietPrism: A large-scale Vietnamese speech and deepfake corpus with diverse dialects and code-switching
We introduce VietPrism, an open, multi-domain corpus that brings these dimensions together at scale: 993.4 hours and 403,941 bona fide utterances from 1,262 verified speakers across 8,388 real-world videos.
PreprintClaims a big stepReal-world useMachine learning of honey bee olfactory behavior identifies repellent odorants in free-flying bees in the field
Additional testing of the top seven candidates using freely foraging honey bees in a field assay confirmed strong repellency, thus predicting a high probability to repel foraging bees from pesticide-treated crops.
Peer-reviewed journalReal-world useMicrorings as programmable temporal kernels enabling photonic AI beyond 100 Gbaud
Here, we redefine MRRs as programmable temporal convolution kernels by exploiting their impulse responses, enabling computation beyond the resonance linewidth.
Peer-reviewed journalBold claims, read criticallyClaims a big stepReal-world useOn a Separate Note: Robust Score-Informed Note Separation with a Two-Stream TFC-TDF U-Net and Adaptive Set Ownership
We present, to our knowledge, the first deep learning approach to score-informed note separation, NoteSep.
PreprintClaims a big stepLow-altitude aircraft will reshape noise exposure across global cities
These findings show that low-altitude aircraft can reshape noise exposure across global cities, requiring route assessment to distinguish newly exposed from already exposed areas and to account for three-dimensional exposure.
PreprintReal-world useYODAS v3: Over 1 Million Hours of High-Bandwidth, Stereophonic, Multilingual Speech
YODAS v3 is not only the largest open speech dataset to date, but also the first truly large-scale speech corpus with high-fidelity stereo audio.
PreprintBold claims, read criticallyClaims a big stepCode availableCinematicVQA: Benchmarking Film-Grammar Reasoning in Large Vision-Language Models
To address this, we introduce CinematicVQA, the first-of-its-kind benchmark for cinematic video understanding that goes beyond technique recognition to evaluate film-grammar reasoning, utilizing our introduced Cinematic Scene Graph (CSG), a structured representation that links filming techniques to their perceptual effects and narrative functions.
PreprintTransformers as Cross-Task Learners: Shared Structure Drives Sample Efficiency in In-Context Learning
To the best of our knowledge, this is the first work to quantify cross-task complexity for general nonlinear task families and explicitly construct a Transformer that exploits their low-dimensional structure to perform ICL.
PreprintClaims a big stepSynthetic Hospital: An Open, Verifiable, Physician-Validated Longitudinal EHR Benchmark
We introduce Synthetic Hospital, an open, fully synthetic, fact-grounded longitudinal EHR benchmark that resolves the open sharing and verifiable ground truth barriers.
PreprintEditVoice: Variable-Length Non-Autoregressive Zero-Shot TTS and Speech Editing with Edit Flows
We introduce EditVoice, to our knowledge the first variable-length NAR zero-shot TTS model, which uses Edit Flows to jointly update speech content and sequence length through insertions, deletions, and substitutions.
PreprintClaims a big stepReal-world useMultiplexed deep visual proteomics resolves spatial heterogeneity and rare endocrine states in human pancreatic islets
Applied to human pancreatic islets, mxDVP segments over 860,000 cells and resolves twelve endocrine subtypes, including rare polyhormonal and intermediate-state populations that exhibit spatial organization patterns, co-expression of INSM1 and SCG3, and hybrid α/β/δ signatures.
Peer-reviewed journalASR ensembling for phoneme intelligibility evaluation of speech anonymizers
Our results show that simple hard-voting ASR metric reaches correlations above 0.9 with human ratings when aggregated by feature, test-type, or condition, provided that multiple ASR models are combined; evaluating stimuli with and without a carrier sentence further improves the correlation at the stimulus level.
PreprintTrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations
We introduce TrackEverything, a 3D point tracker that breaks this trade-off by representing videos as persistent 3D scene tracks in world coordinates.
PreprintBold claims, read criticallyClaims a big stepPHOSA: Photorealistic 3D Sign Avatar Modeling and Benchmark
Building on MVSign, we propose a decoupled sign avatar representation that isolates body, head, and hand components to capture complex articulations, together with a motion-aware sampling strategy to handle motion blur and balance gesture diversity.
PreprintReal-world useComplexSync: High-Fidelity and Real-Time Lip Sync in Complex Scenarios
We present ComplexSync, a unified diffusion-based framework that enables real-time, high-fidelity lip sync under complex conditions.
PreprintReal-world useLabFactory: Building and Evaluating Executable AI Labs
We present LabFactory, a framework in which an AI builder turns a scientific brief into an executable AI lab: a task-specific solver that integrates models, knowledge resources, tools, and a controller behind a fixed interface.
PreprintReal-world useDeep learning extracts MoA-specific signatures from high-throughput images of chemically and genetically perturbed Corynebacteria
Our model robustly classifies MoAs of established antibiotics and recognizes the MoA of previously unseen antibiotics.
Peer-reviewed journalReal-world useThe Entropy Triangle Method (ETM): A novel framework for the prevention of cardiac arrhythmia with a review of more than 10,000 patients
In this article, we introduce two firsts in machine learning and medicine that can predict non-sinus rhythm with over 85% accuracy.
PreprintBold claims, read criticallyClaims a big stepReal-world useOn Growth and Form, and Function: Reusable Regulatory Handles Control Phenotypic Variation
Together, our results provide a computational realization of D'Arcy Thompson's remarkable grid transformations in a 2D NCA---a minimal cybernetic tissue in which variations of fully grown emoji phenotypes can be encoded, combined, and controlled through low-dimensional directions in regulatory weight space.
PreprintBold claims, read criticallyA Living Benchmark for Information Retrieval from Electronic Health Records
We present a scalable framework that automatically generates question--answer pairs from longitudinal EHR notes.
PreprintReal-world useWhat, When, and How: Audio Description as Constrained Global Optimization
When evaluated on REFRAMED, a benchmark for realistic AD of movies, our approach makes better decisions than prompted LLMs about what to describe and when to describe it, establishing a new SOTA on narrative QA and temporally grounded metrics.
PreprintReal-world usePath-specific harm decomposition: A partial identification framework
As a remedy, we develop a novel partial identification framework for direct and indirect FNA.
PreprintThe Last Human Gate: Forward Deployed Engineering for Governance Automation
These results support the technical feasibility of replacing human execution of specified governance-review tasks with agents and software.
PreprintReal-world useCode availablePersonalized single-cell transcriptomics reveals molecular diversity in Alzheimer’s disease
Personalized functional genomics atlas for Alzheimer’s disease that uses knowledge-guided graph neural networks to analyze donor-level functional genomics, identify disease subpopulations and trajectories, and link genetic variants to gene regulation.
Peer-reviewed journalNeural Transport Nested Sampling
We develop a novel sampling algorithm, Neural Transport Nested Sampling (NTNS), which combines the classical strengths of nested sampling with modern neural flow-based methods.
PreprintClaims a big stepDiscovering how ice crystals grow using neural ODEs and symbolic regression
By optimizing against mass growth time series from 290 ice crystals grown in a levitation diffusion chamber, we identify a modified capacitance growth model that more accurately captures observed early-stage growth.
Peer-reviewed journalDeep learning of longitudinal visual fields predicts glaucoma progression rate and identifies fast progressors
On the open-access University of Washington Humphrey Visual Field dataset (4,276 patient-eyes), GLAM achieved an MD-rate mean absolute error of 0.139 dB yr (; 73.5% reduction over a ridge baseline) and an AUC of 0.990 for fast-progressor detection.
PreprintReal-world useImplementation of offset corrected AGAD algorithm on 130-nm CMOS technology–based RRAM array for analog neural network training
We present the first implementation of the analog gradient accumulation with dynamic reference (AGAD), reported as the most advanced and highest-performing version of the TT (Tiki-Taka) algorithm, on an HfO 2 -based resistive random-access memory (RRAM) array for analog neural network training.
Peer-reviewed journalBold claims, read criticallyClaims a big stepBeyond Compression: Training Latent Representations for Stable Long-Horizon Rollout in Neural Surrogate Solvers
Collectively, these interventions reduce long-rollout error by approximately 40% and match or exceed the accuracy of full-resolution models on two physics benchmarks, while requiring 2 orders of magnitude fewer floating point operations and half the GPU memory.
PreprintDAMSEP: Distance-Aware Monaural Source Separation using Multi-RIR Estimation
To address this limitation, we propose Distance-Aware Monaural Source Separation using Multi-RIR Estimation (DAMSEP), the first end-to-end framework that is jointly trained for source separation and multi-source RIR estimation from a single-microphone mixture.
PreprintCode availableScreen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale
We present a hypothesis-driven simulation workflow for screening candidate CX agents before deployment.
PreprintReal-world useTimeBraid: Unifying Time Series and Language for Understanding and Forecasting
We present TimeBraid, a series of unified time-series and language models that align pretrained language models and pretrained time-series foundation models through interleaved global residual attention layers.
PreprintWhere Does Exactly-Once Live? Model, Harness, and Tool-Contract Effects on Duplicate Side Effects in LLM Agents
Across 25,930 episodes spanning nine recent models, three production agent harnesses, two contract variants and fifteen recovery conditions, the answer depends on the fault.
PreprintReal-world useLearning to Discover Interesting Mathematics
Optimizing for our metric creates a model capable of producing more interesting theorems, while also reducing substantial or full overlap with Mathlib from 91.9% to 30.6%, showcasing the creation of more out-of-distribution math.
PreprintBold claims, read criticallySelf-Play Pretraining with Zero Data
We introduce Self-Play Pretraining with Zero Data, an initial proof-of-concept towards realizing this vision.
PreprintShadow Reduction in Ultrasound Imaging Using Differentiable Simulation and Radiance Field Decomposition
We present RFlash, a physics-informed post-processing method that decomposes beamformed ultrasound images into explicit attenuation and scatter-intensity maps using a differentiable radiance-field formulation of image formation.
PreprintReal-world useReachability-Based Formal Verification of Graph Neural Networks with Node and Edge Features
Our results show that GNNV provides tighter robustness guarantees than CORA on graph classification models with ReLU-based activations and, for the first time, delivers edge-aware robustness guarantees for GINE-based PF and OPF models under joint node and edge perturbations.
Preprint with a published versionReal-world useUpTCR: a unified progressive knowledge transfer foundation model for robust T-cell receptor-antigen binding recognition
Here, we present UpTCR, a unified progressive knowledge-transfer foundation model that learns from incomplete data to predict TCR-antigen-HLA binding.
Peer-reviewed journalReal-world useDrGait: Biomechanically Grounded Visual Reasoning for Interpretable Clinical Gait Analysis
By anchoring VLM's reasoning in verifiable geometric and temporal measurements, DrGait reduces hallucinations, achieving competitive diagnostic accuracy while generating transparent and audit-ready clinical reports.
PreprintReal-world useAugur: A Synthetic Decision Lab for Rehearsing Reactions to Product and Policy Changes
Our central finding is methodological and negative: most of the measured gap between frontier cloud models and open-weight models we fine-tune and serve offline is attributable to an under-specified evaluation, not a difference in capability. We show this three ways.
PreprintAgentic Detection of Online Conspiracies
Evaluating our framework on a manually-annotated adversarial dataset, we find that context-aware workflows consistently outperform text-only classification and that the agentic framework performs significantly better than other frameworks and settings, including a non-agentic model exposed to the same contexts available to the agent.
PreprintQwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents
Qwen-Planner-Agent achieves the best overall performance among all evaluated models and systems on MobilePA-Bench, improving over its base model across tool use, memory, skills, and sub-agent coordination.
PreprintReal-world useAICO: Feature significance tests for supervised learning
We introduce AICO (Add-In COvariates), a broadly applicable framework that turns model interpretability into an efficient statistical exercise.
Peer-reviewed journalBold claims, read criticallyReal-world useWildHSR: Metric Feed-Forward 4D People-Scene Reconstruction from a 3D Foundation Model
WildHSR combines both readouts to reconstruct metric cameras, scene and people from monocular video.
PreprintScalarLens: Numerical Embeddings with Stable Coordinates and Contextual Responses for CTR Prediction
We introduce ScalarLens, a numerical embedding that preserves what a value is while adapting how it should be interpreted.
PreprintASIRF: An Agentic Framework for Context-Dependent Sensitive Information Redaction
With only a few dozen expert-authored definitions per domain and no training data, ASIRF's recall exceeds OPF's in 68 of 80 model-domain combinations (85 percent), by at least one of the two architectures, with shortfalls confined mostly to OPF's training-distribution domains.
PreprintReal-world useJust Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures
A single generic question reaches a median AUROC of 0.886 zero-shot and beats supervised baselines on most benchmarks.
PreprintReal-world useCode availableBeyond Spatial Benchmarks: From Spatial Reasoning to Navigation
Our analysis reveals a gap between benchmark-oriented spatial specialization and navigation performance, and shows how aligning spatial supervision with navigation goals, phases, and decision learning improves navigation.
PreprintCode availableTemporal Taxation Compounds Under Post-Training Compression of Whisper Models
We cast the temporal-taxation construct of Choi and Choi (2025) as a quantitative metric, and show that single-snapshot fairness audits on full-precision models do not capture the deployment-time burden that compression places on already-marginalized speakers.
PreprintReal-world useMonitoring Urban Traffic Dynamics at Fine Spatiotemporal Resolution Using Distributed Acoustic Sensing and Deep Learning
Using a DAS deployment along a roadway network in the City of College Station, Texas, USA, this study develops a deep learning-empowered analytical framework that converts raw ground vibration waveforms into spatiotemporal representations, detects vehicle trajectory, and infers traffic states from aggregated traffic volume and speed.
PreprintReal-world useBeyond Surface Style: Aligning Multi-Turn User Simulators with Behavioral Consistency
We propose TRACER, a multi-turn user simulator that explicitly models users' evolving intent and learns to align simulated behavior with real interaction trajectories.
PreprintReal-world useEmpath: Tracing Multi-Level Emotion Dynamics in Crisis Counseling Dialogues
We introduce EMPATH, a framework for understanding affective dynamics in mental health dialogues across three granularities: turn-level labels, transition probabilities, and global conversation archetypes.
PreprintIndustrial Anomaly Detection via Defect-Grounded Reasoning in Visual Latent Space
To address these, we propose Anomaly-LR, a defect-grounded latent reasoning framework that first forms a global understanding of the input and then progressively refines anomaly-relevant representations directly in the visual latent space.
PreprintReal-world useCode availableEnigmaForge: The Question Is Hidden in the Story
Intuition reshuffles the leaderboard: a 22x spread where fact recovery spans 1.6x, the second-best fact-recoverer ranks fourteenth, one model is indifferent to being told the question, and another is significantly better without it.
PreprintLow-Cost Assays for Measuring Model Behavior Across Vendors and Releases
To address these constraints, we present a simple, cheap, scalable, and replicable model for studying model behavior.
PreprintCode availableBeneath the Scores: Rethinking Hallucination Evaluation for Video Understanding Models
Across 60,008 runs on three video-agent architectures, we find that grounding is the dominant source of downstream error, with roughly four times the causal impact of corrupting visual observations.
PreprintWhere Hallucinations Live: A Cross-Architecture Circuit in VQ-Tokenized Vision-Language Models
Using activation patching across twenty-five models spanning eight LLM families, we identify an early-layer () attention routing circuit shared across VQ-tokenized VLMs and propose a three-gate diagnostic that distinguishes the models carrying it from those that do not.
Preprint