Papers nuevos de IA y aprendizaje automático
En los últimos 3 días aparecieron 434 papers nuevos de IA y aprendizaje automático. Pipette los leyó todos y estos son los 60 con más interés, avance real y afirmaciones prudentes. Cada uno muestra la oración de su resumen que dice el resultado principal, tal como la escribieron sus autores.
Modelos de lenguaje 56Agentes y razonamiento 41Visión por computadora 49Generación de imagen, video y audio 17Aprendizaje por refuerzo 23Seguridad, alineación y equidad 26Teoría del aprendizaje y optimización 20Eficiencia y hardware 19IA para ciencia y medicina 48Voz y audio 48Búsqueda y recomendación 19Benchmarks y evaluación 36Otro aprendizaje automático 32
Lo mejor de los últimos 3 días
Double descent without digital computation
Here we demonstrate double descent in a decentralized analog network of self-adjusting resistive elements.
Revista con revisión por paresDice ser un gran avancePUBG Ally: A Conversational Embodied Agent as an AI Teammate
We introduce PUBG Ally, an embodied agent for PUBG: BATTLEGROUNDS that can reason, act autonomously, and play alongside players as a voice-enabled teammate.
PreprintUso en el mundo realSpot, Separate, and Enhance: Fully Generative Approach for Audio Mixing
We introduce Spot, Separate, and Enhance (SSE), the first multimodal, user-guided generative model for audio remixing and enhancement.
PreprintDice ser un gran avanceUso en el mundo realVietPrism: A large-scale Vietnamese speech and deepfake corpus with diverse dialects and code-switching
We introduce VietPrism, an open, multi-domain corpus that brings these dimensions together at scale: 993.4 hours and 403,941 bona fide utterances from 1,262 verified speakers across 8,388 real-world videos.
PreprintDice ser un gran avanceUso en el mundo realMachine learning of honey bee olfactory behavior identifies repellent odorants in free-flying bees in the field
Additional testing of the top seven candidates using freely foraging honey bees in a field assay confirmed strong repellency, thus predicting a high probability to repel foraging bees from pesticide-treated crops.
Revista con revisión por paresUso en el mundo realMicrorings as programmable temporal kernels enabling photonic AI beyond 100 Gbaud
Here, we redefine MRRs as programmable temporal convolution kernels by exploiting their impulse responses, enabling computation beyond the resonance linewidth.
Revista con revisión por paresAfirmaciones fuertes, leer con cuidadoDice ser un gran avanceUso en el mundo realOn a Separate Note: Robust Score-Informed Note Separation with a Two-Stream TFC-TDF U-Net and Adaptive Set Ownership
We present, to our knowledge, the first deep learning approach to score-informed note separation, NoteSep.
PreprintDice ser un gran avanceLow-altitude aircraft will reshape noise exposure across global cities
These findings show that low-altitude aircraft can reshape noise exposure across global cities, requiring route assessment to distinguish newly exposed from already exposed areas and to account for three-dimensional exposure.
PreprintUso en el mundo realYODAS v3: Over 1 Million Hours of High-Bandwidth, Stereophonic, Multilingual Speech
YODAS v3 is not only the largest open speech dataset to date, but also the first truly large-scale speech corpus with high-fidelity stereo audio.
PreprintAfirmaciones fuertes, leer con cuidadoDice ser un gran avanceCódigo disponibleCinematicVQA: Benchmarking Film-Grammar Reasoning in Large Vision-Language Models
To address this, we introduce CinematicVQA, the first-of-its-kind benchmark for cinematic video understanding that goes beyond technique recognition to evaluate film-grammar reasoning, utilizing our introduced Cinematic Scene Graph (CSG), a structured representation that links filming techniques to their perceptual effects and narrative functions.
PreprintTransformers as Cross-Task Learners: Shared Structure Drives Sample Efficiency in In-Context Learning
To the best of our knowledge, this is the first work to quantify cross-task complexity for general nonlinear task families and explicitly construct a Transformer that exploits their low-dimensional structure to perform ICL.
PreprintDice ser un gran avanceSynthetic Hospital: An Open, Verifiable, Physician-Validated Longitudinal EHR Benchmark
We introduce Synthetic Hospital, an open, fully synthetic, fact-grounded longitudinal EHR benchmark that resolves the open sharing and verifiable ground truth barriers.
PreprintEditVoice: Variable-Length Non-Autoregressive Zero-Shot TTS and Speech Editing with Edit Flows
We introduce EditVoice, to our knowledge the first variable-length NAR zero-shot TTS model, which uses Edit Flows to jointly update speech content and sequence length through insertions, deletions, and substitutions.
PreprintDice ser un gran avanceUso en el mundo realMultiplexed deep visual proteomics resolves spatial heterogeneity and rare endocrine states in human pancreatic islets
Applied to human pancreatic islets, mxDVP segments over 860,000 cells and resolves twelve endocrine subtypes, including rare polyhormonal and intermediate-state populations that exhibit spatial organization patterns, co-expression of INSM1 and SCG3, and hybrid α/β/δ signatures.
Revista con revisión por paresASR ensembling for phoneme intelligibility evaluation of speech anonymizers
Our results show that simple hard-voting ASR metric reaches correlations above 0.9 with human ratings when aggregated by feature, test-type, or condition, provided that multiple ASR models are combined; evaluating stimuli with and without a carrier sentence further improves the correlation at the stimulus level.
PreprintTrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations
We introduce TrackEverything, a 3D point tracker that breaks this trade-off by representing videos as persistent 3D scene tracks in world coordinates.
PreprintAfirmaciones fuertes, leer con cuidadoDice ser un gran avancePHOSA: Photorealistic 3D Sign Avatar Modeling and Benchmark
Building on MVSign, we propose a decoupled sign avatar representation that isolates body, head, and hand components to capture complex articulations, together with a motion-aware sampling strategy to handle motion blur and balance gesture diversity.
PreprintUso en el mundo realComplexSync: High-Fidelity and Real-Time Lip Sync in Complex Scenarios
We present ComplexSync, a unified diffusion-based framework that enables real-time, high-fidelity lip sync under complex conditions.
PreprintUso en el mundo realLabFactory: Building and Evaluating Executable AI Labs
We present LabFactory, a framework in which an AI builder turns a scientific brief into an executable AI lab: a task-specific solver that integrates models, knowledge resources, tools, and a controller behind a fixed interface.
PreprintUso en el mundo realDeep learning extracts MoA-specific signatures from high-throughput images of chemically and genetically perturbed Corynebacteria
Our model robustly classifies MoAs of established antibiotics and recognizes the MoA of previously unseen antibiotics.
Revista con revisión por paresUso en el mundo realThe Entropy Triangle Method (ETM): A novel framework for the prevention of cardiac arrhythmia with a review of more than 10,000 patients
In this article, we introduce two firsts in machine learning and medicine that can predict non-sinus rhythm with over 85% accuracy.
PreprintAfirmaciones fuertes, leer con cuidadoDice ser un gran avanceUso en el mundo realOn Growth and Form, and Function: Reusable Regulatory Handles Control Phenotypic Variation
Together, our results provide a computational realization of D'Arcy Thompson's remarkable grid transformations in a 2D NCA---a minimal cybernetic tissue in which variations of fully grown emoji phenotypes can be encoded, combined, and controlled through low-dimensional directions in regulatory weight space.
PreprintAfirmaciones fuertes, leer con cuidadoA Living Benchmark for Information Retrieval from Electronic Health Records
We present a scalable framework that automatically generates question--answer pairs from longitudinal EHR notes.
PreprintUso en el mundo realWhat, When, and How: Audio Description as Constrained Global Optimization
When evaluated on REFRAMED, a benchmark for realistic AD of movies, our approach makes better decisions than prompted LLMs about what to describe and when to describe it, establishing a new SOTA on narrative QA and temporally grounded metrics.
PreprintUso en el mundo realPath-specific harm decomposition: A partial identification framework
As a remedy, we develop a novel partial identification framework for direct and indirect FNA.
PreprintThe Last Human Gate: Forward Deployed Engineering for Governance Automation
These results support the technical feasibility of replacing human execution of specified governance-review tasks with agents and software.
PreprintUso en el mundo realCódigo disponiblePersonalized single-cell transcriptomics reveals molecular diversity in Alzheimer’s disease
Personalized functional genomics atlas for Alzheimer’s disease that uses knowledge-guided graph neural networks to analyze donor-level functional genomics, identify disease subpopulations and trajectories, and link genetic variants to gene regulation.
Revista con revisión por paresNeural Transport Nested Sampling
We develop a novel sampling algorithm, Neural Transport Nested Sampling (NTNS), which combines the classical strengths of nested sampling with modern neural flow-based methods.
PreprintDice ser un gran avanceDiscovering how ice crystals grow using neural ODEs and symbolic regression
By optimizing against mass growth time series from 290 ice crystals grown in a levitation diffusion chamber, we identify a modified capacitance growth model that more accurately captures observed early-stage growth.
Revista con revisión por paresDeep learning of longitudinal visual fields predicts glaucoma progression rate and identifies fast progressors
On the open-access University of Washington Humphrey Visual Field dataset (4,276 patient-eyes), GLAM achieved an MD-rate mean absolute error of 0.139 dB yr (; 73.5% reduction over a ridge baseline) and an AUC of 0.990 for fast-progressor detection.
PreprintUso en el mundo realImplementation of offset corrected AGAD algorithm on 130-nm CMOS technology–based RRAM array for analog neural network training
We present the first implementation of the analog gradient accumulation with dynamic reference (AGAD), reported as the most advanced and highest-performing version of the TT (Tiki-Taka) algorithm, on an HfO 2 -based resistive random-access memory (RRAM) array for analog neural network training.
Revista con revisión por paresAfirmaciones fuertes, leer con cuidadoDice ser un gran avanceBeyond Compression: Training Latent Representations for Stable Long-Horizon Rollout in Neural Surrogate Solvers
Collectively, these interventions reduce long-rollout error by approximately 40% and match or exceed the accuracy of full-resolution models on two physics benchmarks, while requiring 2 orders of magnitude fewer floating point operations and half the GPU memory.
PreprintDAMSEP: Distance-Aware Monaural Source Separation using Multi-RIR Estimation
To address this limitation, we propose Distance-Aware Monaural Source Separation using Multi-RIR Estimation (DAMSEP), the first end-to-end framework that is jointly trained for source separation and multi-source RIR estimation from a single-microphone mixture.
PreprintCódigo disponibleScreen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale
We present a hypothesis-driven simulation workflow for screening candidate CX agents before deployment.
PreprintUso en el mundo realTimeBraid: Unifying Time Series and Language for Understanding and Forecasting
We present TimeBraid, a series of unified time-series and language models that align pretrained language models and pretrained time-series foundation models through interleaved global residual attention layers.
PreprintWhere Does Exactly-Once Live? Model, Harness, and Tool-Contract Effects on Duplicate Side Effects in LLM Agents
Across 25,930 episodes spanning nine recent models, three production agent harnesses, two contract variants and fifteen recovery conditions, the answer depends on the fault.
PreprintUso en el mundo realLearning to Discover Interesting Mathematics
Optimizing for our metric creates a model capable of producing more interesting theorems, while also reducing substantial or full overlap with Mathlib from 91.9% to 30.6%, showcasing the creation of more out-of-distribution math.
PreprintAfirmaciones fuertes, leer con cuidadoSelf-Play Pretraining with Zero Data
We introduce Self-Play Pretraining with Zero Data, an initial proof-of-concept towards realizing this vision.
PreprintShadow Reduction in Ultrasound Imaging Using Differentiable Simulation and Radiance Field Decomposition
We present RFlash, a physics-informed post-processing method that decomposes beamformed ultrasound images into explicit attenuation and scatter-intensity maps using a differentiable radiance-field formulation of image formation.
PreprintUso en el mundo realReachability-Based Formal Verification of Graph Neural Networks with Node and Edge Features
Our results show that GNNV provides tighter robustness guarantees than CORA on graph classification models with ReLU-based activations and, for the first time, delivers edge-aware robustness guarantees for GINE-based PF and OPF models under joint node and edge perturbations.
Preprint con versión publicadaUso en el mundo realUpTCR: a unified progressive knowledge transfer foundation model for robust T-cell receptor-antigen binding recognition
Here, we present UpTCR, a unified progressive knowledge-transfer foundation model that learns from incomplete data to predict TCR-antigen-HLA binding.
Revista con revisión por paresUso en el mundo realDrGait: Biomechanically Grounded Visual Reasoning for Interpretable Clinical Gait Analysis
By anchoring VLM's reasoning in verifiable geometric and temporal measurements, DrGait reduces hallucinations, achieving competitive diagnostic accuracy while generating transparent and audit-ready clinical reports.
PreprintUso en el mundo realAugur: A Synthetic Decision Lab for Rehearsing Reactions to Product and Policy Changes
Our central finding is methodological and negative: most of the measured gap between frontier cloud models and open-weight models we fine-tune and serve offline is attributable to an under-specified evaluation, not a difference in capability. We show this three ways.
PreprintAgentic Detection of Online Conspiracies
Evaluating our framework on a manually-annotated adversarial dataset, we find that context-aware workflows consistently outperform text-only classification and that the agentic framework performs significantly better than other frameworks and settings, including a non-agentic model exposed to the same contexts available to the agent.
PreprintQwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents
Qwen-Planner-Agent achieves the best overall performance among all evaluated models and systems on MobilePA-Bench, improving over its base model across tool use, memory, skills, and sub-agent coordination.
PreprintUso en el mundo realAICO: Feature significance tests for supervised learning
We introduce AICO (Add-In COvariates), a broadly applicable framework that turns model interpretability into an efficient statistical exercise.
Revista con revisión por paresAfirmaciones fuertes, leer con cuidadoUso en el mundo realWildHSR: Metric Feed-Forward 4D People-Scene Reconstruction from a 3D Foundation Model
WildHSR combines both readouts to reconstruct metric cameras, scene and people from monocular video.
PreprintScalarLens: Numerical Embeddings with Stable Coordinates and Contextual Responses for CTR Prediction
We introduce ScalarLens, a numerical embedding that preserves what a value is while adapting how it should be interpreted.
PreprintASIRF: An Agentic Framework for Context-Dependent Sensitive Information Redaction
With only a few dozen expert-authored definitions per domain and no training data, ASIRF's recall exceeds OPF's in 68 of 80 model-domain combinations (85 percent), by at least one of the two architectures, with shortfalls confined mostly to OPF's training-distribution domains.
PreprintUso en el mundo realJust Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures
A single generic question reaches a median AUROC of 0.886 zero-shot and beats supervised baselines on most benchmarks.
PreprintUso en el mundo realCódigo disponibleBeyond Spatial Benchmarks: From Spatial Reasoning to Navigation
Our analysis reveals a gap between benchmark-oriented spatial specialization and navigation performance, and shows how aligning spatial supervision with navigation goals, phases, and decision learning improves navigation.
PreprintCódigo disponibleTemporal Taxation Compounds Under Post-Training Compression of Whisper Models
We cast the temporal-taxation construct of Choi and Choi (2025) as a quantitative metric, and show that single-snapshot fairness audits on full-precision models do not capture the deployment-time burden that compression places on already-marginalized speakers.
PreprintUso en el mundo realMonitoring Urban Traffic Dynamics at Fine Spatiotemporal Resolution Using Distributed Acoustic Sensing and Deep Learning
Using a DAS deployment along a roadway network in the City of College Station, Texas, USA, this study develops a deep learning-empowered analytical framework that converts raw ground vibration waveforms into spatiotemporal representations, detects vehicle trajectory, and infers traffic states from aggregated traffic volume and speed.
PreprintUso en el mundo realBeyond Surface Style: Aligning Multi-Turn User Simulators with Behavioral Consistency
We propose TRACER, a multi-turn user simulator that explicitly models users' evolving intent and learns to align simulated behavior with real interaction trajectories.
PreprintUso en el mundo realEmpath: Tracing Multi-Level Emotion Dynamics in Crisis Counseling Dialogues
We introduce EMPATH, a framework for understanding affective dynamics in mental health dialogues across three granularities: turn-level labels, transition probabilities, and global conversation archetypes.
PreprintIndustrial Anomaly Detection via Defect-Grounded Reasoning in Visual Latent Space
To address these, we propose Anomaly-LR, a defect-grounded latent reasoning framework that first forms a global understanding of the input and then progressively refines anomaly-relevant representations directly in the visual latent space.
PreprintUso en el mundo realCódigo disponibleEnigmaForge: The Question Is Hidden in the Story
Intuition reshuffles the leaderboard: a 22x spread where fact recovery spans 1.6x, the second-best fact-recoverer ranks fourteenth, one model is indifferent to being told the question, and another is significantly better without it.
PreprintLow-Cost Assays for Measuring Model Behavior Across Vendors and Releases
To address these constraints, we present a simple, cheap, scalable, and replicable model for studying model behavior.
PreprintCódigo disponibleBeneath the Scores: Rethinking Hallucination Evaluation for Video Understanding Models
Across 60,008 runs on three video-agent architectures, we find that grounding is the dominant source of downstream error, with roughly four times the causal impact of corrupting visual observations.
PreprintWhere Hallucinations Live: A Cross-Architecture Circuit in VQ-Tokenized Vision-Language Models
Using activation patching across twenty-five models spanning eight LLM families, we identify an early-layer () attention routing circuit shared across VQ-tokenized VLMs and propose a three-gate diagnostic that distinguishes the models carrying it from those that do not.
Preprint