New papers on Speech & audio
154 new papers on speech & audio in the last 7 days, within AI & machine learning. These are the 50 Pipette rates most worth reading, with the main result in the authors' own words.
The best of the week
NAVIR: Neuromorphic Audio-Visual Speech Recognition for Robust Human-Robot Interaction on Edge Hardware
We present NAVIR, an end-to-end AVSR system targeting the BrainChip Akida neuromorphic processor, which natively supports only sequential two-dimensional convolutional inference.
PreprintClaims a big stepReal-world useVietPrism: A large-scale Vietnamese speech and deepfake corpus with diverse dialects and code-switching
We introduce VietPrism, an open, multi-domain corpus that brings these dimensions together at scale: 993.4 hours and 403,941 bona fide utterances from 1,262 verified speakers across 8,388 real-world videos.
PreprintClaims a big stepReal-world useEvoAudio: Recursive Self-Improvement for Audio Understanding
We therefore propose EvoAudio, a recursive self-improvement system for audio understanding.
PreprintClaims a big stepOn a Separate Note: Robust Score-Informed Note Separation with a Two-Stream TFC-TDF U-Net and Adaptive Set Ownership
We present, to our knowledge, the first deep learning approach to score-informed note separation, NoteSep.
PreprintClaims a big stepYODAS v3: Over 1 Million Hours of High-Bandwidth, Stereophonic, Multilingual Speech
YODAS v3 is not only the largest open speech dataset to date, but also the first truly large-scale speech corpus with high-fidelity stereo audio.
PreprintBold claims, read criticallyClaims a big stepCode availableEditVoice: Variable-Length Non-Autoregressive Zero-Shot TTS and Speech Editing with Edit Flows
We introduce EditVoice, to our knowledge the first variable-length NAR zero-shot TTS model, which uses Edit Flows to jointly update speech content and sequence length through insertions, deletions, and substitutions.
PreprintClaims a big stepReal-world useBearings: Self-Supervised Soundfield Embeddings from First-Order Ambisonics
On sound event localization and detection, concatenating our soundfield embeddings with acoustic representations provides the missing spatial information and enables joint detection and localization, raising the location-dependent F-score from below 4 to 50 on TAU-NIGENS 2021 and 39 on STARSS23.
PreprintClaims a big stepLiteCASS: A Lightweight End-to-End Network for Real-Time Stereo Cinematic Audio Source Separation
We present LiteCASS, to our knowledge the first lightweight end-to-end network for real-time stereo CASS.
PreprintReal-world useNemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities
We introduce NemotronLabs VoiceChat, an open full-duplex speech-to-speech model with native tool-calling capabilities.
PreprintReal-world useI'll Keep an Ear Out: Teaching AudioLLMs Proactive Audio Assistance
We introduce proactive audio assistance, where an AudioLLM monitors an audio stream and autonomously decides when to alert the user from a single natural-language intent, motivated by wearable applications for Deaf and Hard of Hearing users.
PreprintBold claims, read criticallyReal-world useTowards Zero-Shot Attribution of Synthetic Speech via Audio-Text Contrastive Retrieval
For generators it has never encountered before, the model reaches a model-level mean reciprocal rank (MRR) of 58.4%.
PreprintReal-world useDAMSEP: Distance-Aware Monaural Source Separation using Multi-RIR Estimation
To address this limitation, we propose Distance-Aware Monaural Source Separation using Multi-RIR Estimation (DAMSEP), the first end-to-end framework that is jointly trained for source separation and multi-source RIR estimation from a single-microphone mixture.
PreprintCode availableTraining Music Sample Identification Models on Real Sample Pairs
In this work, we present SI Embeddings (SIE), an SI model that achieves state-of-the-art results on three benchmarks, including a large-scale test set.
PreprintListen, Critique, and Refine: RL-Based Self-Refinement for Instruction-Following Speech Synthesis
After RL training, the refined two-hop outputs achieve a relative improvement of 7.15% on the InstructTTSEval benchmark, demonstrating the model's reflective ability.
PreprintTTS-Guard: Black-Box Ownership Verification of Text-to-Speech Models via Adaptive Adversarial Speaker-Pair Fingerprints
We present TTS-Guard, a black-box ownership verification framework for TTS models built on adversarial speaker-pair fingerprints.
PreprintReal-world useEnhancing speech representation learning with cross-modal knowledge transfer with HGNN under low resource settings: the case study of Yemba
To our knowledge, this is the first demonstration of explicit cross-modal knowledge transfer for acoustic representation learning using HGNNs, highlighting a promising direction for speech representation in low-resource settings.
PreprintBold claims, read criticallyPhonemizing User-Generated Text: A Benchmark, Taxonomy, and Compositional Approach
We introduce UGTPhon, the first grapheme-to-phoneme (G2P) benchmark for UGT in English, Vietnamese, and Korean, together with an inference-grounded taxonomy for fine-grained diagnosis.
PreprintReal-world useAll I Hear is Noise: Investigating Clever Hans Effects in Clinical Speech Datasets
Across all datasets, silence-only classification frequently matched or exceeded full-audio performance, suggesting that classification performance may be influenced by dataset-specific confounds in addition to disorder-related speech characteristics.
PreprintReal-world useExploring a Single Autoregressive LLM for Unified Target Speech Extraction across Synchronous and Asynchronous Cues
We show that one autoregressive LLM backbone, TSE-Omni, can serve both temporally synchronous cues (lip movements, co-speech gestures) and asynchronous cues (enrollment audio, text).
PreprintBold claims, read criticallyReal-world useCOSED: Setting the Bar for Open-Vocabulary Sound Event Detection
We then introduce COSED, which surpasses prior work on five out of six tasks while staying on par with the best method on the sixth, with margins of 12-33% on three of them.
PreprintPsychoacoustically Aligned Latent Smoothing for Adversarial Robustness of Full-Duplex Speech-to-Speech Dialogue Models
Deployed with no inference-time cost, PALS reduces hijack to 8.3%, mute to 11.2%, and jailbreak to 9.1% at clean quality within 2.3%.
PreprintReal-world useAdaptDuplex: from static to adaptive full-duplex spoken dialogue
We present AdaptDuplex, which upgrades Qwen3-Omni with such a mechanism, co-designed across three layers.
PreprintReal-world useRESTORE: REal-time Steerable Music resTORation and bandwidth Extension via stem disentanglement
Because what constitutes a restored audio signal is subjective, we introduce RESTORE, a framework that formulates audio restoration as a six-source semantic decomposition to allow for real-time interactive user control over the process.
PreprintBold claims, read criticallyReal-world useTowards participatory speech dataset curation: A queer case study and conceptual framework
From this review, we develop a conceptual framework for participatory speech data curation by, for, and with marginalized communities drawing on insights from co-design and knowledge sharing.
PreprintInteractive TTS: Dynamic Speaking Style Adaptation for Expressive Speech Synthesis
To overcome these limitations, we propose Interactive TTS, a dynamic, style-adaptive framework for contextually appropriate and speaker-consistent speech generation.
PreprintReal-world useCOT-TTS: Audio Context-Aware Text-to-Speech with Chain-of-Thought Reasoning
Therefore, we propose COT-TTS, a context-aware, reasoning-based text-to-speech task.
PreprintReal-world useParA-LLM: A Unified Approach to Paralinguistic and Acoustic Speech Understanding
ParA-LLM surpasses state-of-the-art Audio LLMs like GPT-4o-Audio by 7.5% on ParA-Bench, with additional gains of 1.13% on MMAU-Pro Speech and 7.49% on MMAR Speech.
PreprintBenchmarking and Domain Adaptation of Automatic Speech Recognition (ASR) for Adolescent Health Communication in Ghanaian Languages
Across all three stages the evidence converges: for these languages the binding constraint is validated in-domain data, not model capability or computation.
PreprintReal-world useA Harness for Synthesizing Diverse Naturalistic Full-Duplex Conversations
We present a pipeline for synthesizing intent-labeled, two-channel conversational speech from relational event lists.
PreprintReal-world useStructure Before Sampling: Community-Aware Core-Set Selection for Data-Efficient Text-to-Speech
TTS models trained on its 20% core-sets have a significantly lower character error rate (CER) than models trained on equal-duration random or entropy-based subsets in both languages.
PreprintReal-world useBiCFlow-MER: Orchestrating Discriminative and Generative Multimodal Emotion Recognition via Conditional Transport
To address these limitations, BiCFlow-MER (Bidirectional Conditional Flow for Multimodal Emotion Recognition) is proposed as a conditional-flow framework in which audio-text MER is formulated as generative evidence transport within a structured emotion space.
PreprintBold claims, read criticallyThe Vulnerability of Neural Audio Watermarks under Speech Enhancement
Experimental results show that the proposed attack significantly outperforms existing neural re-synthesis methods in watermark removal.
PreprintReal-world useScaling Forced Alignment to End-User Devices
We propose two optimizations to address this issue.
PreprintReal-world useQueer inclusion in speech datasets: An audit and taxonomy of practical tensions
Through an audit of six diverse speech datasets, we find that measurable queer representation is low (0-1.4% of speakers) - insufficient for robust disparity measurement.
PreprintEquiSELD: Efficient training of equivariant sound event localization and detection networks
EquiSELD outperforms prior equivariant networks on both simulated scenes with measured RIRs and recordings of real-world sound scenes at a fraction of the training cost.
PreprintExemplar-Free Analytic Learning for Multi-Label Audio Class-Incremental Learning
Experiments on a 50-class AudioSet-R benchmark across three incremental setups show that the analytic learner substantially outperforms gradient-based methods, and previously learned classes retain nearly unchanged detection performance as new classes are added.
PreprintSPADE: A Multilingual Dataset for Speech Partial Deepfake Detection and Localization
To enable further research in this direction, we propose a multilingual dataset for detection and localization of partially edited speech samples.
PreprintReal-world useCode availableTraining Intelligent Voice Assistant Wakeup with Controllable Synthetic Conversations
This paper introduces a novel wake-up system that extends traditional direct keyword detection with contextual trigger detection.
PreprintReal-world useA Practical Recipe for Semi-Supervised Federated ASR: Online Pseudo-Labels with Server Update Stabilization
These findings yield guidelines for SSFL in ASR training, improving over the strongest prior method on 9 of 11 pairs, by on average in-domain and cross-domain, narrowing the gap to fully-supervised FL.
PreprintReal-world useAURA: Uncertainty-Routed Activation Editing for Acoustic Grounding in Speech Foundation Models
We propose AURA: Activation-editing with Uncertainty-Routed Adaptation, an ultra-efficient representation-editing method that freezes the pretrained model and applies sparse scale-and-shift edits to decoder cross-attention heads.
PreprintReal-world useCycleSpeech: Reciprocal Alignment for Instruction-Controlled Speech Synthesis and Paralinguistic Understanding
We introduce CycleSpeech, a framework that connects generation and understanding through a shared, structured voice profile that serves as a common target for supervision and reciprocal feedback.
PreprintReal-world useA Native-Reference Coordinate Geometry for L2 Pronunciation Deviation Using Self-Supervised Speech Models
We propose a native-reference coordinate geometry in which phone-class averages from native speech define a low-dimensional reference subspace, and L2 speech is evaluated by its distance to matching native phone-class coordinates.
PreprintPartial Accent-Control Editing in Frozen Speech Representations for Accent Conversion
We propose Partial Accent-Control Editing (PACE), an accent conversion framework based upon the editing of frozen WavLM representations without the training of an accent-conditioned generator.
PreprintNo Time to Collapse: Unlocking Robustness and Multiplexed Capacity in Frozen Audio Watermarkers
We argue that this temporal collapse limits both robustness and the recovery of multiple payloads, and that the limitation can be addressed without retraining the underlying watermarker.
PreprintReal-world useNeuMark: Neural Codec Resynthesis-Robust Audio Watermarking in the Codec Latent Space
In this paper, we propose NeuMark, a codec-latent audio watermarking framework that embeds watermark evidence into SpeechTok- enizer acoustic tokens to address this resynthesis threat.
PreprintReal-world useSE-MSB: End-to-End Unpaired Speech Enhancement using Mamba Schr\"odinger Bridges
We propose a fully unpaired SE framework that uses principled Diffusion Schr\"odinger Bridges (DSB) to learn a stochastic transport process between a clean and a degraded speech distribution.
PreprintBold claims, read criticallyReal-world useThe Internet Archive Music Dataset
We introduce the Internet Archive Music Dataset (IAMD), a large-scale collection of captioned music segments derived from the Internet Archive.
Preprint with a published versionCross-Lingual Parkinson's Disease Severity Assessment Using Pre-trained Speech Embeddings: A Multi-Class Evaluation
Our results show that pre-trained speech embeddings enable meaningful cross-lingual transfer, although performance is sensitive to dataset properties, preprocessing, and adaptation strategy.
PreprintReal-world useGenTraceBench: A Benchmark for Tracing Audio Deepfakes Across Pre- and Post-training Stages
DPO and GRPO generally preserve fingerprints, whereas some SFT and pre-training-data changes cause substantial drift; effect sizes vary across three forensic backbones.
PreprintReal-world useSpoken Language Models that Think Aloud
Experiments on spoken reasoning and question-answering benchmarks show that our approach substantially reduces user-audible silence during reasoning while maintaining answer accuracy comparable to that of a serial "think-then-speak" baseline, demonstrating the potential of asynchronous think-aloud for responsive interaction in SLMs.
PreprintReal-world use