pipette
ENEnglish

STAM-ASR: Speaker-Temporal Anchoring with Memory for Multi-Speaker ASR

Victor Tolulope Olufemi, Syeda Faiza Ahmed Sara, Shammur Absar Chowdhury

PreprintUso en el mundo real

En palabras de los autores

Natural conversations make both speech recognition and speaker attribution challenging for ASR, as speakers take turns, overlap, and reappear over time. We propose STAM-ASR, Speaker-Temporal Anchoring with Memory, a lightweight framework that extends an already pretrained AudioLLM for multi-speaker ASR. Without relying on an external diarization system, STAM-ASR learns speaker activity and speaker-aware representations directly from intermediate AudioLLM features. Hence providing explicit who and when cues to modulate the AudioLLM's semantic representation without explicit speech separation. STAM-ASR further maintains fixed-size speaker and conversational memories to carry complementary context across turns. We evaluate STAM-ASR on AMI, ICSI, LibriCSS, and NOTSOFAR-1 across close-talk, far-field, overlapping, and cross-domain conditions. Our reported results shows that speaker-temporal conditioning and memory provide complementary benefits, while the gap between reference and predicted speaker activity identifies robust speaker tracking as a key remaining challenge.

Resultado principalLimitación que admiten los autores

Apareció: viernes, 25 de septiembre. arXiv. Preprint, todavía sin revisión por pares.