pipette
ENEnglish

JASPER: Joint Audio and Speech Pre-trained Encoder Representations

Geeth George, Ameenudeen P E, Hrishikesh H Pillai, Sriram Ganapathy

Preprint

En palabras de los autores

Self-supervised learning (SSL) for speech and audio has largely progressed along separate tracks: speech models emphasise time-domain prediction, whereas audio representation learning has focused on time-frequency patterns. This separation creates a compatibility gap, limiting cross-domain generalization. In this work, we introduce JASPER, Joint Audio and Speech Pre-trained Encoder Representations, a framework that augments speech-pretrained models with time-frequency objectives. Specifically, JASPER performs masked prediction of temporal and spectral targets over long audio segments, enabling spectro-temporal representation learning of speech and audio signals. The proposed method consistently outperforms multiple baselines and existing speech/audio encoders on diverse speech, audio, and music tasks, demonstrating the effectiveness of unified spectro-temporal modeling.

Resultado principalEl resumen no menciona limitaciones.

Apareció: jueves, 24 de septiembre. arXiv. Preprint, todavía sin revisión por pares.