pipette
ESEspañol

JASPER: Joint Audio and Speech Pre-trained Encoder Representations

Geeth George, Ameenudeen P E, Hrishikesh H Pillai, Sriram Ganapathy

Preprint

In the authors' words

Self-supervised learning (SSL) for speech and audio has largely progressed along separate tracks: speech models emphasise time-domain prediction, whereas audio representation learning has focused on time-frequency patterns. This separation creates a compatibility gap, limiting cross-domain generalization. In this work, we introduce JASPER, Joint Audio and Speech Pre-trained Encoder Representations, a framework that augments speech-pretrained models with time-frequency objectives. Specifically, JASPER performs masked prediction of temporal and spectral targets over long audio segments, enabling spectro-temporal representation learning of speech and audio signals. The proposed method consistently outperforms multiple baselines and existing speech/audio encoders on diverse speech, audio, and music tasks, demonstrating the effectiveness of unified spectro-temporal modeling.

Main resultThe abstract does not state a limitation.

Appeared: Thursday, September 24. arXiv. Preprint, not yet peer-reviewed.