pipette
ESEspañol

Attention-Based Adaptive Policies for Simultaneous Speech-to-Text Translation

Filip T\u{a}\c{s}\u{a}dan, Ema Tomanov\'a, Ondrej Lopuch, Pawe{\l} Bilko, Anders S{\o}gaard

PreprintReal-world use

In the authors' words

Simultaneous speech-to-text translation (Simul-S2TT) consists of generating partial translations while the incoming audio frames are processed by the system. However, the streaming nature of this setup creates the challenge of deciding the best moment to perform an accurate translation while minimizing the delay. To address this challenge, we utilize the cross-attention mechanism of the encoder-decoder architecture to find the right alignment between the input speech frames and the target text tokens. In this paper, we propose the Recent Frame Attention Policy (RFAP) and the Dual-Condition Attention Policy (DCAP) that allow offline trained speech-to-text translation models to be used in streaming scenarios without requiring additional training. Results on three different language translation pairs over the CVSS-C corpus show that the RFAP is able to surpass other policies with gains of up to 4.0 BLEU while reducing the translation delay by almost 1 second. Moreover, the DCAP is able to preserve a high translation quality when the latency is very low.

Main resultThe abstract does not state a limitation.

Appeared: Monday, September 28. arXiv. Preprint, not yet peer-reviewed.

Authors' comment: Submitted to ICASSP 2027