pipette
ESEspañol

Training-Free Pronunciation Transcription via Text-Constrained Acoustic Rescoring

Hikaru Asano, Yotaro Kubo, So Kuroki

PreprintReal-world use

In the authors' words

Accurate and efficient pronunciation transcription is essential for preparing text-to-speech training data at scale. Existing approaches have different limitations: grapheme-to-pronunciation (G2P) and speech-to-pronunciation (S2P) methods each capture only partial information, using only text or only speech, while speech-and-text-to-pronunciation (ST2P) methods use both but require costly pronunciation-annotated data. To address this problem, we propose a training-free ST2P pipeline that integrates both lexical and acoustic information at inference time. Lexical resources and G2P tools generate text-constrained candidates, and a left-to-right greedy search selects the best one using whole-sequence negative log-likelihoods from frozen pretrained S2P models. On three Japanese corpora, our method reduces Character Error Rate (CER) from 0.60--1.40% (text-only baseline) to 0.04--0.17% with reference transcripts, and 0.64--1.58% with ASR transcripts. It outperforms all baselines, including a trained ST2P model and commercial multimodal LLMs. Our greedy search method is 3--3.5 faster than beam search at similar CER, and the cascade is 2 faster than direct decoding ensuring the efficiency and accuracy. In Spanish, French, and preliminary English, it also surpasses four open multimodal LLMs and the best traditional methods.

Main resultLimitation the authors admit

Appeared: Monday, September 28. arXiv. Preprint, not yet peer-reviewed.

Authors' comment: 5 pages, 2 figures