pipette
ESEspañol

Tracing and Relearning Detection Evidence in Text-to-Speech Systems

Eunji Shin, Kyudan Jung, Jihwan Kim, Minwoo Lee, Jaegul Choo

PreprintReal-world use

In the authors' words

Recent audio deepfake detectors separate bona fide speech from synthetic speech, yet it remains unclear which stage of a text-to-speech system supplies the detection evidence. We address this with controlled resynthesis and detector adaptation in an F5-TTS-BigVGAN pipeline. Since vocoder reconstruction of a real mel can itself be separable from the source utterance, we fix the vocoder and trace the larger change in detector separation to acoustic generation. Adversarially fine-tuning the acoustic model, with no detector in its objective, raises EER against fixed detectors at comparable quality. However, adapting a detector only on the tuned model's VCTK outputs lowers its LibriSpeech EER from 19.42% to 7.46% and improves detection of unseen base F5-TTS outputs. These results suggest that acoustic-model updates can reduce the detection evidence available to fixed detectors, while detector adaptation keeps the updated outputs detectable in this pipeline.

Main resultThe abstract does not state a limitation.

Appeared: Monday, September 28. arXiv. Preprint, not yet peer-reviewed.

Authors' comment: Submitted to ICASSP 2027