pipette
ESEspañol

Synthetic speech detection in Brazilian Portuguese through accent-related features

Pedro H. L. Leite, Pedro Benevenuto Valadares, Luiz Wagner Pereira Biscainho

PreprintReal-world use

In the authors' words

Leading commercial and open-source Text-to-Speech (TTS) models fail to emulate the regional phonetic diversity of Brazilian Portuguese (pt-BR). By aggregating disparate dialects into a single training distribution, they generate a synthetic "diluted" accent: a phonetic profile attempting to represent all regional distributions simultaneously, but ultimately carrying phonological ambiguity dissociated from natural socio-phonetic realizations. This work introduces a speech deepfake detection methodology combining multilingual phone recognizers with classical signal processing to extract phoneme-level features in consonantal and vocalic realizations with high geographic variance. The analysis reveals that the distributional gap over these features suffices to distinguish natural and synthetic voices through unsupervised Kernel Density Estimation, establishing dialectal inconsistency as a useful and interpretable feature for spoofing detection in pt-BR. Evaluation on pt-BR anti-spoofing datasets shows that these explainable, lightweight, low-dimensional features can boost the performance of foundation models on the task, and show generalization capabilities in a cross-dataset leave-one-out setup.

Main resultThe abstract does not state a limitation.

Appeared: Tuesday, September 22. arXiv. Preprint, not yet peer-reviewed.

Authors' comment: \c{opyright} 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional pur