pipette
ESEspañol

TTS-Guard: Black-Box Ownership Verification of Text-to-Speech Models via Adaptive Adversarial Speaker-Pair Fingerprints

Xubin Yue, Zhenhua Xu, Zhebo Wang, Mengting Li, Zijie Zhou, Wenpeng Xing, Dezhang Kong, Meng Han

PreprintReal-world use

In the authors' words

The rapid maturation of zero-shot Text-to-Speech (TTS) models has turned high-quality voice cloning into a widely available capability, raising acute concerns over unauthorised replication, fine-tuning and resale of proprietary speech models. Yet ownership verification for TTS remains largely open: speech is a continuous waveform whose perturbations are easily destroyed by routine signal processing, and the human auditory system imposes a much tighter perceptual budget than vision. We present TTS-Guard, a black-box ownership verification framework for TTS models built on adversarial speaker-pair fingerprints. TTS-Guard(i) selects key speaker pairs in a dual embedding space for architecture-agnostic stealth;(ii) optimises a perturbation through an adaptive curriculum of shadow models covering fine-tuning, pruning, quantisation and distillation; and (iii) aggregates black-box queries into a calibrated Verification Confidence Score. On five mainstream TTS systems, TTS-Guard reaches an average Fingerprint Success Rate of at a False Positive Rate of , while preserving intelligibility and naturalness. The fingerprint remains effective against ten audio attacks, six model modifications, and two state-of-the-art adversarial purifiers.

Main resultThe abstract does not state a limitation.

Appeared: Tuesday, September 22. arXiv. Preprint, not yet peer-reviewed.