FARE: Forensic Acceptance Region Estimation for Catching Bait-and-Switch Image Generators
In the authors' words
Modern AI image generators are increasingly deployed as opaque APIs, where customers can query the deployed service, but cannot inspect model weights or architecture. This creates a practical challenge: a provider may pass governance certification with one generator and later silently switch to a cheaper and lower-quality one for deployment, compromising public trust or even safety in high-stakes domains. We study integrity auditing at deployment time and propose FARE (Forensic Acceptance Region Estimation). A certified generator is enrolled by training FARE on images sampled from that generator. After deployment, FARE can determine whether a generated image is consistent with the enrolled generator---using only that image. FARE's features are based on image generator-specific artifacts that have been proposed for forensic applications. FARE amplifies these features during training by finding hard samples that tighten the acceptance region and increase sensitivity to subtle changes in the certified generator. Across generator swaps, including substitutions with similar model versions and model variants, FARE is effective at detecting swaps, consistently outperforming existing baselines at strict operating points, and remains effective under the exact-model and decision-only attacks evaluated in this work.
Appeared: Monday, September 28. arXiv. Preprint, not yet peer-reviewed.
Authors' comment: This work has been accepted for publication in the proceedings of The 40th Annual Conference on Neural Information Processing Systems (NeurIPS 2026). 22 pages, including technical appendices. Code: https://github.com/kaikaiyao/FARE