Skip to main navigation Skip to search Skip to main content

Post-training for Deepfake Speech Detection

  • Wanying Ge*
  • , Xin Wang
  • , Xuechen Liu
  • , Junichi Yamagishi
  • *Corresponding author for this work
  • Research Organization of Information and Systems, National Institute of Informatics

Research output: Chapter in Book or Report/Conference proceedingConference Proceedingpeer-review

2 Citations (Scopus)

Abstract

We introduce a post-training approach that adapts self-supervised learning (SSL) models for deepfake speech detection by bridging the gap between general pre-training and domain-specific fine-tuning. We present AntiDeepfake models, a series of post-trained models developed using a large-scale multilingual speech dataset containing over 5 6, 0 0 0 hours of genuine speech and 1 8, 0 0 0 hours of speech with various artifacts in over one hundred languages. Experimental results show that the post-trained models already exhibit strong robustness and generalization to unseen deepfake speech. When they are further fine-tuned on the Deepfake-Eval-2024 dataset, these models consistently surpass existing state-of-the-art detectors that do not leverage post-training. Model checkpoint1 and source code2 are available online.1Zenodo: https://doi.org/10.5281/zenodo.15580542 Hugging Face: https://huggingface.co/nii-yamagishilab2GitHub: https://github.com/nii-yamagishilab/AntiDeepfake

Original languageEnglish
Title of host publicationASRU 2025 - 2025 IEEE Automatic Speech Recognition and Understanding Workshop
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9798331544263
DOIs
Publication statusPublished - 2025
Externally publishedYes
Event2025 IEEE Automatic Speech Recognition and Understanding Workshop, ASRU 2025 - Honolulu, United States
Duration: 6 Dec 202510 Dec 2025

Publication series

NameASRU 2025 - 2025 IEEE Automatic Speech Recognition and Understanding Workshop

Conference

Conference2025 IEEE Automatic Speech Recognition and Understanding Workshop, ASRU 2025
Country/TerritoryUnited States
CityHonolulu
Period6/12/2510/12/25

Keywords

  • deepfake detection
  • post-training
  • speech

Cite this