深度伪造语音检测技术研究综述
survey on deepfake speech detection techniques
Online published: 2026-04-01
Copyright
随着生成式人工智能(Generative Artificial Intelligence,GAI)技术的爆发式发展,深度伪造(Deepfake)语音技术有了显著提升,其生成的语音在音色、韵律和自然度等方面高度逼近真实人类语音,具备极强的欺骗性,给深度伪造语音检测带来了巨大挑战。本文旨在系统地梳理深度伪造语音检测技术的发展脉络、技术方法、面临的挑战以及未来发展趋势。首先,介绍了语音合成和语音转换等深度伪造语音技术的原理和方法;其次,介绍了深度伪造语音检测技术的研究进展,包括基于传统机器学习检测、基于深度学习检测和基于端到端检测等方法,阐述了各类检测方法的原理、特点和性能表现;然后,介绍了深度伪造语音检测的常用数据集和评价指标;最后,总结了深度伪造语音检测技术面临的挑战和未来发展方向。
杨云龙 , 刘权 , 钮艳 , 詹雪东 , 王宁 . 深度伪造语音检测技术研究综述[J]. 网络空间安全科学学报, 2025 , 3(5) : 14 -22 . DOI: 10.20172/j.issn.2097-3136.250502
With the rapid advancement of generative artificial intelligence (GAI), deepfake speech technology has achieved remarkable progress. The synthesized speech now closely mimics authentic human voices in terms of timbre, prosody, and naturalness, exhibiting strong deceptive capabilities and thus posing significant challenges to detection systems. This survey systematically reviews the evolution, technical approaches, current challenges, and future directions of deepfake speech detection. First, we elaborate on the fundamental principles and methodologies of deepfake speech generation, covering speech synthesis and voice conversion (VC). Second, we comprehensively examine detection techniques, classifying them into three paradigms: traditional machine learning–based methods, deep learning–based approaches, and end-to-end detection frameworks. For each paradigm, we provide detailed analysis of its working mechanisms, inherent characteristics, and performance on typical benchmarks. Third, we introduce widely used benchmark datasets and evaluation metrics in this field. Finally, we discuss key challenges—such as poor generalization across unseen forgery types and constraints in real-time deployment—and outline promising future research directions.
表 1 基于深度学习检测技术性能比较Table 1 Detection performance comparison of deep learning-based technologies |
| 文献方法 | 核心模型 | 输入特征 | 关键创新点 | LA攻击场景下 实验结果(EER) | 局限性 |
| Lei等[14] | Siamese CNN | GMM后验 概率 | 结合GMM帧分数与CNN局部关系 建模,具有较高检测性能 | 3.79% | 模型相对简单,对极其逼真 的伪造语音检测可能失效 |
| Rostami等[15] | EfficientNet-A0、 SE-Res2Net50 | 声学特征 | 使用模型融合与SpecAug数据增 强,模型检测性能优越 | 0.86% | 模型复杂度较高,需要消耗 较高计算资源 |
| Wang等[16] | SE-Res2Net-Conformer | 声学特征 | 引入Conformer挖掘局部与全局模 式,探索了时空建模方式 进行检测的价值 | 1.85% | Conformer结构相对复杂, 影响推理速度 |
| Li等[17] | SafeEar | 声学特征 | 内容与声学信息解耦, 保护用户隐私 | 3.10% | 双分支结构增加了模型 复杂度和训练难度 |
| 盖馨怡等[18] | MFF-STViT | MFCC、GFCC、 谱质心、声码器 伪迹特征 | 融合了多种声学特征和STViT 架构,在复杂、长距离伪造 痕迹的检测上表现突出 | 0.41% | STViT模型复杂,参数量大, 训练和推理速度慢 |
表 2 端到端检测技术性能比较Table 2 Performance comparison of end-to-end detection technologies |
| 文献方法 | 核心模型 | 输入特征 | 关键创新点 | LA攻击场景下 实验结果(EER) | 局限性 |
| Tak等[22] | RawNet2 | 原始波形 | 实现SincNet层优化,在ASVspoof 2019数据集 上表现优异,证明了直接处理波形的潜力 | 1.12% | 模型对新型的伪造语音痕迹 不够敏感 |
| Wang等[23] | RawNet2, SimAM | 原始波形 | 结合了元学习与SimAM注意力机制,证明了 注意力机制在聚焦伪造痕迹上的有效性 | 0.99% | 在复杂度增加的情况下,性能 提升有限 |
| Ito等[24] | Wav2Vec2, HuBERT | 原始波形 | 利用自监督学习预训练模型提取深度特征,展示 了大规模无监督数据学习强大通用表征的潜力 | 0.44% | 预训练和推理均需 大量计算资源 |
| Yuan等[25] | MSFA | 原始波形 | 融合了多尺度特征适配与跨层信息,擅长捕获 复杂、多尺度的伪造模式 | 0.36% | 需要消耗大量计算资源和 数据进行训练 |
| 1 |
Li X L, Yu N H, Zhang X P, et al. Overview of digital media forensics technology[J]. Journal of Image and Graphics, 2021, 26 (6): 1216- 1226.
|
| 2 |
Kalchbrenner N, Elsen E, Simonyan K, et al. Efficient neural audio synthesis[C]//Proceedings of the 35th International Conference on Machine Learning. Stockholm, Sweden: PMLR, 2018: 2410-2419.
|
| 3 |
Tachibana H, Uenoyama K, Aihara S. Efficiently trainable text-to-speech system based on deep convolutional networks with guided attention[C]//Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Calgary, Canada: IEEE, 2018: 4784-4788.
|
| 4 |
Chen H, Garner P N. Diffusion transformer for adaptive text-to-speech[C]//Proceedings of the 12th ISCA Speech Synthesis Workshop (SSW 2023). Martigny, Switzerland: IDIAP Research Institute, 2023.
|
| 5 |
Shen J, Ping W, Yan R, et al. Natural TTS synthesis by conditioning WaveNet on Mel spectrogram predictions[C]//Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Calgary, Canada: IEEE, 2018: 4779-4783.
|
| 6 |
Mohammadi S H, Kain A. Voice conversion using deep neural networks with speaker-independent pre-training[C]//Proceedings of the IEEE Spoken Language Technology Workshop (SLT). 2014: 19-23.
|
| 7 |
Kameoka H, Kaneko T, Tanaka K, et al. StarGAN-VC: Non-parallel many-to-many voice conversion using star generative adversarial networks[C]//Proceedings of the 2018 IEEE Spoken Language Technology Workshop (SLT). Athens, Greece: IEEE, 2018: 266-273.
|
| 8 |
Guo H J, Liu C R, Ishi C T, et al. QuickVC: Any-to-many voice conversion using inverse short-time fourier transform for faster conversion [EB/OL]. (2023-06-30). https://arxiv.org/pdf/2302.08296v4.pdf.
|
| 9 |
Wang X, Yamagishi J, A comparative study on recent neural spoofing countermeasures for synthetic speech detection[C]//Proceedings of Interspeech 2021. Brno, Czech Republic: ISCA, 2021: 2396-2400.
|
| 10 |
Kumar A K, Paul D, Pal M, et al. Speech frame selection for spoofing detection with an application to partially spoofed audio-data[J]. International Journal of Speech Technology, 2021, 24 (1): 193- 203.
|
| 11 |
Chang C C, Lin C J. LBSVM: A library for support vector machines [J]. ACM Transactions on Intelligent Systems and Technology, 2011, 2 (3): 27.
|
| 12 |
Li M, Zhang X P, et al. A survey on speech deepfake detection[J]. ACM Computing Surveys, 2025, 57 (7)
|
| 13 |
Pham L, Lam P, Tran D, et al. A comprehensive survey with critical analysis for deepfake speech detection[J]. arXiv preprint, arXiv:, 2409, 15180, 2025.
|
| 14 |
Lei Z C, Yang Y G, Liu C H, et al. Siamese convolutional neural network using Gaussian probability feature for spoofing speech detection[C]//Proceedings of the 21st Interspeech Annual Conference of the International Speech Communication Association. Shanghai, China: ISCA, 2020: 1116-1120.
|
| 15 |
Rostami A M, Homayounpour M M, Nickabadi A. Efficient attention branch network with combined loss function for automatic speaker verification spoof detection[J]. Circuits, Systems, and Signal Processing, 2021, 42 (7): 4252- 4270.
|
| 16 |
Wang L, Yeoh B, Ng J W. Synthetic voice detection and audio splicing detection using SE-Res2Net-Conformer architecture[C]//Proceedings of the 13th International Symposium on Chinese Spoken Language Processing (ISCSLP). Singapore: IEEE, 2022: 115-119.
|
| 17 |
Li X F, Li K, Zheng Y F, et al. SafeEar: Content privacy-preserving audio deepfake detection[C]//Proceedings of the ACM SIGSAC Conference on Computer and Communications Security (CCS). 2024: arXiv: 2409.09272.
|
| 18 |
盖馨怡, 涂国庆, 刘树波, 等. 基于特征融合的音频伪造检测方法[J]. 计算机应用研究, 2025, 42 (7): 2109- 2115.
Gai X Y, Tu G Q, Liu S B, et al. Audio forgery detection method based on feature fusion[J]. Application Research of Computers, 2025, 42 (7): 2109- 2115.
|
| 19 |
Yadav A K S, Bartusiak E R, Bhagtani K, et al. Synthetic speech attribution using self supervised audio spectrogram transformer[J]. Electronic Imaging, 2023, 35, 1- 11.
|
| 20 |
Xie Y, Cheng H, Wang Y, et al. Domain generalization via aggregation and separation for audio deepfake detection[J]. IEEE Transactions on Information Forensics and Security, 2024, 19, 344- 358.
|
| 21 |
Wang C, He J, Yi J, et al. Multi-scale permutation entropy for audio deepfake detection[C]//International Conference on Acoustics, Speech, and Signal Processing. IEEE, 2024: 1406-1410.
|
| 22 |
Tak H, Patino J, Todisco M, et al. End-to-end anti-spoofing with RawNet2[C]//Proceedings of the International Conference on Acoustics, Speech and Signal Processing (ICASSP). 2021: 6369-6373.
|
| 23 |
Wang Z Y, Hansen J H L. Audio anti-spoofing using simple attention module and joint optimization based on additive angular margin loss and meta-learning[C]//Proceedings of the 23rd Interspeech Annual Conference of the International Speech Communication Association. Incheon, ISCA, 2022: 376-380.
|
| 24 |
Ito A, Horiguchi S. Spoofing attacker also benefits from self-supervised pretrained model[C]//Proceedings of Interspeech 2023. Dublin, Ireland: ISCA, 2023: 5346-5350.
|
| 25 |
Yuan H Y, Zhang L J, Niu B N, et al. A spoofing speech detection method combining multi-scale features and cross-layer information[J]. Information, 2025, 16 (3): 194.
|
| 26 |
Shin H, Heo J, Kim J, et al. Hm-conformer: A conformer-based audio deepfake detection system with hierarchical pooling and multi-level classification token aggregation methods[C]//International Conference on Acoustics, Speech, and Signal Processing, IEEE, 2024: 10581-10585.
|
| 27 |
Cuccovillo L, Gerhardt M, Aichroth P. Audio transformer for synthetic speech detection via formant magnitude and phase analysis[C]//International Conference on Acoustics, Speech, and Signal Processing. IEEE, 2024: 4805-4809.
|
| 28 |
Lu J, Zhang Y, Wang W, et al. One-class knowledge distillation for spoofing speech detection[C]//International Conference on Acoustics, Speech, and Signal Processing. IEEE, 2024: 11251-11255.
|
| 29 |
Wang X, Yamagishi J. Can large-scale vocoded spoofed data improve speech spoofing countermeasure with a self-supervised front end[C]//International Conference on Acoustics, Speech, and Signal Processing. IEEE, 2024: 10311-10315.
|
| 30 |
Ren Y, Peng H, Li L, et al. Lightweight voice spoofing detection using improved one-class learning and knowledge distillation[J]. IEEE Transactions on Multimedia, 2023, 26, 4360- 4374.
|
| 31 |
Lavrentyeva G, Novosyolov S, Tseren A, et al. STC anti-spoofing systems for the ASVspoof2019 challenge[C]//Proceedings of Interspeech, 2019: 1033-103.
|
| 32 |
Deng J, Ren Y, Zhang T, et al. VFD-Net: Vocoder fingerprints detection for fake audio[C]//International Conference on Acoustics, Speech, and Signal Processing. IEEE, 2024: 12151-12155.
|
| 33 |
Zhu Y, Koppisetti S, Tran T, et al. SLIM: Style-linguistics mismatch model for generalized audio deepfake detection[J]. arXiv preprint, arXiv:, 2407, 18517, 2024.
|
/
| 〈 |
|
〉 |