基于U-Net和离散小波变换的鲁棒音频水印方法
网络出版日期: 2025-08-20
基金资助
国家自然科学基金(61962019);湖北省高等学校优秀中青年科技创新团队计划项目(T2023013);湖北省自然科学基金联合基金重点项目(2023AFD061);湖北民族大学2025年研究生科研创新项目(MYK2025033)
版权
Robust audio watermarking method based on U-Net and discrete wavelet transform
Online published: 2025-08-20
Copyright
近年来,生成式人工智能技术迅速发展,使数字内容的生成和传播速度大幅提升。其中,高逼真的深度伪造音频给真伪媒体内容的识别带来了巨大挑战。音频水印能有效增强深度伪造音频溯源,保护数字内容的真实性和完整性。然而现有的音频水印方法主要使用传统技术,难以在复杂环境下保证鲁棒性。针对鲁棒音频水印,提出了一种基于U-Net和离散小波变换的鲁棒音频水印方法,利用了U-Net在多尺度特征提取及跳跃连接中的优势,以及离散小波变换在频率分量分解和低频特征稳定性的特点。通过联合训练的嵌入网络和提取网络分别实现了水印信息的嵌入与提取。在嵌入网络中,通过U-Net将水印信息嵌入到低频系数中;在提取网络中,解码器从低频系数中提取水印信息。同时,在迭代训练过程中引入噪声层以进一步提升鲁棒性。实验结果表明,该方法能有效地进行音频水印的嵌入和提取,并且具有良好的鲁棒性和不可感知性。
文双兵 , 张齐山 , 谭寒钟 , 胡涛 , 李军 . 基于U-Net和离散小波变换的鲁棒音频水印方法[J]. 网络空间安全科学学报, 2025 , 3(3) : 91 -101 . DOI: 10.20172/j.issn.2097-3136.250307
In recent years, the rapid advancement of generative artificial intelligence has significantly accelerated the generation and dissemination of digital content. Among these developments, highly realistic deepfake audio poses substantial challenges for identifying authentic media content. Audio watermarking can effectively enhance the traceability of deepfake audio, preserving the authenticity and integrity of digital content. However, existing audio watermarking methods primarily rely on traditional techniques, which struggle to maintain robustness in complex environments. Addressing the need for robust audio watermarking, this study proposes a method based on U-Net and Discrete Wavelet Transform (DWT). This approach leverages U-Net’s advantages in multi-scale feature extraction and skip connections, as well as the stability of DWT in frequency component decomposition and low-frequency features. The embedding and extraction of watermark information are achieved through jointly trained embedding and extraction networks. In the embedding network, U-Net embeds watermark information into low-frequency coefficients. In the extraction network, the decoder retrieves watermark information from these low-frequency coefficients. Additionally, a noise layer is introduced during iterative training to improve robustness further. Experimental results demonstrate that this method effectively embeds and extracts audio watermarks, achieving good robustness and imperceptibility.
表 1 本文方法与基准方法在FMA数据集上面对单一攻击的SNR与平均恢复位精度比较Table 1 Comparison of SNR and average bit recovery accuracy between the proposed method and the baseline method under single attacks on the FMA dataset |
| 攻击类型 | FDLM | FSVC | DeAR | UDAW (本文) |
| SNR | ||||
| AGWN15dB | ||||
| AGWN20dB | ||||
| AENca | ||||
| AENke | ||||
| AENla | ||||
| AENsi | ||||
| AENwa | ||||
| AENra | ||||
| AFD 0.5/L | ||||
| AFD 0.5/EX | ||||
| AFD 0.5/QS | ||||
| AFD 0.5/HS | ||||
| AFD 0.5/LG | ||||
| AEC 1K/0.5 | ||||
| AEC 2K/0.5 | ||||
| ATS 16k | ||||
| ATS 32k | ||||
| AVC 0.1 | ||||
| AVC 0.5 | ||||
| ARS | ||||
| AMF | ||||
| AER | ||||
| ABPF | ||||
| ALPF |
表 2 本文方法与基准方法在GTZAN和MagnaTagATune数据集上面对单一攻击的SNR与平均恢复位精度比较Table 2 Comparison of SNR and average bit recovery accuracy between the proposed method and baseline methods under single attacks on the GTZAN and MagnaTagATune datasets |
| 攻击类型 | GTZAN | MagnaTagATune | |||||||
| FDLM | FSVC | DeAR | UDAW (本文) | FDLM | FSVC | DeAR | UDAW (本文) | ||
| SNR | 34.167 7 | 28.817 8 | 22.451 6 | 26.667 0 | 35.575 8 | 36.142 5 | 25.142 7 | 28.635 0 | |
| AGWN15dB | 0.725 9 | 0.801 7 | 0.971 8 | 0.990 3 | 0.651 2 | 0.664 5 | 0.907 3 | 0.979 4 | |
| AGWN20dB | 0.749 2 | 0.844 2 | 0.991 4 | 0.997 1 | 0.680 2 | 0.700 1 | 0.954 4 | 0.992 9 | |
| AENca | 0.816 7 | 0.927 3 | 0.999 3 | 0.999 9 | 0.762 8 | 0.808 3 | 0.999 3 | 1.000 0 | |
| AENke | 0.810 9 | 0.820 3 | 1.000 0 | 0.998 4 | 0.766 1 | 0.735 8 | 1.000 0 | 0.996 7 | |
| AENla | 0.789 8 | 0.970 9 | 1.000 0 | 0.999 9 | 0.727 0 | 0.905 6 | 1.000 0 | 1.000 0 | |
| AENsi | 0.798 5 | 0.965 0 | 0.994 6 | 0.999 4 | 0.753 8 | 0.902 3 | 0.966 8 | 0.998 4 | |
| AENwa | 0.848 3 | 0.868 2 | 0.999 9 | 0.999 9 | 0.795 4 | 0.766 4 | 0.999 9 | 1.000 0 | |
| AENra | 0.825 5 | 0.825 9 | 0.999 9 | 0.999 9 | 0.778 9 | 0.688 5 | 0.999 9 | 1.000 0 | |
| AFD 0.5/L | 0.712 5 | 0.998 3 | 0.963 6 | 0.999 9 | 0.726 5 | 0.999 8 | 0.985 7 | 1.000 0 | |
| AFD 0.5/EX | 0.696 4 | 0.998 1 | 0.884 1 | 0.999 9 | 0.710 2 | 0.999 7 | 0.897 4 | 1.000 0 | |
| AFD 0.5/QS | 0.729 1 | 0.998 9 | 0.993 0 | 0.999 9 | 0.741 1 | 0.999 5 | 0.998 0 | 1.000 0 | |
| AFD 0.5/HS | 0.702 0 | 0.990 8 | 0.860 6 | 0.996 1 | 0.714 2 | 0.999 3 | 0.860 2 | 0.997 3 | |
| AFD 0.5/LG | 0.749 5 | 0.999 0 | 0.999 9 | 0.999 9 | 0.757 5 | 0.999 4 | 0.999 9 | 1.000 0 | |
| AEC 1K/0.5 | 0.999 2 | 0.999 6 | 0.999 9 | 0.999 9 | 0.999 3 | 0.999 7 | 0.999 9 | 1.000 0 | |
| AEC 2K/0.5 | 0.999 2 | 0.999 4 | 0.999 9 | 0.999 9 | 0.999 3 | 0.999 3 | 0.999 9 | 1.000 0 | |
| ATS 16k | 0.798 2 | 0.526 7 | 0.999 5 | 0.497 5 | 0.788 8 | 0.494 8 | 0.999 4 | 0.502 4 | |
| ATS 32k | 0.789 3 | 0.520 3 | 0.421 3 | 0.494 8 | 0.784 3 | 0.501 8 | 0.420 3 | 0.488 5 | |
| AVC 0.1 | 0.999 2 | 0.999 8 | 0.999 8 | 0.999 9 | 0.999 3 | 0.999 7 | 0.999 9 | 1.000 0 | |
| AVC 0.5 | 0.999 2 | 0.999 5 | 0.999 9 | 0.999 9 | 0.999 3 | 0.999 3 | 0.999 9 | 1.000 0 | |
| ARS | 0.997 6 | 0.999 5 | 0.999 9 | 0.999 9 | 0.584 7 | 0.480 7 | 0.999 9 | 1.000 0 | |
| AMF | 0.901 3 | 0.999 3 | 0.999 9 | 0.999 9 | 0.906 7 | 0.998 2 | 0.999 9 | 1.000 0 | |
| AER | 0.974 2 | 0.979 3 | 0.981 4 | 0.973 0 | 0.979 2 | 0.983 8 | 0.989 4 | 0.973 0 | |
| ABPF | 0.988 1 | 0.999 7 | 0.999 5 | 0.955 2 | 0.987 7 | 0.997 9 | 0.999 8 | 0.882 7 | |
| ALPF | 0.999 5 | 0.999 4 | 0.999 7 | 0.995 0 | 0.999 4 | 0.999 6 | 0.999 9 | 0.971 9 | |
表 3 本文方法与基准方法面对5种组合攻击的平均恢复位精度比较Table 3 Comparison of average bit recovery accuracy between the proposed method and baseline methods against five types of combined attacks |
| 方法 | AGWN20dB+AFD 0.5/HS | AGWN20dB+AER | ABPF+AER | AGWN20dB+ABPF | AGWN20dB+ABPF+AER |
| UDAW(本文) | 0.964 2 | 0.989 8 | 0.905 3 | 0.986 0 | 0.960 7 |
| FDLM | 0.666 3 | 0.751 9 | 0.948 9 | 0.751 9 | 0.741 6 |
| FSVC | 0.755 1 | 0.803 1 | 0.951 1 | 0.806 2 | 0.773 7 |
| DeAR | 0.708 2 | 0.954 8 | 0.987 4 | 0.942 3 | 0.922 7 |
表 4 UDAW与DeAR面对组合攻击的平均恢复位精度比较Table 4 Comparison of average bit recovery accuracy between UDAW and DeAR against combined attacks |
| 方法 | AGWN20dB | AFD 0.5/HS | AEC2K/0.5 | ABPF | AER | |||||||||
| UDAW(本文) | DeAR | UDAW(本文) | DeAR | UDAW(本文) | DeAR | UDAW(本文) | DeAR | UDAW(本文) | DeAR | |||||
| AGWN20dB | 0.989 6 | 0.953 4 | 0.989 5 | 0.776 3 | 0.994 5 | 0.955 4 | 0.995 1 | 0.959 4 | 0.980 2 | 0.920 5 | ||||
| AFD 0.5/HS | 0.991 3 | 0.769 2 | 0.999 2 | 0.975 4 | 0.998 8 | 0.911 1 | 0.990 1 | 0.882 9 | 0.984 6 | 0.897 4 | ||||
| AEC 2K/0.5 | 0.994 8 | 0.949 6 | 0.997 6 | 0.912 8 | 1.000 0 | 1.000 0 | 0.999 9 | 0.999 0 | 0.985 8 | 0.990 0 | ||||
| ABPF | 0.995 0 | 0.955 5 | 0.998 4 | 0.884 7 | 1.000 0 | 0.999 0 | 0.929 5 | 0.998 0 | 0.989 8 | 0.981 4 | ||||
| AER | 0.982 6 | 0.931 7 | 0.988 7 | 0.896 7 | 0.994 0 | 0.986 9 | 0.982 5 | 0.984 4 | 0.969 0 | 0.982 0 | ||||
表 5 本文方法与不同网络结构的SNR与平均恢复位精度比较Table 5 Comparison of SNR and average bit recovery accuracy of the proposed method with different network architectures |
| 方法 | CNN | ResNet | UDAW |
| SNR | 26.673 7 | 28.209 7 | 30.087 3 |
| AGWN20dB | 0.963 0 | 0.935 2 | 0.991 1 |
| AFD 0.5/ EX | 0.884 2 | 0.938 7 | 0.999 9 |
| AGWN+ABPF | 0.955 9 | 0.926 1 | 0.986 8 |
| AFD+AER | 0.870 3 | 0.926 2 | 0.975 6 |
| 1 |
LV Z. Generative artificial intelligence in the metaverse era[J]. Cognitive Robotics, 2023, 3, 208- 217.
|
| 2 |
RAMASAMY R, ARUMUGAM V. Digital watermarking-A tutorial[J]. IEEE Potentials, 2022, 41 (4): 43- 48.
|
| 3 |
UDDIN M S, HASAN M, SHIMAMURA T. Audio watermarking: A comprehensive review[J]. International Journal of Advanced Computer Science & Applications, 2024, 15(5): 1410-1418.
|
| 4 |
DEEBA F, KUN S, DHAREJO F A, et al. Digital image watermarking based on ANN and least significant bit[J]. Information Security Journal: A Global Perspective, 2020, 29 (1): 30- 39.
|
| 5 |
WANG S, YUAN W, UNOKI M. Multi-subspace echo hiding based on time-frequency similarities of audio signals[J]. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2020, 28, 2349- 2363.
|
| 6 |
NEJAD M Y, MOSLEH M, HEIKALABAD S R. A blind quantum audio watermarking based on quantum discrete cosine transform[J]. Journal of Information Security and Applications, 2020, 55, 102495.
|
| 7 |
ABODENA O, ALASHTIR A. High hiding capacity audio watermarking method based on discrete cosine transform[J]. Internation Journal of Advance Research and Innovative Ideas in Education, 2021, 7, 677- 684.
|
| 8 |
SALAH E, AMINE K, REDOUANE K, et al. A fourier transform based audio watermarking algorithm[J]. Applied Acoustics, 2021, 172, 107652.
|
| 9 |
TANG Y, WANG C, XIANG S, et al. A robust reversible watermarking scheme using attack-simulation-based adaptive normalization and embedding[J]. IEEE Transactions on Information Forensics and Security, 2024, 19, 4114- 4129.
|
| 10 |
ZHU J, KAPLAN R, JOHNSON J, et al. Hidden: Hiding data with deep networks[C]//Proceedings of the European Conference on Computer Vision (ECCV). Springer, 2018: 657-672.
|
| 11 |
MUN S M, NAM S H, JANG H, et al. Finding robust domain from attacks: A learning framework for blind watermarking[J]. Neurocomputing, 2019, 337, 191- 202.
|
| 12 |
AHMADI M, NOROUZI A, KARIMI N, et al. ReDMark: Framework for residual diffusion watermarking based on deep networks[J]. Expert Systems with Applications, 2020, 146, 113157.
|
| 13 |
LIU Y, GUO M, ZHANG J, et al. A novel two-stage separable deep learning framework for practical blind watermarking[C]//Proceedings of the 27th ACM International conference on multimedia. ACM, 2019: 1509-1517.
|
| 14 |
HUANG J, LUO T, LI L, et al. ARWGAN: Attention-guided robust image watermarking model based on GAN[J]. IEEE Transactions on Instrumentation and Measurement, 2023, 72, 1- 17.
|
| 15 |
ZHU L, FANG Y, ZHAO Y, et al. Lite localization network and DUE-based watermarking for color image copyright protection[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2024, 34 (10): 9311- 9325.
|
| 16 |
PAVLOVIĆ K, KOVAČEVIĆ S, DJUROVIĆ I, et al. Robust speech watermarking by a jointly trained embedder and detector using a DNN[J]. Digital Signal Processing, 2022, 122, 103381.
|
| 17 |
LIU C, ZHANG J, FANG H, et al. Dear: A deep-learning-based audio re-recording resilient watermarking[C]//Proceedings of the AAAI Conference on Artificial Intelligence. AAAI, 2023, 37(11): 13201-13209.
|
| 18 |
AZAD R, AGHDAM E K, RAULAND A, et al. Medical image segmentation review: The success of U-net[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46 (12): 10076- 10095.
|
| 19 |
DUAN X, JIA K, LI B, et al. Reversible image steganography scheme based on a U-Net structure[J]. IEEE Access, 2019, 7, 9314- 9323.
|
| 20 |
LIU L, MENG L, PENG Y, et al. A data hiding scheme based on U-Net and wavelet transform[J]. Knowledge-Based Systems, 2021, 223, 107022.
|
| 21 |
GUIMARÃES H R, NAGANO H, SILVA D W. Monaural speech enhancement through deep wave-U-net[J]. Expert Systems with Applications, 2020, 158, 113582.
|
| 22 |
LIU C, ZHANG J, ZHANG T, et al. Detecting voice cloning attacks via timbre watermarking[J]. arXiv preprint, arXiv: 2312. 03410.
|
| 23 |
CHEN G, WU Y, LIU S, et al. Wavmark: Watermarking for audio generation[J]. arXiv preprint, arXiv:, 2308, 12770, 2023.
|
| 24 |
LI B, CHEN J, XU Y, et al. DRAW: Dual-decoder-based robust audio watermarking against desynchronization and replay attacks[J]. IEEE Transactions on Information Forensics and Security, 2024, 19, 6529- 6544.
|
| 25 |
LIU W, LI Y, LIN D, et al. GROOT: Generating robust watermark for diffusion-model-based audio synthesis[C]//Proceedings of the 32nd ACM International Conference on Multimedia. ACM, 2024: 3294-3302.
|
| 26 |
RONNEBERGER O, FISCHER P, BROX T. U-net: Convolutional networks for biomedical image segmentation[C] //International Conference on Medical Image Computing and Computer-Assisted Intervention. Cham: Springer International Publishing, 2015: 234-241.
|
| 27 |
GURROLA-RAMOS J, DALMAU O, ALARCÓN T. U-Net based neural network for fringe pattern denoising[J]. Optics and Lasers in Engineering, 2022, 149, 106829.
|
| 28 |
ZHANG J, NIU Y, SHANGGUAN Z, et al. A novel denoising method for CT images based on U-net and multi-attention[J]. Computers in Biology and Medicine, 2023, 152, 106387.
|
| 29 |
KE J, WANG L. DF-UDetector: An effective method towards robust deepfake detection via feature restoration[J]. Neural Networks, 2023, 160, 216- 226.
|
| 30 |
ZHANG J, CHEN D, LIAO J, et al. Robust model watermarking for image processing networks via structure consistency[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46 (10): 6985- 6992.
|
| 31 |
LI Q, WANG X, MA B, et al. Concealed attack for robust watermarking based on generative model and perceptual loss[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2021, 32 (8): 5695- 5706.
|
| 32 |
ZHANG J, CHEN D, LIAO J, et al. Model watermarking for image processing networks[C]//Proceedings of the AAAI Conference on Artificial Intelligence. AAAI, 2020, 34(7): 12805-12812.
|
| 33 |
HUANG H N, CHEN S T, LIN M S, et al. Optimization-based embedding for wavelet-domain audio watermarking[J]. Journal of Signal Processing Systems, 2015, 80, 197- 208.
|
| 34 |
POURHASHEMI S M, MOSLEH M, ERFANI Y. A novel audio watermarking scheme using ensemble-based watermark detector and discrete wavelet transform[J]. Neural Computing and Applications, 2021, 33 (11): 6161- 6181.
|
| 35 |
DEFFERRARD M, BENZI K, VANDERGHEYNST P, et al. FMA: A dataset for music analysis[J]. arXiv preprint, arXiv:, 1612, 01840, 2016.
|
| 36 |
STURM B L. The GTZAN dataset: Its contents, its faults, their effects on evaluation, and its future use[J]. arXiv preprint, arXiv:, 1306, 1461, 2013.
|
| 37 |
LAW E, WEST K, MANDEL M I, et al. Evaluation of algorithms using games: The case of music tagging[C]//Proceedings of 10th International Society for Music Information Retrieval Conference (ISMIR 2009). International Society for Music Information Retrieval, 2009: 387-392.
|
| 38 |
LIU Z, HUANG Y, HUANG J. Patchwork-based audio watermarking robust against de-synchronization and recapturing attacks[J]. IEEE Transactions on Information Forensics and Security, 2018, 14 (5): 1171- 1180.
|
| 39 |
ZHAO J, ZONG T, XIANG Y, et al. Desynchronization attacks resilient watermarking method based on frequency singular value coefficient modification[J]. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2021, 29, 2282- 2295.
|
| 40 |
PICZAK K J. ESC: Dataset for environmental sound classification[C]//Proceedings of the 23rd ACM international conference on Multimedia. ACM, 2015: 1015-1018.
|
/
| 〈 |
|
〉 |