基于精确扩散反演的生成式图像内生水印方法
网络出版日期: 2024-05-18
基金资助
国家重点研发计划(2023YFF0905000);国家自然科学基金(U23B2023,62371278,62302286,62376148);中国博士后科学基金面上项目(2023M742207)
版权
Generative image endogenous watermarking method based on exact diffusion inversion
Online published: 2024-05-18
Supported by
Natural Key R&D Program of China under Grant 2023YFF0905000, Natural Science Foundation of China under Grant U23B2023, 62371278,62302286,62376148, and China Postdoctoral Science Foundation under the grant No. 2023M742207
Copyright
扩散模型在图像生成方面取得了显著成就,但生成的图像真假难辨,因此滥用扩散模型将引发隐私安全、法律伦理等社会问题。对生成模型的输出添加水印可以追踪生成内容版权,防止人工智能生成内容造成潜在危害。对于去噪扩散模型,在初始噪声向量中添加水印的内生水印方法可直接生成含水印图像,版权验证时通过反向扩散重建初始向量以提取水印。但扩散模型中的采样过程并不是严格可逆,重建的噪声向量与原始噪声存在较大误差,很难保证水印的准确提取。通过引入基于耦合变换的精确反向扩散,可以更加准确地重建初始噪声向量,提升水印提取的准确性。通过实验验证了引入基于耦合变换的精确反向扩散对于生成式图像内生水印的性能提升,实验结果表明,内生水印可以在生成图像中嵌入不可见水印,嵌入的水印可通过精确反向扩散被准确提取,并具有一定的稳健性。
李莉 , 张新鹏 , 王子驰 , 吴德阳 , 吴汉舟 . 基于精确扩散反演的生成式图像内生水印方法[J]. 网络空间安全科学学报, 2024 , 2(1) : 92 -100 . DOI: 10.20172/j.issn.2097-3136.240108
The diffusion model has achieved significant success in image generation, but it is difficult to distinguish the authenticity of the generated images. Therefore, abusing the diffusion model will lead to social issues such as privacy and security, legal ethics, and so on. Adding watermarks to the output of the generated model can track the copyright of the generated content and prevent potential harm caused by artificial intelligence-generated content. For the diffusion model, the endogenous watermarking method of adding watermarks to the initial noise vector can directly generate watermarked images. During copyright verification, the initial vector is reconstructed through reverse diffusion to extract the watermark. However, the sampling process in the diffusion model is not strictly reversible, and there is a significant error between the reconstructed noise vector and the original noise, making it difficult to ensure accurate watermark extraction. By introducing Exact Diffusion Inversion via Coupled Transformations (EDICT), the initial noise vector can be reconstructed more accurately, improving the accuracy of watermark extraction. The performance improvement of generative image endogenous watermarking by introducing EDICT has been verified through experiments. The experimental results show that endogenous watermarking can embed invisible watermarks in generated images, and the embedded watermarks can be accurately extracted through precise backdiffusion and have a certain degree of robustness.
表 1 不同嵌入量下的水印提取准确率Table 1 Accuracy of watermark extraction under different embedding payloads |
| 嵌入量/bit | 水印提取准确率 | TPR@1%FPR |
| 1 | 1.000 | 1.000 |
| 2 | 1.000 | 1.000 |
| 4 | 1.000 | 1.000 |
| 8 | 1.000 | 1.000 |
| 16 | 1.000 | 1.000 |
表 2 不同攻击下的水印提取准确率Table 2 Accuracy of watermark extraction under different attacks |
| 攻击方法 | 水印提取准确率 | TPR@1%FPR |
| 高斯噪声 | 1.000 | 1.000 |
| 旋转 | 0.990 | 0.990 |
| 裁剪10% | 0.960 | 0.980 |
| JPEG压缩(QF=85%) | 1.000 | 1.000 |
| 高斯噪声+旋转 | 0.980 | 0.980 |
| 旋转+裁剪10% | 0.960 | 0.970 |
| 旋转+压缩(QF=85%) | 0.970 | 0.980 |
| 高斯噪声+旋转+裁剪10% +压缩(QF=85%) | 0.940 | 0.950 |
表 3 不同方法的性能比较Table 3 Performance comparison of different methods |
| 方法 | 误比特率 | 水印提取 准确率 | TPR@ 1%FPR | FID 得分 | CLIP 得分 |
| DwtDct | 0.010 | 0.985 | 0.632 | 25.10 | 0.362 |
| DwtDctSvd | 0.009 | 1.000 | 1.000 | 25.01 | 0.359 |
| HiDDeN | 0.050 | 1.000 | 1.000 | 24.51 | 0.361 |
| Tree-ring | 0.180 | 1.000 | 1.000 | 25.93 | 0.364 |
| 本文 | 0.090 | 1.000 | 1.000 | 25.86 | 0.365 |
表 4 不同方法的鲁棒性比较Table 4 Robustness comparison of different methods |
| 方法 | 水印提取准确率 | |||
| 高斯噪声 | 旋转 | 裁剪10% | JPEG压缩 | |
| DwtDct | 1.000 | 0.830 | 0.800 | 0.810 |
| DwtDctSvd | 1.000 | 0.880 | 0.840 | 0.820 |
| HiDDeN | 1.000 | 0.990 | 0.970 | 0.960 |
| Tree-ring | 1.000 | 0.820 | 0.830 | 0.890 |
| 本文 | 1.000 | 0.990 | 0.960 | 1.000 |
表 5 DDIM和EDICT的水印提取误比特率 |
| 嵌入量/bit | 误比特率(原始水印) | |
| DDIM | EDICT | |
| 16 | 0.210 | 0.070 |
| 36 | 0.190 | 0.080 |
| 64 | 0.180 | 0.060 |
| 81 | 0.230 | 0.080 |
| 225 | 0.240 | 0.090 |
表 6 不同设置下的性能Table 6 Performances under different settings |
| 水印 半径 | 引导 强度 | 生成步数 | 水印提取 准确率 | TPR@ 1%FPR | FID 得分 | CLIP 得分 |
| 5 | 7.5 | 50 | 1.000 | 0.850 | 25.60 | 0.810 |
| 10 | 7.5 | 50 | 1.000 | 0.830 | 25.10 | 0.800 |
| 10 | 5.0 | 50 | 1.000 | 0.840 | 26.01 | 0.790 |
| 10 | 7.5 | 100 | 1.000 | 0.890 | 24.68 | 0.830 |
| 1 |
FATIMA N, IMRAN A S, KASTRATI Z, et al. A systematic literature review on text generation using deep neural network models[J]. IEEE Access, 2022, 10, 53490- 53503.
|
| 2 |
GUO B, WANG H, DING Y, et al. Conditional text generation for harmonious human-machine interaction[J]. ACM Transactions on Intelligent Systems and Technology, 2021, 12 (2): 1- 50.
|
| 3 |
ZHANG H, SONG H, LI S, et al. A survey of controllable text generation using transformer-based pre-trained language models[J]. ACM Computing Surveys, 2023, 56 (3): 1- 37.
|
| 4 |
LI Y,LIU H,WU Q,et al. Gligen:Open-set grounded text-to-image generation[C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,2023:22511-22521.
|
| 5 |
BIANCHI F,KALLURI P,DURMUS E,et al. Easily accessible text-to-image generation amplifies demographic stereotypes at large scale[C]//Proceedings of the 2023 ACM Conference on Fairness,Accountability and Transparency,2023:1493-1504.
|
| 6 |
ZHANG J,CHEN K,QIN C,et al. AAS:Automatic virtual data augmentation for deep image steganalysis[J]. IEEE Transactions on Dependable and Secure Computing,2023,doi:10.1109/TDSC.2023.3333913.
|
| 7 |
YANG D, YU J, WANG H, et al. Diffsound: Discrete diffusion model for text-to-sound generation[J]. IEEE/ACM Transactions on Audio, Speech and Language Processing,, 2023, 31, 1720- 1733.
|
| 8 |
BORSOS Z, MARINIER R, VINCENT D, et al. Audiolm: A language modeling approach to audio generation[J]. IEEE/ACM Transactions on Audio, Speech and Language Processing,, 2023, 31, 2523- 2533.
|
| 9 |
ESSER P,CHIU J,ATIGHEHCHIAN P,et al. Structure and content-guided video synthesis with diffusion models[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision,2023:7346-7356.
|
| 10 |
WU J Z,GE Y,WANG X,et al. Tune-a-video:One-shot tuning of image diffusion models for text-to-video generation[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision,2023:7623-7633.
|
| 11 |
KARNEWAR A,MITRA N J,VEDALDI A,et al. Holofusion:Towards photo-realistic 3D generative modeling[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision,2023:22976-22985.
|
| 12 |
WANG H,DU X,LI J,et al. Score jacobian chaining:Lifting pretrained 2D diffusion models for 3d generation[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,2023:12619-12629.
|
| 13 |
FROMER J C, COLEY C W. Computer-aided multi-objective optimization in small molecule discovery[J]. Patterns, 2023, 4 (2): 1- 17.
|
| 14 |
VASWANI A,SHAZEER N,PARMAR N,et al. Attention is all you need[J]. Advances in Neural Information Processing Systems,2017,30:1-11.
|
| 15 |
BROWN T, MANN B, RYDER N, et al. Language models are few-shot learners[J]. Advances in Neural Information Processing Systems, 2020, 33, 1877- 1901.
|
| 16 |
RAY P P. ChatGPT:a comprehensive review on background,applications,key challenges,bias,ethics,limitations and future scope[J]. Internet of Things and Cyber-Physical Systems,2023,doi:10.1016/j.iotcps.2023.04.003.
|
| 17 |
LI L,ZHANG W,BARNI M. Covert task embedding:Turning a DNN into an insider agent leaking out private information[J]. IEEE Transactions on Neural Networks and Learning Systems,2022,doi:10.1109/TNNLS.2022.3216010.
|
| 18 |
FANG H, QIU Y, QIN G, et al. DP2: Dataset protection by data poisoning[J]. IEEE Transactions on Dependable and Secure Computing, 2024, 21 (2): 636- 649.
|
| 19 |
ZHAO Y, PANG T, DU C, et al. A recipe for watermarking diffusion models[J]. arXiv preprint arXiv:, 2303, 10137, 2023.
|
| 20 |
FERNANDEZ P, COUAIRON G, JÉGOU H, et al. The Stable Signature: rooting watermarks in latent diffusion models[J]. arXiv preprint arXiv:, 2303, 15435, 2023.
|
| 21 |
XIONG C,QIN C,FENG G,et al. Flexible and secure watermarking for latent diffusion model[C]//Proceedings of the 31st ACM International Conference on Multimedia,2023:1668-1676.
|
| 22 |
WEN Y, KIRCHENBAUER J, GEIPING J, et al. Tree-ring watermarks: fingerprints for diffusion images that are invisible and robust[J]. arXiv preprint arXiv:, 2305, 20030, 2023.
|
| 23 |
ROMBACH R,BLATTMANN A,LORENZ D,et al. High-resolution image synthesis with latent diffusion models[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,2022:10684-10695.
|
| 24 |
HO J, JAIN A, ABBEEL P. Denoising diffusion probabilistic models[J]. Advances in Neural Information Processing Systems, 2020, 33, 6840- 6851.
|
| 25 |
SONG J, MENG C, ERMON S. Denoising diffusion implicit models[J]. arXiv preprint arXiv:, 2010, 02502, 2020.
|
| 26 |
BEGUM M, UDDIN M S. Digital image watermarking techniques: A review[J]. Information, 2020, 11 (2): 110.
|
| 27 |
GHOUTI L, BOURIDANE A, IBRAHIM M K, et al. Digital image watermarking using balanced multiwavelets[J]. IEEE Transactions on Signal Processing, 2006, 54 (4): 1519- 1536.
|
| 28 |
MAHTO D K, SINGH A K. A survey of color image watermarking: State-of-the-art and research directions[J]. Computers & Electrical Engineering, 2021, 93, 107255.
|
| 29 |
WAN W, WANG J, ZHANG Y, et al. A comprehensive survey on robust image watermarking[J]. Neurocomputing, 2022, 488, 226- 247.
|
| 30 |
ZHU T, QU W, CAO W. An optimized image watermarking algorithm based on SVD and IWT[J]. The Journal of Supercomputing, 2022, 78 (1): 222- 237.
|
| 31 |
WU D, LI L, WANG J, et al. Robust zero-watermarking scheme using DTCWT and improved differential entropy for color medical images[J]. Journal of King Saud University-Computer and Information Sciences, 2023, 35 (8): 101708.
|
| 32 |
ZHU J,KAPLAN R,JOHNSON J,et al. Hidden:Hiding data with deep networks[C]//Proceedings of the European Conference on Computer Vision (ECCV),2018:657-672.
|
| 33 |
MELLIMI S, RAJPUT V, ANSARI I A, et al. A fast and efficient image watermarking scheme based on deep neural network[J]. Pattern Recognition Letters, 2021, 151, 222- 228.
|
| 34 |
ZHONG X, HUANG P C, MASTORAKIS S, et al. An automated and robust image watermarking scheme based on deep neural networks[J]. IEEE Transactions on Multimedia, 2020, 23, 1951- 1961.
|
| 35 |
LIU G,SI Y,QIAN Z,et al. WRAP:Watermarking approach robust against film-coating upon printed photographs[C]//Proceedings of the 31st ACM International Conference on Multimedia,2023:7274-7282.
|
| 36 |
FANG H,CHEN K,QIU Y,et al. DeNoL:A few-shot-sample-based decoupling noise layer for cross-channel watermarking robustness[C]//Proceedings of the 31st ACM International Conference on Multimedia,2023:7345-7353.
|
| 37 |
LI L, ZHANG W, BARNI M. Universal blackmarks: Key-image-free blackbox multi-bit watermarking of deep neural networks[J]. IEEE Signal Processing Letters, 2023, 30, 36- 40.
|
/
| 〈 |
|
〉 |