Generative image endogenous watermarking method based on exact diffusion inversion
Online published: 2024-05-18
Supported by
Natural Key R&D Program of China under Grant 2023YFF0905000, Natural Science Foundation of China under Grant U23B2023, 62371278,62302286,62376148, and China Postdoctoral Science Foundation under the grant No. 2023M742207
Copyright
The diffusion model has achieved significant success in image generation, but it is difficult to distinguish the authenticity of the generated images. Therefore, abusing the diffusion model will lead to social issues such as privacy and security, legal ethics, and so on. Adding watermarks to the output of the generated model can track the copyright of the generated content and prevent potential harm caused by artificial intelligence-generated content. For the diffusion model, the endogenous watermarking method of adding watermarks to the initial noise vector can directly generate watermarked images. During copyright verification, the initial vector is reconstructed through reverse diffusion to extract the watermark. However, the sampling process in the diffusion model is not strictly reversible, and there is a significant error between the reconstructed noise vector and the original noise, making it difficult to ensure accurate watermark extraction. By introducing Exact Diffusion Inversion via Coupled Transformations (EDICT), the initial noise vector can be reconstructed more accurately, improving the accuracy of watermark extraction. The performance improvement of generative image endogenous watermarking by introducing EDICT has been verified through experiments. The experimental results show that endogenous watermarking can embed invisible watermarks in generated images, and the embedded watermarks can be accurately extracted through precise backdiffusion and have a certain degree of robustness.
LI Li , ZHANG Xinpeng , WANG Zichi , WU Deyang , WU Hanzhou . Generative image endogenous watermarking method based on exact diffusion inversion[J]. Journal of Cybersecurity, 2024 , 2(1) : 92 -100 . DOI: 10.20172/j.issn.2097-3136.240108
表 1 不同嵌入量下的水印提取准确率Table 1 Accuracy of watermark extraction under different embedding payloads |
| 嵌入量/bit | 水印提取准确率 | TPR@1%FPR |
| 1 | 1.000 | 1.000 |
| 2 | 1.000 | 1.000 |
| 4 | 1.000 | 1.000 |
| 8 | 1.000 | 1.000 |
| 16 | 1.000 | 1.000 |
表 2 不同攻击下的水印提取准确率Table 2 Accuracy of watermark extraction under different attacks |
| 攻击方法 | 水印提取准确率 | TPR@1%FPR |
| 高斯噪声 | 1.000 | 1.000 |
| 旋转 | 0.990 | 0.990 |
| 裁剪10% | 0.960 | 0.980 |
| JPEG压缩(QF=85%) | 1.000 | 1.000 |
| 高斯噪声+旋转 | 0.980 | 0.980 |
| 旋转+裁剪10% | 0.960 | 0.970 |
| 旋转+压缩(QF=85%) | 0.970 | 0.980 |
| 高斯噪声+旋转+裁剪10% +压缩(QF=85%) | 0.940 | 0.950 |
表 3 不同方法的性能比较Table 3 Performance comparison of different methods |
| 方法 | 误比特率 | 水印提取 准确率 | TPR@ 1%FPR | FID 得分 | CLIP 得分 |
| DwtDct | 0.010 | 0.985 | 0.632 | 25.10 | 0.362 |
| DwtDctSvd | 0.009 | 1.000 | 1.000 | 25.01 | 0.359 |
| HiDDeN | 0.050 | 1.000 | 1.000 | 24.51 | 0.361 |
| Tree-ring | 0.180 | 1.000 | 1.000 | 25.93 | 0.364 |
| 本文 | 0.090 | 1.000 | 1.000 | 25.86 | 0.365 |
表 4 不同方法的鲁棒性比较Table 4 Robustness comparison of different methods |
| 方法 | 水印提取准确率 | |||
| 高斯噪声 | 旋转 | 裁剪10% | JPEG压缩 | |
| DwtDct | 1.000 | 0.830 | 0.800 | 0.810 |
| DwtDctSvd | 1.000 | 0.880 | 0.840 | 0.820 |
| HiDDeN | 1.000 | 0.990 | 0.970 | 0.960 |
| Tree-ring | 1.000 | 0.820 | 0.830 | 0.890 |
| 本文 | 1.000 | 0.990 | 0.960 | 1.000 |
表 5 DDIM和EDICT的水印提取误比特率 |
| 嵌入量/bit | 误比特率(原始水印) | |
| DDIM | EDICT | |
| 16 | 0.210 | 0.070 |
| 36 | 0.190 | 0.080 |
| 64 | 0.180 | 0.060 |
| 81 | 0.230 | 0.080 |
| 225 | 0.240 | 0.090 |
表 6 不同设置下的性能Table 6 Performances under different settings |
| 水印 半径 | 引导 强度 | 生成步数 | 水印提取 准确率 | TPR@ 1%FPR | FID 得分 | CLIP 得分 |
| 5 | 7.5 | 50 | 1.000 | 0.850 | 25.60 | 0.810 |
| 10 | 7.5 | 50 | 1.000 | 0.830 | 25.10 | 0.800 |
| 10 | 5.0 | 50 | 1.000 | 0.840 | 26.01 | 0.790 |
| 10 | 7.5 | 100 | 1.000 | 0.890 | 24.68 | 0.830 |
| 1 |
FATIMA N, IMRAN A S, KASTRATI Z, et al. A systematic literature review on text generation using deep neural network models[J]. IEEE Access, 2022, 10, 53490- 53503.
|
| 2 |
GUO B, WANG H, DING Y, et al. Conditional text generation for harmonious human-machine interaction[J]. ACM Transactions on Intelligent Systems and Technology, 2021, 12 (2): 1- 50.
|
| 3 |
ZHANG H, SONG H, LI S, et al. A survey of controllable text generation using transformer-based pre-trained language models[J]. ACM Computing Surveys, 2023, 56 (3): 1- 37.
|
| 4 |
LI Y,LIU H,WU Q,et al. Gligen:Open-set grounded text-to-image generation[C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,2023:22511-22521.
|
| 5 |
BIANCHI F,KALLURI P,DURMUS E,et al. Easily accessible text-to-image generation amplifies demographic stereotypes at large scale[C]//Proceedings of the 2023 ACM Conference on Fairness,Accountability and Transparency,2023:1493-1504.
|
| 6 |
ZHANG J,CHEN K,QIN C,et al. AAS:Automatic virtual data augmentation for deep image steganalysis[J]. IEEE Transactions on Dependable and Secure Computing,2023,doi:10.1109/TDSC.2023.3333913.
|
| 7 |
YANG D, YU J, WANG H, et al. Diffsound: Discrete diffusion model for text-to-sound generation[J]. IEEE/ACM Transactions on Audio, Speech and Language Processing,, 2023, 31, 1720- 1733.
|
| 8 |
BORSOS Z, MARINIER R, VINCENT D, et al. Audiolm: A language modeling approach to audio generation[J]. IEEE/ACM Transactions on Audio, Speech and Language Processing,, 2023, 31, 2523- 2533.
|
| 9 |
ESSER P,CHIU J,ATIGHEHCHIAN P,et al. Structure and content-guided video synthesis with diffusion models[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision,2023:7346-7356.
|
| 10 |
WU J Z,GE Y,WANG X,et al. Tune-a-video:One-shot tuning of image diffusion models for text-to-video generation[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision,2023:7623-7633.
|
| 11 |
KARNEWAR A,MITRA N J,VEDALDI A,et al. Holofusion:Towards photo-realistic 3D generative modeling[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision,2023:22976-22985.
|
| 12 |
WANG H,DU X,LI J,et al. Score jacobian chaining:Lifting pretrained 2D diffusion models for 3d generation[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,2023:12619-12629.
|
| 13 |
FROMER J C, COLEY C W. Computer-aided multi-objective optimization in small molecule discovery[J]. Patterns, 2023, 4 (2): 1- 17.
|
| 14 |
VASWANI A,SHAZEER N,PARMAR N,et al. Attention is all you need[J]. Advances in Neural Information Processing Systems,2017,30:1-11.
|
| 15 |
BROWN T, MANN B, RYDER N, et al. Language models are few-shot learners[J]. Advances in Neural Information Processing Systems, 2020, 33, 1877- 1901.
|
| 16 |
RAY P P. ChatGPT:a comprehensive review on background,applications,key challenges,bias,ethics,limitations and future scope[J]. Internet of Things and Cyber-Physical Systems,2023,doi:10.1016/j.iotcps.2023.04.003.
|
| 17 |
LI L,ZHANG W,BARNI M. Covert task embedding:Turning a DNN into an insider agent leaking out private information[J]. IEEE Transactions on Neural Networks and Learning Systems,2022,doi:10.1109/TNNLS.2022.3216010.
|
| 18 |
FANG H, QIU Y, QIN G, et al. DP2: Dataset protection by data poisoning[J]. IEEE Transactions on Dependable and Secure Computing, 2024, 21 (2): 636- 649.
|
| 19 |
ZHAO Y, PANG T, DU C, et al. A recipe for watermarking diffusion models[J]. arXiv preprint arXiv:, 2303, 10137, 2023.
|
| 20 |
FERNANDEZ P, COUAIRON G, JÉGOU H, et al. The Stable Signature: rooting watermarks in latent diffusion models[J]. arXiv preprint arXiv:, 2303, 15435, 2023.
|
| 21 |
XIONG C,QIN C,FENG G,et al. Flexible and secure watermarking for latent diffusion model[C]//Proceedings of the 31st ACM International Conference on Multimedia,2023:1668-1676.
|
| 22 |
WEN Y, KIRCHENBAUER J, GEIPING J, et al. Tree-ring watermarks: fingerprints for diffusion images that are invisible and robust[J]. arXiv preprint arXiv:, 2305, 20030, 2023.
|
| 23 |
ROMBACH R,BLATTMANN A,LORENZ D,et al. High-resolution image synthesis with latent diffusion models[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,2022:10684-10695.
|
| 24 |
HO J, JAIN A, ABBEEL P. Denoising diffusion probabilistic models[J]. Advances in Neural Information Processing Systems, 2020, 33, 6840- 6851.
|
| 25 |
SONG J, MENG C, ERMON S. Denoising diffusion implicit models[J]. arXiv preprint arXiv:, 2010, 02502, 2020.
|
| 26 |
BEGUM M, UDDIN M S. Digital image watermarking techniques: A review[J]. Information, 2020, 11 (2): 110.
|
| 27 |
GHOUTI L, BOURIDANE A, IBRAHIM M K, et al. Digital image watermarking using balanced multiwavelets[J]. IEEE Transactions on Signal Processing, 2006, 54 (4): 1519- 1536.
|
| 28 |
MAHTO D K, SINGH A K. A survey of color image watermarking: State-of-the-art and research directions[J]. Computers & Electrical Engineering, 2021, 93, 107255.
|
| 29 |
WAN W, WANG J, ZHANG Y, et al. A comprehensive survey on robust image watermarking[J]. Neurocomputing, 2022, 488, 226- 247.
|
| 30 |
ZHU T, QU W, CAO W. An optimized image watermarking algorithm based on SVD and IWT[J]. The Journal of Supercomputing, 2022, 78 (1): 222- 237.
|
| 31 |
WU D, LI L, WANG J, et al. Robust zero-watermarking scheme using DTCWT and improved differential entropy for color medical images[J]. Journal of King Saud University-Computer and Information Sciences, 2023, 35 (8): 101708.
|
| 32 |
ZHU J,KAPLAN R,JOHNSON J,et al. Hidden:Hiding data with deep networks[C]//Proceedings of the European Conference on Computer Vision (ECCV),2018:657-672.
|
| 33 |
MELLIMI S, RAJPUT V, ANSARI I A, et al. A fast and efficient image watermarking scheme based on deep neural network[J]. Pattern Recognition Letters, 2021, 151, 222- 228.
|
| 34 |
ZHONG X, HUANG P C, MASTORAKIS S, et al. An automated and robust image watermarking scheme based on deep neural networks[J]. IEEE Transactions on Multimedia, 2020, 23, 1951- 1961.
|
| 35 |
LIU G,SI Y,QIAN Z,et al. WRAP:Watermarking approach robust against film-coating upon printed photographs[C]//Proceedings of the 31st ACM International Conference on Multimedia,2023:7274-7282.
|
| 36 |
FANG H,CHEN K,QIU Y,et al. DeNoL:A few-shot-sample-based decoupling noise layer for cross-channel watermarking robustness[C]//Proceedings of the 31st ACM International Conference on Multimedia,2023:7345-7353.
|
| 37 |
LI L, ZHANG W, BARNI M. Universal blackmarks: Key-image-free blackbox multi-bit watermarking of deep neural networks[J]. IEEE Signal Processing Letters, 2023, 30, 36- 40.
|
/
| 〈 |
|
〉 |