Robust model watermarking based on adversarial simulation
Online published: 2025-01-25
Supported by
National Natural Science Foundation of China (No. 62261160653, No. U23A20305, No. 62172435)、Henan Academy of Sciences Science and Technology Open Cooperation Project (220907015)
Copyright
The development of the Internet of Things (IoT) has significantly improved people’s lives and productivity. In this process, deep neural network models play a crucial role in data processing and intelligence within IoT. To prevent unauthorized use of models, model watermarking has been emerged as an effective means of copyright protection. Model owners can embed specific watermark behaviors into the models before release and detect the watermark behaviors to identify potential pirated models. However, adversaries can use low-cost methods to remove watermarks with minimal impact on model performance, thus evading copyright verification. To address this problem, an innovative robust model watermarking method based on adversarial simulation was proposed. The method optimized a set of watermark samples to ensure that the watermark samples could trigger the watermark behaviors even after undergoing watermark removal attacks. Specifically, by analyzing the common characteristics of watermark removal attacks, a watermark removal simulator was constructed to mimic these attacks and a clean model simulator was constructed to emulate the model’s performance without watermarks. These simulators were used together to guide the optimization of the watermark samples. Experiments were conducted on CIFAR-10 and CIFAR-100 datasets. The results show that the proposed robust model watermarking method exhibits strong resistance to various watermark removal attacks, demonstrating its effectiveness and practicality.
XI Zuping , QU Zuomin , LU Wei , ZHANG Wei , LUO Xiangyang , XIAO Hongtao . Robust model watermarking based on adversarial simulation[J]. Journal of Cybersecurity, 2024 , 2(5) : 67 -77 . DOI: 10.20172/j.issn.2097-3136.240506
表 1 不同水印方法的保真性与有效性评估Table 1 Fidelity and effectiveness evaluation of different watermarking methods |
| 数据集 | 方法 | 干净模型 | 含水印模型 | |||
| CIFAR10 | Content | 92.23 | 0.52 | 92.34 | 100.00 | |
| RN | 1.56 | 92.29 | 100.00 | |||
| Unrelated | 0.00 | 92.52 | 100.00 | |||
| 本文 | 13.54 | 91.38 | 100.00 | |||
| CIFAR100 | Content | 70.35 | 0.00 | 69.94 | 100.00 | |
| RN | 0.00 | 70.20 | 100.00 | |||
| Unrelated | 0.00 | 70.23 | 100.00 | |||
| 本文 | 5.98 | 70.47 | 100.00 | |||
表 2 不同水印方法抵抗水印移除攻击的鲁棒性评估Table 2 Robustness evaluation of different watermarking methods against watermark removal attacks |
| 数据集 | 方法 | 平均下降 率(%) | ||||
| FT | FP | NAD | ANP | |||
| CIFAR10 | Content | 23.96 | 45.83 | 23.00 | 0.60 | 76.65 |
| RN | 59.89 | 82.81 | 64.8 | 41.4 | 37.78 | |
| Unrelated | 30.73 | 64.84 | 35.40 | 0.80 | 67.06 | |
| 本文 | 90.89 | 92.19 | 87.15 | 64.48 | 16.32 | |
| CIFAR100 | Content | 42.71 | 5.2 | 25.37 | 46.80 | 69.98 |
| RN | 44.79 | 17.19 | 26.53 | 93.00 | 54.62 | |
| Unrelated | 53.91 | 0.26 | 5.4 | 72.60 | 66.96 | |
| 本文 | 76.82 | 32.03 | 61.76 | 82.77 | 36.66 | |
表 3 水印移除模型上的测试集准确率Table 3 Test set accuracy on watermark-removed models |
| 数据集 | 方法 | ||||
| FT | FP | NAD | ANP | ||
| CIFAR10 | Content | 91.64 | 92.40 | 90.70 | 86.87 |
| RN | 92.20 | 92.45 | 90.98 | 61.92 | |
| Unrelated | 91.97 | 92.60 | 90.50 | 87.54 | |
| 本文 | 91.73 | 91.95 | 89.67 | 87.44 | |
| CIFAR100 | Content | 69.19 | 67.70 | 63.72 | 60.14 |
| RN | 69.13 | 67.50 | 63.73 | 60.66 | |
| Unrelated | 68.94 | 66.70 | 65.33 | 60.38 | |
| 本文 | 68.21 | 66.50 | 63.25 | 61.01 | |
表 4 不同扰动幅度预算 |
| 含水印模型 | 水印移除模型 | 平均下降 率(%) | ||||
| FT | FP | NAD | ANP | |||
| 0.01 | 100 | 57.29 | 85.16 | 72.20 | 40.25 | 36.28 |
| 0.02 | 100 | 90.89 | 92.19 | 87.15 | 64.48 | 16.32 |
| 0.03 | 100 | 91.15 | 94.01 | 82.80 | 75.35 | 14.17 |
| 0.04 | 100 | 93.49 | 88.28 | 66.53 | 92.02 | 14.92 |
表 5 不同扰动幅度预算 |
| 含水印模型 | 水印移除模型 | ||||
| FT | FP | NAD | ANP | ||
| 0.01 | 91.22 | 91.54 | 92.25 | 89.26 | 87.64 |
| 0.02 | 91.38 | 91.73 | 91.95 | 89.67 | 87.44 |
| 0.03 | 91.21 | 91.43 | 91.60 | 89.38 | 86.96 |
| 0.04 | 91.46 | 91.59 | 91.65 | 89.73 | 87.77 |
表 6 不同网络结构下本文水印方法的有效性、保真性及抵抗水印移除攻击的鲁棒性评估Table 6 Effectiveness, fidelity, and robustness evaluations against watermark removal attacks of this watermark method under different network architectures |
| 网络结构 | 方法 | 含水印模型 | FT | ||
| SENet | Content | 92.31 | 100 | 91.99 | 19.79 |
| RN | 92.33 | 100 | 92.17 | 48.43 | |
| Unrelated | 92.17 | 99.74 | 91.67 | 42.19 | |
| 本文 | 91.42 | 100 | 91.50 | 83.07 | |
| MobileNetV2 | Content | 91.08 | 100 | 90.86 | 12.24 |
| RN | 91.24 | 100 | 90.72 | 34.75 | |
| Unrelated | 91.48 | 100 | 90.17 | 22.13 | |
| 本文 | 90.72 | 100 | 90.28 | 71.61 | |
| 1 |
HINTON G, DENG L, YU D, et al. Deep neural networks for acoustic modeling in speech recognition: the shared views of four research groups[J]. IEEE Signal Processing Magazine, 2012, 29 (6): 82- 97.
|
| 2 |
HE K,ZHANG X,REN S,et al. Deep residual learning for image recognition[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 2016:770-778.
|
| 3 |
吴蜜. 人工智能与物联网技术在智慧城市中的应用[J]. 集成电路应用, 2024, 41 (2): 362- 364.
WU M. Application of artificial intelligence and Internet of Things technology in smart cities[J]. Application of IC, 2024, 41 (2): 362- 364.
|
| 4 |
柴天佑. 工业人工智能与工业互联网协同实现生产过程智能化及其未来展望[J]. 控制工程, 2023, 30 (8): 1378- 1388.
CHAI T Y. Industrial AI and industrial internet collaboratively achieving production process intelligence and its future perspectives[J]. Control Engineering of China, 2023, 30 (8): 1378- 1388.
|
| 5 |
前瞻产业研究院. 中国物联网行业应用领域市场需求与投资预测分析报告[EB/OL]. (2023-12-05)[2024-03-15]. https://bg.qianzhan.com/report/detail/300/231205-fc0151e7.html.
Prospective Industry Research Institute. Report of application tield market demand and investment forecast on China Internet of Things industry[EB/OL]. (2023-12-05)[2024-03-15]. https://bg.qianzhan.com/report/detail/300/231205-fc0151e7.html.
|
| 6 |
吴汉舟, 张杰, 李越, 等. 人工智能模型水印研究进展[J]. 中国图象图形学报, 2023, 28 (6): 1792- 1810.
WU H Z, ZHANG J, LI Y, et al. Overview of artificial intelligence model watermarking[J]. Journal of Image and Graphics, 2023, 28 (6): 1792- 1810.
|
| 7 |
王馨雅, 华光, 江昊, 等. 深度学习模型的版权保护研究综述[J]. 网络与信息安全学报, 2022, 8 (2): 1- 14.
WANG X Y, HUA G, JIANG H, et al. survey on intellectual property protection for deep learning model[J]. Chinese Journal of Network and Information Security, 2022, 8 (2): 1- 14.
|
| 8 |
CHIDA Y,NAGAI Y,SAKAZAWA S,et al. Embedding watermarks into deep neural networks[C]//Proceedings of the ACM on International Conference on Multimedia Retrieval. ACM, 2017:269-277.
|
| 9 |
LIU H,WENG Z,ZHU Y. Watermarking deep neural networks with Greedy residuals[C]// Proceedings of the International Conference on Machine Learning. 2021:6978-6988.
|
| 10 |
张准, 李佳睿, 岳鹏, 等. 基于后门水印的联邦模型授权方案[J]. 网络空间安全科学学报, 2024, 2 (1): 113- 122.
ZHANG Z, LI J R, YUE P, et al. Federated model authorization scheme based on backdoor watermarking[J]. Journal of Cybersecurity, 2024, 2 (1): 113- 122.
|
| 11 |
陈可江, 李帅, 张卫明, 等. 基于知识注入的大语言模型水印[J]. 网络空间安全科学学报, 2024, 2 (1): 63- 71.
CHEN K J, LI S, ZHANG W M, et al. Watermarking for large language models based on knowledge injection[J]. Journal of Cybersecurity, 2024, 2 (1): 63- 71.
|
| 12 |
ZHANG J,GU Z,JANG J,et al. Protecting intellectual property of deep neural networks with watermarking[C]//Proceedings of the 2018 on Asia Conference on Computer and Communications Security. 2018:159-172.
|
| 13 |
LUKAS N,JIANG E,LI X,et al. Sok:how robust is image classification deep neural network watermarking?[C]//2022 IEEE Symposium on Security and Privacy (SP). IEEE,2022:787-804.
|
| 14 |
SHAFIEINEJAD M,LUKAS N,WANG J,et al. On the robustness of backdoor-based watermarking in deep neural networks[C]//Proceedings of the 2021 ACM Workshop on Information Hiding and Multimedia Security. ACM, 2021:177-188.
|
| 15 |
ZHAO P, CHEN P Y, DAS P, et al. Bridging mode connectivity in loss landscapes and adversarial robustness[J]. arXiv preprint, arXiv:, 2005, 00060, 2020.
|
| 16 |
GAN G,LI Y,WU D,et al. Towards robust model watermark via reducing parametric vulnerability[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. IEEE, CVF, 2023:4751-4761.
|
| 17 |
LI Y, LYU X, KOREN N, et al. Neural attention distillation: Erasing backdoor triggers from deep neural networks[J]. arXiv preprint, arXiv:, 2101, 05930, 2021.
|
| 18 |
LIU K,DOLAN-GAVITT B,GARG S. Fine-pruning:defending against backdooring attacks on deep neural networks[C]//International Symposium on Research in Attacks,Intrusions,and Defenses. Cham:Springer International Publishing,2018:273-294.
|
| 19 |
WU D X, XIA S T, WANG Y S. Adversarial weight perturbation helps robust generalization[J]. Advances in Neural Information Processing Systems, 2020, 33, 2958- 2969.
|
| 20 |
KRIZHEVSKY A,HINTON G. Learning multiple layers of features from tiny images[M/OL]. (2009-04-08)[2009-04-08]. https://www.cs.utoronto.ca/~kriz/learning-features-2009-TR.pdf.
|
| 21 |
WU H, LIU G, YAO Y, et al. Watermarking neural networks with watermarked images[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2020, 31 (7): 2591- 2601.
|
| 22 |
KINGMA D P. ADAM: a method for stochastic optimization[J]. arXiv preprint, arXiv, 1412, 6980, 2014.
|
| 23 |
NETZER Y,WANG T,COATES A,et al. Reading digits in natural images with unsupervised feature learning[C]//NIPS Workshop on Deep Learning and Unsupervised Feature Learning. 2011,2011(2):4.
|
| 24 |
HU J,SHEN L,SUN G. Squeeze-and-excitation networks[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 2018:7132-7141.
|
| 25 |
SANDLER M,HOWARD A,ZHU M,et al. Mobilenetv2:inverted residuals and linear bottlenecks[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 2018:4510-4520.
|
/
| 〈 |
|
〉 |