Detection of AI drag-style edited images based on dual representation of pixel and latent space
Online published: 2026-01-04
Copyright
In recent years, the rapid development of generative artificial intelligence has made the digital image manipulation significantly easier. If left unchecked, these fake images could pose serious threats to society. With AI technology, users can effortlessly generate fake images on a global scale or use tools like drag-style image editing for localized manipulation. Currently, most research in digital image forensics focuses on identifying the globally AI generated images, with less emphasis on detecting locally tampered images. To address the security risks posed by such localized editing, a detection method for AI drag-style edited images was proposed. By analyzing the process behind drag-style editing, distinctive traces left by such operation were identified. Based on this, a detection method combining pixel-level characterization and latent space representations of images was designed, which could effectively improve the detection accuracy of AI drag-style edited ones. To support this research, four typical drag-style editing techniques based on diffusion models to construct a dataset of 9 806 locally drag-style edited images. Extensive experiments were conducted on this dataset, the real-image dataset, and the globally generated image dataset, and the results showed that the proposed method achieved excellent detection performance.
ZOU Xinping , LI Haodong . Detection of AI drag-style edited images based on dual representation of pixel and latent space[J]. Journal of Cybersecurity, 2025 , 3(4) : 111 -124 . DOI: 10.20172/j.issn.2097-3136.250418
图 1 由4种不同的AI拖拽式编辑技术产生的图像。在每个例子中,由左到右分别为输入图像、用户指令、编辑结果。对于一张给定的输入图像,用户通过拖拽点击和涂画编辑区域给出编辑指令,AI模型根据指令完成对输入图像的修改Fig.1 Diagram of four different AI drag-style editing techniques. In each example, the sequence from left to right consists of the input image, user instructions, and the edited result. For a given input image, users provide editing instructions by dragging, clicking, and painting on the areas to be edited, and the AI model completes the modification of the input image based on these instructions |
表 1 扩散模型全图生成和拖拽式编辑总结Table 1 Summary of full map generation and drag-editing of diffusion models |
| 特征 | 扩散模型全图生成 | 扩散模型拖拽式编辑 |
| 输入 | 文本或随机噪声或真实图像 | 原图+拖拽指令 |
| 作用位置 | 文本通过跨模态融合机制在反向扩散阶段起作用 | 拖拽指令主要在局部噪声优化阶段起作用 |
| 核心操作 | 全局去噪解码成新图像 | 局部优化+去噪解码 |
| 关键步骤 | 前向扩散(加噪)、反向扩散 (去噪) | 潜空间反演、局部约束优化、去噪解码 |
| 输出特性 | 完全新图像 | 原图局部修改(保留全局 结构) |
| 检测差异 | 高频噪声全局分布 | 局部高频伪影、掩码边界不一致性 |
表 2 实验数据集组成Table 2 Composition of experimental dataset |
| 类别 | 名称 | 来源 | 数量 | 生成器种类 | |
| 真实 | DragBench | 公开数据集 | 175 | — | |
| ImageNet | 公开数据集 | 2 000 | — | ||
| Download_SRC | 网络收集 | 1 631 | — | ||
| 全局 生成 | GenImage | 公开数据集 | 1 331 167 | ADM、SDv1.4、SDv1.5、Wukong、Glide、VQDM、 Midjourney | |
| Fake2M | 公开数据集 | 2 300 000 | SDv1.5、IF_v1.0、Midjourney | ||
| 拖拽 编辑 | DragDiffusion | 本工作构建 | 1 736 | SDv1.5 | |
| DiffEditor | 本工作构建 | 4 000 | SDv1.5 | ||
| DragNoise | 本工作构建 | 1 734 | SDv1.5 | ||
| FreeDrag | 本工作构建 | 2 336 | SDv1.5 | ||
表 3 三分类任务实验结果(准确率/%)Table 3 Experimental results of three-classification task (accuracy/%) |
| 方法 | 测试/预测类型 | 真实 | 生成 | 拖拽 | 所有平均 |
| ResNet | 真实 | 80.3 | 26.2 | 7.5 | 79.3 |
| 生成 | 16.3 | 70.3 | 5.1 | ||
| 拖拽 | 3.4 | 3.5 | 87.4 | ||
| CF | 真实 | 87.3 | 5.4 | 6.6 | 78.7 |
| 生成 | 16.3 | 84.6 | 29.1 | ||
| 拖拽 | 3.4 | 10.0 | 64.3 | ||
| CNNSpot | 真实 | 87.1 | 4.6 | 9.2 | 88.7 |
| 生成 | 12.3 | 94.3 | 6.0 | ||
| 拖拽 | 0.6 | 1.1 | 84.8 | ||
| DIRE | 真实 | 89.8 | 14.9 | 2.7 | 89.1 |
| 生成 | 8.3 | 83.2 | 3.0 | ||
| 拖拽 | 1.9 | 1.9 | 94.3 | ||
| PSCC | 真实 | 96.9 | 17.4 | 0 | 90.5 |
| 生成 | 2.9 | 76.2 | 1.5 | ||
| 拖拽 | 0.2 | 6.4 | 98.5 | ||
| UnivFD | 真实 | 97.7 | 1.6 | 12.8 | 89.2 |
| 生成 | 1.0 | 90.1 | 7.4 | ||
| 拖拽 | 1.3 | 8.3 | 79.8 | ||
| NPR | 真实 | 91.4 | 7.6 | 1.0 | 86.5 |
| 生成 | 16.0 | 78.8 | 5.2 | ||
| 拖拽 | 4.4 | 9.2 | 89.2 | ||
| CFLNet | 真实 | 92.2 | 7.2 | 0.6 | 89.4 |
| 生成 | 16.7 | 82.8 | 0.5 | ||
| 拖拽 | 4.5 | 2.2 | 93.3 | ||
| 所提方法 | 真实 | 93.4 | 13.1 | 2.9 | 91.4 |
| 生成 | 5.5 | 85.8 | 2.2 | ||
| 拖拽 | 1.1 | 1.1 | 94.9 |
表 4 二分类任务实验结果(准确率/%)Table 4 Experimental results of binary classification task (accuracy/%) |
| 测试数据 | ResNet | CNNSpot | PSCC | UnivFD | 所提方法 |
| 真实图像 | 88.2 | 99.9 | 96.8 | 98.7 | 98.8 |
| 拖拽图像 | 87.4 | 32.2 | 100 | 20.1 | 96.3 |
| 以上平均 | 87.8 | 66.1 | 98.4 | 59.4 | 97.6 |
表 5 泛化性实验结果(准确率/%)Table 5 Generalization experimental results (accuracy/%) |
| 测试数据集 | ResNet | CNNSpot | DIRE | PSCC | 所提方法 |
| DragNoise | 91.4 | 85.0 | 81.1 | 96.1 | 92.9 |
| WuKong | 90.7 | 90.8 | 88.8 | 83.8 | 94.3 |
| IF | 21.8 | 25.8 | 24.4 | 15.6 | 22.9 |
| 平均值 | 73.9 | 71.7 | 68.9 | 65.2 | 75.8 |
表 6 不同变体实验结果(准确率/%)Table 6 Experimental results for different variants (accuracy/%) |
| 测试数据集 | 基线模型 | 变体1 | 变体2 | 变体3 | 变体4 | 所提方法 |
| 拖拽图像 | 90.1 | 93.1 | 81.9 | 81.4 | 90.6 | 94.9 |
| 真实图像 | 56.4 | 69.4 | 67.3 | 64.4 | 77.1 | 93.4 |
| 生成图像 | 69.2 | 76.2 | 70.8 | 74.0 | 70.7 | 85.8 |
| 平均值 | 71.9 | 79.5 | 73.3 | 73.3 | 81.8 | 91.4 |
表 7 不同损失参数实验结果(准确率/%)Table 7 Experimental results with different loss parameters (accuracy/%) |
| 测试数据集 | (1,1) | (1,2) | (2,1) |
| 拖拽图像 | 94.9 | 93.9 | 93.2 |
| 真实图像 | 93.4 | 91.4 | 91.9 |
| 生成图像 | 85.8 | 87.5 | 85.8 |
| 平均值 | 91.4 | 90.9 | 90.3 |
图 5 网络中不同卷积层的Grad-CAM热图。(a)从左至右:原图、编辑指示图、编辑图、编辑图与原图的差值图、编辑图与重建图的差值图;(b)、(c)分别是在空域分支和潜空间分支不同网络层中获得的热图,从左至右网络变深;(d)是在特征融合模块不同网络层中获得的热图,从上至下网络变深Fig.5 Grad-CAM heatmaps of different convolutional layers in the network. (a) From left to right: original image, editing instruction map, edited image, difference map between the edited image and the original image, and the difference map between the edited image and the reconstructed image; (b) and (c) are the heatmaps obtained from different network layers in the spatial domain branch and the latent space branch, respectively, with the network becoming deeper from left to right; (d) are the heatmaps obtained from different network layers within the feature fusion module, with the network becoming deeper from top to bottom |
①
②
| 1 |
WANG S, WANG O, ZHANG R, et al. CNN-generated images are surprisingly easy to spot. . for now[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. San Francisco, CA, USA: IEEE, 2020: 8695-8704.
|
| 2 |
WANG Z, BAO J, ZHOU W, et al. Dire for diffusion-generated image detection[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. Paris, France: IEEE, 2023: 22445-22455.
|
| 3 |
AMOROSO R, MORELLI D, CORNIA M, et al. Parents and children: Distinguishing multimodal deepfakes from natural images[J]. ACM Transactions on Multimedia Computing, Communications and Applications, 2024, 21 (1): 1- 23.
|
| 4 |
SHI Y, XUE C, LIEW J H, et al. DragDiffusion: Harnessing diffusion models for interactive point-based image editing[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle Convention Center, USA: IEEE, 2024: 8839-8849.
|
| 5 |
PAN X, TEWARI A, LEIMKÜHLER T, et al. Drag your gan: Interactive point-based manipulation on the generative image manifold[C]//ACM SIGGRAPH 2023 Conference Proceedings. Los Angeles, CA, USA: ACM, 2023: 1-11.
|
| 6 |
MAO J, WANG X, AIZAWA K. Guided image synthesis via initial image editing in diffusion model[C]//Proceedings of the 31st ACM International Conference on Multimedia. Ottawa, Canada: ACM, 2023: 5321-5329.
|
| 7 |
RONNEBERGER O, FISCHER P, BROX T. U-Net: Convolutional networks for biomedical image segmentation[C]//Medical Image Computing and Computer-Assisted Intervention-MICCAI 2015: 18th International Conference. Springer International Publishing, 2015: 234-241.
|
| 8 |
MOU C, WANG X, SONG J, et al. DiffEditor: Boosting accuracy and flexibility on diffusion-based image editing[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle Convention Center, USA: IEEE, 2024: 8488-8497.
|
| 9 |
LIU H, XU C, YANG Y, et al. Drag your noise: Interactive point-based editing via diffusion semantic propagation[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle Convention Center, USA: IEEE, 2024: 6743-6752.
|
| 10 |
LING P, CHEN L, ZHANG P, et al. FreeDrag: Feature dragging for reliable point-based image editing[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle Convention Center, USA: IEEE, 2024: 6860-6870.
|
| 11 |
ZHANG Z, LIU H, CHEN J, et al. GoodDrag: Towards good practices for drag editing with diffusion models[J]. arXiv preprint, arXiv:, 2404, 07206, 2024.
|
| 12 |
ZHAO X, GUAN J, FAN C, et al. FastDrag: Manipulate anything in one step[J]. Advances in Neural Information Processing Systems, 2025, 37, 74439- 74460.
|
| 13 |
LI H, LI B, TAN S, et al. Identification of deep network generated images using disparities in color components[J]. Signal Processing, 2020, 174, 107616.
|
| 14 |
QIAN Y, YIN G, SHENG L, et al. Thinking in frequency: Face forgery detection by mining frequency-aware clues[C]//Proceedings of the European Conference on Computer Vision. Scottish Event Campus, Glasgow: Springer, 2020: 86-103.
|
| 15 |
LI Y, CHANG M, LYU S. In Ictu Oculi: Exposing AI created fake videos by detecting eye blinking[C]//Proceedings of the Workshop on Information Forensics and Security. Hong Kong, China: IEEE, 2018: 1-7.
|
| 16 |
LI L, BAO J, ZHANG T, et al. Face X-ray for more general face forgery detection[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. San Francisco, CA, USA: IEEE, 2020: 5001-5010.
|
| 17 |
HE K, ZHANG X, REN S, et al. Deep residual learning for image recognition[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Las Vegas, NV, USA: IEEE, 2016: 770-778.
|
| 18 |
LIU X, LIU Y, CHEN J, et al. PSCC-Net: Progressive spatio-channel correlation network for image manipulation detection and localization[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2022, 32 (11): 7505- 7517.
|
| 19 |
NILOY F F, BHAUMIK K K, WOO S S. CFL-Net: image forgery localization using contrastive learning[C]//Proceedings of the IEEE/CVF winter conference on applications of computer vision. Waikoloa, HI: IEEE, 2023: 4642-4651.
|
| 20 |
OJHA U, LI Y, LEE Y J. Towards universal fake image detectors that generalize across generative models[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Los Alamitos, CA, USA: IEEE, 2023: 24480-24489.
|
| 21 |
TAN C, ZHAO Y, WEI S, et al. Rethinking the up-sampling operations in CNN-based generative network for generalizable deepfake detection[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle Convention Center, USA: IEEE, 2024: 28130-28139.
|
| 22 |
SUN Z, FANG H, CAO J, et al. Rethinking image editing detection in the era of generative AI revolution[C]//Proceedings of the ACM Multimedia 2024. Melbourne, Australia: ACM, 2024: 2023.
|
| 23 |
KIRILLOV A, MINTUN E, RAVI N, et al. Segment anything[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. Paris, France: IEEE, 2023: 4015-4026.
|
| 24 |
LI J, LI D, SAVARESE S, et al. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models[C]//International Conference on Machine Learning. San Diego, CA, USA: JMLR , 2023: 19730-19742.
|
| 25 |
SONG J, MENG C, ERMON S. Denoising diffusion implicit models[J]. arXiv preprint, arXiv, 2010, 02502, 2010.
|
| 26 |
MAHFOUDI GA, TAJINI B, RETRAINT F, et al. DEFACTO: Image and face manipulation dataset[C]//Proceedings of the European Signal Processing Conference. Coruña, Spain: IEEE, 2019: 1-5.
|
| 27 |
KARRAS T, LAINE S, AITTALA M, et al. Analyzing and improving the image quality of stylegan[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. San Francisco, CA, USA: IEEE, 2020: 8110-8119.
|
| 28 |
ROMBACH R, BLATTMANN A, Lorenz D, et al. High-resolution image synthesis with latent diffusion models[C]//Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. New Orleans, LA, USA: IEEE, 2022: 10684-10695.
|
| 29 |
DENG J, DONG W, SOCHER R, et al. ImageNet: A large-scale hierarchical image database[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Miami, FL, USA: IEEE, 2009: 248-255.
|
| 30 |
ZHU M, CHEN H, YAN Q, et al. GenImage: A million-scale benchmark for detecting ai-generated image[C]//Advances in Neural Information Processing Systems. La Jolla, California: NIPS, 2023, 36.
|
| 31 |
LU Z, HUANG D, BAI L, et al. Seeing is not always believing: Benchmarking human and model perception of AI-generated images[C]//Advances in Neural Information Processing Systems. New Orleans, LA, USA: IEEE, 2023: 36.
|
| 32 |
RAMPRASAATH R, MICHAEL C, ABHISHEK D, et al. Grad-CAM: Visual explanations from deep networks via gradient-based localization[C]// Proceedings of the IEEE International Conference on Computer Vision. Dordrecht, Netherlands: Springer, 2017: 618-626.
|
/
| 〈 |
|
〉 |