生成式图像水印在潜在频率掩蔽下的弱点探索
Exploring Weaknesses of Generative Image Watermarks against Latent Frequency Masking
查看机构详情
- MSU AI Institute(莫斯科国立大学人工智能研究所)
- Trusted AI Research Center RAS(俄罗斯科学院可信人工智能研究中心)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究提出潜在频率掩蔽攻击,通过替换潜在表示中的傅里叶系数擦除水印,在六种扩散水印方法上验证了其有效性与高效性,并指出该攻击面需纳入鲁棒性评估。
中文摘要 AI 辅助
不可见水印已成为追踪AI生成图像的核心工具,但其对自适应移除攻击的鲁棒性仍是一个未解决的安全问题。我们提出潜在频率掩蔽(Latent Frequency Masking)攻击,通过替换含水印图像潜在表示中选定的傅里叶系数来擦除水印证据。替换系数可从高斯噪声中采样以提高效率,或通过扩散再生推导以更好地保留图像质量。我们提供了一个理论失真界,将重建的对抗图像变化与掩蔽的潜在频率扰动联系起来。我们在DiffusionDB和MS-COCO提示生成的图像上,针对六种扩散水印方法评估了所提出的攻击。潜在频率掩蔽能够移除或显著削弱多种水印,同时保持感知质量,并与现有攻击相比具有更优的运行时间。这些结果揭示了潜在频率操作是一个实际的攻击面,并强调了在生成式图像水印的鲁棒性评估中需纳入此类攻击。
英文摘要
Invisible watermarking has become a central tool for tracing AI-generated images, but its robustness against adaptive removal attacks remains an open security question. We introduce Latent Frequency Masking, an attack that erases watermark evidence by replacing selected Fourier coefficients in the latent representation of a watermarked image. The replacement can be sampled from Gaussian noise for efficiency or derived from diffusion regeneration for improved image preservation. We provide a theoretical distortion bound relating the change between the reconstructed adversarial image and the masked latent-frequency perturbation. We evaluate the proposed attack against six diffusion watermarking methods on images generated from DiffusionDB and MS-COCO prompts. Latent Frequency Masking removes or substantially weakens several watermarks while preserving perceptual quality and achieving favorable runtime compared with existing attacks. These results identify latent-frequency manipulation as a practical attack surface and highlight the need to include such attacks in robustness evaluations of generative image watermarking.