arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10166cs.CRcs.AI

MarkNull:通过流形上潜在空间操纵实现AI生成图像的模型无关水印去除

MarkNull: Model-Agnostic Watermark Removal in AI-Generated Images via On-Manifold Latent Manipulation

Jie Cao, Qi Li, Zelin Zhang, Xiaodong Wu, Lingshuang Liu, Xiangman Li, Jianbing Ni

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出MarkNull及变体MarkNull-A,通过流形上潜在空间操纵实现AI生成图像的模型无关水印去除,可降低水印比特准确率至近随机猜测水平,还能破解谷歌SynthID-Image系统,为防御提供攻击检测机制。

中文摘要 AI 辅助

数字水印已成为AI生成图像来源追踪和版权归属的关键技术,但针对现实场景中模型无关去除攻击的鲁棒性仍缺乏充分探索。现有攻击要么仅针对特定生成模型有效,要么在去除水印时伴随严重的视觉质量下降。本文提出MarkNull,一种通过流形上潜在空间操纵实现的模型无关水印去除攻击。MarkNull基于一项关键观察:带水印图像的生成潜在表示与嵌入的初始噪声之间存在强统计依赖关系。为量化该依赖关系,我们引入噪声-潜在对齐分数(NLAS),并构建优化目标,在保持语义保真度的同时选择性地使潜在表示与嵌入水印去相关。对不同类型水印范式(包括事后方案、微调方案和初始噪声方案)的广泛评估显示,MarkNull将平均比特准确率降至53.14%,接近随机猜测水平(50%),且无明显视觉质量下降。为进一步提升可扩展性,我们提出MarkNull-A,一种免优化的摊销变体,将攻击提炼为单次前向传播,每张图像耗时0.50秒,计算开销适中。值得注意的是,我们的攻击成功破解了谷歌的SynthID-Image系统,同时保持高视觉质量,并能有效迁移至视频水印。最后,我们提出一种攻击检测机制作为MarkNull和MarkNull-A的防御对应方案,强调开发能抵御模型无关潜在空间攻击的水印设计的必要性。

英文摘要

Digital watermarking has emerged as a critical technique for provenance and copyright attribution in AI-generated imagery, yet its robustness against realistic, model-agnostic removal attacks remains poorly explored. Existing attacks either succeed only against specific generative models or achieve removal at the cost of severe visual degradation. In this paper, we propose MarkNull, a model-agnostic watermark removal attack via on-manifold latent manipulation. MarkNull is grounded in a key observation: watermarked images exhibit a strong statistical dependency between the generated latent representation and the embedded initial noise. To quantify this dependency, we introduce the Noise-Latent Alignment Score (NLAS) and formulate an optimization objective that selectively decorrelates the latent representation from the embedded watermark while preserving semantic fidelity. Extensive evaluations across different categories of watermarking paradigms, including post-hoc, fine-tuning-based, and initial-noise-based schemes, demonstrate that MarkNull reduces average bit accuracy to 53.14%, approaching random-guessing (50%), without perceptible image degradation. To further improve scalability, we propose MarkNull-A, an amortized, optimization-free variant that distills the attack into a single forward pass, achieving 0.50 s/image with modest computational overhead. Notably, our attacks successfully compromise Google's SynthID-Image system while preserving high visual quality and transfer effectively to video watermarking. Finally, we present an attack detection mechanism as a defensive counterpart to MarkNull and MarkNull-A, highlighting the necessity of developing watermark designs resilient to model-agnostic latent-space attacks.

发表机构

  • University of Waterloo(滑铁卢大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑