arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.08213cs.CRcs.LG

DRIFT:通过偏转生成轨迹去除扩散水印

DRIFT: Removing Diffusion Watermarks by Deflecting the Generative Trajectory

Rui Bao, Zheng Gao, Xiaoyu Li, Xiaoyan Feng, Yang Song, Jiaojiao Jiang

首次发表
浏览论文内容

中文总结 AI 辅助

DRIFT通过结合部分前向扩散与随机反向重采样偏转生成轨迹,以黑盒方式去除扩散水印,在九种水印上实现98-100%攻击成功率并保持最佳图像质量。

中文摘要 AI 辅助

扩散水印在生成过程中嵌入可验证信号,通常通过恢复依赖于轨迹的证据来进行验证,这使得水印对传统的像素域失真具有鲁棒性。现有的去除攻击要么沿确定性轨迹重新生成,这往往会保留带有水印的潜在结构,要么对每张图像分别进行优化。我们识别出对可恢复生成轨迹的依赖是我们研究的方案中常见的攻击面。基于这一观察,我们提出了DRIFT,一种黑盒攻击,它结合了部分前向扩散与随机反向重采样。前向重新加噪限制了固定深度恢复流程可用的源信息,而随机反向提供了替代的噪声驱动路径,我们通过匹配的采样器比较来隔离其去除效益。自适应DRIFT为每张图像搜索所选阶梯上第一个被验证器拒绝的梯级,并优化保真度,同时仅保留被同一验证器拒绝的更新。在固定深度下,我们推导了信息论和Wasserstein源依赖界限;在实现阶梯单调性下,第一个被拒绝的梯级是该阶梯上被拒绝梯级中失真最小的,且验证器门控的细化保持了拒绝状态。在跨越三种范式的九种水印中,DRIFT实现了98-100%的攻击成功率和所比较攻击中最佳的图像质量,无需秘密密钥、验证器内部信息或逐图像梯度优化。

英文摘要

Diffusion watermarking embeds verifiable signals into the generative process and commonly verifies them by recovering trajectory-dependent evidence, making the marks robust to conventional pixel-space distortions. Existing removal attacks either regenerate along deterministic trajectories, which often preserve the watermark-bearing latent structure, or optimize every image separately. We identify the reliance on a recoverable generative trajectory as a common attack surface among the schemes we study. Based on this observation, we propose DRIFT, a black-box attack that combines partial forward diffusion with stochastic reverse resampling. Forward re-noising limits source information available to a fixed-depth recovery pipeline, while stochastic reversal supplies alternative noise-driven paths whose removal benefit we isolate through matched sampler comparisons. Adaptive DRIFT searches a selected ladder for each image's first verifier-rejected rung and refines fidelity while retaining only updates rejected by the same verifier. At fixed depth, we derive information-theoretic and Wasserstein source-dependence bounds; under realized-ladder monotonicity, the first rejected rung is least distorted among rejected rungs on that ladder, and verifier-gated refinement preserves rejection. Across nine watermarks spanning three paradigms, DRIFT achieves 98-100% attack success and the best image quality among the compared attacks, without secret keys, verifier internals, or per-image gradient optimization.

发表机构

  • University of New South Wales(新南威尔士大学)
  • Griffith University(格里菲斯大学)

机构由 AI 辅助整理,请以论文原文为准。

↑