arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一个提示就足够:通过基础图像模型进行水印洗白

One Prompt Is Enough: Watermark Laundering Through Foundation Image Models

Jidong Yang, Qi Li, Wei Zong, Yang-Wai Chow, Willy Susilo, Huaike Yu, Chunpeng Wang, Suo Gao

arXiv 2609.01249首次发表:更新:

发表机构

Qilu University of Technology (Shandong Academy of Sciences); University of Wollongong; Dalian Polytechnic University(齐鲁工业大学(山东省科学院); 卧龙岗大学; 大连工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究揭示了基础图像模型可通过单步重建实现不可见水印洗白,通过多模型多方案实验明确其机制与独特性,建议将其作为水印评估的新鲁棒性条件。

AI 中文摘要

不可见水印通常针对压缩、模糊、噪声、裁剪和去噪等预定义扰动进行评估。公共基础图像模型带来了一种独特的威胁:攻击者可提交带有不可见水印的图像,仅用单个重建提示就能获得视觉保真的输出,而该输出无法可靠解码出原始不可见水印。我们将这种失效模式形式化为“水印洗白”,并结合误码率(BER)与视觉、语义保留的联合载荷-保真度曲线对其进行评估。在6个OpenAI和Google图像编辑模型、3种代表性水印方案及1800个重建输出中,我们确定了两种互补的洗白机制:OpenAI模型对所有评估方案产生最强的载荷破坏,而Nano Banana 2显示DwtDct在高保真重建下仍存在漏洞。提示消融实验表明,载荷破坏无需单一的去水印指令,说明该效应主要由重建路径而非显式攻击措辞诱导。与传统攻击的对比进一步显示,提示条件重建构成了一种独特的操作攻击接口。这些发现促使将基础模型重建作为不可见水印评估中缺失的鲁棒性条件。

英文摘要

Invisible watermarks are typically evaluated against predefined perturbations such as compression, blur, noise, cropping, and denoising. Public foundation image models expose a distinct threat: an attacker can submit a watermarked image with a single reconstruction prompt and obtain a visually faithful output from which the invisible watermark can no longer be decoded reliably. We formalize this failure mode as watermark laundering and evaluate it using a joint payload-fidelity profile that combines bit error rate (BER) with visual and semantic preservation. Across six OpenAI and Google image editing models, three representative watermarking schemes, and 1,800 reconstructed outputs, we identify two complementary laundering regimes: OpenAI models produce the strongest payload disruption across the evaluated schemes, whereas Nano Banana 2 shows that DwtDct remains vulnerable under high-fidelity reconstruction. Prompt ablations show that no single removal-oriented instruction is necessary for payload disruption, indicating that the effect is primarily induced by the reconstruction pathway rather than by explicit attack wording. Comparisons with conventional attacks further show that prompt-conditioned reconstruction constitutes a distinct operational attack interface. These findings motivate foundation-model reconstruction as a missing robustness condition in invisible watermark evaluation.

Comments11 pages, 4 figures, and 2 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑