AI 中文总结
该研究复现了针对Tree-Ring语义水印的Reprompt伪造攻击,在低内存GPU上验证其有效性,并揭示了检测器统计量恢复、半精度兼容性及预测矛盾等关键发现。
AI 中文摘要
语义水印方案(如Tree-Ring)在扩散模型的初始噪声潜变量中隐藏可检测的模式。近期研究表明,这些水印不仅可被移除,还可被伪造:从未见过水印密钥的攻击者仍能生成被真实检测器接受的图像。我们使用作者发布的代码,在免费层级的双T4 GPU(每设备可用内存14.6 GB,远低于原始研究中使用的A40硬件每GPU内存)上,复现了Müller等人针对Stable Diffusion XL上Tree-Ring的Reprompt伪造攻击。攻击成功复现。在六次试验、三个组别中,我们检测到真实图像6/6,干净图像0/6,伪造图像5/6,每次攻击耗时325-332秒。在约束条件下运行还产生了另外三项结果。发布的检测器计算非中心χ²统计量并仅返回其CDF,因此我们恢复了被丢弃的统计量;我们的恢复结果与发布的检测器完全一致,基于该统计量构建的两个自然分数在相同的十八个观测值上区分伪造组与干净零假设的AUC分别为0.861和0.972。以半精度运行SDXL需要对流水线的直接自编码器调用进行修补,受控探针确认修补后的路径不改变检测器统计量。最后,我们报告了一项基于阅读检测器源代码所做的预测,而我们的测量结果与之相矛盾。笔记本、固定分支及所有测量工件均随论文发布。
英文摘要
Semantic watermarking schemes such as Tree-Ring hide a detectable pattern in the initial noise latent of a diffusion model. Recent work shows these watermarks are not only removable but forgeable: an attacker who never sees the watermarking key can still produce images the genuine detector accepts. We reproduce the Reprompt forgery attack of Müller et al. against Tree-Ring on Stable Diffusion XL, using the authors' released code, on free-tier dual T4 GPUs with 14.6 GB of usable memory per device, substantially less per-GPU memory than the A40 hardware used in the original study. The attack reproduces. Over six trials of three arms we detect genuine images 6/6, clean images 0/6, and forged images 5/6, at 325-332 s per attack. Three further results came out of running it under constraint. The released detector computes a non-central $χ^2$ statistic and hands back only its CDF, so we recovered the discarded statistic; our recovery matches the released detector exactly, and two natural scores built from it separate the forged arm from the clean null at AUC 0.861 and 0.972 on the same eighteen observations. Running SDXL in half precision requires patching the pipeline's direct autoencoder calls, and a controlled probe confirms the patched path leaves the detector statistic unchanged. Finally, we report a prediction we made from reading the detector source that our measurements then contradicted. The notebook, the pinned fork and every measurement artifact are released with the paper.
Comments11 pages, 6 figures, 4 tables. Code and artifacts: https://github.com/sr-tamim/watermark-forgery-research