发表机构
University of Edinburgh(爱丁堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究被动图像溯源在对抗性编辑下的鲁棒性,给出仅像素验证器的最优统计极限,并揭示公共验证器因信息泄露而提前失效,提出应分别评估统计上限与验证器信息泄露。
AI 中文摘要
被动图像溯源问题探讨的是:仅凭像素信息能否揭示图像的来源——是来自人类、某一聚合AI类别,还是某个特定的生成器。一旦源图像在验证器查看之前被编辑,该问题便转化为一个鲁棒性问题。我们将此问题建模为对抗性分布偏移下的源-目标验证。我们的第一个结果给出了任何仅基于图像的验证器的最优极限:最大的鲁棒目标接受差距等于目标分布与受攻击源分布集合之间的最小全变差距离。该量取决于源、目标和编辑类别,而非验证器架构。我们的第二个结果解释了为何已部署的公共验证器可能在此统计极限之前就失效。如果验证器在攻击区域上可以被模拟至误差ε,则替代黑盒攻击能达到目标接受率,其误差在2ε加上白盒最优的优化误差之内;基于公共特征的可揭示评分的逻辑斯蒂和softmax头部是可识别的,且近似评分访问能给出稳定的恢复界限。一个有限状态实验检验了双方均可计算的最小最大恒等式。在相同提示的真实/扩散基准上,评估的公共CLIP验证器在定向像素攻击下失效,而ResNet-18受害者表现出部分假到真的迁移。带有弃权(不执行)的二元反馈降低了实测攻击成功率,但正的实证差距上界并不能确立鲁棒性。这些结果促使分别评估源-目标统计上限和已部署验证器所泄露的信息。
英文摘要
Passive image provenance asks whether pixels alone can reveal where an image came from: a human, an aggregate AI class, or a particular generator. This becomes a robustness problem once a source image can be edited before the verifier sees it. We study the problem as source--target verification under adversarial distribution shift. Our first result gives the exact best-case limit for any image-only verifier: the largest robust target-acceptance gap equals the minimum total-variation distance between the target distribution and the set of attacked source distributions. This quantity depends on the source, target, and edit class, not on the verifier architecture. Our second result explains why deployed public verifiers can fail before this statistical limit is reached. If the verifier can be emulated on the attack region to error $\varepsilon$, then a surrogate black-box attack reaches target acceptance within $2\varepsilon$ plus optimization error of the white-box optimum; score-revealing logistic and softmax heads over public features are identifiable, and approximate score access gives stable recovery bounds. A finite-state experiment checks the minimax identity where both sides are computable. On same-prompt real/diffusion benchmarks, the evaluated public CLIP verifiers fail under targeted pixel attacks, while a ResNet-18 victim exhibits partial fake-to-real transfer. Binary feedback with abstention reduces measured attack success, but positive empirical gap upper bounds do not establish robustness. These results motivate separate evaluation of the source--target statistical ceiling and the information released by a deployed verifier.
CommentsAccepted at the 40th Annual Conference on Neural Information Processing Systems (NeurIPS 2026). 29 pages, including technical appendices. Code: https://github.com/kaikaiyao/pixels-alone-provenance