arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.05870cs.CV

真实图像的校准内容认证

Certification of Real Images through Calibrated Content Authentication

Sarim Hashmi, Abdelrahman Elsayed, Mohammed Talha Alam, Samuele Poppi, Nils Lukas

首次发表
浏览论文内容

中文总结 AI 辅助

针对深度伪造检测器准确率下降问题,提出基于忠实重建的校准内容认证方法,可校准至最多1%错误认证,优于基线,并揭示事后可验证性随时间侵蚀。

中文摘要 AI 辅助

生成模型可以合成高质量的伪造多媒体内容,这些内容已经在规模上被滥用。我们评估了二十个深度伪造检测器,针对过去四年发布的十个生成器,发现准确率随时间下降,从接近完美的99.5%降至76%。对抗性扰动进一步将每个基线检测器的准确率降至2%以下,实际上逆转了检测器分配的标签。我们认为这种不可靠性反映了一个根本的模糊性:生成器可以精确地再现真实内容(例如,通过记忆),因此仅凭内容无法揭示真实的来源。因此,生成器产生的内容必须允许由同一生成器进行忠实重建,找到这样的重建使得合成来源具有合理性,真实性也具有合理的可否认性。因此,我们提出并评估了一种检测范式,该范式输出一个校准的预测,即真实性是否合理可否认:由任何已知生成器进行的忠实重建建立了合理的可否认性,而校准则限制了已知生成器产生的内容无法被重建的频率。我们的评估表明:(i)我们的检测器可以被校准,使得最多1%的生成内容被错误认证,在这个操作点上,大多数基线检测器的召回率接近零,包括最强的准确率为93%的检测器;(ii)在攻击样本上校准更严格的安全阈值,在评估的有界扰动攻击空间内,可以保持该界限对抗自适应攻击者,其扰动破坏了所有基线,但不覆盖任意的对抗性变换;(iii)事后可验证性正在侵蚀,因为3,000张Reddit图像中有1,116张抵抗2022年生成器的重建,但只有55到79张抵抗2024年生成器的重建。

英文摘要

Generative models can synthesize high-quality inauthentic multimedia content that is already being misused at scale. We evaluate twenty deepfake detectors against ten generators released in the last four years and find accuracy decreasing over time, from near-perfect 99.5% to 76%. Adversarial perturbations further reduce every baseline detector to below 2% accuracy, effectively inverting the detector's assigned label. We argue that this unreliability reflects a fundamental ambiguity: generators can reproduce authentic content exactly (e.g., through memorization), so content alone cannot reveal the true provenance label. For this reason, content produced by a generator must admit a faithful reconstruction by that same generator, and finding such a reconstruction makes synthetic provenance plausible and authenticity plausibly deniable. We therefore propose and evaluate a detection paradigm that outputs a calibrated prediction of whether authenticity is plausibly deniable: a faithful reconstruction by any known generator establishes plausible deniability, while calibration bounds how often content from known generators fails to be reproduced. Our evaluation shows that (i) our detector can be calibrated so that at most 1% of generated content is wrongly certified, an operating point at which most baseline detectors reach near-zero recall, including the strongest with 93% accuracy; (ii) calibrating a stricter security threshold on attacked samples preserves this bound against adaptive adversaries within the evaluated bounded-perturbation attack space, whose perturbations break every baseline, but does not cover arbitrary adversarial transformations; and (iii) post-hoc verifiability is eroding, as 1,116 of 3,000 Reddit images resist reproduction by a 2022 generator, but only 55 to 79 resist reproduction by 2024 generators.

发表机构

  • Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

↑