发表机构
Original Pictures Technologies, Inc.(Original Pictures Technologies 公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该基准在开放许可模型下比较水印与指纹两类内容来源软绑定方法,覆盖图像、音频和视频,发现PixelSeal和AudioSeal分别最优,并揭示指纹在图像副本检测上更有效,而水印在视频上更有效。
AI 中文摘要
诸如C2PA等内容来源标准允许平台通过软绑定恢复被剥离的清单:从内容中读取的不可见水印,或在注册表中查找的指纹。我们在一个协议下对这两类方法进行了基准测试,仅限于我们审计过许可证的公开可用模型,在公共媒体上进行,假匹配率在保留的负样本上校准,性能估计采用源级自助法区间。对于水印,我们评估了来自12种方法的25种图像、7种音频和7种视频配置,涉及感知质量、鲁棒性、假阳性和成本;对于指纹,评估了35种方法,注册表规模高达98,985张图像,涵盖部分编辑和对抗性攻击。PixelSeal在图像和视频水印方面给出了最佳平衡,AudioSeal在音频方面最佳,但每个TrustMark变体的纠错检测器在5.9%至15.3%的未标记图像上触发,因此验证者应测试预期载荷。在指纹中,基于非商业数据训练的副本检测器在配对级假匹配率为10^{-7}时检测到高达75.5%的变换图像,而DINOv2(最佳许可宽松方法)检测到65.0%;产品目录中的近似副本主导了假匹配,在评估的图像流水线中,在匹配校准的假绑定约束下,未观察到几何验证带来的检测改进。在相同的受攻击副本上,这两类方法失败方式不同:水印和指纹成功的并集覆盖了69%的图像副本;指纹覆盖了更多音频查询,而预期密钥水印验证覆盖了更多视频查询。嵌入水印使62.0%的图像的ISCC码越过其匹配阈值,平台相关的颜色转换在x86和ARM主机之间改变了水印位。
英文摘要
Content-provenance standards such as C2PA let a platform recover a stripped manifest through a soft binding: an invisible watermark read from the content, or a fingerprint looked up in a registry. We benchmarked both families under one protocol, restricted to openly available models whose licences we audited, on public media, with false-match rates calibrated on held-out negatives and source-level bootstrap intervals for performance estimates. For watermarking we evaluated 25 image, 7 audio and 7 video configurations from 12 methods on perceptual quality, robustness, false positives and cost; for fingerprinting, 35 methods on registries of up to 98,985 images, partial edits and adversarial attacks. PixelSeal gave the best balance for image and video watermarks and AudioSeal for audio, but the error-correcting detector of every TrustMark variant fired on 5.9 to 15.3% of unmarked images, so a verifier should test the expected payload. Among fingerprints, copy detectors trained on non-commercial data detected up to 75.5% of transformed images at a pair-level false-match rate of $10^{-7}$ and DINOv2, the best permissively licensed method, 65.0%; near-copies in a product catalogue dominated the false matches, and no detection improvement from geometric verification was observed under a matched calibration false-binding constraint in the evaluated image pipelines. On the same attacked copies the two families failed differently: the union of watermark and fingerprint successes covered 69% of image copies; fingerprints covered more audio queries, whereas expected-key watermark verification covered more video queries. Embedding a watermark moved the ISCC code of 62.0% of images past its match threshold, and platform-dependent colour conversion changed watermark bits between x86 and ARM hosts.
Comments59 pages. Code, recorded results and manuscript source: https://github.com/Original-Pictures/soft-bindings-benchmark (v1.0)