arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

无验证的水印:欧盟人工智能法案之后的人工智能文本水印

Watermarks Without Verification: AI Text Watermarking After the EU AI Act

Alexander Nemecek, Vipin Chaudhary, Erman Ayday

arXiv 2609.09604首次发表:更新:

发表机构

Case Western Reserve University(凯斯西储大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文指出欧盟AI法案下文本水印的反对与保证均无法验证,通过开源实现评估发现质量影响有限,并提出五项治理要求以解决不可验证性。

AI 中文摘要

2026年8月2日,欧盟人工智能法案第50条的义务生效,要求生成式人工智能提供商标记其系统生成的内容,并确保这些内容能够被检测为人工智能生成。几天后,Anthropic披露,在该日期之后发布的每个Claude模型都在所有生成的文本中嵌入了基于SynthID-Text的水印,默认启用且用户无法选择退出;Google自2024年起已在Gemini中部署了SynthID-Text。用户反对称,水印降低了质量,尤其是对代码而言,它秘密编码了识别信息,并且(相互矛盾地)既容易被移除又无法逃脱;供应商则以质量不变、无识别信息以及对轻度编辑具有鲁棒性的保证作为回应。在这项工作中,我们认为,无论是反对意见还是保证,目前都无法得到验证,而这种不可验证性(而非水印本身)才是实质性的治理失败。我们根据解决每项争议所需的条件对争议性主张进行分类,并在两个开放权重模型上评估开源SynthID-Text实现,因为没有公开工具可以测试已部署的系统。在散文方面,水印的测量效果不超过更改采样种子的效果。在代码方面,一个模型的正确性成本为三个百分点,另一个模型则低于测量阈值,而检测仍接近随机水平,这是可检测性的限制而非质量限制。其余差距可追溯至访问受限或机构缺失,我们将每项差距映射到一项要求:发布匹配输出、配置披露、认可审计、共享评估协议以及互操作检测。

英文摘要

On August 2, 2026, the obligations of Article 50 of the EU AI Act took effect, requiring generative AI providers to mark the content their systems produce and ensure it can be detected as AI-generated. Days later, Anthropic disclosed that every Claude model released after that date embeds a watermark based on SynthID-Text in all generated text, enabled by default with no user opt-out; Google has deployed SynthID-Text in Gemini since 2024. Users objected that the watermark degrades quality, particularly for code, that it secretly encodes identifying information, and, in mutual contradiction, that it is easily removable and inescapable; the vendor answered with assurances of unchanged quality, no identifying information, and robustness to light editing. In this work, we argue that neither the objections nor the assurances can currently be verified and that this unverifiability, rather than watermarking itself, is the substantive governance failure. We sort the contested assertions by what it would take to settle each and evaluate the open-source SynthID-Text implementation on two open-weight models, because no public tool can test the deployed systems. On prose, the measured effect of the watermark does not exceed that of changing the sampling seed. On code, the cost is three points of correctness on one model and below measurement on the other, while detection remains near chance, a limitation of detectability rather than quality. The remaining gaps trace to withheld access or missing institutions and we map each to a requirement: release of matched outputs, configuration disclosure, accredited audits, a shared evaluation protocol, and interoperable detection.

Comments10 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑