arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

人工智能水印证据无法满足法医鉴定要求:一项实证评估

AI Watermark Evidence Fails Forensic Readiness: An Empirical Evaluation

Saifur Rahman Tamim, Amir Labib Khan

arXiv 2607.16010首次发表:更新:

发表机构

Northern University Bangladesh(孟加拉国北方大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对政府要求大语言模型生成内容带水印这一规定所基于的假设进行检验,通过FRS框架评估三种水印方法,以意义保留释义为攻击向量,发现这些方法存在诸多问题,不符合法庭证据标准,FRS评分系统也有局限性。

AI 中文摘要

政府越来越多地要求大语言模型生成的内容带有水印。欧盟人工智能法案要求水印“足够可靠和强大”。加利福尼亚州的SB 942要求水印披露是“永久性的或极难去除的”。这两项规定都基于一个未经检验的假设:水印检测能产生足够可靠的证据供法庭使用。本文直接检验这一假设。我们根据多伯特可采性标准和美国国家标准与技术研究院(NIST)SP 800-86数字取证流程,评估了三种有代表性的大语言模型水印方法——KGW、单字模型和SynthID-Text的MarkLLM实现。为构建此次评估,我们提出了一个法医准备分数(FRS)框架,它有12条标准、三个强制关卡和一个60分的评分系统。我们将意义保留释义作为攻击向量,因为它既符合法律现实,又难以被视为证据篡改而不予理会。结果引发了严重的证据问题。在每种方法针对15个不同提示进行的846次有效释义运行中,最初检测到的每个KGW和单字模型文本在释义后都失去了水印——100%的条件去除率。SynthID的情况稍好,为98.3%。甚至在任何攻击之前,误报率就已经很高:KGW为70%,单字模型为83%,SynthID为80%。SynthID配置还将5.4%的释义后的人工撰写对照标记为人工智能生成,并显示出18.6%的矛盾率,其自身原始水印输出的80%落在不确定区间。这三种方法中没有一种能满足多伯特五个因素中的两个以上。我们还发现,尽管FRS基于点的评分系统按设计运行,但不能完全捕捉法医无用性——这是未来框架设计中值得注意的一个限制。这些配置经测试不符合法庭要求的证据标准。

英文摘要

Governments are increasingly mandating that LLM-generated content carry watermarks. The EU AI Act calls for markings that are "sufficiently reliable and robust." California's SB 942 requires disclosure that is "permanent or extraordinarily difficult to remove." Both mandates rest on an untested assumption: that watermark detection yields evidence reliable enough for courts. This paper tests that assumption directly. We evaluate three representative LLM watermarking methods -- KGW, Unigram, and the MarkLLM implementation of SynthID-Text -- against the Daubert admissibility criteria and the NIST SP 800-86 digital forensic process. To structure this evaluation, we propose a Forensic Readiness Score (FRS) framework with 12 criteria, three mandatory gates, and a 60-point scoring system. We focus on meaning-preserving paraphrase as the attack vector, since it is both legally realistic and difficult to dismiss as evidence tampering. The results raise serious evidentiary concerns. Out of 846 valid paraphrase runs across 15 diverse prompts per method, every single initially-detected KGW and Unigram text lost its watermark after paraphrasing -- 100% conditional removal. SynthID fared only slightly better at 98.3%. Even before any attack, false-negative rates were already high: 70% for KGW, 83% for Unigram, 80% for SynthID. The SynthID configuration also flagged 5.4% of paraphrased human-written controls as AI-generated and showed an 18.6% paradox rate, with 80% of its own pristine watermarked output landing in the uncertainty deadband. None of the three methods satisfy more than two of five Daubert factors. We also find that the FRS point-based scoring system, despite working as designed, cannot fully capture forensic uselessness -- a limitation worth noting for future framework design. These configurations, as tested, do not meet the evidentiary bar that courts require.

Comments9 pages, 4 figures. A version of this paper was submitted to the AAAI/ACM Conference on AI, Ethics, and Society (AIES) 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑