AI 中文总结
研究针对合成文本标记需求,提出ChainMark无模型大语言模型水印方法,通过特定方式划分词汇表状态并强制马尔可夫转移,检测器无需访问语言模型,推导相关闭式,证明鲁棒性阈值,在多模型和领域中优于其他方法,可恢复目标误报率。
AI 中文摘要
欧盟人工智能法案等监管制度要求对合成文本进行机器可读标记,但现有的水印检测器依赖于生成语言模型和启发式阈值,没有闭式校准。我们引入了ChainMark,一种主动水印,它通过带密钥的SHA-256将词汇表划分为S个状态,并在一部分rho的位置上强制进行硬马尔可夫转移;检测器通过O(n)次哈希操作从相同密钥重放分区,无需访问语言模型。我们推导了一个闭式S*(n, rho, alpha),将目标误报率、文本长度和预算映射到最小状态数(定理1),证明了一个通用的鲁棒性阈值delta* = 1 - 1/sqrt(2)约为29.3%,在(S, rho, n)中不变(定理2),并将两者推广到任何k正则转移拓扑(定理3)。在三个指令微调的大语言模型和四个领域中,在匹配预算下,ChainMark在翻译和随机替换攻击下严格优于KGW和SWEET;一次语料库经验重新校准可在自然语言文本上恢复1%的目标误报率。
英文摘要
Regulatory regimes such as the EU AI Act mandate machine-readable marking of synthetic text, but existing watermark detectors rely on the generating LM and on heuristic thresholds with no closed-form calibration. We introduce ChainMark, an active watermark that partitions the vocabulary into S states via keyed SHA-256 and forces a hard Markov transition on a fraction rho of positions; the detector replays the partition from the same key in O(n) hash operations, with no LM access. We derive a closed-form S*(n, rho, alpha) mapping a target FPR, text length, and budget to the minimum state count (Theorem 1), prove a universal robustness threshold delta* = 1 - 1/sqrt(2) approximately 29.3% that is invariant in (S, rho, n) (Theorem 2), and generalise both to any k-regular transition topology (Theorem 3). Across three instruction-tuned LLMs and four domains, ChainMark strictly dominates KGW and SWEET under translation and random-substitution attacks at matched budget; a one-corpus empirical recalibration restores the 1% target FPR on natural-language text.