AI 中文总结
DHMark是一种LLM生成文本的公钥水印框架,通过Diffie-Hellman引导拒绝采样分离载荷授权与文本证据,在多类攻击下保持高有效率且负控制接受率为0。
AI 中文摘要
大语言模型(LLM)水印为追踪生成文本的来源提供了重要机制。现有统计水印通常有效且鲁棒,但大多依赖私有检测密钥,这会使验证中心化并增加公开审计的复杂度。近期的公开或可公开验证的水印方案改进了密钥管理,但许多方案依赖嵌入密码字符串的精确恢复,使其在token编辑、截断、复制粘贴及低熵生成下脆弱。本文提出DHMark,一种用于LLM生成文本的公钥水印框架。核心思路是将载荷授权与噪声文本证据分离:签发方对绑定公开上下文的短注册载荷签名,载荷被扩展为多个1比特方程;生成期间,Diffie-Hellman引导的token标记接口为每个候选token分配公开方程投票,采样器柔和或选择性提升投票与授权载荷一致的候选;验证时,第三方验证者使用公开信息提取token投票,将其聚合为方程级证据,仅对已签名的注册记录评分。该设计避免了长嵌入签名的精确恢复,将水印检测视为注册辅助的统计证据聚合。我们形式化了可公开验证的设置,分析了标签伪随机性、注册支持的可靠性及采样失真,并在截断、替换、复制粘贴、错误上下文及纯生成攻击下评估了原型。在默认32位配置下,DHMark在8种编辑条件下维持至少0.967的有效率,同时在3种负控制下产生0.000的接受率。
英文摘要
Large language model (LLM) watermarking provides an important mechanism for tracing the provenance of generated text. Existing statistical watermarks are often effective and robust, but most of them rely on private detection keys, which centralizes verification and complicates public auditing. Recent public or publicly verifiable watermarking schemes improve key management, yet many of them rely on exact recovery of embedded cryptographic strings, making them fragile under token edits, truncation, copy-paste, and low-entropy generation. This paper introduces DHMark, a public-key watermarking framework for LLM-generated text. The key idea is to separate payload authorization from noisy textual evidence. An issuer signs a short registry payload bound to a public context, and the payload is expanded into many one-bit equations. During generation, a Diffie-Hellman-guided token-labeling interface assigns each candidate token a public equation vote, and the sampler softly or selectively promotes candidates whose votes agree with the authorized payload. During verification, third-party verifiers use public information to extract token votes, aggregate them into equation-level evidence, and score only signed registry records. This design avoids exact recovery of a long embedded signature and instead treats watermark detection as registry-aided statistical evidence aggregation. We formalize the public-verification setting, analyze label pseudorandomness, registry-backed soundness, and sampling distortion, and evaluate a prototype under truncation, substitution, copy-paste, wrong-context, and plain-generation attacks. In the default 32-bit configuration, DHMark maintains at least a 0.967 valid rate across eight edit conditions while yielding a 0.000 acceptance rate on three negative controls.