语言模型水印检测的预测似然比
Predictive Likelihood Ratios for Language Model Watermark Detection
浏览论文内容
中文总结 AI 辅助
本研究提出预测似然比方法用于语言模型水印检测,通过混合先验平均不确定性,在固定显著性下最大化先验功效,并验证了其作为检验鞅的任意时间有效性,实验表明在多种备择下具有稳健性。
中文摘要 AI 辅助
密钥水印检测检验观测到的令牌与从密钥重建的伪随机变量之间的依赖性。基于Li等人(2025)的枢轴框架,我们构建了预测似然比,对不确定的概率赤字和残差尾部分布进行平均。目标是在不同备择设定下实现稳健的检测功效,而无需单一信号强度调优。混合先验结合了尾部形状和有效宽度;分层扩展允许文档内赤字或宽度的变化。该检验在固定显著性水平下最大化先验平均功效,但并非普遍一致最优或极小极大。在精确条件枢轴零假设下,在每次观测前选择的归一化预测备择产生贝叶斯因子,该因子也是检验鞅:第一类错误控制不受备择误设影响,并在可选停止下保持有效。此保证不涵盖条件零假设的违反,且插值实现没有认证的任意时间保证。Gumbel边际似然通过固定求积评估。在评估的尾部形状和尾部宽度备择以及三个时间范围中,联合尾部混合的最大观测第二类错误遗憾为0.0080,而等尾混合为0.0962,相对于最佳测试规则。在来自两个开放模型的温度匹配输出上,它在所有八个非饱和模型-温度单元中相对于等尾基线提高了AUC,尽管领先参考分数通常具有更高的AUC。补充实验显示在独立零类似替换下保持功效,以及分层依赖建模的较小变化。证据支持在评估的备择中的稳健性,而非均匀功效保证或对任意文本编辑的抵抗。
英文摘要
Keyed watermark detection tests dependence between observed tokens and pseudorandom variables reconstructed from a secret key. Building on the pivotal framework of Li et al. (2025), we construct predictive likelihood ratios that average over uncertain probability deficits and residual-tail distributions. The aim is robust detection power across alternative specifications without requiring a single signal-strength tuning. A mixture prior combines tail shape and effective width; hierarchical extensions allow within-document variation in deficit or width. The test maximizes prior-averaged power at a fixed size, but is not generally uniformly most powerful or minimax. Under the exact conditional pivot null, normalized predictive alternatives selected before each observation yield a Bayes factor that is also a test martingale: Type I error control is unaffected by alternative misspecification and remains valid under optional stopping. This guarantee does not cover violations of the conditional null, and the interpolated implementation has no certified anytime guarantee. Gumbel marginal likelihoods are evaluated by fixed quadrature. Across the evaluated tail-shape and tail-width alternatives and three horizons, the union-tail mixture has maximum observed Type II error regret .0080, compared with .0962 for the equal-tail mixture, relative to the best tested rule. On temperature-matched outputs from two open models, it improves AUC over the equal-tail baseline in all eight non-saturated model-temperature cells, although the leading reference score generally has higher AUC. Supplementary experiments show retained power under independent null-like replacement and smaller changes from hierarchical dependence modeling. The evidence supports robustness across the evaluated alternatives, not uniform power guarantees or resistance to arbitrary text edits.
发表机构
- University of Chicago(芝加哥大学)
机构由 AI 辅助整理,请以论文原文为准。