发表机构
Stanford University; University of Pennsylvania(斯坦福大学; 宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对混合源大语言模型文本的水印定位问题,将其建模为token级多重检验,提出自适应阈值方法,实现最优发现边界,经模拟和实验验证了理论与实际性能。
AI 中文摘要
水印为大语言模型(LLM)生成的文本提供了一种原则性的认证方式。但在实践中,最终文本可能是混合源的,经过重写、插入、删除或paraphrasing后,水印证据仅保留在部分token位置。尽管已有研究探讨了水印信号的全局检测,但这类信号能否被定位仍不明确。我们将水印定位建模为基于枢轴统计量的token级多重检验问题,其中包含一个潜在指示器,记录每个位置是否保留了水印相关性。在由信号稀疏性、下一个token集中度和有效词汇量增长的指数所索引的渐近 regime下,我们推导了全局检测的清晰边界,以及基于坐标枢轴的定位规则类中发现和分类的相变。我们证明,在该类参数 regime内,发现比检测更难,且无法实现一致分类。随后,我们开发了一种自适应阈值方法,该方法无需知晓指数或随时间变化的下一个token分布,而是使用数据驱动的存活水印比例估计。该方法相对于同质枢轴规则达到了最优发现边界和接近最优的发现 power。模拟验证了理论相变,而在模型生成文本上的实验证明了常见编辑机制下的实际定位性能。
英文摘要
Watermarking provides a principled way to authenticate text generated by large language models (LLMs). In practice, however, the final text may be mixed-source, with watermark evidence surviving at only a subset of token positions after rewriting, insertion, deletion, or paraphrasing. Although prior work has studied global detection of watermark signals, when such signals can be localized remains unclear. We formulate watermark localization as a token-level multiple-testing problem based on pivotal statistics, with a latent indicator recording whether watermark dependence survives at each position. Under an asymptotic regime indexed by exponents for signal sparsity, next-token concentration, and effective-vocabulary growth, we derive a sharp boundary for global detection and phase transitions for discovery and classification within the class of coordinatewise pivot-based localization rules. We show that discovery is strictly harder than detection and that consistent classification is impossible across the parameter regime within this class. We then develop an adaptive thresholding method that does not require knowledge of the exponents or time-varying next-token distributions, but uses a data-driven estimate of the surviving watermark fraction. The method attains the optimal discovery boundary and near-optimal discovery power relative to homogeneous pivot-based rules. Simulations support the theoretical phase transitions, while experiments on model-generated texts demonstrate practical localization performance under common edit mechanisms.
Comments66 pages, 13 figures