发表机构
School of Mathematical Sciences, Shanghai Jiao Tong University; School of Management, Fudan University; Department of Industrial Engineering & Management, Shanghai Jiao Tong University(上海交通大学数学科学学院; 复旦大学管理学院; 上海交通大学工业工程系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FissionGuard通过数据拆分技术,在控制错误发现率(FDR)的前提下,实现了对文本局部水印比例的有效测试,提升了水印token覆盖率。
AI 中文摘要
水印检测器即使在文本中只有一小部分token带有水印时也能标记该文本。我们研究序列审计以识别局部段落,其水印比例超过预先指定的阈值。每个警报确定候选段落的终点,而起点则从更早的token中选择。重复使用相同的证据来定位和测试段落可能会使推断失效。我们提出FissionGuard,它使用数据拆分从每个token级p值构造警报值和验证值。警报值指导警报和定位,验证值通过部分合取产生段落级p值,排除尽可能多的最强信号,排除数量由比例原假设允许的比例决定。这些p值用于在线多重测试程序分配的水平上测试相应的假设。在所述水印生成模型下,允许水印位置处的token级p值之间存在依赖关系,我们证明在每个固定有限时间范围内,报告段落中存在有限样本错误发现率(FDR)控制。我们还推导了测试非零候选的功效的有限样本下界。在合成流以及Mistral和Qwen生成的文本上,针对四种标记布局的实验表明,与样本拆分方法相比,水印token覆盖率有所提高,且经验FDR低于目标值。
英文摘要
A watermark detector can flag a text even when only a small fraction of its tokens are watermarked. We study sequential auditing to identify local passages whose watermark proportions exceed a prespecified threshold. Each alarm fixes a candidate passage's endpoint, while its start is chosen from earlier tokens. Reusing the same evidence to locate and test a passage can invalidate inference. We propose FissionGuard, which uses data fission to construct an alarm value and a validation value from each token-level $p$-value. The alarm values guide alarms and localization. The validation values yield passage-level $p$-values through partial conjunction, excluding as many of the strongest signals as the proportion null allows. These $p$-values are used to test the corresponding hypotheses at levels assigned by an online multiple-testing procedure. Under the stated watermark generation model, allowing dependence among token-level $p$-values at watermarked positions, we prove finite-sample false discovery rate (FDR) control among reported passages at every fixed finite horizon. We also derive a finite-sample lower bound on power for testing nonnull candidates. Experiments on synthetic streams and Mistral and Qwen generations across four marking layouts show gains in watermarked-token coverage over sample splitting, with empirical FDR below the target.