发表机构
School of Mathematics and Statistics, Xi’an Jiaotong University; Institute of Statistics and Big Data, Renmin University of China(西安交通大学数学与统计学院; 中国人民大学统计与大数据研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出参考零校准框架,利用独立零样本构造更尖锐的拒绝阈值,在保持I型错误控制下提升E-过程推断效率,并在共形鞅和LLM水印检测中验证了性能增益。
AI 中文摘要
E-过程为随时有效推断提供了一个灵活的框架,传统的拒绝边界 $1/\alpha$ 通常由Ville不等式证明其合理性。这种边界具有普遍性,但可能过于保守,因为它没有利用关于零分布的其他信息。我们提出了一种参考零校准框架,利用独立的零样本来构造更尖锐的拒绝阈值,同时保持I型错误控制。为了研究更尖锐阈值带来的统计增益,我们将该框架专门应用于共形鞅。基于直方图投注,我们引入了Krichevsky--Trofimov平滑,并开发了一种用于分布偏移检测的重启混合构造。我们建立了检测能力和检测延迟的定量结果,并刻画了相对于传统Ville边界,拒绝边界的缩减如何转化为能力和延迟的增益。数值实验在多种分布变化下证实了理论发现。最后,我们将所提出的校准策略应用于现有的在线LLM水印检测E-过程,表明该方法可以在不修改底层E-过程的情况下提高顺序检测性能。这些结果表明,参考零校准提供了一种通用且模块化的方式来提高基于E-过程的顺序推断的效率。
英文摘要
E-processes provide a flexible framework for anytime-valid inference, with the conventional rejection boundary $1/α$ typically justified by Ville's inequality. Such a boundary is universal but can be conservative, as it does not exploit additional information about the null distribution. We propose a reference-null calibration framework that uses independent null samples to construct sharper rejection thresholds while preserving type-I error control. To study the statistical gain from sharper thresholds, we specialize the framework to conformal martingales. Building on histogram-based betting, we incorporate Krichevsky--Trofimov smoothing and develop a restart-mixture construction for distribution shift detection. We establish quantitative results for detection power and detection delay and characterize how the reduction in the rejection boundary translates into power and delay gains relative to the conventional Ville boundary. Numerical experiments corroborate the theoretical findings across a range of distributional changes. Finally, we apply the proposed calibration strategy to an existing e-process for online LLM watermark detection, showing that the method can improve sequential detection performance without modifying the underlying e-process. These results demonstrate that reference-null calibration provides a general and modular way to enhance the efficiency of e-process-based sequential inference.
Comments45 pages, 8 figures