发表机构
Columbia University(哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究小校准预算下思维链验证器的无分布保证,发现通过弃权(不执行)实现有效性,并提出改进的固定序列证书以提高覆盖率。
AI 中文摘要
预测思维链(CoT)轨迹是否正确的信号通常通过AUC进行比较,但部署此类信号需要一个具有保证的阈值。我们探讨在现实校准预算(从几十到几百个标记问题)下,无分布选择性保证能为CoT验证器提供什么,使用了七个开放模型、五种验证器信号和37,000条分级轨迹。核心观察是“通过弃权(不执行)实现有效性”:一个$(\alpha,\delta)$-有效的过程以概率$P_{\rm fire}$发出证书,其将已发出证书的失败概率限制为$\delta/P_{\rm fire}$,因此一个很少触发的证书可能有效,但每次使用时都可能是错误的。在已知风险的模拟中,标准证书在最多0.3%的校准抽样中失败,但在其触发的抽样中,失败率高达69%。认证下限和Benjamini-Hochberg共形选择的格条件解释了为何这些预算下证书会弃权(不执行),数据也证实了这一点:标准证书要么不返回任何内容,要么返回一个大的接受集,且一个不可读的残差流探针比可读信号提供两到三倍的覆盖率,而交叉拟合重建无法从可读特征中线性恢复这种优势。随后,我们给出一个从下限开始的固定序列证书,该证书在无单调性假设下有效,在每个模型-信号对上覆盖范围都超过Bonferroni证书,并将非空目标$0.75\pi_0$处的覆盖率从0.05提高到0.16,尽管下限使得绝对覆盖率仍然较小。最后,证书无法看到部署后重要的事情:在基准偏移下,接受轨迹中的错误率跟随新任务的基础错误率;在针对验证器的最佳$n$选择下,错误率上升超过目标,而经验失败频率保持在$\delta$以下,因为弃权(不执行)吸收了失败。
英文摘要
Signals that predict whether a chain-of-thought (CoT) trace is correct are compared by AUC, but deploying one requires a threshold with a guarantee. We ask what distribution-free selective guarantees deliver for CoT verifiers at realistic calibration budgets of tens to a few hundred labelled problems, using seven open models, five verifier signals and 37,000 graded traces. The central observation is validity by abstention: an $(α,δ)$-valid procedure that issues a certificate with probability $P_{\rm fire}$ bounds the failure probability of an issued certificate only by $δ/P_{\rm fire}$, so a certificate that rarely fires can be valid and wrong every time it is used. In a simulation with known risk the standard certificate fails in at most 0.3% of calibration draws but in up to 69% of those in which it fires. A certification floor and a lattice condition for Benjamini-Hochberg conformal selection explain why certificates abstain at these budgets, and the data bear them out: the standard certificate returns nothing or a large accepted set, and an unreadable residual-stream probe buys two to three times the coverage of the readable signals, an edge a cross-fitted reconstruction cannot recover linearly from the readable features. We then give a floor-started fixed-sequence certificate, valid without monotonicity assumptions, that covers more than the Bonferroni certificate on every model-signal pair and raises coverage at the non-vacuous target $0.75π_0$ from 0.05 to 0.16, although the floor keeps absolute coverage small. Finally, a certificate cannot see what matters after deployment: under benchmark shift the error among accepted traces tracks the new task's base error, and under best-of-$n$ selection against the verifier it rises past the target while the empirical failure frequency stays below $δ$, because abstention absorbs the failures.
Comments22 pages, 7 figures, 12 tables. Under submission at AISTATS 2027