arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

绘制廉价,信任昂贵:认证测试时扩展曲线

Cheap to Draw, Expensive to Trust: Certifying Test-Time Scaling Curves

Sohail, Sarkar, Shakuntala Baichoo

arXiv 2609.40190首次发表:更新:

发表机构

PMCC AI Lab, Peter Munk Cardiac Centre; University Health Network, Toronto, Ontario, Canada(PMCC AI实验室,彼得·芒克心脏中心; 大学健康网络,多伦多,安大略省,加拿大)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种配对审计方法,以最小最大成本认证测试时采样扩展曲线,通过区分问题间与问题内方差,显著降低所需答案数量,并适用于pass@k和多数投票。

AI 中文摘要

采样多个答案并保留验证器评分最高的一个,是在测试时购买准确性的最简单方法之一。其效果以扩展曲线的形式报告:准确率与采样答案数量$k$的函数关系。这条曲线绘制廉价,但信任昂贵。从曲线上读取的预算是在查看所有点之后选择的,因此只有同时覆盖所有预算的置信带才能保护该选择,而在一个100道题的基准上,一个固定的精确二项式设计需要生成192,000个答案,才能以95%的置信度认证64个预算,误差在$\pm1/32$以内。大部分成本用于支付错误的的不确定性。基准是固定的问题列表;在预算64时,所选答案正确性的方差中约有四分之三存在于问题之间,而重新访问每个问题的审计无需为此付费。我们推导了认证整条曲线的最小最大成本,精确到对数因子。它由三部分组成:校准分数分布的尾部、区分问题,以及沿曲线累加的问题内噪声。在单个基准上,最后一部分在答案到问题的最佳分配下锐化为单个答案影响力的方差,这是每个有效审计都必须支付的,而学习分配的审计在精度增长时(精确到对数)可以达到该方差。基于同一问题两次独立抽取的指数不等式构建的配对审计无需预试验。在185个保留的分数池上,它在64个预算时使用的答案数量是最便宜的竞争认证审计的0.74倍,在1024个预算时为0.53倍;在一项新生成的MMLU-Pro研究中,它用79,133个答案认证了曲线,与事先根据研究的问题内方差拟合的成本定律预测值相差在0.6%以内。同样的路径也适用于pass@$k$和多数投票的认证,且置信带可扩展到问题总体以及依赖于先前答案的答案。

英文摘要

Sampling several answers and keeping the one a verifier scores highest is one of the simplest ways to buy accuracy at test time. Its effect is reported as a scaling curve: accuracy against the number $k$ of sampled answers. The curve is cheap to draw and expensive to trust. A budget read off it is chosen after looking at every point, so only a band that covers all budgets at once protects the choice, and on a 100-question benchmark a fixed exact-binomial design needs 192,000 generated answers to certify 64 budgets to within $\pm1/32$ at 95%. Most of that cost pays for the wrong uncertainty. A benchmark is a fixed list of questions; at budget 64, about three quarters of the variance of a selected answer's correctness lies between questions, and an audit that revisits every question need not pay for it. We derive the minimax cost of certifying the whole curve, up to logarithmic factors. It has three parts: calibrating the tail of the score distribution, telling the questions apart, and within-question noise summed along the curve. At a single benchmark the last part sharpens to the variance of one answer's influence under the best allocation of answers to questions, which every valid audit pays and an audit that learns the allocation attains, up to a logarithm, as the precision grows. A paired audit built on an exponential inequality for two independent draws at the same question needs no pilot. On 185 held-out score pools it uses 0.74 times the answers of the cheapest competing certified audit at 64 budgets and 0.53 times at 1,024, and on a newly generated MMLU-Pro study it certified the curve with 79,133 answers, within 0.6% of what a cost law fitted beforehand predicted from the study's within-question variance. The same paths certify pass@$k$ and majority voting, and the bands extend to populations of questions and to answers that depend on earlier ones.

Comments32 pages, 10 figures, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑