硬预算重复评估中诚实不确定性的尖锐界限
Sharp Limits for Honest Uncertainty in Hard-Budget Repeated Evaluation
浏览论文内容
中文总结 AI 辅助
本研究提出硬预算下重复评估的尖锐理论界限,证明任务覆盖设计显著降低估计误差和置信区间宽度,使复制与任务覆盖成为高效评估的关键设计变量。
中文摘要 AI 辅助
重复评估可以准确估计基准分数,但仍需要复制来证明狭窄的不确定性。我们在固定网格上刻画了这一要求,该网格包含$M$个任务,每个任务有$L$条二元路径,在硬预算$(M+t)K$下,每条路径的成本至多为$K$次响应或回合。对于固定的$L \ge 3$和$0 < \alpha \le 1/12$,当每个任务都被观测时,最差纯队列上的最优期望宽度为$\Theta_{\alpha,L}([M(t+1)]^{-1/2})$;当允许省略时,则为$\Theta_{\alpha,L}([M(t+\sqrt{M})]^{-1/2})$。下界覆盖了自适应硬预算策略,而固定的随机子集设计通过不一致性证书达到两种速率。联合均值/不一致性区间将任务覆盖法则转化为实用的有限预算推断。在LiveCodeBench的等预算重放中,使用16个模型、880个任务和每个任务5个输出,任务覆盖设计相对于池化均匀采样将中位数点估计MSE降低了87.0%,而联合证书在15/16个面板中产生更窄的置信区间,并将中位数区间宽度降低了30.6%。有限区域分析确定任务覆盖在评估规模上是有效选择,并刻画了队列规模和任务内一致性如何决定有用操作区域。总之,尖锐法则和固定预算证据使复制和任务覆盖成为信息高效重复评估的明确设计变量。
英文摘要
Repeated evaluation can estimate a benchmark score accurately while still requiring replication to certify narrow uncertainty. We characterize that requirement on a fixed grid of $M$ tasks with $L$ binary paths per task under the hard budget $(M+t)K$, where each path costs at most $K$ responses or episodes. For fixed $L \ge 3$ and $0 < α\le 1/12$, the optimal expected width on the worst pure cohort is $Θ_{α,L}([M(t+1)]^{-1/2})$ when every task is observed and $Θ_{α,L}([M(t+\sqrt{M})]^{-1/2})$ when omission is allowed. The lower bounds cover adaptive hard-budget policies, and fixed random-subset designs attain both rates through disagreement certificates. A joint mean/disagreement interval turns the task-covering law into practical finite-budget inference. In an equal-budget LiveCodeBench replay with 16 models, 880 tasks, and five outputs per task, the task-covering design reduces median point-estimation MSE by 87.0\% relative to pooled uniform sampling, while the Joint certificate produces narrower confidence intervals in 15/16 panels and reduces median interval width by 30.6\%. Finite-regime analyses identify task coverage as the effective choice at the evaluated scale and characterize how cohort size and within-task agreement determine the useful operating region. Together, the sharp laws and fixed-budget evidence make replication and task coverage explicit design variables for information-efficient repeated evaluation.
发表机构
- Northwestern University(西北大学)
- Pinterest, Inc.(Pinterest 公司)
机构由 AI 辅助整理,请以论文原文为准。