arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.09245stat.MLcs.AIcs.LG

固定采样次数 pass@k 评估能够识别什么

What Fixed-Rollout pass@k Evaluations Can Identify

Pranav Singh, Prashant Singh

首次发表
浏览论文内容

中文总结 AI 辅助

本文证明在固定采样次数下,pass@k 外推(k>n)不可识别,并给出识别区间与报告标准,为推理时缩放定律提供非参数基线。

中文摘要 AI 辅助

重复采样评估日益将 pass@k 外推到每个问题收集的样本数 n 之外。我们证明,在合并/随机任务条件二项模型中,固定 n 的成功计数仅能识别潜在每任务成功分布的 n 个自由矩。因此,对于 k ≤ n,直接 pass@k 是可识别的,但对于 k > n,一般外推的 pass@k、尾部指数和尾部常数是不可识别的,即使在相同的采样预算下拥有任意多的可交换任务。这比通常的估计量在 n 之外未定义这一观察更强:它刻画了固定深度计数律实验中缺失的信息。我们给出了保持计数律但外推不相容的精确构造,陈述了唯一的例外扩展情况,并通过 Hausdorff 主表示计算了尖锐的总体识别区间。在 Brown 等人公开的每个问题 10,000 次采样的数据上,反事实的 n = 16 评估使得 k = 1000 时的失败率在四个 MATH/GSM8K/CodeContests 配置中模糊性因子从 1.5 到超过 2,600。校准表明,仅中间规模的失败份额并不能决定区间宽度。我们的结果并不拒绝参数化的推理时缩放定律;它提供了非参数基线,用以评估这些定律的假设。我们给出了一个精确、保守的单坐标有限任务置信度证书,以及一个区分直接估计、识别集和模型条件预测的报告标准。

英文摘要

Repeated-sampling evaluations increasingly extrapolate pass@k far beyond the number n of samples collected per problem. We show that, in the pooled/random-task conditional-Binomial model, fixed-n success counts identify only the n free moments of the latent per-task success distribution. Consequently, direct pass@k is identified for k <= n, but generic extrapolated pass@k, tail exponents, and tail constants are not identified for k > n, even with arbitrarily many exchangeable tasks at the same rollout budget. This is stronger than the observation that the usual estimator is undefined beyond n: it characterizes the information missing from the fixed-depth count-law experiment. We give exact count-law-preserving constructions with incompatible extrapolations, state the exceptional unique-extension case, and compute sharp population identified intervals through Hausdorff principal representations. On the public 10,000-rollout-per-problem release of Brown et al., counterfactual n = 16 evaluations leave failure at k = 1000 ambiguous by factors from 1.5 to over 2,600 across four MATH/GSM8K/CodeContests configurations. The calibration shows that intermediate-scale failure share alone does not determine width. Our result does not reject parametric inference-time scaling laws; it supplies the nonparametric baseline against which their assumptions can be evaluated. We give an exact, conservative one-coordinate finite-task confidence certificate and a reporting standard separating direct estimates, identified sets, and model-conditioned forecasts.

发表机构

  • Indian Institute of Technology Ropar(印度理工学院鲁尔基分校(罗帕尔分校))

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑