最佳采样温度何时随预算升高?Pass@k的充分条件
When Does the Best Sampling Temperature Rise with the Budget? Sufficient Conditions for Pass@k
AI总结:
该研究针对使pass@k最大化的采样温度随预算升高的经验模式,给出了总体层面的形式化充分条件,推导了相关相图与核表示,构建了该现象的条件理论。
AI中文摘要:
使pass@k最大化的温度,在采样预算较小时通常较低,而在预算较大时更高。从Codex到近期的多样本推理研究都报告了这一模式。但这并非pass@k的代数性质:正如Slocum等人(ICLR 2025)所观察到的,对于单个固定任务,最大化温度与k无关。基于该固定任务的观察及难易任务的解释,我们为该聚合模式给出了形式化的总体层面充分条件。对于任务X,设p_t(X)为温度t下的单样本成功概率,并定义条件对数成功响应m_t(u)=E[ṗ_t(X)|p_t(X)=u]/u。若m_t(u)随当前成功概率非递增,则聚合pass@k的标准化温度导数随k非递减。因此,导数符号在各预算间呈嵌套关系;若每条温度-性能曲线严格单峰,其唯一最大化点随k非递减。该证明将机制识别为朝向低成功任务的单调似然比幂倾斜。我们推导了闭式双分层相图,包含上升与下降 regime,并表明边际温度导数具有精确的Beta(2,k)核表示,其核集中于约1/k的单样本成功处。将该尺度解释为任务层面定位,还需要在零附近存在规则、非消失的密度响应因子。带符号矩表示给出了诊断形状约束,而简短附录记录了现有多配置分配公式的精确离散细化。本研究未训练任何语言模型,也未使用任何模型查询作为实验测量:其贡献是针对已确立的经验现象的条件理论,所做假设可在未来工作中进行检验。
英文摘要:
The temperature that maximizes pass@$k$ is often low for a small sampling budget and higher for a large budget. This pattern has been reported from Codex through recent multi-sample inference studies. It is not an algebraic property of pass@$k$: as Slocum et al. (ICLR 2025) observe, for one fixed task the maximizing temperature is independent of $k$. Building on that fixed-task observation and the hard/easy-task explanation, we give a formal population-level sufficient condition for the aggregate pattern. For task $X$, let $p_t(X)$ be one-sample success probability at temperature $t$, and define the conditional log-success response $m_t(u)=\mathbb{E}[\dot p_t(X)\mid p_t(X)=u]/u$. If $m_t(u)$ is nonincreasing in current success probability, then the normalized temperature derivative of aggregate pass@$k$ is nondecreasing in $k$. Consequently, derivative signs are nested across budgets; if each temperature-performance curve is strictly single-peaked, its unique maximizer is nondecreasing in $k$. The proof identifies the mechanism as a monotone-likelihood-ratio power tilt toward lower-success tasks. We derive a closed-form two-stratum phase diagram, including upward and downward regimes, and show that the marginal temperature derivative admits an exact $\mathrm{Beta}(2,k)$ kernel representation whose kernel concentrates at one-sample success of order $1/k$. Interpreting that scale as task-level localization additionally requires a regular, nonvanishing density-response factor near zero. A signed-moment representation yields diagnostic shape restrictions, while a short appendix records exact discrete refinements of the existing multi-configuration allocation formulation. No language model is trained, and no model query is used as an experimental measurement: the contribution is a conditional theory of an established empirical phenomenon, with assumptions that can be tested in future work.