发表机构
LMU Munich & MCML; Carnegie Mellon University; Meta(慕尼黑大学与慕尼黑机器学习中心; 卡内基梅隆大学; Meta)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出一种仅基于样本的框架,利用顺序统计结构估计BoN策略值,开发双稳健估计器并给出两种预算选择规则,在合成实验和GSM8K上验证了准确性与有效性。
AI 中文摘要
Best-of-N (BoN) 是一种常见的推理时对齐方法,它从参考模型生成的N个样本中选择得分最高的响应。在仅能访问样本的情况下,从日志数据中评估BoN策略具有挑战性,因为标准的离线策略估计器需要依赖于不可用的响应似然度的密度比。在本文中,我们提出了一种仅基于样本的框架,用于在无需访问这些似然度的情况下评估和选择BoN策略。我们证明,BoN的顺序统计结构使得所需的密度比可以通过得分秩概率来表达,而这些概率仅从样本中即可估计。随后,我们开发了一种BoN策略值的双稳健估计器(BoN-DR),该估计器能有效地跨候选预算共享辅助样本池。即使在奖励估计器误设的情况下,我们也建立了有效的渐近推断,并证明了我们的BoN-DR估计器的效率。由于较大的预算可能放大得分函数中的误差并导致奖励过度优化,我们推导了两种选择规则:(i)最大化估计的策略值,以及(ii)最大化相对于参考策略改进的置信下界,该下界考虑了估计不确定性并提供了无损害保证。在合成实验和GSM8K(使用多个参考模型和奖励模型)中,我们的框架能准确估计BoN策略值并选择有效的采样预算。
英文摘要
Best-of-N (BoN) is a common inference-time alignment method that selects the highest-scoring response among N samples from a reference model. Evaluating BoN policies from logged data is challenging under sample-only access because standard off-policy estimators require density ratios that depend on unavailable response likelihoods. In this paper, we propose a sample-only framework for evaluating and selecting BoN policies without access to these likelihoods. We show that the order-statistic structure of BoN allows the required density ratios to be expressed through score-rank probabilities that are estimable from samples alone. We then develop a doubly robust estimator of the BoN policy value (BoN-DR) that efficiently reuses a shared auxiliary sample pool across candidate budgets. We establish valid asymptotic inference even under reward estimator misspecification and prove the efficiency of our BoN-DR estimator. Since larger budgets can amplify errors in the score function and lead to reward overoptimization, we derive two selection rules: (i) maximizing the estimated policy value and (ii) maximizing a lower confidence bound on the improvement over the reference policy, which accounts for estimation uncertainty and provides a no-harm guarantee. Across synthetic experiments and GSM8K with multiple reference and reward models, our framework accurately estimates BoN policy values and selects effective sampling budgets.
CommentsPreprint