arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21899cs.LGcs.ITmath.IT

ExpBoN:用于高效测试时大语言模型对齐的指数噪声最佳n采样

ExpBoN: Exponential-Noise Best-of-$n$ for Efficient Test-Time LLM Alignment

Yanxiao Liu, Sicheng Wan, Deniz Gündüz

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出ExpBoN,一种基于指数噪声机制的软BoN采样方法,实现指数级快速收敛,并集成到GSI框架形成ExpGSI,在保持精度的同时大幅降低测试时对齐的计算成本。

中文摘要 AI 辅助

最佳n采样(Best-of-$n$, BoN)是一种简单而有效的推理时对齐方法,但硬最大化仅能对奖励与分布偏移之间的权衡提供粗略控制。软最佳n采样(Soft Best-of-$n$, Verdun等人,2025)提供了更平滑的控制,并收敛到与KL正则化奖励最大化相关的最优分布。在本文中,我们引入了ExpBoN,一种基于指数噪声报告-无噪声最大机制(exponential-noise report-noisy-max mechanism)的替代性软BoN方法。它允许精确的有限n分解,从而在总变差距离、期望奖励以及KL散度的两个方向上都实现指数级快速收敛。我们对其收敛性和遗憾行为提供了全面的理论分析。我们进一步将ExpBoN集成到引导式推测推理(guided speculative inference, GSI)框架(Geuter, Mroueh, 和 AlvarezMelis,2025)中,形成ExpGSI,用于高效的奖励引导的大语言模型对齐。ExpGSI在保持相当准确性的同时,大幅降低了计算成本。在MATH500、MMLU-STEM和Minerva Math数据集上,使用Qwen2.5-Math和Qwen3模型家族的实验表明,对于Qwen2.5-Math,ExpGSI在候选预算范围内将估计计算量减少了14%至39%;对于Qwen3,在n=16时减少高达45%。总体而言,我们的结果为指数噪声BoN和高效的测试时大语言模型对齐提供了理论和算法基础。

英文摘要

Best-of-$n$ (BoN) sampling is a simple yet effective inference-time alignment method, but hard maximization provides only coarse control over the trade-off between reward and distribution shift. Soft Best-of-$n$ (Verdun et al. 2025) provides smoother control and converges to the optimal distribution associated with KL-regularized reward maximization. In this paper, we introduce ExpBoN, an alternative soft BoN method based on the exponential-noise report-noisy-max mechanism. It admits an exact finite-$n$ decomposition, which yields exponentially fast convergence in total variation, expected reward, and both directions of KL divergence. We provide comprehensive theoretical analyses of its convergence and regret behavior. We further integrate ExpBoN into the guided speculative inference (GSI) framework (Geuter, Mroueh, and AlvarezMelis 2025), resulting in ExpGSI, for efficient reward-guided LLM alignment. ExpGSI yields substantial reductions in computational cost while maintaining comparable accuracy. Experiments on MATH500, MMLU-STEM, and Minerva Math with the Qwen2.5-Math and Qwen3 model families show that ExpGSI reduces estimated computation by $14\%$-$39\%$ across candidate budgets for Qwen2.5-Math and by up to $45\%$ at $n=16$ for Qwen3. Overall, our results provide a theoretical and algorithmic foundation for exponential-noise BoN and efficient test-time LLM alignment.

发表机构

  • Imperial College London(伦敦帝国理工学院)
  • University of Washington(华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

↑