AI 中文总结
本文证明在解码时策略组合中,Metropolis-Hastings 校正方法在任意 rollout 预算下均优于重要性重采样,并给出改进下界与渐近误差分析,实验验证其有效性。
AI 中文摘要
大型语言模型(LLM)的后训练通常需要探索多个奖励之间的权衡,但针对每个权衡进行重新训练代价高昂。解码时策略组合允许在推理时通过组合奖励特定策略来调整这些权衡。这种组合针对完整响应上策略概率的加权乘积,但标准实现组合了它们的下一词元概率,通常会引入采样偏差。我们分析了一种基于独立 Metropolis-Hastings(MH)的已知迭代校正方法。我们的主要结果表明,对于每个 rollout 预算,MH 产生的输出分布与目标的接近程度至少与使用相同预算的采样重要性重采样(SIR)相当,这是通过每个凸 f-散度来衡量的。我们还推导了 MH 相对于未校正解码器在共识目标中的改进下界,该目标衡量与所提供策略的一致性。我们进一步刻画了校正采样误差在两个渐近状态下的表现:当奖励特定策略趋于一致时,以及当目标与未校正解码器概率之间的对数比率波动越来越剧烈时(这在长响应中可能发生)。我们通过可枚举和 LLM 规模设置中的实验补充了我们的分析。
英文摘要
Post-training a large language model (LLM) often requires exploring trade-offs between multiple rewards, but retraining for each trade-off is expensive. Decoding-time policy composition allows these trade-offs to be adjusted by combining reward-specific policies at inference time. This composition targets a weighted product of the policies' probabilities over complete responses, but standard implementations combine their next-token probabilities, generally introducing sampling bias. We analyze a known iterative correction based on independence Metropolis-Hastings (MH). Our main result shows that, for every rollout budget, MH produces an output distribution at least as close to the target as sampling-importance-resampling (SIR) with the same budget, as measured by every convex f-divergence. We also derive a lower bound on MH's improvement over the uncorrected decoder in a consensus objective measuring agreement with the supplied policies. We further characterize the correction's sampling error in two asymptotic regimes: when the reward-specific policies approach agreement, and when the log ratio between target and uncorrected-decoder probabilities fluctuates increasingly widely, as can happen for long responses. We complement our analysis with experiments in enumerable and LLM-scale settings.
Comments56 pages, 6 figures