arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

偏好优化的归一化奖励

Normalized Rewards for Preference Optimization

Shawn Im, Federico Danieli, Skyler Seto, Barry-John Theobald, Katherine Metcalf

arXiv 2607.16240首次发表:更新:

AI 中文总结

研究针对直接对齐算法过度优化隐式奖励模型的问题,提出添加正则化项的方法,通过研究似然变化分布理解过度优化,应用该正则化于相关方法,实现生成质量与基准能力权衡及奖励建模改进,提升了模型性能。

AI 中文摘要

直接对齐算法(如DPO)已成为训练后将大语言模型与人类偏好对齐的常用方法。然而,观察到这些算法会过度优化其隐式奖励模型,降低偏好响应的可能性,导致偏好数据集中响应的总可能性降低,可能产生不良行为。为抵消此副作用,研究了使用添加正则化项的目标来维持所选和拒绝响应的总长度归一化概率的效果。通过研究有无正则化时响应似然变化在令牌上的分布来更好理解过度优化。发现很大一部分似然变化归因于一小部分异常令牌。将正则化应用于基于参考和无参考方法,取得了生成质量与通用基准能力之间更好的权衡,以及跨数据集奖励建模的改进,如在Llama-3.1-8B-Instruct上,AlpacaEval2分数相对提高>20%,通用基准上相对性能提升>9%,还有效减轻了偏好响应中的位移量。

英文摘要

Direct Alignment Algorithms (DAAs) such as DPO have become a common way to post-train and align LLMs with human preferences. However, DAAs have been observed to over-optimize their implicit reward model and decrease the likelihood of preferred responses. This results in a decrease in the total likelihood assigned to responses seen in the preference dataset, potentially resulting in undesirable behavior. To counteract this undesired side-effect of DAAs, we examine the effect of using objectives that add a regularization term to maintain the total length-normalized probabilities of the chosen and rejected responses. To better understand over-optimization, we investigate how response likelihood changes are distributed over the tokens with and without regularization. We find that a significant portion of the likelihood changes are due to a small set of outlier tokens, which explains how DAAs improve generation quality despite decreasing the likelihoods of chosen responses. We apply the proposed regularization to reference-based (DPO) and reference-free (SimPO) methods and find (1) improved trade-offs between generation quality and general benchmark capability and (2) improvements in reward modeling across datasets. For example, on Llama-3.1-8B-Instruct, we see both a >20% relative increase in AlpacaEval2 scores and >9% relative performance gains on general benchmarks. Additionally, we find that the added regularization term effectively mitigates the amount of displacement within preferred responses overall, and for the outlier tokens specifically, by utilizing low-likelihood tokens.

CommentsICML 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑