arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

不确定性归一化边际用于直接偏好优化

Uncertainty-Normalized Margins for Direct Preference Optimization

Sadegh Khorasani, Petrus Mikkola, Matthias Grossglauser

arXiv 2609.38647首次发表:更新:

发表机构

EPFL; University of Helsinki(洛桑联邦理工学院; 赫尔辛基大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出不确定性归一化边际DPO(UNM-DPO),结合强度相关边际与学习提示尺度,通过两种训练目标及长度归一化提升偏好优化性能,在HelpSteer和AlpacaEval上显著优于DPO和SimPO。

AI 中文摘要

直接偏好优化(DPO)通过具有共同噪声尺度的Bradley-Terry模型对二元偏好进行建模,没有明确考虑偏好强度或来自人类反馈的提示相关不确定性。我们引入了不确定性归一化边际DPO(UNM-DPO),它将强度相关边际与学习的提示尺度相结合。受异方差Bradley-Terry模型的启发,我们开发了两种训练目标。两者都比较首选和拒绝响应的隐式奖励,这些奖励源自响应对数概率比率相对于参考策略。仅优势(AO)在减去边际之前将奖励差异除以提示尺度;整体残差(WR)在除以尺度之前减去边际。对于WR比较模型,我们建立了已知边际使提示尺度可识别的必要且充分条件。我们提出了一种学习尺度的实用程序。基于WR,我们引入了ULNM-DPO-WR,它通过长度对每个响应的隐式奖励进行归一化。我们在HelpSteer2和HelpSteer3上使用Skywork奖励模型作为评判者,将我们的方法与DPO及相关基线进行了评估。使用Llama-3.1-8B-Instruct,ULNM-DPO-WR在评估面板上对匹配的DPO实现了68.00%和65.31%的平局调整胜率,具有更高的平均奖励和平均更短的响应。在AlpacaEval上使用GPT-4.1评判者和GPT-4-Turbo参考答案,相同的8B策略实现了21.62%的长度控制胜率,而DPO为16.39%,SimPO为15.30%。这些结果证明了将偏好强度边际、学习提示尺度和长度归一化相结合用于策略优化的潜力。

英文摘要

Direct preference optimization (DPO) models binary preferences through a Bradley-Terry model with a common noise scale, without explicitly accounting for preference strength or prompt-dependent uncertainty from human feedback. We introduce uncertainty-normalized margin DPO (UNM-DPO), which combines strength-dependent margins with a learned prompt scale. Motivated by a heteroskedastic Bradley-Terry model, we develop two training objectives. Both compare the implicit rewards of preferred and rejected responses, derived from response log-probability ratios to a reference policy. Advantage-only (AO) divides this reward difference by the prompt scale before subtracting the margin; whole-residual (WR) subtracts the margin before dividing by the scale. For the WR comparison model, we establish a necessary and sufficient condition under which known margins make the prompt scale identifiable. We introduce a practical procedure for learning the scale. Building on WR, we introduce ULNM-DPO-WR, which normalizes each response's implicit reward by its length. We evaluate our methods against DPO and related baselines on HelpSteer2 and HelpSteer3, using the Skywork reward model as a judge. With Llama-3.1-8B-Instruct, ULNM-DPO-WR achieves tie-adjusted win rates against matched DPO of 68.00% and 65.31% on evaluation panels, with higher mean rewards and shorter responses on average. On AlpacaEval with a GPT-4.1 judge and GPT-4-Turbo reference answers, the same 8B policy achieves a length-controlled win rate of 21.62%, compared with 16.39% for DPO and 15.30% for SimPO. These results demonstrate the potential of combining preference-strength margins, learned prompt scales, and length normalization for policy optimization.

Comments36 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑