arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.01511cs.CL

GAW-PO:基于梯度对齐令牌权重的偏好优化

GAW-PO: Preference Optimization with Gradient-Aligned Token Weights

  • National University of Science and Technology POLITEHNICA Bucharest(布加勒斯特国立科技理工大学)
  • NVIDIA(英伟达)

机构由 AI 辅助整理,请以论文原文为准。

Andreea Dutulescu, Stefan Ruseti, Mihai Masala, Traian Rebedea, Mihai Dascalu

AI总结:

提出GAW-PO方法,通过梯度对齐的令牌重加权,在DPO中区分被拒绝令牌的惩罚强度,提升偏好优化性能,在11个基准上平均优于标准DPO 0.97分,且对激进优化更鲁棒。

AI中文摘要:

大多数偏好优化方法,如直接偏好优化(DPO),在响应级别应用偏好监督,尽管自回归语言模型是逐令牌优化的。因此,被拒绝响应中的所有令牌都会对负面训练信号做出贡献,包括可能编码对偏好响应有用的行为的令牌。我们提出了GAW-PO,一种用于DPO的梯度对齐令牌重新加权方法,该方法估计对于每个被拒绝的令牌,惩罚它是否会干扰偏好更新方向。梯度与偏好行为强对齐的令牌获得较弱的负面贡献,而冲突的令牌则保留更强的惩罚。我们的方法在评估的偏好优化方法中取得了最高的平均性能,在涵盖数学、推理、编码和问答的11个基准上,比标准DPO提高了0.97分,比最强的竞争基线提高了0.65分。我们进一步表明,梯度对齐加权对激进的偏好优化具有更强的鲁棒性:随着DPO正则化参数$\eta$的减小,标准DPO性能急剧下降,而GAW-PO继续改进。这些结果表明,考虑被拒绝令牌更新与偏好行为之间的相互作用,为偏好优化提供了一种有效的令牌级信用分配形式。

英文摘要:

Most preference optimization methods, such as Direct Preference Optimization (DPO), apply preference supervision at the response level, although autoregressive language models are optimized token by token. As a result, all tokens in a rejected response contribute to the negative training signal, including tokens that may encode behavior that is useful for the preferred response. We introduce GAW-PO, a gradient-aligned token reweighting method for DPO that estimates, for each rejected token, whether penalizing it would interfere with the preferred update directions. Tokens whose gradients are strongly aligned with the preferred behavior receive a weaker negative contribution, while conflicting tokens retain a stronger penalty. Our method achieves the highest average performance among the evaluated preference-optimization methods, improving by 0.97 points over standard DPO and 0.65 points over the strongest competing baseline across 11 benchmarks spanning mathematics, reasoning, coding, and question answering. We further show that gradient-aligned weighting is substantially more robust to aggressive preference optimization: as the DPO regularization parameter $β$ decreases, standard DPO degrades sharply, whereas GAW-PO continues to improve. These results suggest that accounting for the interaction between rejected-token updates and preferred behavior provides an effective form of token-level credit assignment for preference optimization.

↑