arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Se-DPO:用于直接偏好优化的自进化令牌信用

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization

Wenxiao Zhao, Shu Wang, Ying Nian Wu

arXiv 2608.09568首次发表:更新:

发表机构

University of California, Los Angeles; Shanghai AI Laboratory(加州大学洛杉矶分校; 上海人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Se-DPO是一种在DPO训练中基于模型自身内部信号动态校准令牌信用的方法,无需外部模型,在AlpacaEval 2和Arena-Hard上较DPO分别提升9.8和12.2个百分点。

AI 中文摘要

直接偏好优化(DPO)通过均匀求和聚合令牌级对数概率比,隐含地将所有令牌视为对偏好信号的贡献相同,但单个令牌对偏好信号的贡献存在差异。我们引入令牌信用,该信用会根据每个令牌对偏好结果的贡献来调整其KL正则化。我们推导得出,有效令牌信用与每个令牌的隐式奖励大小成正比,且观察到该量在训练过程中会发生显著变化,这意味着静态令牌信用会随着训练推进而日益失配。在本研究中,我们提出Se-DPO(用于DPO的自进化令牌信用),这是一种在DPO训练期间从模型自身不断演变的内部信号中推导令牌信用的实时机制。由于奖励信号在不同位置的可靠性存在差异,Se-DPO会根据每个令牌贡献的强度和置信度来校准令牌信用。Se-DPO无需外部模型,仅添加一个计算开销极小的轻量级校准网络。实验表明,Se-DPO在AlpacaEval 2上较DPO提升了9.8个百分点,在Arena-Hard上提升了12.2个百分点。

英文摘要

Direct Preference Optimization (DPO) aggregates token-level log-probability ratios via uniform summation, implicitly treating all tokens as contributing equally to the preference signal. However, the contribution of individual tokens to the preference signal varies. We introduce token credit, which modulates each token's KL regularization based on its contribution to the preference outcome. We derive that effective token credit is proportional to the magnitude of each token's implicit reward, and observe that this quantity evolves substantially during training. This implies that static token credit becomes increasingly misaligned as training progresses. In this work, we propose Se-DPO (Self-Evolving Token Credit for DPO), a live mechanism that derives token credit from the model's own evolving internal signals during DPO training. Since the reward signal varies in reliability across positions, Se-DPO calibrates token credit based on both the strength and the confidence of each token's contribution. Se-DPO requires no external models, adding only a lightweight calibration network with minimal computational overhead. Experiments show that Se-DPO improves over DPO by up to 9.8 points on AlpacaEval~2 and 12.2 points on Arena-Hard.

Comments16 pages, 2 figures, COLM2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑