arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06893cs.CLcs.LG

弥合离线与迭代对齐差距:基于偏好蒸馏的方法

Towards Bridging the Gap Between Offline and Iterative Alignment via Preference Distillation

Wenbo Zhang, Wenzhuo Zhou, Hengrui Cai, Zhengling Qi

首次发表
浏览论文内容

中文总结 AI 辅助

本文通过受控实验发现显式偏好模型是迭代DPO优于离线方法的关键,据此提出DP3O算法,将显式偏好知识蒸馏到策略优化中,在性能上媲美迭代DPO且训练时间减少42%。

中文摘要 AI 辅助

直接偏好优化(DPO)是一种有前景的离线对齐方法,用于对齐大型语言模型(LLMs),因其简单性、计算效率和对人类偏好的隐式建模而受到关注。有趣的是,DPO的迭代扩展在学术基准上取得了更强的性能,这引发了两个关键问题:(i)为什么迭代方法通常优于离线方法?(ii)它们的优势能否被纳入离线对齐中?为回答第一个问题,我们的受控实验揭示,迭代过程中额外引入的显式偏好模型是其优于离线方法的关键因素。这一洞察使我们肯定地回答第二个问题,并提出蒸馏偏好概率策略优化(DP3O),一种有效且高效的离线对齐算法。DP3O首先使用辅助类LLM学习一个显式偏好模型,然后将其知识蒸馏到策略优化中。理论上,我们证明显式偏好建模比隐式公式具有更好的估计误差控制,且DP3O通过方差缩减实现了比硬标签DPO更紧的泛化界。实证上,我们在广泛的基于聊天的和下游任务上评估DP3O,表明它优于最先进的离线方法,达到与迭代DPO相当的性能,并将训练时间减少约42%,展示了其有效性和效率。

英文摘要

Direct preference optimization DPO is a promising offline approach for aligning large language models (LLMs) due to its simplicity, computational efficiency, and implicit modeling of human preferences. Interestingly, iterative extensions of DPO have achieved stronger performance on academic benchmarks, raising two key questions: (i) Why do iterative methods generally outperform offline ones? (ii) Can their advantages be incorporated into offline alignment? To answer the first question, our controlled experiments reveal that the explicit preference model, additionally introduced in the iterative procedure, is a key factor behind its superiority over offline methods. This insight leads us to answer the second question affirmatively and propose Distilled Preference Probability Policy Optimization (DP3O), an effective and efficient offline alignment algorithm. DP3O first learns an explicit preference model using a helper class of LLMs and then distills its knowledge into policy optimization. Theoretically, we show that explicit preference modeling admits better estimation error control than implicit formulations, and that DP3O achieves a tighter generalization bound than hard-label DPO through variance reduction. Empirically, we evaluate DP3O on a wide range of chat-based and downstream tasks and show that it outperforms state-of-the-art offline methods, achieves performance comparable to iterative DPO, and reduces training time by about $42\%$, demonstrating both its effectiveness and efficiency.

发表机构

  • University of California, Irvine(加利福尼亚大学欧文分校)
  • George Washington University(乔治华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑