在DPO中解耦优化尺度与偏好尺度
Disentangling Optimization Scale from Preference Scale in DPO
AI总结:
本文针对DPO中β同时控制逆偏好噪声尺度与优化步长的纠缠问题,提出居中-软plus重构方法,使两种效应可独立调整,消除损失值跨β不可比性。
AI中文摘要:
直接偏好优化(Direct Preference Optimization,DPO)是一种广泛用于基于偏好数据对齐语言模型的目标函数,系数β通常被解释为控制对参考策略的KL散度约束。本文表明β纠缠了两种不同的作用:它既控制有效逆偏好噪声尺度,同时又重新缩放优化动态,将该尺度与有效步长耦合。因此,在固定学习率下,所获得的策略偏差随β呈非单调变化:在小β时消失于“死区”,在中间值达到峰值,而在更大的β时再次减小。此外,标准DPO的损失值在不同β之间不具有可比性:损失曲线几乎相同的运行,与参考模型的KL散度可能相差数倍。这种纠缠模糊了β的作用,增加了对超参数选择的敏感性,并使学习率调度复杂化。本文提出了一种居中-软plus(centered-softplus)重构,对于β>0,该重构与DPO等价于argmin,同时使逆偏好噪声尺度和学习率的效应显式化且可独立调整。归一化的居中-软plus目标还允许连续的β→0端点,该端点简化为线性偏好间隔目标。
英文摘要:
Direct Preference Optimization (DPO) is a widely used objective for aligning language models from preference data, with the coefficient $β$ commonly interpreted as controlling the KL constraint to a reference policy. We show that $β$ entangles two distinct roles: it governs the effective inverse preference-noise scale and simultaneously rescales the optimization dynamics, coupling this scale with the effective step size. As a consequence, at a fixed learning rate the achieved policy deviation is non-monotone in $β$: it vanishes in a dead zone at small $β$, reaches a peak at an intermediate value, and decreases again for larger $β$. Moreover, standard DPO loss values are not comparable across $β$: runs with nearly identical loss curves can differ several-fold in KL divergence from the reference model. This entanglement obscures the role of $β$, increases sensitivity to hyperparameter choices, and complicates learning-rate scheduling. We propose a centered-softplus reformulation that is argmin-equivalent to DPO for $β>0$, while making the inverse preference-noise-scale and learning-rate effects explicit and independently tunable. The normalized centered-softplus objective also admits a continuous $β\to0$ endpoint that reduces to a linear preference-margin objective.