AI 中文总结
针对DPO中高频词元对称出现导致梯度纠缠、偏好信号被稀释的问题,提出各向异性DPO(ADPO)及频率硬DPO,通过固定掩码置零高频词元隐式奖励,在多个基准上超越标准DPO,实现零开销的稳健偏好对齐。
AI 中文摘要
直接偏好优化(DPO)通过优化序列级词元级隐式奖励差异之和来对齐语言模型。然而,我们在此公式中发现了一个普遍存在的病理现象:一小部分不成比例的高频词元类型主导了累积序列得分,同时对称地出现在偏好响应和不偏好响应中。具体而言,在Anthropic HH-RLHF上使用规范的Qwen分词时,仅69种词元类型就占所有响应词元的55.1%,占配对内共享词元质量的85.9%,其偏好侧特异性显著低于词汇表中的其余部分。这种对称的普遍性导致梯度纠缠,并稀释了通过目标传播的判别性偏好信号。为解决此问题,我们引入了各向异性DPO(ADPO)及其规范实现——频率硬DPO。使用固定的、与标签无关的词汇表掩码,我们的方法将高频响应词元的隐式奖励贡献置零,同时为信息性位置分配单位权重,从而在不修改偏好对、不丢弃上下文或不引入学习参数的情况下抑制梯度干扰。此处,“各向异性”指非均匀的词元级目标加权,而非表征几何。在AlpacaEval、MT-Bench和Arena-Hard上的广泛实证评估表明,频率硬DPO在Qwen-2.5-7B-Instruct和Llama-3-8B-Instruct上始终优于标准DPO,证明选择性地掩蔽共享高频词元提供了一种有效、零开销的稳健偏好对齐机制。
英文摘要
Direct Preference Optimization (DPO) aligns language models by optimizing over sequence-level sums of token-wise implicit reward differences. However, we identify a pervasive pathology in this formulation: a disproportionately small subset of high-frequency token types dominates cumulative sequence scores while appearing symmetrically across both preferred and dispreferred responses. Specifically, under canonical Qwen tokenization on Anthropic HH-RLHF, merely 69 token types account for $55.1\%$ of all response tokens and $85.9\%$ of within-pair shared token mass, exhibiting substantially lower preference-side specificity than the remaining vocabulary. This symmetric ubiquity induces gradient entanglement and dilutes the discriminative preference signal propagated through the objective. To resolve this issue, we introduce \emph{Anisotropic DPO} (\textsf{ADPO}) and its canonical realization, \emph{Frequency-Hard DPO}. Using a fixed, label-agnostic vocabulary mask, our method zeroes the implicit reward contribution of high-frequency response tokens while assigning unit weight to informative positions, thereby suppressing gradient interference without modifying preference pairs, discarding context, or introducing learned parameters. Here, \emph{anisotropy} designates non-uniform token-level objective weighting rather than representational geometry. Extensive empirical evaluations on AlpacaEval, MT-Bench, and Arena-Hard demonstrate that Frequency-Hard DPO consistently outperforms standard DPO across Qwen-2.5-7B-Instruct and Llama-3-8B-Instruct, establishing that selectively masking shared high-frequency tokens offers an effective, zero-overhead mechanism for robust preference alignment.