arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32579cs.LG

DimPO:使用偏好优化进行注意力降维

DimPO: Dimensionality Reduction for Attention using Preference Optimization

Vojtěch Lanz, Yufei Cui, Prasanna Parthasarathi

AI总结:

DimPO通过偏好优化与top-k交叉熵结合,在注意力降维中保留关键键的排序和集中度,在长上下文任务中优于KL散度方法,实现更好的下游性能。

AI中文摘要:

线性投影可以在不更新预训练模型的情况下降低查询向量和键向量的维度,但尚不清楚哪种训练目标最能保持模型行为。我们探究,在长上下文场景中,对键的偏好以及最高权重键上的注意力质量是否比使用KL散度匹配完整注意力分布提供更好的信号。我们引入了DimPO,它将列表式偏好优化与轻量级top-k交叉熵项相结合,以实现头部保真度。DimPO从冻结语言模型的注意力模式中离线训练,每层一个映射,查询和键共享,并与其他层分开训练。在LLaMA3.2-3B、LLaMA3.1-8B、Qwen2.5-7B和Qwen3-4B Instruct模型上,成对偏好目标优于三元组基线,并在将最后40%的层投影到一半维度时,在短上下文任务中保留了原始分数的98%。随着投影层数增加或在长上下文RULER上,它们迅速退化。相比之下,KL和DimPO在训练期间使用每个键,在8B模型上投影多达50%的层时,保留了原始RULER 4k分数的大约95%。基于KL的投影保持更接近原始注意力分布和注意力输出,但DimPO实现了更好的下游性能。在超过50%的投影层时,DimPO在SQuAD、常见词提取、频繁词提取和变量跟踪等任务上越来越优于KL。这些结果表明,在降维下,保留任务相关注意力的排序和集中度可能比重现完整注意力分布更重要。

英文摘要:

A linear projection can reduce the dimension of query and key vectors without updating the pretrained model, but it remains unclear which training objective best preserves model behavior. We ask whether preferences over keys and attention mass on the highest-weighted keys provide a better signal than matching the full attention distribution with KL divergence, especially in long-context settings. We introduce DimPO, which combines listwise preference optimization with a lightweight top-k cross-entropy term for head-fidelity. DimPO is trained offline from the attention patterns of a frozen language model, with one map per layer, shared by the query and the keys and trained separately from the other layers. Across LLaMA3.2-3B, LLaMA3.1-8B, Qwen2.5-7B, and Qwen3-4B Instruct models, pairwise preference objectives outperform the triplet baseline and retain 98% of the original score on short-context tasks when projecting to half the dimension on the last 40% of the layers. With more projected layers or on long-context RULER, they degrade rapidly. In contrast, KL and DimPO, which use every key during training, retain about 95% of the original RULER 4k score on the 8B model when projecting up to 50% of the layers. KL-based projections remain closer to the original attention distribution and attention output, yet DimPO achieves better downstream performance. Beyond 50% of projected layers, DimPO increasingly outperforms KL on tasks including SQuAD, common-word extraction, frequent-word extraction, and variable tracking. These results suggest that under dimensionality reduction, preserving the ordering and concentration of task-relevant attention can matter more than reproducing the full attention distribution.

↑