arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

缓解大语言模型对齐税的偏好数据选择

Preference Data Selection for Mitigating the Alignment Tax in Large Language Models

Minsu Kim, Jianxun Lian, Xing Xie, Steven Euijong Whang

arXiv 2608.24192首次发表:更新:

发表机构

Microsoft Research Asia(微软亚洲研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出BALIGN策略,通过筛选偏好数据的三个关键特征,缓解大语言模型对齐税,在保留通用能力的同时提升对齐效果,达到最优帕累托前沿。

AI 中文摘要

将大语言模型与人类偏好对齐对实际部署至关重要,但常产生对齐税,导致预训练通用能力的灾难性遗忘。以往研究主要将此问题视为优化或架构挑战,而驱动这种退化的偏好数据固有特征却未被充分探索。本文提出BALIGN,一种平衡数据选择策略,在优化对齐效果的同时明确缓解灾难性遗忘。通过对偏好优化梯度的理论与实证分析,我们确定了决定参数漂移的三个关键数据中心特征:参考模型的对数概率边际、选中与拒绝响应的词元长度差,以及与通用能力语料库的TF-IDF相似度。通过将这些正交特征聚合成统一的复合风险评分,BALIGN系统地过滤掉会扰乱模型固有参数或对齐效用极低的高风险偏好样本。在标准人类偏好数据集上的大量实验表明,BALIGN能在不损害对齐收益的前提下有力保留基础能力,以极低的计算开销始终达到最优帕累托前沿。

英文摘要

Aligning large language models to human preferences is crucial for real-world deployment but frequently incurs an alignment tax, leading to the catastrophic forgetting of pre-trained general capabilities. While previous works primarily frame this problem as an optimization or architectural challenge, the inherent characteristics of preference data that drive this degradation remain largely underexplored. In this paper, we propose BALIGN, a balanced data selection strategy that explicitly mitigates catastrophic forgetting while optimizing alignment efficacy. Through theoretical and empirical analyses of the preference optimization gradient, we identify three key data-centric features that dictate parameter drift: the reference model's log-probability margin, the token length difference between chosen and rejected responses, and the TF-IDF similarity to general capability corpora. By aggregating these orthogonal features into a unified composite risk score, BALIGN systematically filters out high-risk preference samples that disrupt intrinsic model parameters or provide minimal alignment utility. Extensive experiments on standard human preference datasets demonstrate that BALIGN strongly preserves foundational capabilities without compromising alignment gains, consistently achieving the optimal Pareto frontier with minimal computational overhead.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑