arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AlignDiff:利用模型内在信息进行更好的偏好数据选择

AlignDiff: Exploiting Model-Intrinsic Information for Better Preference Data Selection

Peng Lai, He Zhu, Zhiwen Ruan, Dongdong Zhang, Yun Chen, Peng Li, Furu Wei, Yang Liu, Guanhua Chen

arXiv 2609.05899首次发表:更新:

发表机构

Southern University of Science and Technology; Peking University; MSRA; Shanghai University of Finance and Economics; Tsinghua University(南方科技大学; 北京大学; 微软亚洲研究院; 上海财经大学; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

AlignDiff利用模型内在信号过滤偏好数据,优先选择清晰且具挑战性的样本,在LLaMA和Qwen上超越七个基线,提升对齐性能。

AI 中文摘要

将大型语言模型与人类偏好对齐仍然是一个挑战,这主要是由于偏好数据质量在有效对齐中起着关键作用。现有数据集经常受到固有噪声和分布偏移的困扰,这从根本上限制了模型性能。为了弥合这一差距,我们提出了AlignDiff,一个由模型内在信号驱动的偏好数据过滤框架。AlignDiff首先利用正信号和逆信号识别具有清晰偏好的样本,然后基于平均负对数似然差距优先选择更具挑战性的样本,鼓励模型从中学习更丰富的信息。AlignDiff在两个广泛使用的模型家族(LLaMA和Qwen)以及对齐社区广泛采用的三个基准(AlpacaEval 2.0、Arena-Hard和MT-Bench)上进行了评估。在所有设置中,它始终优于七个强基线。我们进行了全面的消融研究以验证AlignDiff的有效性,并进一步表明基于难度的课程学习可以提高模型性能。

英文摘要

Aligning large language models with human preferences remains a challenge, primarily due to the critical role of preference data quality in effective alignment. Existing datasets are frequently plagued by inherent noise and distribution shifts, which inherently limit model performance. To bridge this gap, we propose AlignDiff, a preference data filtering framework driven by intrinsic model signals. AlignDiff first identifies samples with clear preferences using both positive and inverse signals, then prioritizes the more challenging samples based on the average negative log-likelihood gap, encouraging the model to learn richer information from them. AlignDiff is evaluated on two widely used model families (LLaMA and Qwen) and three benchmarks widely adopted in the alignment community (AlpacaEval 2.0, Arena-Hard, and MT-Bench). Across all settings, it consistently outperforms seven strong baselines. We conduct comprehensive ablation studies to validate the effectiveness of AlignDiff, and further show that difficulty-based curriculum learning improves model performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑