Data-DPO:面向大语言模型后训练中目标模型数据选择的直接偏好优化
Data-DPO: Direct Preference Optimization for Target Model Data Selection in LLM Post-Training
浏览论文内容
中文总结 AI 辅助
针对现有数据选择方法忽略数据与目标模型能力分布兼容性的问题,提出Data-DPO方法,结合目标模型偏好、外部质量评分与边际多样性筛选数据,在Vision-Flan和LLaVA-CoT上性能优于基线及全数据训练。
中文摘要 AI 辅助
监督微调(SFT)中的数据选择旨在从大规模候选数据中筛选出一小部分有效样本,以降低训练成本同时保留模型性能。然而,现有方法通常将数据价值视为相对静态的属性,对数据与目标模型能力分布之间的兼容性关注有限。为解决该问题,我们提出Data-DPO,一种面向目标模型的SFT数据选择方法。Data-DPO通过一步探测观测目标模型在不同样本上的局部训练反馈,将样本间的激活差异转化为成对数据偏好,并训练一个轻量奖励模型以学习感知目标模型的数据偏好。在最终选择阶段,Data-DPO进一步结合目标模型偏好、外部质量评分和边际多样性,构建更稳定有效的训练子集。在Vision-Flan和LLaVA-CoT上的实验结果表明,Data-DPO在多种数据预算下均优于现有数据选择基线,且稳定超越全数据训练的性能。
英文摘要
Data selection in supervised fine-tuning aims to select a small set of effective samples from large-scale candidate data, reducing training cost while preserving model performance. However, existing methods usually treat data value as a relatively static property, and pay limited attention to the compatibility between data and the capability distribution of the target model. To address this issue, we propose Data-DPO, a target model-oriented SFT data selection method. Data-DPO observes the local training feedback of the target model on different samples through one-step probing, transforms activation differences among samples into pairwise data preferences, and trains a lightweight reward model to learn target-model-aware data preferences. In the final selection stage, Data-DPO further combines target model preference, external quality scores, and marginal diversity to construct a more stable and effective training subset. Experimental results on Vision-Flan and LLaVA-CoT show that Data-DPO consistently outperforms existing data selection baselines under multiple data budgets and stably surpasses full data training performance.