发表机构
McGill University(麦吉尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究如何从语言建模到推理进行任务和预算感知的数据选择,提出PPL-Factory框架,结合任务感知困惑度得分与数据预算感知标准,实验表明该框架在GSM8K等任务上仅用少量训练集就能超越其他方法,有效提升微调效率。
AI 中文摘要
并非所有训练样本对大语言模型微调的贡献都相同。选择信息丰富的训练样本可降低计算成本并保持下游性能。现有数据选择方法多依赖间接启发式,效果因任务而异。基于困惑度的数据选择可估计样本难度,但现有方法忽略语言建模和推理任务学习目标差异。本文提出PPL-Factory,结合任务感知困惑度得分和数据预算感知选择标准。实验表明,PPL-Factory在GSM8K上仅用1%训练集就超越其他方法,用10%数据时在GSM8K和MATH上均有出色表现,证明该方法对高效微调有效且适用。
英文摘要
Not all training samples contribute equally to large language model fine-tuning. Selecting informative training samples can reduce the computational cost while preserving downstream performance. Many existing data selection methods rely on indirect heuristics, such as data quality, diversity or reasoning trace length. However, the effectiveness of these fixed criteria is task-dependent and difficult to generalize across diverse downstream tasks. Perplexity-based data selection provides a simple and model-aware solution to estimate the sample difficulty, but existing approaches typically score the entire training sequence and ignore the difference in learning objectives of language modeling and reasoning tasks. In this paper, we propose PPL-Factory, a simple and interpretable data selection framework that combines task-aware perplexity-based scores and data budget-aware selection criteria. Experiments on GSM8K demonstrate that PPL-Factory outperforms other state-of-the-art data selection methods using only $1\%$ of the training set. With $10\%$ of the data, PPL-Factory exceeds full-data fine-tuning accuracy by 0.9 on GSM8K and 4.8 on MATH. Overall, our results demonstrate that task-aware and budget-aware perplexity-based selection provides an effective and applicable approach for efficient fine-tuning.
Comments13 pages, 4 figures, 5 tables