发表机构
Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对合成数据整理,指出个体级信号不足,需采用群体级信号以提升模型生成能力,并提出廉价诊断方法以优化计算预算下的数据选择。
AI 中文摘要
合成数据现已成为大语言模型训练中不可或缺的部分,用于增强自主执行和长周期任务执行等高级能力。然而,近期研究表明,大规模使用合成数据进行训练可能会降低模型生成质量,因此决定哪些合成数据值得用于训练变得尤为重要。当前的数据整理实践通常使用个体级信号(即孤立地评估每个数据样本的训练效用),但在预训练和后训练场景中,我们发现这种方法对合成数据而言并不充分,而群体级信号(即考虑数据样本间相互作用的效用估计)对于有效的数据整理是必要的。首先,我们证明个体级信号无法察觉样本如何共同影响训练:不同组成的合成数据集在个体级影响下可能难以区分,但在群体级影响下却差异显著,且基于群体级信号进行整理能带来更好的下游性能,尤其在生成能力方面。其次,我们发现随着训练流程越来越依赖合成数据,群体级信号的重要性日益凸显:在广泛使用的数据整理方法中,只有那些纳入群体级信号的方法能超越基线,且当放大捕捉样本间关系的权重时,性能提升更为明显。最后,我们将这些发现付诸实践——对于受计算预算限制的模型开发者,我们提供一种廉价诊断方法,用于优先确定哪些合成数据组最需要群体级估计,以极小的计算成本即可获得大部分完整群体级评分的收益。
英文摘要
Synthetic data now is essential to LLM training, used to strengthen advanced capabilities such as autonomous and long-horizon task execution. Yet recent work shows that training on it at scale can degrade model generation, making it important to decide what synthetic data is worth training on. While current data curation practices do so with individual-level signals (i.e., estimates of each data sample's training utility in isolation), across pre-training and post-training settings we show that this is insufficient for synthetic data, and that group-level signals (i.e., estimates of utility that account for interactions among data samples) are necessary for effective data curation. First, we show that individual-level signals are blind to how samples jointly affect training: synthetic datasets with different compositions can be indistinguishable under individual-level influence yet differ sharply under group-level influence, and curating by the latter yields better downstream performance, particularly in generative capability. Second, we find that group-level signals matter more as training pipelines become increasingly synthetic: among widely used data curation methods, only those incorporating them improve over baseline, with gains increasing when weights capturing relations among samples are amplified. Finally, we translate these findings into practice -- for model developers under a compute budget, we offer a cheap diagnostic that prioritizes which groups of synthetic data most need group-level estimation, recovering much of the benefit of full group-level scoring at a fraction of the compute cost.