发表机构
University of Pennsylvania; University of Texas at Austin(宾夕法尼亚大学; 德克萨斯大学奥斯汀分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LESSER利用输出层梯度近似全参数梯度进行训练后数据选择,仅需前向传播,将SFT和RL的FLOP成本分别降低9.7倍和3.0倍,同时保持下游性能。
AI 中文摘要
大型语言模型的训练后数据选择对下游性能有显著影响。基于梯度的数据选择是一种常用方法,它通过训练数据梯度与小验证集梯度的对齐程度对数据进行排序。然而,使用全参数梯度进行排序需要对每个样本执行昂贵的反向传播,使得计算在大规模候选池中变得不可行。这引出一个自然问题:我们能否以一小部分成本近似全梯度特征?我们方便地发现,输出层梯度足以进行有效的数据选择,且仅需更廉价的前向传播。我们将其实现为LESSER,一种即插即用的选择方法包装器,可将SFT的特征提取FLOP成本降低9.7倍,将RL基准的特征提取FLOP成本降低3.0倍,同时在下游任务上保持与全梯度相当的性能。实验发现,即使输出层梯度与全梯度对单个样本的排序不同,它们选择的批次仍具有对齐的梯度。
英文摘要
The choice of post-training data for large language models substantially affects downstream performance. Gradient-based data selection is a popular approach that ranks training data by how well their gradients align with those of a small validation set. However, ranking with full-parameter gradients requires an expensive backward pass on every sample, making computation intractable for large candidate pools. This raises a natural question: can we approximate full-gradient features at a fraction of the cost? Conveniently, we find that output-layer gradients suffice for effective data selection, yet require only the cheaper forward pass. We implement this as LESSER, a drop-in wrapper for selection methods that reduces the feature-extraction FLOP cost by $9.7\times$ for SFT and $3.0\times$ for RL benchmarks, while tracking full-gradient performance on downstream tasks. Empirically, we find that even when output-layer and full gradients rank individual samples differently, they select batches with aligned gradients.