少刷新,优选择:复用陈旧梯度特征实现高效基于影响力的数据选择
Refreshing Less, Selecting Better: Reusing Stale Gradient Features for Efficient Influence-Based Data Selection
AI总结:
针对基于梯度的数据选择中重复计算梯度特征成本高的问题,提出CDIS方法,通过复用陈旧缓存特征并仅刷新高排名部分,在保持选择质量的同时将选择成本降低3.4倍。
AI中文摘要:
基于梯度的数据选择方法(如LESS)通过评估每个候选样本的梯度与目标验证梯度之间的对齐程度来为其打分,而在每个新的检查点重新计算每个样本的梯度特征主导了其计算成本。在三个选择种子、两个模型家族、两个候选池和两个目标任务中,在预热后检查点缓存的特征与新鲜验证梯度配对使用,在40个优化步骤后仍能保持排序,Spearman相关系数在0.952至0.991之间,而它们所诱导的前10%子集相比完全重新计算所选择的子集,遗漏了10%至22%的样本。因此,我们提出了缓存多样化影响力选择(CDIS),该方法在陈旧分数下为排名前p比例(p)的候选样本重新计算梯度特征,在重新计算的样本上拟合仿射校准,并在源和长度配额下选择最终子集。当校准后的陈旧分数具有有界误差时,刷新比例达到或超过选择比例即可恢复精确的前k子集,且预算曲线遵循此规则:在p=0.3时,恢复的前k子集在所有设置下与完全重新计算一致,按层分配相同预算可恢复分层子集(匹配度0.92至1.00),梯度阶段墙钟时间降低3.5至3.6倍。在四个检查点上迭代缓存,使前k重叠保持在0.98或更高,每个样本仅需1.9个梯度特征,而完全重新计算需要4个,且刷新样本上陈旧到重新计算的一致性为不安全复用提供了免费检查。在下游任务中,无约束的前k选择会坍缩到单一数据源,在GSM8K上比随机选择低13分。CDIS比随机选择高12分,配对置信区间排除零,落后完全重新计算4.3分(一个训练运行标准差),而选择成本降低3.4倍。
英文摘要:
Gradient-based data selection methods such as LESS score each candidate by the alignment between its gradient and a target validation gradient, and recomputing per-example gradient features at every new checkpoint dominates their cost. Across three selection seeds, two model families, two candidate pools, and two target tasks, features cached at a post-warmup checkpoint and paired with fresh validation gradients preserve the ranking 40 optimizer steps later with Spearman correlation from 0.952 to 0.991, while the top-10% subset they induce misses 10 to 22% of the examples that full recomputation selects. We therefore propose Cached Diverse Influence Selection (CDIS), which recomputes gradient features for the top-ranked fraction $p$ of candidates under the stale scores, fits an affine calibration on the recomputed examples, and selects the final subset under source and length quotas. A refresh fraction at or above the selection fraction recovers the exact top-$k$ subset whenever the calibrated stale scores have bounded error, and the budget curves follow this rule: at $p=0.3$ the recovered top-$k$ subset coincides with full recomputation in every setting, allocating the same budget per stratum recovers the stratified subset at 0.92 to 1.00, and gradient-stage wall-clock drops 3.5 to 3.6 times. Iterating the cache over four checkpoints keeps top-$k$ overlap at 0.98 or higher at 1.9 gradient features per example against 4 for full recomputation, and stale-to-recomputed agreement on the refreshed examples provides a free check for unsafe reuse. Downstream, unconstrained top-$k$ selection collapses to a single data source and scores 13 points below random selection on GSM8K. CDIS scores 12 points above random selection with paired confidence intervals that exclude zero and trails full recomputation by 4.3 points, one training-run standard deviation, at 3.4 times lower selection cost.