发表机构
Sichuan University; University of Electronic Science and Technology of China(四川大学; 电子科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对在线批次选择中样本内梯度匹配的偏差问题,提出留一法梯度匹配(LOOM),通过移除Gram矩阵对角线实现无偏估计,在多个微调任务上显著优于全批次训练及其他选择器。
AI 中文摘要
在线批次选择在候选批次中最有用的部分上对语言模型进行微调。匹配候选批次梯度的选择器很有吸引力,因为它们不需要保留数据,但很少能超越在整个批次上训练的效果。我们展示了原因。样本内梯度匹配将每个示例作为其自身目标的一部分,因此其目标将每个示例的梯度噪声归因于自身。这是使训练误差乐观的协方差惩罚,现在位于梯度Gram矩阵的对角线上:它将选择导向噪声最大的示例,并使完整批次成为目标能达到的最佳解。这一修复无需任何成本。对于每个示例,其他候选构成了数据分布的独立样本,因此移除对角线将匹配目标转化为对更新误差相对于总体梯度的无偏估计。这一留一法目标的最小化器根据每个示例的梯度信噪比(SNR)对示例进行加权,并且每当每个示例的SNR异质性足够强时,半个批次产生的更新误差低于整个批次;我们给出了确切的条件。我们基于这一原理构建了\method{}。它在普通反向传播过程中以Adam预条件子的度量计算Gram矩阵,以$(1-e^{-\gamma})$的保证贪婪地选择加权子集,并且不使用保留数据。在四个微调任务和从1.5B到8B参数的七个骨干网络上,LOOM在Llama-3.1-8B和Qwen2.5-7B上比全批次训练提高了2.3和2.4个百分点,超过每个样本内梯度匹配器2.4个百分点,超过验证引导的GREATS和OPUS 1.6--2.0个百分点,并以低于其基础率五分之一的比例选择注入的标签噪声。
英文摘要
Online batch selection fine-tunes a language model on the most useful part of each candidate batch. Selectors that match the gradient of the candidate batch are attractive because they need no held-out data, yet they rarely beat training on the whole batch. We show why. In-sample gradient matching uses every example as part of its own target, so its objective credits each example with its own gradient noise. This is the covariance penalty that makes training error optimistic, now sitting on the diagonal of the gradient Gram matrix: it steers selection toward the noisiest examples and makes the full batch the best solution the objective can reach. The fix costs nothing. For each example, the other candidates form an independent sample of the data distribution, so removing the diagonal turns the matching objective into an unbiased estimate of the update's error with respect to the population gradient. The minimizer of this leave-one-out objective weights examples by their gradient signal-to-noise ratio (SNR), and whenever per-example SNR is heterogeneous enough, half of a batch yields a lower-error update than the whole batch; we give the exact condition. We build \method{} on this principle. It computes the Gram matrix in the metric of the Adam preconditioner during the ordinary backward pass, selects a weighted subset greedily with a $(1-e^{-γ})$ guarantee, and uses no held-out data. Across four fine-tuning tasks and seven backbones from 1.5B to 8B parameters, LOOM improves on full-batch training by 2.3 and 2.4 points on Llama-3.1-8B and Qwen2.5-7B, exceeds every in-sample gradient matcher by 2.4 points and the validation-guided GREATS and OPUS by 1.6--2.0, and selects injected label noise at under a fifth of its base rate.