发表机构
University of Illinois Chicago; Cisco Research(伊利诺伊大学芝加哥分校; 思科研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出Fisher引导的子模数据选择方法,通过分解梯度为锚定和前沿分量,在持续预训练中提升目标领域质量并控制遗忘,实现10倍令牌效率提升。
AI 中文摘要
数据选择已成为大型语言模型训练中的一个核心瓶颈,因为网络规模的语料库噪声大且令牌预算有限。在持续预训练(CPT)中,这变成了一个遗忘控制问题:选择不当的目标领域语料库可能会覆盖预训练检查点中编码的能力。现有的CPT实践要么使用与参数无关的标量(如困惑度)对候选数据进行评分,要么通过花费大量额外的通用领域回放令牌来缓解遗忘。这两种策略都没有直接询问在候选数据上训练将如何移动模型参数。我们证明,基于损失的选择会导致CPT后的Fisher对角线在预训练模型已承诺的高Fisher坐标上精确地向下漂移,而低Fisher坐标则基本不受影响。这种不对称性揭示了灾难性遗忘的参数空间机制。受此观察启发,我们提出了一种Fisher感知的CPT选择器,它将每个候选数据的梯度分解为锚定分量(衡量沿已承诺参数方向的扰动)和前沿分量(衡量在无约束低Fisher子空间中的更新能力)。我们使用对数行列式子模目标聚合这些信号,并通过可扩展的流式数据选择流水线在一次遍历中优化它。在TinyLlama-1.1B和Llama-3.1-8B的医学数据CPT上,我们的选择器提高了目标领域质量,同时限制了在保留的预训练基准上的遗忘。最重要的是,它比遗忘感知回放具有显著更高的令牌效率。1B个选定令牌在适应性和遗忘方面已经优于使用10B令牌训练的回放策略,提供了10倍的令牌效率优势。
英文摘要
Data selection is already a central bottleneck in large-language-model training, where web-scale corpora are noisy and token budgets are finite. In continual pre-training (CPT), it becomes a forgetting-control problem: a poorly chosen target-domain corpus can overwrite capabilities encoded in the pretrained checkpoint. Existing CPT practice either scores candidates with parameter-agnostic scalars such as perplexity, or mitigates forgetting by spending many extra general-domain replay tokens. Neither strategy directly asks how training on a candidate will move the model parameters. We show that loss-based selection causes the post-CPT Fisher diagonal to drift downward on exactly the high-Fisher coordinates the pretrained model had committed to, while leaving low-Fisher coordinates largely untouched. This asymmetry exposes a parameter-space mechanism for catastrophic forgetting. Motivated by this observation, we propose a Fisher-aware CPT selector that decomposes each candidate's gradient into an anchor component, which measures perturbation along committed parameter directions, and a frontier component, which measures update capacity in unconstrained low-Fisher subspaces. We aggregate these signals with a log-determinant submodular objective and optimize it in a single pass using a scalable streaming data selection pipeline. On TinyLlama-1.1B and Llama-3.1-8B CPT over medical data, our selector improves target-domain quality while bounding forgetting on held-out pretraining benchmarks. Most importantly, it is substantially more token-efficient than forgetting-aware replay. 1B selected tokens already outperform the replay strategy trained with 10B tokens on both adaptation and forgetting, giving a 10x token-efficiency advantage.