发表机构
University of Illinois at Urbana-Champaign; Google(伊利诺伊大学厄巴纳-尚佩恩分校; 谷歌)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
dattri-LLM是一个统一高效的训练数据归因库,通过紧凑梯度表示和成本模型路由,实现高效、兼容和可扩展的梯度归因,平均吞吐量提升3.2倍,支持110B参数模型。
AI 中文摘要
训练数据归因(TDA)估计单个训练示例对模型输出的贡献。大多数可扩展的TDA方法依赖于每个示例的梯度,在LLM规模下,这些梯度的计算和使用在效率、兼容性和可扩展性方面构成挑战。我们引入了dattri-LLM,一个使基于梯度的归因在大规模下更加实用的TDA库。在效率方面,dattri-LLM使用紧凑的梯度表示,并基于成本模型动态路由梯度操作。在兼容性方面,其捕获机制从调用backward()的现有训练循环中收集每个示例的梯度,无需更改循环或其配置。这包括使用DDP和FSDP的分布式训练,以及使用HuggingFace Transformers、TRL和OLMo构建的流水线。在可扩展性方面,dattri-LLM提供了可重用的梯度操作和训练时回调,用于实现归因方法和应用。这些接口支持多种归因方法,包括梯度相似性、基于曲率的影响和基于轨迹的方法,以及在训练期间对梯度进行操作的应用,如在线数据选择。在相同的硬件和工作负载下,dattri-LLM平均实现了最快竞争库3.2倍的吞吐量,将多种归因方法扩展到110B参数模型,跨越四个H200 GPU,并在不同模型系列和规模的一系列模型中提供了优越的归因保真度-成本权衡。
英文摘要
Training data attribution (TDA) estimates the contribution of individual training examples to model outputs. Most scalable TDA methods rely on per-example gradients, whose computation and use at LLM scale pose challenges in efficiency, compatibility, and extensibility. We introduce dattri-LLM, a TDA library that makes gradient-based attribution more practical at scale. For efficiency, dattri-LLM uses compact gradient representations and dynamically routes gradient operations based on a cost model. For compatibility, its capture mechanism collects per-example gradients from existing training loops that call backward(), without requiring changes to the loop or its configuration. This includes distributed training with DDP and FSDP and pipelines built with HuggingFace Transformers, TRL, and OLMo. For extensibility, dattri-LLM exposes reusable gradient operations and training-time callbacks for implementing attribution methods and applications. These interfaces support a variety of attribution methods, including gradient similarity, curvature-based influence, and trajectory-based methods, as well as applications that act on gradients during training, such as online data selection. On the same hardware and workload, dattri-LLM achieves 3.2x the throughput of the fastest competing library on average, scales multiple attribution methods to 110B-parameter models across four H200 GPUs, and offers superior attribution fidelity-cost trade-offs across a range of models with different model families and scales. The source code of dattri-LLM is available at https://github.com/TRAIS-Lab/dattri-llm.