发表机构
DatologyAI; ETH Zürich; ITU Copenhagen; Stanford University; Arcee AI; CMU(DatologyAI; 苏黎世联邦理工学院; 哥本哈根IT大学; 斯坦福大学; Arcee AI; 卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Zephon提出一种支持在线有状态管道的基础模型数据加载器,通过拓扑无关通道和序列化排序实现弹性确定性,并高效恢复检查点,在文本和视觉-语言任务上达到有竞争力吞吐量。
AI 中文摘要
确定性数据加载对基础模型开发至关重要:模型研究人员需要确信,他们在昂贵的消融实验中观察到的差异是由他们更改的参数引起的,而非训练数据序列中的非确定性所致。数据加载器必须提供弹性确定性,即尽管跨运行的GPU拓扑发生变化(例如由于GPU稀缺)、频繁的检查点恢复周期以及不同的数据处理执行后端,仍能提供全局训练数据批次的确定性序列。实现这一目标很困难,因为现代基础模型数据管道在线进行分词、打包和混合样本,引入了破坏样本索引的有状态n对m转换。现有数据加载器大多假设可索引的1对1管道,而常见的离线物化解决方法成本高昂,且对于某些模态(如视频)不可行。我们提出了Zephon,一种支持在线、有状态管道的基础模型数据加载器,同时提供弹性确定性和从检查点的高效恢复。它将全局流划分为与拓扑无关的通道,在并行化可互换后端上的无状态工作的同时序列化排序决策,并且仅检查点有界在途状态,使恢复成本不随训练进度增长。我们在文本和视觉-语言工作负载上评估了Zephon,结果表明它在提供现有加载器无法为在线、有状态管道提供的组合保证的同时,实现了有竞争力的吞吐量。
英文摘要
Deterministic data loading is important for foundation model development: model researchers need confidence that differences they observe across costly ablations are caused by the parameter they changed rather than non-determinism in the training data sequence. The data loader must provide elastic determinism, i.e., a deterministic sequence of global training data batches despite changes to the GPU topology across runs (e.g., due to GPU scarcity), frequent checkpoint-resume cycles, and different data processing execution backends. Achieving this is difficult because modern foundation model data pipelines tokenize, pack, and mix samples online, introducing stateful n-to-m transformations that break sample indexing. Existing data loaders largely assume indexable 1-to-1 pipelines, and the common workaround of offline materialization is expensive and, for some modalities such as video, infeasible. We present Zephon, a data loader for foundation models that supports online, stateful pipelines while providing elastic determinism and efficient resumption from checkpoints. It partitions the global stream into topology-independent lanes, serializes ordering decisions while parallelizing stateless work on interchangeable backends, and checkpoints only bounded in-flight state so recovery cost does not grow with training progress. We evaluate Zephon on text and vision-language workloads and show that it achieves competitive throughput while providing a combination of guarantees that no existing loader offers for online, stateful pipelines.
Commentspreprint; currently under revision at VLDB'27