arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FastCI:面向LLM训练框架的高效GPU密集型持续集成

FastCI: Efficient GPU-Intensive CI for LLM Training Frameworks

Tianshuo Qiao, Naiqian Zheng, Xiaopeng Liu, Shuguang Wang, Diandian Gu, Xuanzhe Liu, Xin Jin

arXiv 2610.01967首次发表:更新:

发表机构

Peking University; ByteDance Seed(北京大学; 字节跳动Seed)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FastCI通过运行时证据选择受影响测试、剪除等价上下文测试、优先高风险测试并优化工作负载,将LLM训练框架CI延迟降低77.5%,GPU资源减少63.9%,并提高覆盖率保持率。

AI 中文摘要

随着大型语言模型(LLM)在规模和复杂性上持续增长,其训练框架也在快速演进。因此,持续集成(CI)对于维护这些框架的质量和稳定性至关重要。然而,与传统软件不同,LLM训练框架的CI依赖于GPU密集型测试,这些测试通常涉及完整的模型训练或评估。这导致CI本身成为快速开发的新瓶颈。在本文中,我们介绍了FastCI,一个提高LLM训练框架CI效率的框架。FastCI利用运行时证据来选择受影响的测试,并剪除在等价上下文中执行更改代码的测试。然后,FastCI优先处理高风险测试,以更早暴露潜在失败,并沿着每个测试预期验证范围之外的维度优化测试工作负载。在我们LLM训练框架的CI工作负载上评估,与当前部署的CI流水线相比,FastCI将CI延迟降低了77.5%,GPU资源使用减少了63.9%,同时将修改代码覆盖率保持率提高了3.2%。FastCI现已集成到字节跳动LLM训练框架的CI流水线中。

英文摘要

As large language models (LLMs) keep growing in size and complexity, their training frameworks evolve at a rapid pace as well. Therefore, continuous integration (CI) is critical for maintaining the quality and stability of these frameworks. However, unlike traditional software, CI for LLM training frameworks relies on GPU-intensive tests, which usually involve complete model training or evaluation. This leads CI itself to become a new bottleneck for fast-paced development. In this paper, we introduce FastCI, a framework that improves the efficiency of CI for LLM training frameworks. FastCI leverages runtime evidence to select affected tests and prune tests that execute changed code in equivalent contexts. Then FastCI prioritizes high-risk tests to expose potential failures earlier, and optimizes test workloads along dimensions outside the intended validation scope of each test. Evaluated on the CI workload of our LLM training framework, FastCI reduces the CI latency by 77.5% and the GPU resource usage by 63.9%, while improving the modified code coverage retention by 3.2%, compared with the currently deployed CI pipelines. FastCI has now been integrated into the CI pipelines of our LLM training framework at ByteDance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑