发表机构
Tsinghua University; Harbin Institute of Technology; ModelBest Inc.; Peking University(清华大学; 哈尔滨工业大学; 模智科技有限公司; 北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ForgeTrain提出Forge Engineering范式,为每个场景从头构建专用训练框架,以可信框架为黄金参考,逐步放宽等价性,实验显示MFU提升4.7%-33.2%,是首个由AI端到端锻造的生产级训练框架。
AI 中文摘要
训练大型模型仍然依赖于诸如Megatron-LM之类的通用框架,这种通用性税限制了针对特定场景的优化,并通过累积的抽象增加了运行时开销。AI代码生成降低了构建框架的成本,使得为每个场景打造一个专用框架变得可行。我们提出了Forge Engineering:为每个场景从头构建专用实现,并在正确性和可用性约束下迭代优化以达到峰值性能。专用实现不继承抽象边界,因此它们可以跨栈集成优化并达到更高的性能上限。我们将这一范式实例化为训练框架ForgeTrain,它持有一个可信框架作为黄金参考,并单调地放宽等价性从逐位等价到超越。跨多个模型-硬件配置的实验表明,ForgeTrain始终生成正确的训练引擎,并将已建立训练框架的MFU提高了4.7%至33.2%。据我们所知,这是第一个由AI端到端锻造的生产级训练框架,能够匹配或超越其人类参考。
英文摘要
Training large models still relies on general-purpose frameworks such as Megatron-LM, whose generality tax constrains scenario-specific optimization and adds runtime overhead through accumulated abstraction. AI code generation reduces the cost of building a framework, and makes it affordable to forge one per scenario. We propose Forge Engineering: building a dedicated implementation from scratch for each scenario and iteratively optimizing it toward peak performance under correctness and usability constraints. Dedicated implementations inherit no abstraction boundaries, so they can integrate optimizations across the stack and reach a higher performance ceiling. We instantiate this paradigm for training frameworks as ForgeTrain, which holds a trusted framework as a golden reference and relaxes equivalence monotonically from Bit-for-Bit to Surpass. Experiments across multiple model--hardware configurations show that ForgeTrain consistently produces correct training engines and improves MFU over established training frameworks by 4.7--33.2%. To our knowledge this is the first production-grade training framework forged end-to-end by AI to match or surpass its human reference.
Comments21 pages, 9 figures