arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FlashBoot:机架级大模型的亚秒级权重加载方案

FlashBoot: Sub-Second Weight Loading for Large Models at Rack Scale

Issac Zhu, Hscos Zhang, Keith Jiang, Jack Li, Hugh Yin, Jason Zhao

arXiv 2608.08482首次发表:更新:

AI 中文总结

FlashBoot是基于SGLang设计的机架级大模型权重加载子系统,通过FabricArena布局等技术解决现有方案的结构性缺陷,将单节点及机架级权重加载速度大幅提升,性能远超现有最先进方案。

AI 中文摘要

旗舰级混合专家(MoE)模型正同时在总参数数量和专家数量两个维度快速增长。在弹性部署场景中,多个节点上的大量GPU必须快速达到服务就绪状态,而这种增长使得权重加载成为延迟预算中不可忽视的部分。即使在NVIDIA的GB300 NVL72平台上,当前最先进的加载器仍未充分利用大部分带宽,损失源于三个结构性问题:(C1)权重内存被碎片化为数万个单张量对象,导致传输速率远低于链路带宽;(C2)跨节点复制受NCCL通信器设置的限制,在传输单个权重字节前需要耗费10-110秒;(C3)现有的跨节点GPU到GPU克隆路径是串行的,在并发多节点启动时扩展性很差。我们提出FlashBoot,这是一个基于SGLang设计的、面向硬件与框架工作流协同的权重加载子系统,其核心是FabricArena,一种连续、可导出且可跨节点寻址的张量内存布局。在此基础上,FlashLoad以单次批量零拷贝传输的方式从CPU加载权重,FlashClone则通过消除NCCL设置的远程映射机制从远程GPU复制驻留模型。在NVL72平台上使用DeepSeek-V4-Pro和DeepSeek-V4-Flash进行的实验显示,FlashClone映射远程权重内存仅需约10毫秒(而NCCL需10-110秒),每个克隆操作的持续速率不低于700 GB/s。与现有最先进方案相比,FlashBoot将单节点权重加载速度提升了最高50倍(从20.1秒降至0.4秒),并发机架级权重加载速度提升了超过270倍(从87秒降至0.32秒)。我们的代码将公开提供。

英文摘要

Flagship Mixture-of-Experts (MoE) models are growing fast along two axes at once: total parameter count and the number of experts. In elastic deployment scenarios, many GPUs across many nodes must become serving-ready quickly, and this growth makes weight loading a noticeable part of the latency budget. Even on NVIDIA's GB300 NVL72, today's state-of-the-art loaders leave most of that bandwidth unused. The losses are structural: (C1) weight memory is fragmented into tens of thousands of per-tensor objects, so transfers run far below link bandwidth; (C2) cross-node replication is gated by NCCL communicator setup, which costs 10-110 s before a single weight byte moves; and (C3) the existing cross-node GPU->GPU clone path is serial and scales poorly to concurrent multi-node bring-up. We present FlashBoot, a hardware-friendly, framework-workflow co-designed weight-loading subsystem built on SGLang. At its core is FabricArena, a contiguous, exportable and inter-node addressable tensor memory layout. On top of it, FlashLoad loads from CPU as a single bulk, zero-copy transfer, and FlashClone replicates a resident model from a remote GPU via a remote-mapping mechanism that removes NCCL setup. In experiments on NVL72 with DeepSeek-V4-Pro and DeepSeek-V4-Flash, FlashClone maps remote weight memory in ~10 ms (versus 10-110 s for NCCL) and sustains >=700 GB/s per clone. Against the state of the art, FlashBoot accelerates single-node weight loading by up to 50x (from 20.1 s to 0.4 s) and concurrent rack-level weight loading by >270x (from 87 s to 0.32 s). Our code will be made publicly available.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑