发表机构
NVIDIA; NVIDIA Research(英伟达; 英伟达研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM推理服务中单点故障导致恢复慢的问题,提出基于快照和GPU内存服务解耦设备内存生命周期的方法,实现7秒内恢复,比热重启快13-29倍,可回收79%的GPU小时损失。
AI 中文摘要
大语言模型(LLM)推理副本运行在紧密耦合的 GPU 上,并持续数周不间断地服务流量。因此,硬件和软件故障不可避免,一个工作节点的故障可能中断整个副本。恢复需要重新初始化引擎,即使权重和编译产物已缓存,也需要数分钟时间。生产部署会过度配置服务容量以掩盖这一窗口期。我们认为,主要的成本是就绪服务容量的损失,而非请求进度,因此恢复应保留已初始化的引擎状态,而非重建它。我们基于这一原则提出了 Dynamo 的快速恢复方案。快照捕获一次已初始化的引擎,并在恢复时恢复它,而不是重新初始化。对 Dynamo 集群 18 周故障的分析表明,大多数故障是设备保留型:引擎进程失败,而 GPU 及其驻留分配保持完好。我们的关键见解是,独立的引擎进程可以重用相同的 GPU 驻留状态,同时保持可变执行状态私有。GPU 内存服务(GMS)将设备内存所有权与引擎进程解耦,使引擎能够共享和重新附加幸存的分配,而无需复制它们。GMS 保留模型权重,并在替换引擎和影子引擎之间以只读方式共享它们,避免权重重新加载。在同一 GPU 上运行第二个已初始化的运行时,可将恢复简化为提升。在 vLLM 和 SGLang 上的四个模型中,这些机制在 7 秒内恢复失败的副本,比热重启快 13-29 倍,每个 GPU 使用固定的 4-8 GiB 设备内存,与模型大小无关。重放生产轨迹,我们估计它们将回收因恢复而损失的 GPU 小时数的 79%。
英文摘要
Large language model (LLM) inference replicas run across tightly coupled GPUs and serve traffic continuously for weeks. Hardware and software failures are therefore inevitable, and one worker failure can disrupt an entire replica. Recovery requires reinitializing the engine, taking minutes even when weights and compilation artifacts are cached. Production deployments overprovision serving capacity to mask this window. We argue that the dominant cost is loss of ready serving capacity, not request progress, so recovery should preserve initialized engine state rather than reconstruct it. We present fast recovery for Dynamo based on this principle. Snapshots capture an initialized engine once and restore it instead of reinitializing it. Analysis of 18 weeks of failures from the Dynamo cluster shows that most failures are device-preserving: the engine process fails while the GPU and its resident allocations remain intact. Our key insight is that independent engine processes can reuse the same GPU-resident state while keeping mutable execution state private. The GPU Memory Service (GMS) decouples device-memory ownership from engine processes, enabling engines to share and reattach surviving allocations without copying them. GMS preserves model weights and shares them read-only between replacement and Shadow Engines, avoiding weight reloads. A second initialized runtime on the same GPUs reduces recovery to promotion. Across four models on vLLM and SGLang, these mechanisms recover a failed replica in under 7 seconds, 13-29 times faster than a warm restart, using a fixed 4-8 GiB of device memory per GPU independent of model size. Replaying the production trace, we estimate they would reclaim 79% of GPU-hours lost to recovery.
Comments16 pages, 12 figures. Open-source implementations: https://github.com/ai-dynamo/dynamo and https://github.com/ai-dynamo/snapshot