发表机构
NingboTech University; Dalian Ocean University(宁波诺丁汉大学; 大连海洋大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对分布式训练中网络故障导致的检查点开销问题,提出AccelPact运行时,通过动态依赖重绑定实现零I/O内存故障恢复,在Mistral-7B训练中提升吞吐量达1.698倍。
AI 中文摘要
大规模分布式模型训练经常因瞬时网络故障而中断,传统上迫使集群管理器终止所有进程并回滚到最新检查点。虽然周期性检查点提供了持久性,但频繁快照引入了严重的存储反压:我们在一个1.216B参数的解码器上的测量表明,每次更新的异步检查点会导致高达+656.7%的延迟开销,消耗32.5 GiB的主机内存,并产生3.39 TB/小时的存储流量。为了消除这一开销,我们提出了AccelPact,一个并行运行时,为分片分布式训练实现零I/O的内存中故障恢复。当在已提交的优化器步骤中通信失败时,设备内存保持静止且未损坏,但继续执行失败,因为像PyTorch FSDP这样的框架在模块包装器和参数层次结构中缓存内部通信句柄。AccelPact通过由带外Gloo共识协议协调的非侵入式引用重绑定机制解决了这种依赖失效问题。在16个NVIDIA RTX 5880 GPU上训练全参数Mistral-7B时,AccelPact消除了检查点重放,在检查点年龄为5时,相比冷重启实现了1.197x的全程吞吐量提升,相比NVRx检查点恢复实现了1.194x的提升,在年龄为18时提升至1.698x。在连续十次故障注入中,所有16个秩保持比特相同的参数状态,零数值漂移。在4到16个GPU集群拓扑中,引用重绑定以恒定时间(0.493-0.518毫秒)执行。直接操作原生通信器实例,AccelPact需要零应用程序代码修改,并避免在此http URL下的编译器图断裂,为弹性深度学习提供了高效基础。
英文摘要
Distributed model training at scale is frequently interrupted by transient network failures, conventionally forcing cluster managers to abort all processes and roll back to the latest checkpoint. While periodic checkpointing provides durability, frequent snapshotting introduces severe storage backpressure: our measurements on a 1.216B-parameter decoder reveal that per-update asynchronous checkpointing incurs up to a +656.7% latency overhead, consumes 32.5 GiB of host memory, and generates 3.39 TB/hour of storage traffic. To eliminate this overhead, we present AccelPact, a parallel runtime enabling zero-I/O in-memory fault recovery for sharded distributed training. When communication fails at a committed optimizer step, device memory remains quiescent and uncorrupted, yet continuation fails because frameworks like PyTorch FSDP cache internal communication handles across module wrappers and parameter hierarchies. AccelPact resolves this dependency invalidation via a non-invasive reference-rebinding mechanism coordinated by an out-of-band Gloo consensus protocol. On 16 NVIDIA RTX 5880 GPUs training full-parameter Mistral-7B, AccelPact eliminates checkpoint replay, yielding a 1.197x whole-run goodput improvement over cold restart and 1.194x over NVRx checkpoint restoration at checkpoint age 5, rising to 1.698x at age 18. Across ten successive fault injections, all 16 ranks maintain bit-identical parameter state with zero numerical drift. Across 4-to-16 GPU cluster topologies, reference rebinding executes in constant time (0.493-0.518 ms). Operating directly on native communicator instances, AccelPact requires zero application-code modifications and avoids compiler graph breaks under torch.compile, providing an efficient foundation for resilient deep learning.
Comments8 pages, 4 figures, 10 tables