arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FailBench:评估分布式训练架构中的容错能力

FailBench: Evaluating Fault Tolerance Across Distributed Training Architectures

Khaled Aljbab, Amine Barrak

arXiv 2610.07688首次发表:更新:

发表机构

Oakland University(奥克兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FailBench提出统一评估框架,覆盖七种分布式训练架构和八种容错机制,在8xV100集群上评估142种组合,发现无普遍最优机制,机制成本依赖架构,并发布决策框架与开源工件。

AI 中文摘要

分布式深度学习依赖于数据、流水线、张量和混合并行,然而容错机制通常仅在其所设计的架构上进行评估。这使得从业者在跨架构选择机制时缺乏指导。FailBench提供了一个统一的评估框架,涵盖七种分布式训练架构、八种崩溃容错机制以及一个无容错基线,并支持单一、并发和级联的故障停止故障。我们在一个8xV100集群上评估了142种(架构、机制、轨迹)组合,每个秩的检查点状态大小从205 MB到2.7 GB不等。研究得出三个发现。首先,没有一种机制是普遍最优的:在A2上,磁盘检查点的稳态开销最低(0.5%),内存复制恢复最快(约17毫秒),而即时检查点避免了周期性稳态检查点成本,但在故障时产生约0.9秒的开销。当进程组重新形成需要数秒时,机制在运行时开销上的差异大于恢复速度。其次,八卦训练在丢失一个工作节点后样本吞吐量提高了17.9%,但在损失进展上与匹配的无故障基线相比没有可检测的改善。第三,机制成本强烈依赖于架构:内存复制的开销范围从3.7%到176%。我们将这些发现转化为一个用于选择容错机制的决策框架,并将FailBench作为开放工件发布。

英文摘要

Distributed deep learning relies on data, pipeline, tensor, and hybrid parallelism, yet fault-tolerance mechanisms are typically evaluated only on the architecture for which they were designed. This leaves practitioners with little guidance when choosing mechanisms across architectures. FailBench provides a unified evaluation harness covering seven distributed training architectures, eight crash-fault-tolerance mechanisms and a no-FT baseline, and single, concurrent, and cascading fail-stop failures. We evaluate 142 (architecture, mechanism, trace) combinations on an 8xV100 cluster, with per-rank checkpoint states ranging from 205 MB to 2.7 GB. Three findings emerge. First, no mechanism is universally best: on A2, disk checkpointing has the lowest steady-state overhead (0.5%), in-memory replication restores fastest (~17 ms), and just-in-time checkpointing avoids periodic steady-state checkpoint cost but incurs ~0.9 s upon failure. When process-group re-formation takes seconds, mechanisms differ more in runtime overhead than restore speed. Second, gossip training increases sample throughput by 17.9% after losing a worker, yet shows no detectable improvement in loss progress over a matched no-fault baseline. Third, mechanism cost depends strongly on architecture: in-memory replication overhead ranges from 3.7% to 176%. We translate these findings into a decision framework for selecting fault-tolerance mechanisms and release FailBench as an open artifact.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑