arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2504.02211cs.DCcs.AIcs.LG

FT-Transformer:具有端到端容错注意力的弹性与可靠Transformer

FT-Transformer: Resilient and Reliable Transformer with End-to-End Fault Tolerant Attention

  • University of California, Riverside(加州大学河滨分校)

机构由 AI 辅助整理,请以论文原文为准。

Huangliang Dai, Shixun Wu, Jiajun Huang, Zizhe Jian, Yue Zhu, Haiyang Hu, Zizhong Chen

更新

AI总结:

本文提出FT-Transformer,通过端到端容错注意力实现模块级保护,解决传统操作级容错的开销问题,在非线性计算中提供全面覆盖,对线性模块设计避免线程通信的ABFT算法,实现7.56倍加速和13.9%平均开销。

AI中文摘要:

Transformer模型依赖高性能计算(HPC)资源进行推理,而软错误在大规模系统中不可避免,这使得模型的可靠性尤为关键。现有的Transformer容错框架在操作层面设计,缺乏架构优化,导致显著的计算和内存开销,进而降低保护效率并限制向更大模型的扩展。本文通过将注意力模块内的操作视为单一内核并应用端到端容错,实现了Transformer的模块级保护。该方法在多步计算中提供统一保护,同时实现对非线性计算中潜在错误的全面覆盖。对于线性模块,我们设计了一种避免线程间通信的跨步算法容错(ABFT)。实验结果表明,我们的端到端容错相比传统方法实现高达7.56倍加速,平均容错开销为13.9%。

英文摘要:

Transformer models rely on High-Performance Computing (HPC) resources for inference, where soft errors are inevitable in large-scale systems, making the reliability of the model particularly critical. Existing fault tolerance frameworks for Transformers are designed at the operation level without architectural optimization, leading to significant computational and memory overhead, which in turn reduces protection efficiency and limits scalability to larger models. In this paper, we implement module-level protection for Transformers by treating the operations within the attention module as a single kernel and applying end-to-end fault tolerance. This method provides unified protection across multi-step computations, while achieving comprehensive coverage of potential errors in the nonlinear computations. For linear modules, we design a strided algorithm-based fault tolerance (ABFT) that avoids inter-thread communication. Experimental results show that our end-to-end fault tolerance achieves up to 7.56x speedup over traditional methods with an average fault tolerance overhead of 13.9%.

补充信息

↑