发表机构
Zhejiang University; Huawei(浙江大学; 华为)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对LLM训练的SDC问题,系统表征其在Transformer训练中的脆弱性,提出TrainSDC框架,经实验验证可有效缓解SDC且运行时开销较低。
AI 中文摘要
大型语言模型(LLM)训练日益易受静默数据损坏(SDC)影响,但现有保护方法大多对Transformer计算采取统一处理,因为其脆弱性仍未被充分理解。我们首次系统表征了Transformer训练前向和反向传播中主要计算接口的SDC脆弱性,分析揭示两种不同的误差传播机制:前向传播脆弱性高度依赖位置,Q/K路径上的故障会产生持续的训练偏差;而反向传播脆弱性主要由梯度指数分布决定,而非计算位置。基于这些发现,我们提出TrainSDC,这是一个由Q/K路径重计算、残差增益监测和指数感知梯度缩放构成的表征引导保护框架。在Llama 3.2-1B和Qwen3-0.6B上的实验表明,TrainSDC在稀疏和密集故障注入下均能保持接近无故障执行的训练行为,仅引入1.65%-6.76%的运行时开销。
英文摘要
LLM training is increasingly vulnerable to silent data corruption (SDC), yet existing protection methods largely treat Transformer computations uniformly because their vulnerability remains poorly understood. We present the first systematic characterization of SDC vulnerability across major computation interfaces in both the forward and backward passes of Transformer training. Our analysis reveals two distinct error propagation mechanisms: forward-pass vulnerability is highly location dependent, with faults on the Q/K path producing persistent training deviations, whereas backward-pass vulnerability is largely governed by gradient exponent distributions rather than computation locations. Motivated by these observations, we propose TrainSDC, a characterization-guided protection framework consisting of Q/K-path recomputation, residual-gain monitoring, and exponent-aware gradient scaling. Experiments on Llama 3.2-1B and Qwen3-0.6B show that TrainSDC maintains training behavior close to fault-free execution under both sparse and dense fault injection while introducing only 1.65%-6.76% runtime overhead.
Comments12 pages, 5 figures, and 6 tables. Includes an appendix with additional experiments