草绘误差而非乘积:半精度GPU矩阵乘法的事后故障恢复
Sketching the Error, Not the Product: Post Hoc Fault Recovery for Half Precision GPU Matrix Multiplication
浏览论文内容
中文总结 AI 辅助
针对半精度GPU矩阵乘法中的静默数据损坏,提出事后验证器FP-Sketch,通过和草图与哈希一阶矩草图定位故障,按实测噪声调整桶数,在Transformer和Llama-2-7B上实现高恢复率并显著降低困惑度损害。
中文摘要 AI 辅助
由缺陷加速器引起的静默数据损坏(SDC)现已中断大规模训练,然而现有的缓解措施作用于整个节点。针对单个GEMM的基于算法的容错(ABFT)必须融合到内核中或对操作数进行编码,并且每个校验和最多只能定位一个错误。我们提出FP-Sketch,一种验证器,它在未修改的张量核心GEMM之后运行,该GEMM的半精度操作数以FP32累加并交付。一个和草图检测每次调用的损坏。由独立重计算确认的哈希一阶矩草图,随后定位多个损坏条目,且构造上无误报,每个故障产生一个坐标和一个幅度,用于集群诊断。在浮点运算中,限制定位的是草图噪声而非桶冲突。我们测量了该噪声,发现其常数取决于BLAS和操作数格式,并且对于方阵乘积,桶数必须按$n^{2.57}$增长。根据测量噪声而非拟合的$n$的幂来调整桶数,将八种Transformer形状的恢复率从0.402提高到1.000,并且在运行时测量噪声使桶数适应内核和模型。使用NVBit的指令级注入表明,活跃累加器中的翻转通常仅为典型条目的2%至9%,这是输出侧注入无法产生的人群。输出侧注入可恢复每个故障,而在NVBit下,为典型幅度故障设计的同一引擎恢复率为0.550,而为测量幅度设计则恢复至1.000。在Llama-2-7B上,保护MLP下投影可消除由2048位翻转引起的困惑度损害的99.4%(BF16)和99.9%(FP16),且干净路径探针耗时0.78至3.06毫秒,而GEMM耗时0.35至12.47毫秒。
英文摘要
Silent data corruption (SDC) from defective accelerators now interrupts large scale training, yet deployed mitigations act on whole nodes. Algorithm based fault tolerance (ABFT) for a single GEMM has to be fused into the kernel or encode the operands, and it localizes at most one error per checksum. We present FP-Sketch, a verifier that runs after an unmodified tensor core GEMM whose half precision operands are accumulated and delivered at FP32. A sum sketch detects corruption on every call. Hashed first moment sketches, confirmed by independent recomputation, then localize several corrupted entries with no false positives by construction, and each fault yields a coordinate and a magnitude for fleet diagnosis. In floating point, sketch noise rather than bucket collisions limits localization. We measure that noise and find that its constant depends on the BLAS and the operand format and that the bucket count must grow as $n^{2.57}$ for a square product. Sizing the bucket count by measured noise rather than by a fitted power of $n$ raises recovery on eight transformer shapes from 0.402 to 1.000, and measuring the noise at run time adapts the bucket count to the kernel and the model. Instruction level injection with NVBit shows that upsets in a live accumulator are often only 2 to 9% of a typical entry, a population that output side injection cannot produce. Output side injection recovers every fault, while under NVBit the same engine sized for faults of typical magnitude recovers 0.550, and sizing for the measured magnitudes restores 1.000. On Llama-2-7B, guarding the MLP down projections removes 99.4% (BF16) and 99.9% (FP16) of the perplexity damage caused by 2048 bit flips, and the clean path probe costs 0.78 to 3.06 ms against GEMMs of 0.35 to 12.47 ms.
发表机构
- National Institute of Technology Warangal(印度理工学院瓦朗加尔分校)
- AthenaHealth(阿西娜健康)
机构由 AI 辅助整理,请以论文原文为准。