arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19743cs.DC

量化整数GPU算术中静默数据损坏的综合征解码

Syndrome Decoding for Silent Data Corruption in Quantized Integer GPU Arithmetic

  • National Institute of Technology Warangal(瓦朗加尔国立理工学院)
  • Athenahealth(阿西娜健康)

机构由 AI 辅助整理,请以论文原文为准。

Pranav Napolean, Vikas Srivastava, Napolean Periathambi

AI总结:

针对量化整数GPU算术中的静默数据损坏,提出SProbe验证内核,利用随机化Freivalds门和Reed-Solomon综合征解码,实现高概率检测与精确修复,在H100上全面检测故障,代价可控。

AI中文摘要:

量化神经网络推理在GPU张量核心上执行整数矩阵乘法,这些核心内部的INT32累加器既没有奇偶校验也没有ECC。该数据路径中的瞬态故障会返回一个有效但错误的整数,且不会引发中断。基于校验和的算法级容错(ABFT)能够检测此类静默数据损坏(SDC),但其判定是二元的。它无法识别损坏的元素或其幅度,并且未加权的行和列校验和在构造上对在两个轴上相互抵消的错误是盲目的。我们提出了SProbe,一个尾随验证内核,它读取未经修改的厂商GEMM的输出。一个在61位素数域中具有三个独立评估点的随机化Freivalds门以至多$2^{-141}$的概率漏检非零错误。当门触发时,每行在三个素数上的幂和综合征通过Berlekamp-Massey、Chien搜索和Forney的Reed-Solomon链进行解码,恢复每行最多四个碰撞错误的列和精确幅度。SProbe随后就地修复累加器或重新计算GEMM。在NVIDIA H100上,SProbe在七种故障类别和四种矩阵大小下检测到每一个注入的故障,包括TR-ABFT从未检测到的构造模式和加权网格码检测到但无法纠正的模式。该门在N=16384时占cuBLASLt GEMM时间的49%,在N=65536时占11%。在INT8医学LLM中,保护以30%的吞吐量成本消除了所有观察到的静默损坏。我们的测量还表明,在我们测试的每种配置中,重新计算比就地恢复更快,诊断而非修复主导了恢复成本,并且我们报告了在验证验证器本身时发现的缺陷。

英文摘要:

Quantized neural network inference runs integer matrix multiplications on GPU tensor cores, and the INT32 accumulators inside those cores have neither parity nor ECC. A transient fault in this datapath returns a valid but wrong integer and raises no interrupt. Checksum based Algorithm Based Fault Tolerance (ABFT) can detect such silent data corruptions (SDCs), but its verdict is binary. It cannot identify the corrupted element or its magnitude, and unweighted row and column checksums are blind by construction to errors that cancel on both axes. We present SProbe, a trailing verification kernel that reads the output of an unmodified vendor GEMM. A randomized Freivalds gate with three independent evaluation points in a 61 bit prime field misses a nonzero error with probability at most $2^{-141}$. When the gate fires, per row power sum syndromes over three primes are decoded with the Reed Solomon chain of Berlekamp Massey, Chien search, and Forney, recovering the column and exact magnitude of up to four colliding errors per row. SProbe then repairs the accumulator in place or recomputes the GEMM. On an NVIDIA H100, SProbe detects every injected fault across seven fault classes and four matrix sizes, including constructed patterns that TR-ABFT never detects and patterns that a weighted grid code detects but cannot correct. The gate costs 49% of the cuBLASLt GEMM time at N=16384 and 11% at N=65536. In an INT8 medical LLM, protection eliminates all observed silent corruptions at a 30% throughput cost. Our measurements also show that recomputation is faster than in place recovery in every configuration we tested, that diagnosis rather than repair dominates recovery cost, and we report the defects we found while validating the verifier itself.

↑