AI 中文总结
本研究校准了 bfloat16 GEMM 上单侧 ABFT 的误报阈值,发现其检测率远低于预期,且可检测性由位位置而非收缩长度决定,否定了原假设机制。
AI 中文摘要
基于算法的容错技术通过将乘积的校验和与从因子重新计算的校验和进行比较来检测损坏的矩阵乘法,当残差超过阈值时标记故障。经典阈值为 tau = Ku,源自点积舍入界限,该界限要求 Ku <= 0.01。在 bfloat16 且收缩长度 K = 4096 时,Ku 约为 32,因此该界限不成立,阈值未定义。先前的工作因此仅在 float32 下评估校验和 ABFT。我们提供了缺失的校准。本文是对输出级、GEMM 后扰动模型下检测器的表征;它不是对硅片中故障的表征,此处报告的任何比率均不估计部署硬件中的故障发生率。我们将广义帕累托分布拟合到 7.99e7 个干净的逐行残差上,并外推得到每 GEMM 误报率为 1e-6 时的阈值,得到 3.1445e-6,该值介于我们观察到的两个干净样本最大值之间。每 GEMM 转换背后的独立性假设是通过经验而非断言来检验的,其方差比为 0.985,与二项式参考相比。针对该阈值,检测器捕获了 400 次单指数位翻转中的 304 次(76%),未达到在收集任何数据之前注册的 90% 的标准,并且捕获了 600 次低尾数位翻转中的 0 次。可检测性在 K 的 16 倍范围内不变,尽管干净噪声底限随 K^-0.4999 下降:信号和底限一起缩放,因此位位置决定可检测性,而收缩长度不决定。这否定了该研究围绕设计的机制。我们报告了四项失败的注册预测,包括上述那项,以及四项对我们早期发现的撤回,其中两项由出版前审计而非预测发现。代码、数据和带时间戳的预注册文件已发布。
英文摘要
Algorithm-based fault tolerance detects corrupted matrix multiplications by comparing a checksum of the product against a checksum recomputed from the factors, flagging a fault when the residual exceeds a threshold. The classical threshold is tau = Ku, derived from a dot-product rounding bound that requires Ku <= 0.01. At bfloat16 with contraction length K = 4096, Ku is approximately 32, so the bound does not hold and the threshold is not defined. Prior work therefore evaluated checksum ABFT only at float32. We supply the missing calibration. This paper is a characterization of a detector under an output-level, post-GEMM perturbation model; it is not a characterization of faults in silicon, and no rate reported here estimates fault incidence in deployed hardware. We fit a Generalized Pareto distribution to 7.99e7 clean per-row residuals and extrapolate a threshold at a per-GEMM false-positive rate of 1e-6, obtaining 3.1445e-6, which sits between the two clean sample maxima we observed. The independence assumption behind the per-GEMM conversion is checked empirically rather than asserted, at a variance ratio of 0.985 against a binomial reference. Against this threshold the detector catches 304 of 400 single exponent-bit flips (76%), failing a bar of 90% that was registered before any data was collected, and catches 0 of 600 low-mantissa flips. Detectability is invariant in K across a 16x range even though the clean noise floor falls as K^-0.4999: signal and floor scale together, so bit position governs detectability and contraction length does not. This falsified the mechanism the study was designed around. We report four registered predictions that failed, including that one, and four retractions of our own earlier findings, two of which were caught by pre-publication audit rather than by prediction. Code, data, and timestamped pre-registration documents are released.
Comments15 pages. Code, data, and pre-registration documents at https://github.com/GautamTalksDev/assay-gpu (tag v0.1.1-paper). Also available at doi:10.5281/zenodo.22054179