发表机构
Krixvon(克里克斯冯)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究量化GEMM内核一致性检查套件的检测极限,发现基于1个间距容忍度的套件无法检测多数尾声故障,将权重尺度重新量化为二的幂次可提升CUTLASS与Triton的一致性及生成序列的字节级匹配度。
AI 中文摘要
量化通用矩阵乘(GEMM)内核的一致性检查套件用于判断两种实现是否在容忍度内一致,本文对该套件的检测能力进行了测量。针对Qwen3-1.7B的8232个层-故障-工况单元,在参考INT8流水线中注入9种故障,结果显示5种尾声故障(尺度精度、双舍入、乘法顺序、输出截断、融合顺序)中,每一种故障都会使输出移动至多1个bfloat16间距,且在5880个单元中,只要发生移动就恰好为1个间距。因此,1个间距的容忍度从结构上就对该类故障完全不可察觉:5种故障中有4种未被套件中的任何检查检测到,第5种仅在二的幂次尺度下被检测到。违反累加器精确性前提条件或破坏操作数共享的故障则会被全部检测到,且空故障从未触发。因此,这类基于容忍度的套件所确立的一致性范围比可互换性更窄,仅能确认前提条件成立、操作数共享且差异保持在1个间距内。可用于检测1种已检测故障的二的幂次约束也具备部署可行性:将所有权重尺度重新量化为最接近的二的幂次,可使CUTLASS和Triton在所有线性层上逐位一致(权重尺度为检查点自身尺度时,CUTLASS为8/196、Triton为10/252,重新量化后分别为196/196、252/252),并在1.7B、8B、14B模型上生成字节级相同的token序列(8/8提示,三种模型均达到该比例)。观测到的困惑度点估计分别为+0.32%、-0.28%、+0.48%;90%置信区间在较小的两个模型上覆盖零值,但在14B模型上未覆盖,分别为+0.71%和+0.76%。此前报道的该干预措施导致困惑度增加157%的结果是探针的伪影,该探针在重写尺度时未对权重进行重新量化;分离效应后,99.8%的困惑度变化源于权重-尺度不匹配,而非二的幂次约束本身。
英文摘要
Conformance suites for quantized GEMM kernels ask whether two implementations agree within a tolerance. We measure what such a suite can detect. Injecting nine faults into a reference INT8 pipeline over 8,232 layer--fault--regime cells of Qwen3-1.7B, we find that every one of five epilogue faults -- scale precision, double rounding, multiplication order, output truncation, fused ordering -- moves the output by at most a single bfloat16 spacing, and by exactly one whenever it moves it at all, across 5,880 cells. A tolerance of one spacing is therefore blind to the entire class by construction: four of the five faults are detected by no check in the suite, and the fifth only under power-of-two scales. Faults that violate the accumulator's exactness preconditions, or that break operand sharing, are detected without exception, and a null fault never fires. What a tolerance-based suite of this shape establishes is therefore narrower than interchangeability: that the preconditions hold, that operands are shared, and that differences stay within one spacing. The power-of-two constraint that exposes the one detected fault is also deployable. Requantizing every weight scale to its nearest power of two makes CUTLASS and Triton agree bitwise at every linear layer (196/196 and 252/252, against 8/196 and 10/252 under the checkpoints' own scales) and yields byte-identical generated token sequences at 1.7B, 8B and 14B (8/8 prompts, against 0/8 at all three). Observed perplexity point estimates are +0.32%, -0.28% and +0.48%; the 90% intervals cover zero at the two smaller sizes but not at 14B, reaching +0.71% and +0.76%. A previously reported +157% perplexity for this intervention was an artifact of a probe that rewrote scales without requantizing the weights; separating the effects attributes 99.8% of it to the resulting weight--scale mismatch rather than to the power-of-two constraint itself.
Comments9 pages (IEEEtran two-column), 3 figures, 5 tables. Fault injection against a two-stage-pinned prediction matrix (8,232 scored cells; 63-cell pre-data core, corrected and re-pinned after a disclosed smoke run); power-of-two requantization measured at 1.7B/8B/14B. Companion to arXiv:2608.13756, whose open questions on check sensitivity and deployability this paper answers