arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

整数不在场证明:定位INT8量化大语言模型推理中的跨内核差异

The Integer Alibi: Localizing Cross-Kernel Divergence in INT8-Quantized LLM Inference

Teng-Ruei Chen

arXiv 2608.13756首次发表:更新:

发表机构

Krixvon(克里克斯冯公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究验证了vLLM中INT8线性内核(CUTLASS与Triton)的可互换性,定位到差异源于累加器后的缩放应用与输出舍入,提出的干预措施可恢复端到端逐位一致,还揭示了FP8 GEMM差异的不同特征并发布相关一致性检查流程。

AI 中文摘要

通常认为,实现相同缩放INT8通用矩阵乘法(GEMM)接口的两个GPU内核是可互换的。我们验证这一假设:固定检查点、提示词、硬件、推理引擎、解码及量化配置,仅替换vLLM内部的INT8线性内核(CUTLASS与Triton)。对于17亿参数模型,两次冷重启间的输出逐位一致,但在所有端到端对比中,两个内核对任意序列的预测均不一致(0/8、0/16及0/64)。这一差异不止是基准测试偏差,更源于整数不在场证明:在经验证的无溢出边界下,共享INT8操作数对应的INT32点积是精确且与顺序无关的,因此累加器不可能是差异来源。向Qwen3-1.7B和8B模型的所有线性层(分别为196和252层)提供两个内核的相同操作数,我们发现在2的幂次缩放因子下输出逐位一致,确认1.7B模型的196/196个序列的预测列表被固定,8B模型的252/252个序列的预测列表被固定但非盲预测;在检查点的实际缩放因子下,观测到的差异最多为1个bfloat16间距。这将差异定位到精确累加器后的缩放应用及输出舍入。作为探针检查点应用时,相同干预可恢复端到端逐位一致(8/8和16/16序列)。跨实现的FP8 GEMM呈现不同特征:差异的发生率和幅度随约简深度增加,而INT8差异的占比保持在百万分之几,且在K的64倍范围内不超过1个间距。教师强制重放将层与标记关联:翻转集中在小对数几率边缘,其ROC-AUC达0.94,可在16384个位置上预测翻转风险。我们将发布预注册信息、每层预测结果、含内核选择证据的清单,以及将这些控制措施转化为内核互换性具体检查的一致性流程。

英文摘要

Two GPU kernels implementing the same scaled INT8 GEMM interface are usually treated as interchangeable. We test that assumption: holding the checkpoint, prompts, hardware, inference engine, decoding, and quantization configuration fixed, we swap only the INT8 linear kernel (CUTLASS versus Triton) inside vLLM. At 1.7B each arm reproduces itself bit-for-bit across cold restarts, yet the arms agree on no sequence in any end-to-end comparison we ran (0/8, 0/16, and 0/64). What makes this more than a benchmark discrepancy is an integer alibi: for shared INT8 operands under a verified no-overflow bound, the INT32 dot product is exact and order-independent, so the accumulator cannot be the source of any difference. Feeding both kernels identical operands from every linear layer of Qwen3-1.7B and 8B (196 and 252 layers), we find bit-identical outputs under power-of-two scales, confirming a pinned prediction list 196/196 and 252/252 (pre-registered at 1.7B, pinned but not blind at 8B), and observed differences of at most one bfloat16 spacing under the checkpoints' real scales. This localizes the divergence to scale application and output rounding after the exact accumulator. Applied as a probe checkpoint, the same intervention restores end-to-end bitwise agreement (8/8 and 16/16 sequences). Cross-implementation FP8 GEMM shows a different signature: both the prevalence and the magnitude of differences grow with reduction depth, while the INT8 fraction stays at parts per million and within one spacing over a 64x range of K. Teacher-forced replay ties layers to tokens: flips concentrate at small logit margins, which predict flip risk with ROC-AUC 0.94 on 16,384 positions. We will release the pre-registration, per-layer predictions, manifests with kernel-selection evidence, and a conformance procedure that turns these controls into a concrete check for kernel interchangeability.

Comments10 pages (IEEEtran two-column), 3 figures, 4 tables. Pre-registered protocol with append-only amendments. Companion to arXiv:2608.11693. v2: corrects the product-bound attribution (weight side, not activation side; 16256 is exact) and distinguishes the layer-level uniform pow2 regime from the per-channel probe; no result changes. Figure count corrected from v1

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑