arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

规格表并非内核:对NVIDIA Blackwell Ultra上INT8可用性的ISA与源代码级审计

Spec Sheets Are Not Kernels: An ISA- and Source-Level Audit of INT8 Availability on NVIDIA Blackwell Ultra

Teng-Ruei Chen

arXiv 2608.11693首次发表:更新:

AI 中文总结

本文审计发现,NVIDIA Blackwell Ultra(B300)的INT8支持在ISA、CUTLASS内核库、vLLM和SGLang四层栈中被分层撤回,名义存在的INT8格式实际无法部署,凸显量化格式可用性是全栈属性。

AI 中文摘要

NVIDIA公布的规格显示,Blackwell Ultra GPU(B300)的FP8与INT8张量核心吞吐量的密集计算比约为30:1,其前代产品H200和B200均提供1:1的比例。我们通过追踪INT8 W8A8支持在四层栈中的情况,审计这种优先级降低在实际中的含义:公布的规格、PTX ISA、NVIDIA的CUTLASS内核库,以及两个主要的开源LLM服务引擎(vLLM和SGLang)。我们发现了一致的分层式撤回:(i)即使同一PTX修订版将FP4种类扩展到该目标,PTX ISA也从未在sm_103a上公开第五代张量核心整数路径(相关内容见此链接,带有.kind::i8),使得传统的 warp 级IMMA成为B300上唯一架构合法的整数张量核心路径;(ii)CUTLASS的内核生成器会为任何针对103a的构建明确跳过INT8 UMMA的生成,同时无条件生成FP8;(iii)vLLM未提供适用于Blackwell的INT8 GEMM,且在模型加载后的第一次前向传播时会因严重运行时错误而失败;(iv)SGLang的提前编译INT8 GEMM在Sm90处停止,而其FP8调优配置已覆盖B200。我们记录了一个逃逸通道(通过环境变量将vLLM的INT8路径重定向到JIT编译的Triton后端)、在sm_103上检测“原生INT8”的明显分析器方法中的假阴性陷阱,以及使朴素测试成本高昂的实际失败语义。这些发现共同表明,量化格式的可用性是整个栈的属性,而非模型或规格表的属性。四个不同的层,其中三层属于NVIDIA自身,以相互一致的方式撤回了INT8支持,且数据表上名义上存在的格式默认情况下在该硬件上无法部署。

英文摘要

NVIDIA's published specifications give the Blackwell Ultra GPU (B300) a dense-compute ratio of roughly 30:1 between FP8 and INT8 tensor-core throughput; its predecessors, H200 and B200, both provide 1:1. We audit what this deprioritization means in practice by tracing INT8 W8A8 support through four layers of the stack: the published specifications, the PTX ISA, NVIDIA's CUTLASS kernel library, and the two major open-source LLM serving engines (vLLM and SGLang). We find a consistent, layered withdrawal: (i) the PTX ISA never exposes the fifth-generation tensor-core integer path (tcgen05.mma with .kind::i8) on sm_103a, even though the same PTX revision extends the FP4 kinds to that target, leaving legacy warp-level IMMA as the only architecturally legal integer tensor-core path on B300; (ii) CUTLASS's kernel generator explicitly skips INT8 UMMA generation for any build targeting 103a, while generating FP8 unconditionally; (iii) vLLM ships no INT8 GEMM for Blackwell and fails with a hard runtime error at the first forward pass, after the model has loaded; and (iv) SGLang's ahead-of-time INT8 GEMM stops at Sm90, while its FP8 tuning configurations already cover B200. We document an escape hatch (rerouting vLLM's INT8 path to a JIT-compiled Triton backend via an environment variable), a false-negative trap in the obvious profiler methodology for detecting "native INT8" on sm_103, and the practical failure semantics that make naive testing expensive. Together, these findings show that a quantization format's availability is a property of the whole stack rather than of the model or the spec sheet. Four distinct layers, three of them NVIDIA's own, withdrew INT8 support in mutually consistent ways, and a format that is nominally present on the datasheet is, by default, undeployable on this hardware.

Comments8 pages (IEEEtran two-column), 2 figures, 3 tables. All sources pinned to commits/digests with access dates; no performance measurements (companion measurement study in preparation). v2: bibliography only -- 15 refs recited to their peer-reviewed venues, author dashes suppressed, artifact URL added; no technical changes

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑