arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

剖析Nvidia Hopper上LLM推理的GPU利用率

Dissecting GPU Utilization for LLM Inference on Nvidia Hopper

Mohammad Siavashi, Gerald Q. Maguire, Dejan Kostic, Marco Chiesa

arXiv 2609.12923首次发表:更新:

发表机构

KTH Royal Institute of Technology(瑞典皇家理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM推理中单一SM利用率掩盖真实性能瓶颈的问题,本文在H100上剖析vLLM,提出八个计数器验证的视图,将利用率差距映射到片段填充、占用率等具体机制。

AI 中文摘要

单一的SM利用率百分比可能使LLM推理工作负载看起来计算饱和,同时掩盖了实际完成的有用工作量。问题不在于计数器有误,而在于它将多种不同机制压缩成一个数字。这在解码阶段最为严重,因为每个请求仅贡献一个新token,密集投影GEMM变成小行矩阵乘法。在Hopper上,bfloat16 GMMA路径以固定的64行矩阵片段执行这些操作,因此小批量解码只能用真实token行填充每个片段的很小一部分。本文在H100 NVL上,使用FlashAttention-3和cuBLASLt对vLLM进行剖析,覆盖冷预填充、热预填充和解码阶段,并扫描序列长度和批量大小。我们用八个经计数器验证的视图替代通常的单一利用率数字,这些视图源自原始Nsight Compute报告,每个视图对应一个NCU计数器或显式公式。这些视图共同将利用率差距映射到具体机制——片段填充、占用率限制、停顿特征、波量化及内核选择——涵盖四个生产模型和六种逐层内核角色。

英文摘要

A single SM utilization percentage can make an LLM inference workload look compute-saturated while hiding how much useful work is being done. The problem is not that the counter is wrong, but that it collapses several different mechanisms into one number. This is most severe during decode, where each request contributes only one new token and dense projection GEMMs become small-row matrix multiplications. On Hopper, the bfloat16 GMMA path executes these operations in fixed 64-row matrix fragments, so small-batch decode can fill only a small fraction of each fragment with real token rows. In this paper, we profile vLLM with FlashAttention-3 and cuBLASLt on an H100 NVL across cold prefill, warm prefill, and decode, sweeping sequence length and batch size. We replace the usual single utilization number with eight counter-validated views derived from raw Nsight Compute reports, each pinned to an NCU counter or explicit formula. Together, these views map utilization gaps to concrete mechanisms - fragment fill, occupancy limits, stall signatures, wave quantization, and kernel selection - across four production models and six per-layer kernel roles.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑