arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

NVIDIA H100上普通全局负载的并发响应

Concurrency Response of Plain Global Loads on the NVIDIA H100

Somashekar Manjunath, Rahul Ramachandra M

arXiv 2608.15764首次发表:更新:

AI 中文总结

本文通过微基准测试探究NVIDIA H100上普通全局负载的并发响应,发现其LDG带宽在每线程负载K≈2时达峰后下降约35%,排除分配大小依赖,还对比了异步复制与普通加载的性能。

AI 中文摘要

内存受限的GPU内核维持的带宽由其在途字节数决定,本文将利特尔法则(Little's Law)用作吞吐量核算工具,而非测量硬件池。CUDA在Hopper架构上通过普通加载(plain loads)和异步复制(asynchronous copies)等路径填充该预算;我们在3颗H100 SXM5芯片上通过独立的微基准测试表征这些路径的并发响应。主要结果关于普通加载路径:在我们的主要配置下,达到的LDG带宽在每线程提供的负载K约为2时达到峰值,随后从K=2到K=8下降约35%。该下降在固定工作控制(匹配不同K值下总发出的逻辑加载量、升序和反向扫描顺序)以及在两颗芯片上用相同仪器重复测试时依然存在,下降幅度分别为-35.0%和-35.2%。单独分析的计数器显示,在K=2到8之间,DRAM字节数几乎恒定,而L2扇区流量上升;40倍标称分配大小扫描(512 MB至20 GB,均高于约50 MB的L2,无地址跟踪)基本未改变该下降,排除了简单的分配大小依赖性。由于L2命中率在每个分配下随K上升,聚合请求流确实随K变化;我们将K报告为提供的软件指令级并行度(ILP),硬件机制留待后续研究。初步调查补充了匹配的异步复制与普通加载的对比(高提供深度下为2.1-2.9倍,两颗芯片)、B芯片上同一块线程束(CTA)的双流观测(C芯片的对应检查结果不同,未纳入汇总)以及跨芯片原语基线。

英文摘要

The bandwidth a memory-bound GPU kernel sustains is set by how many bytes it keeps in flight. We use Little's Law here as throughput accounting, not as a measured hardware pool. CUDA fills that budget on Hopper through plain loads (ld.global) and asynchronous copies (cp.async), among other paths; we characterize their concurrency response with clean-room microbenchmarks on three H100 SXM5 dies. Our main result concerns the plain-load path: attained LDG bandwidth peaks at a small offered per-thread load (K ~ 2) and then declines, by about 35% from K=2 to K=8 at our primary configuration. The decline survives a fixed-work control matching total issued logical loads across K, ascending and reversed sweep orders, and replication on two dies with the same instrument (-35.0% and -35.2%). Separately profiled counters show DRAM bytes nearly constant over K=2->8 while L2-sector traffic rises, and a 40x nominal allocation-size sweep (512 MB to 20 GB, all above the ~50 MB L2; no address trace) leaves the decline essentially unchanged, disfavoring a simple allocation-size dependence. Because the L2 hit-rate nonetheless rises with K at every allocation, the aggregate request stream does change with K; we report K as offered software ILP and leave the hardware mechanism open. A preliminary survey adds a matched cp.async-versus-plain-load comparison (2.1-2.9x at high offered depth, two dies), a die-B same-CTA two-stream observation whose companion die-C check differs and is not pooled, and a cross-die primitive baseline.

Comments12 pages, 8 figures. Measurements on three NVIDIA H100 80GB HBM3 SXM5 dies

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑