arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从帕斯卡到布莱克威尔架构的 warp 发散特性研究

Characterizing Warp Divergence from Pascal to Blackwell

Alpin Dale

arXiv 2607.23402首次发表:更新:

AI 中文总结

研究从帕斯卡到布莱克威尔架构的 warp 发散特性,通过多种测试方法区分稳定行为与架构变化,发现发散路径线性序列化,执行效率与路径数量有关,不同架构重新收敛机制有变化,新屏障类无明显运行时效果,发散性能成本稳定可预测。

AI 中文摘要

自伏特架构引入独立线程调度(ITS)以来,人们普遍认为英伟达 GPU 以固定方式处理 warp 发散。本文以 ITS 之前的帕斯卡架构为基线,对安培、霍珀以及数据中心和消费级的布莱克威尔 GPU 进行测试。通过循环精确微基准测试、硬件计数器和对编译器生成的 SASS 的静态分析,区分稳定行为和架构变化。结果表明,所有测试代中,发散路径随路径数量\(k\)线性序列化,执行效率随\(32/k\)下降,惩罚与占用率无关,预测可消除序列化成本。帕斯卡架构也有类似行为,但其编译器发出的重新收敛机制有很大变化,布莱克威尔架构引入了新特性,如两级收敛屏障分类等。控制位翻转实验表明新屏障类在测试中无明显运行时效果。所以,即使英伟达的控制流 ISA 和重新收敛机制不断发展,发散仍保持稳定且可预测的性能成本。

英文摘要

Since Volta introduced Independent Thread Scheduling (ITS), NVIDIA GPUs have been widely assumed to handle warp divergence in a fixed manner. We test this assumption across Ampere, Hopper, and datacenter and consumer Blackwell GPUs, using pre-ITS Pascal as a baseline. Combining cycle-accurate microbenchmarks, hardware counters, and static analysis of compiler-generated SASS, we separate stable behavior from architectural change. Across all tested generations, divergent paths serialize linearly with the number of paths $k$, following $T(k) \approx sk$ with no super-linear reconvergence penalty. Warp execution efficiency falls as $32/k$, the penalty is independent of occupancy, and predication removes the serialization cost. The same behavior appears on Pascal, showing that this programmer-visible cost model predates ITS. The compiler-emitted reconvergence machinery, however, has changed substantially. Pascal uses a per-warp SSY/SYNC instruction stack, whereas later generations use barrier-register instructions. Deferred reconvergence beyond the immediate post-dominator falls from 29 cases on Ampere to 2 on Blackwell. Blackwell also introduces a two-tier convergence-barrier classification, uniform-branch instructions, and explicit partial-mask warp synchronization, none of which appear on Ampere or Hopper. Controlled bit-flip experiments indicate that the new barrier class is a static compiler classification with no observable runtime effect in our tests. Thus, divergence retains a stable and predictable performance cost even as NVIDIA's control-flow ISA and reconvergence mechanisms continue to evolve.

Comments6 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑