arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16389cs.PL

Exo-GPU:面向张量核心的安全、命令式、用户可调度的编程

Exo-GPU: Safe, Imperative, User-schedulable Programming for Tensor Cores

  • MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

David Zhao Akeley, Yuka Ikarashi, Jonathan Ragan-Kelley

AI总结:

Exo-GPU是一种低级命令式语言,将并行与同步视为注释,验证顺序-并行等价性,用于编写H100 GEMM内核,性能超理论峰值80%以上,部分优于CUBLAS。

AI中文摘要:

现代GPU不仅需要SIMT风格的并行性,还需要计算与数据移动之间的软件管理并发,以达到最大性能。性能工程师必须考虑将工作细分为计算资源层级(线程、线程束、线程束组、线程块、线程块簇),并且在许多情况下,还必须在内存层级的不同级别(寄存器、张量核心累加器、共享内存、全局内存)使用异步张量核心和内存拷贝指令。与CPU不同,CPU中的乱序执行由硬件管理并对程序员隐藏,而GPU通过异步指令向软件显式暴露指令重排序。成熟的GPU编程语言通常要么提供无安全保证的直接低级控制(例如,CUDA C++内联汇编或内部函数),要么提供更易于分析的高级抽象(例如,Triton的基于块的模型),这些抽象在编译器后端隐藏异步指令,这可能阻碍性能工程师通过调整关键细节来最大化性能。我们提出Exo-GPU,一种命令式低级语言,在CUDA之上创建最小抽象。我们的关键思想是将并行性和同步视为顺序代码上的简单注释,而不是基本的控制流原语,从而能够验证这些构造不会改变程序语义。其好处有两方面:程序员可以在没有隐藏控制流或修改的情况下推理代码,同时允许Exo-GPU编译器验证顺序-并行等价性——保证并行执行在功能上等同于其顺序解释。我们使用Exo-GPU为H100 GPU编写了GEMM内核,使用了wgmma、TMA和split-k。在大型问题上,我们的内核达到了理论峰值的80%以上,在某些情况下优于供应商提供的CUBLAS库。

英文摘要:

Modern GPUs require not only SIMT-style parallelism but also software-managed concurrency between compute and data movement to reach maximum performance. Performance engineers must reason about subdividing work into the hierarchy of computation resources (threads, warps, warpgroups, blocks, clusters), and, in many cases, also must use asynchronous tensor core and memcpy instructions on different levels of the memory hierarchy (registers, tensor core accumulators, shared memory, global memory). Unlike CPUs, where out-of-order execution is managed by hardware and hidden from programmers, GPUs expose explicit instruction reordering to software through these asynchronous instructions. Well-established GPU programming languages generally offer either direct low-level control without safety guarantees (e.g., CUDA C++ inline assembly or intrinsics) or easier-to-analyze, high-level abstractions (e.g., Triton's tile-based model) that hide asynchronous instructions in the compiler backend, which may prevent performance engineers from maximizing performance by tuning critical details. We propose Exo-GPU, an imperative, low-level language that creates minimal abstraction over CUDA. Our key idea is to treat parallelism and synchronization as mere annotations on sequential code rather than as fundamental control flow primitives, enabling verification that these constructs do not alter the program semantics. The benefit is twofold: programmers can reason about code without hidden control flow or mutation, while allowing the Exo-GPU compiler to verify sequential-parallel equivalence--guaranteeing that parallel execution is functionally equivalent to its sequential interpretation. We used Exo-GPU to author GEMM kernels for the H100 GPU, using wgmma, TMA, and split-k. Our kernels achieved over 80% of theoretical peak on large problem sizes, in some cases outperforming the vendor-provided CUBLAS library.

补充信息

↑