发表机构
Stanford University; EPFL(斯坦福大学; 瑞士洛桑联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探索用大型语言模型智能体直接将 Triton 内核编译为 PTX,替代传统编译器后端,在多种 GPU 上实现最高 3.34 倍性能提升,并扩展验证器支持新架构。
AI 中文摘要
编译器后端在编程模型、工作负载和加速器不断演进的过程中,其构建和维护成本高昂。我们研究大型语言模型能否取代传统的优化和降级流水线,这一过程我们称之为 AI 降级(AI lowering)。我们研究了从 Triton 到 NVIDIA PTX 的 AI 降级:一个 LLM 智能体将 Triton 内核直接翻译为 PTX。我们构建了一个评估候选 PTX 的环境,以及一个智能体框架,其中 LLM 将 Triton 内核翻译为 PTX。在 Ada、Hopper 和 Blackwell GPU 上的十二个常见内核以及来自近期 ML 论文的十个内核上,AI 降级达到了自动调优 Triton 性能的 0.83 倍至 3.34 倍。最大的性能提升来自于 Triton 降级流水线未执行的转换,例如将打包的二进制权重直接解码为 Tensor Core 操作数(在 BitDelta 上为 3.34 倍),在张量存储器中为每个线程分配完整的 softmax 行(在 FlashAttention 上为 1.37 倍),以及重用重叠的卷积窗口(最高 2.23 倍)。这些结果依赖于一个具有全面验证支持的稳健评估框架。我们基于现有的 PTX 验证器 Volta,并大幅扩展其以支持现代 GPU 架构,引入了对 Blackwell 的 tcgen05 Tensor Core 接口的支持。这需要对三个架构特性进行建模:托管张量存储器、基于描述符的操作数布局,以及通过提交、等待、内存屏障和代理栅栏协调的异步执行。我们讨论了将这些特性形式化所面临的挑战,以及当前的局限性。我们的结果表明了一个新兴的未来:AI 编译器取代自定义编写的中间表示和检查器,减少为新通用和定制芯片启动软件所需的时间和工程工作量。
英文摘要
Compiler backends are expensive to build and maintain as programming models, workloads, and accelerators evolve. We investigate whether large language models can replace the conventional optimizing and lowering pipeline, a process that we call AI lowering. We study AI lowering from Triton to NVIDIA PTX: an LLM agent translates Triton kernels directly into PTX. We build an environment that evaluates candidate PTX, and an agentic harness in which an LLM translates Triton kernels into PTX. Across twelve common kernels on Ada, Hopper, and Blackwell GPUs and ten kernels from recent ML papers, AI lowering achieves 0.83x-3.34x the performance of autotuned Triton. The largest gains come from transformations that Triton's lowering pipeline does not perform, such as decoding packed binary weights directly into Tensor Core operands (3.34x on BitDelta), assigning each thread a complete softmax row in tensor memory (1.37x on FlashAttention), and reusing overlapping convolution windows (up to 2.23x). These results rely on a robust evaluation harness with comprehensive verification support. We build on Volta, an existing PTX verifier, and substantially extend it to support modern GPU architectures by introducing support for Blackwell's tcgen05 Tensor Core interface. This requires modeling three architectural features: managed tensor memory, descriptor-based operand layouts, and asynchronous execution coordinated through commits, waits, memory barriers, and proxy fences. We discuss the challenges involved in formalizing them, as well as the current limitations. Our results suggest an emerging future in which AI compilers replace custom-written intermediate representations and checkers, reducing the time and engineering effort required to bring up software for new general-purpose and custom chips.