KernelOPT:面向GPU内核优化的调度感知智能体搜索
KernelOPT: Dispatch-Aware Agentic Search for GPU Kernel Optimization
浏览论文内容
中文总结 AI 辅助
KernelOPT提出多智能体系统,通过保留库调用并优化Triton子内核,结合四门验证级联,在250个KernelBench问题上实现最高1.40倍加速。
中文摘要 AI 辅助
深度学习推理和训练性能关键取决于GPU内核效率。现代编译器(如PyTorch Inductor)能从高层模型代码自动生成GPU内核,但通常与专家编写的实现相比性能差距悬殊。近期基于大语言模型(LLM)的内核优化器可以缩小这一差距,但仅针对独立内核,将编译后的模型视为黑盒,通常优化单个独立内核而不考虑编译器的结构决策或验证模型端到端性能。我们提出KernelOPT,一个将编译后模型视为结构化工件的多智能体系统。它保留厂商库调用(cuBLAS、cuDNN),并专门针对生成的Triton子内核,使用五个基于性能剖析的LLM智能体。一个四门验证级联(静态验证、多随机种子正确性、模型级float64回退验证和性能门控)在优化过程中筛选候选,并端到端验证重新拼接的模型。如果没有候选通过所有四个门,系统保留编译器基线。该系统接受PyTorch模型、独立Triton内核和Helion内核。在250个KernelBench问题上评估,KernelOPT相对于PyTorch Inductor实现了几何平均加速比:Level 1(51/100)为1.40倍,Level 2(31/100)为1.15倍,Level 3(12/50)为1.07倍。
英文摘要
Deep learning inference and training performance depends critically on GPU kernel efficiency. Modern compilers such as PyTorch Inductor automatically generate GPU kernels from high-level model code, but frequently underperform expert-written implementations by wide margins. Recent LLM-assisted kernel optimizers can close this gap for standalone kernels, yet treat compiled models as black boxes, generally optimizing individual standalone kernels without respecting the compiler's structural decisions or verifying the model end-to-end. We present KernelOPT, a multi-agent system that treats compiled models as structured artifacts. It preserves vendor library calls (cuBLAS, cuDNN) and exclusively targets generated Triton sub-kernels using five profiling-guided LLM agents. A four-gate verification cascade applies static validation, multi-seed correctness checking, model-level float64-fallback verification, and performance gating ($γ{=}1.03$) to filter candidates and verify the re-stitched model end-to-end. When candidates fail verification, the system preserves the compiler baseline. The system accepts PyTorch nn Modules, standalone Triton kernels, and Helion kernels. Evaluated on 250 KernelBench problems (100 Level 1, 100 Level 2 and 50 Level 3) on NVIDIA H200, KernelOPT achieves geometric mean speedups over torch compile of 1.40$\times$ (L1), 1.15$\times$ (L2), and 1.07$\times$ (L3) across all kernels, including fallback cases. Optimized-only geomeans (excluding cases where verification gates preserve the compiler baseline) are substantially higher: 2.54$\times$ (L1: 36/100), 1.84$\times$ (L2: 23/100), and 1.37$\times$ (L3: 11/50), reflecting where the optimizer achieves meaningful leverage.
发表机构
- Red Hat(红帽公司)
机构由 AI 辅助整理,请以论文原文为准。