用于大语言模型驱动的GPU内核生成的工具工程
Harness Engineering for LLM-Driven GPU Kernel Generation
浏览论文内容
中文总结 AI 辅助
在MLSys 2026 FlashInfer竞赛中,针对NVIDIA Blackwell B200 GPU,提出以工具为中心的大语言模型驱动的GPU内核优化系统,分离评估工具与优化控制器,利用Codex等生成候选内核,实验显示优化后平均延迟加速显著,且代理辅助内核效果更佳。
中文摘要 AI 辅助
大语言模型可辅助GPU内核生成,但其实际效果取决于生成代码能否可靠地受限、验证、分析和选择。本文在MLSys 2026 FlashInfer AI内核生成竞赛中,针对NVIDIA Blackwell B200 GPU提出了一个以工具为中心的大语言模型驱动的GPU内核优化系统。该系统将评估工具与基于分析的优化控制器分离,工具负责编译、正确性、官方对齐的计时和工件存档,控制器利用分析器和工作负载证据做出有界候选生成决策。通过五个运算符定义,保留的官方对齐工件相对于提供的FlashInfer基线,平均延迟加速分别为1.62倍、18.05倍、29.68倍、1.12倍和13.70倍。在评估定义中,代理辅助内核优于完全由代理生成的工件,表明专家提供的优化方向、高质量参考和工作负载上下文对于可靠的人工智能驱动的内核优化仍然至关重要。
英文摘要
Large language models (LLMs) can assist GPU kernel generation, but their practical effectiveness depends on whether generated code can be reliably constrained, validated, profiled, and selected. This paper presents a harness-centered system for LLM-driven GPU kernel optimization in the MLSys 2026 FlashInfer AI Kernel Generation Contest on NVIDIA Blackwell B200 GPUs. The system separates an evaluation harness from a profile-backed optimization controller: the harness enforces compilation, correctness, official-aligned timing, and artifact archival, while the controller turns profiler and workload evidence into bounded candidate-generation decisions. Human-authored skills capture operator constraints, references, profiling procedures, and promotion rules, while Codex and Claude Code agents generate candidate kernels inside those constraints. Across five operator definitions, the retained official-aligned artifacts achieved mean-latency speedups over supplied FlashInfer baselines of 1.62x, 18.05x, 29.68x, 1.12x, and 13.70x. The Agent-Assisted kernels outperform the Full-Agent artifacts across the evaluated definitions, indicating that expert-provided optimization directions, high-quality references, and workload context remain critical for reliable AI-driven kernel optimization.
发表机构
- Baidu, Inc.(百度公司)
机构由 AI 辅助整理,请以论文原文为准。