arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

智能体内核优化:无需手动编写CUDA即可生成最先进的GPU内核

Agentic Kernel Optimization: Generating State-of-the-Art GPU Kernels Without Hand-Written CUDA

Mao Luo, Hongbin Li, Feng Lin, Hanling Yi, Zhe Huang

arXiv 2608.14560首次发表:更新:

AI 中文总结

该研究构建多智能体编排框架Houmao的GPU内核优化工作流,利用通用代码智能体在无需手动CUDA代码的情况下,生成的GPU内核在多个任务上实现远超PyTorch和FlashInfer基线的加速,在竞赛中取得最优结果,证明代码智能体可作为GPU内核开发的有效自主优化器。

AI 中文摘要

我们研究通用型代码智能体能否在无需任何手动编写的CUDA代码的情况下生成最先进的GPU内核。我们使用FlashInfer-Bench中的代表性工作负载来探究这一问题,重点关注Fused MoE、DSA TopK索引器和DSA稀疏注意力,并在NVIDIA B200 GPU上通过正确性门控的FlashInfer-Bench协议评估所有生成的内核。从PyTorch实现、工作负载定义、基准测试命令以及一组紧凑的CUDA优化技巧出发,我们在Houmao(一个用于异构编码智能体的多智能体编排框架)中构建了一个内核优化工作流,用于生成、调试、分析和优化内核。人类仅担任编排角色:定义工作流、执行正确性与反破解约束、提供关键参考资料、在进展停滞时重定向搜索,无需审核或编辑内核代码本身。在约19亿个智能体令牌的处理过程中,生成的内核相较于PyTorch参考实现,在Fused MoE上实现了92.68倍的加速,在DSA TopK索引器上实现了1101.02倍的加速,在DSA稀疏注意力上实现了181.35倍的加速,同时也显著优于对应的FlashInfer基线。在MLSys 2026 FlashInfer AI内核生成竞赛的官方评估中,我们生成的Fused MoE内核相较于FlashInfer基线实现了1.71倍的加速,超过了Fused MoE智能体辅助赛道的最高结果(1.68倍加速)。这些结果表明,在严谨的正确性优先工作流下,代码智能体可作为现代GPU内核开发的有效自主优化器。

英文摘要

We study whether general-purpose code agents can produce state-of-the-art GPU kernels without any manually written CUDA code. We investigate this question using representative workloads from FlashInfer-Bench, focusing on the Fused MoE, DSA TopK Indexer, and DSA Sparse Attention, and evaluate all generated kernels under the correctness-gated FlashInfer-Bench protocol on NVIDIA B200 GPUs. Starting from the PyTorch implementations, workload definitions, benchmark commands, and a compact set of CUDA optimization skills, we build a kernel optimization workflow in Houmao, a multi-agent orchestration framework for heterogeneous coding agents, to generate, debug, profile, and optimize the kernels. Humans remain strictly in an orchestration role: defining the workflow, enforcing correctness and anti-hacking constraints, supplying key references, and redirecting the search when progress stalls, without reviewing or editing the kernel code itself. Across roughly 1.9 billion agent tokens, the resulting kernels achieve speedups of 92.68x on Fused MoE, 1101.02x on DSA TopK Indexer, and 181.35x on DSA Sparse Attention relative to the PyTorch reference implementations, while also significantly outperforming the corresponding FlashInfer baselines. In the official evaluation of the MLSys 2026 FlashInfer AI Kernel Generation Contest, our generated Fused MoE kernel achieves a 1.71x speedup over the FlashInfer baseline, exceeding the top result of the Fused MoE agent-assisted track, which reports a 1.68x speedup. These results suggest that, under a disciplined correctness-first workflow, code agents can serve as effective autonomous optimizers for modern GPU kernel development.

CommentsTechnical report on AI code generation for practical GPU kernels on NVIDIA Blackwell GPUs

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑