发表机构
Beijing Normal University; Beijing Jiaotong University; Institute of Automation, CAS; Peking University; BAAI(北京师范大学; 北京交通大学; 中国科学院自动化研究所; 北京大学; 北京智源人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出KernelGenBench基准,评估LLM与智能体生成的Triton内核,发现智能体方法优于纯LLM采样,cuBLAS算子最具挑战,跨平台性能退化严重,自主生成成本远高于简单采样。
AI 中文摘要
大语言模型(LLM)大幅提升了对高效加速器内核的需求,但内核开发仍是高度专业化且劳动密集型的任务。近期LLM与智能体框架的兴起为自动内核生成提供了可行路径。不过,尽管进展迅速,目前仍缺乏全面的基准测试来跨多样算子源或异构硬件平台严格评估LLM生成的内核。我们提出KernelGenBench,这是一个用于系统评估基于LLM和智能体生成的Triton内核的统一基准,覆盖多样算子源与异构硬件平台。它包含两个互补的子基准:KernelGenBench-MS(多源),评估210个来自三个源的算子,超出标准PyTorch中心任务范畴;KernelGenBench-MC(多芯片),使用110个算子子集测量跨6个异构硬件平台的性能可移植性。我们的大规模评估消耗超150亿token,结果显示:(1)基于智能体的方法始终优于纯LLM采样方法,而cuBLAS算子对所有方法而言最具挑战性;(2)不同硬件平台的生成性能差异显著,即使是近期的内核专用智能体也存在严重的跨平台性能退化(如AutoKernel在NVIDIA上的准确率为87%,在平台E上降至25%);(3)自主内核生成的成本仍很高,专用智能体方法每个成功算子平均消耗511万token(AKO4all达519万),比简单LLM采样方法高出数个数量级。
英文摘要
Modern AI systems depend on specialized accelerator kernels, whose development is complicated by increasingly diverse operators and hardware. LLMs and agentic systems promise to automate this work, but existing evaluations do not show whether their performance transfers across operator sources and hardware platforms, or what such transfer costs. We present KernelGenBench, the first unified multi-source and multi-chip infrastructure for evaluating LLM- and agent-generated Triton kernels. With a common Triton target spanning six hardware platforms, it provides the broadest cross-vendor hardware coverage among existing kernel-generation benchmarks. We report two controlled analytical views: KernelGenBench-MS (Multi-Source) covers 210 operators from PyTorch ATen, production vLLM operators, and proprietary cuBLAS routines, while KernelGenBench-MC (Multi-Chip) evaluates a semantically stable 110-operator subset across six hardware platforms. Our evaluation consumed over 15 billion tokens. Agentic execution improved correctness, but no method dominated across sources and platforms: vLLM posed the strongest correctness challenge, cuBLAS set the highest performance ceiling, and AutoKernel accuracy fell from 87% on NVIDIA to 25% on Iluvatar CoreX. These improvements were costly: specialized agents averaged 4.99 million tokens per successful operator, rising to 6.25 million for CUDA Optimized Skill. The results establish operator source, hardware platform, and agentic scaffold as distinct dimensions of kernel-generation capability, and show that success in a familiar source-hardware setting is not a reliable proxy for deployment readiness.
Comments9 pages, 3 figures. Code and data are publicly available at https://github.com/flagos-ai/KernelGenBench