arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

D2K-Bench:LLM智能体能否将专家设计转化为高效的GPU内核?

D2K-Bench: Can LLM Agents Turn Expert Designs into Efficient GPU Kernels?

Daifeng Li, Huiqiang Jiang, Chengruidong Zhang, Wei Wu, Xudong Guo, Jianhong Tu, Jianwei Zhang, Binhang Yuan, Dayiheng Liu

arXiv 2610.03226首次发表:更新:

发表机构

HKUST; Alibaba Group; USTC(香港科技大学; 阿里巴巴集团; 中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

D2K-Bench通过26个任务和85个工作负载诊断基准,评估LLM智能体在专家设计指导下生成GPU内核的效率,实验表明指导显著提升正确率与性能,并揭示未实现的设计属性。

AI 中文摘要

大语言模型(LLM)智能体生成的GPU内核在效率上可能仍不及专家实现,但仅凭运行时性能无法揭示这一差距与设计发现和实现之间的关系。我们提出了D2K-Bench,一个包含26个任务和85个工作负载的诊断基准,用于衡量智能体将专家设计指导转化为高效GPU内核的有效程度。该指导涵盖L1:高层算法见解,L2:数据流设计,以及L3:底层优化技巧,并包括这些层级之间的依赖关系。有无指导的成对运行共享任务描述、工作负载、工具、硬件以及350轮次的预算。补充评估考察独立提出的设计以及生成代码中实现的设计属性。在NVIDIA B200 GPU上对五个模型进行的实验中,指导将130个模型-任务对上的正确率从93.1%提升至98.5%,并将所有26个任务的性能得分从1.46提高到1.95。对于在两次运行中均在所有26个任务上提交正确的三个前沿模型(GPT-6-Astra、Claude-Opus-4.8和GPT-5.6-Sol),几何平均加速比从1.69倍增加到2.49倍。在所有五个模型中,平均综合实现得分从100分中的57分提高到70分。这些结果显示了专家设计指导的价值,同时识别出仍未实现的设计属性。

英文摘要

GPU kernels generated by large language model (LLM) agents can remain less efficient than expert implementations, but runtime alone does not reveal how the gap relates to design discovery and implementation. We introduce D2K-Bench, a diagnostic benchmark of 26 tasks and 85 workloads that measures how effectively agents translate expert design guidance into efficient GPU kernels. The guidance covers L1: high-level algorithmic insights, L2: dataflow design, and L3: low-level optimization tricks, including dependencies among these levels. Pairwise runs with and without guidance share task descriptions, workloads, tools, hardware, and a 350-turn budget. Complementary assessments examine independently proposed designs and the design properties implemented in generated code. Across five models on NVIDIA B200 GPUs, guidance raises correctness over 130 model-task pairs from 93.1% to 98.5% and increases the Performance Score over all 26 tasks from 1.46 to 1.95. For the three frontier models with correct submissions on all 26 tasks in both runs (GPT-6-Astra, Claude-Opus-4.8, and GPT-5.6-Sol), geometric mean speedup increases from $1.69\times$ to $2.49\times$. Across all five models, the mean combined implementation score increases from 57 to 70 out of 100. These results show the value of expert design guidance while identifying design properties that remain unimplemented.

Comments30 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑