arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RealisticTritonBench:面向真实世界AI框架的Triton内核生成基准测试

RealisticTritonBench: A Benchmark for Triton-Kernel Generation in Real-World AI Frameworks

Jinjun Huang, Zhongzhen Wen, Tongtong Xu, Meng Yan, Xin Xia, Zhongxin Liu

arXiv 2608.12004首次发表:更新:

发表机构

State Key Lab for Novel Software Technology, Nanjing University; Huawei; The School of Big Data and Software Engineering, Chongqing University(南京大学计算机软件新技术国家重点实验室; 华为; 重庆大学大数据与软件学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出RealisticTritonBench基准,解决现有Triton内核生成基准的局限,评估发现领先LLM在真实世界Triton内核生成任务上仍表现不佳。

AI 中文摘要

在现代AI框架中,GPU内核是决定系统整体性能的关键。Triton兼具易用性、可移植性以及接近手工编写CUDA的性能,因此被广泛用于实现GPU内核。近期研究进展显示,大型语言模型(LLM)具备自动生成Triton内核的潜力,可减少专业内核开发人员的手动工作量。已有若干基准测试用于评估LLM生成的Triton内核,但这些基准存在三个关键局限:其一,它们将任务限制在PyTorch到Triton的转换,无法反映真实世界Triton任务的多样性与复杂性;其二,它们仅评估单个内核的性能,而非端到端性能,而端到端性能才是AI框架实际部署的核心标准;其三,它们依赖手动编写的单个内核评估脚本,这些脚本可能存在缺陷,导致模型可利用缺陷绕过正确性检查并获得虚高分数。为解决这些局限,我们推出RealisticTritonBench,这是首个从流行AI框架的真实世界拉取请求(PR)中衍生Triton内核生成任务的基准测试,可实现逼真的、类似生产环境的评估。RealisticTritonBench系统地从流行的开源AI框架中提取修改Triton内核的PR,并将其转换为具有具体工程上下文的生成任务。每个任务以自然语言需求作为输入,要求生成对应的Triton内核实现,并配备完整且可复现的评估环境。与以往专注于孤立内核性能的基准不同,RealisticTritonBench将生成的内核集成到其原始框架中,并使用端到端测试对其进行评估,从而实现更忠实的评估。我们在RealisticTritonBench上评估了领先的LLM,发现它们在真实世界的Triton内核生成任务上仍然表现不佳。

英文摘要

In modern AI frameworks, GPU kernels are key to overall system performance. Combining usability, portability, and near-handwritten CUDA performance, Triton is widely adopted for implementing GPU kernels. Recent advances show the potential of large language models (LLMs) to automatically generate Triton kernels, reducing the manual effort required from expert kernel developers. Several benchmarks evaluate LLM-generated Triton kernels. However, they suffer from three key limitations: (1) they restrict tasks to PyTorch-to-Triton translation, failing to reflect the diversity and complexity of real-world Triton tasks; (2) they evaluate only individual-kernel performance rather than end-to-end performance, the core criterion for real-world deployment in AI frameworks; and (3) they rely on manually written evaluation scripts for individual kernels, which may contain flaws that models can exploit to bypass correctness checks and obtain inflated scores. To address these limitations, we introduce RealisticTritonBench, the first benchmark to derive Triton kernel generation tasks from real-world pull requests in popular AI frameworks, enabling realistic, production-like evaluation. RealisticTritonBench systematically extracts PRs that modify Triton kernels from popular open-source AI frameworks and transforms them into generation tasks with concrete engineering contexts. Each task takes a natural language requirement as input and requires a corresponding Triton kernel implementation, with a complete and reproducible evaluation environment. Unlike prior benchmarks focused on isolated kernel performance, RealisticTritonBench integrates generated kernels into their original frameworks and evaluates them using end-to-end tests, enabling a more faithful assessment. We evaluate leading LLMs on RealisticTritonBench and find that they still struggle with real-world Triton kernel generation tasks.

CommentsAccepted by ASE 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑