arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12471cs.CL

AMDKernelVault:面向AMD GPU内核优化的大规模数据集与智能体训练

AMDKernelVault: Large-Scale Datasets and Agentic Training for AMD GPU Kernel Optimization

Ji Liu, Saptarshi Majumder, Yiqing Huang, Wenwen Ouyang, Umang Pandey, Zeping Li, Chushi Chen, Zihao An, Puyuan Yang, Zekai Li, Sina Rafati, Ziqiong Liu, Pratik… 展开作者

Ji Liu, Saptarshi Majumder, Yiqing Huang, Wenwen Ouyang, Umang Pandey, Zeping Li, Chushi Chen, Zihao An, Puyuan Yang, Zekai Li, Sina Rafati, Ziqiong Liu, Pratik Prabhanjan Brahma, Dong Li, Zicheng Liu, Sharon Zhou, Emad Barsoum

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出AMDKernelVault,一个包含HIP和Triton内核的大规模语料库及智能体训练框架,通过HIPKernelGen和TritonKernelGen流水线生成并验证内核,训练Qwen3-8B在多个基准上取得最高正确率,弥补了AMD GPU内核优化的空白。

中文摘要 AI 辅助

我们推出了AMDKernelVault,这是一个面向近期AMD CDNA GPU的开源HIP和Triton内核语料库及训练框架。现有的基于LLM的内核智能体主要围绕CUDA/NVIDIA,且常常依赖反复调用前沿LLM进行生成、反思和优化。为弥补这一空白,我们开发了HIPKernelGen和TritonKernelGen,这两个智能体驱动的流水线将PyTorch参考实现转换为HIP或Triton内核,在ROCm下编译并验证候选内核,并在AMD硬件上进行延迟性能分析。该语料库包含62,153个经执行验证的HIP内核样本、2,377条基于生产环境的ROCm库问答条目,以及39,893个Triton内核。我们进一步使用监督微调和执行感知的强化学习训练Qwen3-8B,以展示该语料库的实用性。在固定的评估预算下,该模型在PyTorch到HIP转换(34.0% Pass@1)、TritonBench-G(33.2% Corr@3)和ROCmBench(41.94% Corr@3)上取得了对比模型中最高的正确率,但在编译或速度指标上并未全面领先。语料库和文档可在该https URL获取,相关的训练和内核生成代码可在该https URL获取。

英文摘要

We introduce AMDKernelVault, an open HIP and Triton kernel corpus and training framework for recent AMD CDNA GPUs. Existing LLM-based kernel agents are largely CUDA/NVIDIA-centric and often depend on repeated frontier-LLM calls for generation, reflection, and optimization. To address this gap, we develop HIPKernelGen and TritonKernelGen, agent-driven pipelines that transform PyTorch references into HIP or Triton kernels, compile and validate candidates under ROCm, and latency-profile them on AMD hardware. The corpus contains 62,153 execution-verified HIP kernel samples, 2,377 production-grounded ROCm Libraries QA entries, and 39,893 Triton kernels. We further train Qwen3-8B with supervised fine-tuning and execution-aware reinforcement learning as a demonstration of the corpus's utility. Under fixed evaluation budgets, it achieves the highest correctness among the compared models on PyTorch-to-HIP (34.0% Pass@1), TritonBench-G (33.2% Corr@3), and ROCmBench (41.94% Corr@3), but does not uniformly lead compilation or speed metrics. The corpus and documentation are available at https://huggingface.co/datasets/amd/AIG-Datasets, and the associated training and kernel-generation code is available at https://github.com/AMD-AGI/hip_kernel_llm_lab.

发表机构

  • Advanced Micro Devices, Inc.(超威半导体公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑