arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24762cs.AIcs.PF

Kernel Forge:用于基于大语言模型的CUDA内核生成与优化的代理框架

Kernel Forge: An Agent Harness for LLM-based Generation and Optimization of CUDA Kernels

Joshua Brodsky, Dhravid Kumar, Savini Kashmira, Jayanaka Danatanarayana, Jason Mars, Krisztian Flautner, Lingjia Tang

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对机器学习模型中计算内核优化需专家手写代码的问题,提出开源端到端代理框架Kernel Forge,它支持多种工作负载,用蒙特卡洛树搜索优化路径,经实验评估,优化后的内核性能优于PyTorch即时模式。

中文摘要 AI 辅助

机器学习模型越来越多地嵌入日常软件,其大部分运行时间花在矩阵乘法、卷积和归一化等少量计算内核上。优化这些内核是降低延迟和成本的直接方法,但传统上需要专家工程师手写低级GPU代码。基于大语言模型的代理系统现在可以用更少人力生成和优化内核,但现有工具存在诸多问题。我们提出Kernel Forge,一个开源的端到端代理框架,支持多种工作负载,使用蒙特卡洛树搜索探索优化路径,还有图形用户界面。我们在四个PyTorch模型上评估,结果显示其优化后的内核性能优于PyTorch即时模式。

英文摘要

Machine learning models are increasingly embedded in everyday software, and most of their runtime is spent in a small set of compute kernels such as matrix multiplication, convolution, and normalization. Optimizing these kernels is one of the most direct ways to reduce latency and cost, but it has traditionally required expert engineers to hand-write low-level GPU code. Agentic systems built on large language models (LLMs) can now generate and optimize kernels with far less human effort, yet existing tools are largely evaluated on randomly generated tensors and isolated kernels, emit standalone CUDA code that developers must manually reintegrate, mostly target only LLM PyTorch models, and offer limited support for inspecting and debugging results. We present Kernel Forge, an open-source, end-to-end agentic harness that accepts any unmodified PyTorch model in place. Kernel Forge supports vision, diffusion, and LLM workloads, uses Monte Carlo Tree Search (MCTS) to explore multiple optimization paths rather than a single linear refinement chain, and ships with a graphical user interface for monitoring progress, inspecting candidate kernels, and debugging failures. We evaluate Kernel Forge on four PyTorch models spanning vision, diffusion, and LLM workloads on an NVIDIA DGX Spark with GB10 GPU. With only 50 optimization iterations per kernel, it optimizes 14 kernels to outperform PyTorch eager mode, reaching $1.52\times$ on adaptive\_avgpool2d in ResNet-50, $1.70\times$ on group\_norm in Stable Diffusion 3.5 Medium, $2.83\times$ on softmax in Gemma 4 E2B, and $1.54\times$ on softmax in Qwen 3.5 35B-A3B.

发表机构

  • University of Michigan(密歇根大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑