发表机构
Google; Google DeepMind(谷歌; 谷歌DeepMind)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出面向TPU的多智能体系统MaxKernel,含三种内核开发范式,在JaxBench及真实工作负载上生成的优化内核性能匹配专家基线,已开源。
AI 中文摘要
为加速器设计和编写高性能自定义内核是一项复杂任务,需要深厚的硬件级专业知识。大型语言模型(LLM)可与实时编译器反馈结合,构建用于内核生成的智能体系统。本研究提出MaxKernel,这是一个多智能体系统,实现了三种不同的TPU内核开发范式:(1)用于协作式逐步设计的Human-in-the-Loop(HITL,人在回路中)智能体;(2)执行完全自动化、基于指标/跟踪的优化循环的Autonomous(Auto,自主)智能体;(3)将Auto智能体扩展到设计空间全局探索的基于图的自主搜索。三种范式均利用一组共享的专用子智能体处理规划、实现、自调试、测试和硬件分析。我们在JaxBench(一套针对TPU的包含50项不同内核任务的综合套件)以及来自最先进开源模型的复杂真实工作负载上评估MaxKernel。结果表明,MaxKernel始终能生成高度优化的实现,与专家手动调优的基线匹配,并在基准测试中实现显著性能提升。该智能体已开源,可通过此URL获取。
英文摘要
Designing and authoring high-performance custom kernels for accelerators is a complex task that requires deep hardware-level expertise. Large Language Models (LLM) can be leveraged together with real-time compiler feedback to build agentic systems for kernel generation. In this work, we present MaxKernel, a multi-agent system that implements three distinct paradigms for TPU kernel development: (1) a Human-in-the-Loop (HITL) agent for collaborative, step-by-step design; (2) an Autonomous (Auto) agent that executes a fully automated, metric/trace-driven optimization loop; and (3) a Graph-Based Autonomous Search that scales the Auto agent for global exploration of the design space. All three paradigms leverage a shared pool of specialized sub-agents to handle planning, implementation, self-debugging, testing, and hardware profiling. We evaluate MaxKernel on JaxBench, a comprehensive suite of 50 diverse kernel tasks for TPUs, alongside complex, real-world workloads from state-of-the-art open-source models. We demonstrate that MaxKernel consistently generates highly optimized implementations, matching expert hand-tuned baselines and delivering significant performance across the benchmark. Our agent is open-sourced and available https://github.com/AI-Hypercomputer/accelerator-agents/tree/main/MaxKernel.
Comments14 pages, 6 figures, 4 tables