KernelArc:面向GPU内核优化的多智能体框架
KernelArc: A Multi-Agent Framework for GPU Kernel Optimization
- Interuniversity Microelectronics Centre (IMEC)(校际微电子中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出KernelArc多智能体框架,在NVIDIA H100等GPU上针对SOL-ExecBench工作负载优化出多款高性能内核,相关成果在对应基准任务中排名第一,验证了共享多智能体搜索的优化价值。
AI中文摘要:
我们提出KernelArc,一种面向异构工作负载的自主GPU内核优化多智能体框架。策略专业化智能体并行运行,仅通过结论型共享内存、确定性基准守卫、只读跨智能体状态及 plateau 触发的起草机制进行协调。我们在NVIDIA H100与B200 GPU上,采用具类别代表性的SOL-ExecBench工作负载评估KernelArc。所得实现涵盖自定义BF16 GEMM、静态cuBLASLt Expert-API配置表、融合型混合专家反向传播、形状门控解码器层融合、原生NVFP4分组查询注意力及分页预填充注意力。在2026年7月30日记录的公开SOL-ExecBench排行榜快照中,这些提交结果在具代表性的L1、L2、量化及FlashInfer任务上排名第一。上述轨迹印证了本文核心动机:共享多智能体搜索可在固定候选预算内拓展探索范围并达成更优的现有方案,而各协调特征的价值取决于内核与优化阶段。
英文摘要:
We present KernelArc, a multi-agent framework for autonomous GPU kernel optimization across heterogeneous workloads. Strategy-specialized agents run in parallel and coordinate through conclusions-only shared memory, a deterministic benchmark guard, and read-only cross-agent state with plateau-triggered drafting. We evaluate KernelArc on NVIDIA H100 and B200 GPUs using category-representative SOL-ExecBench workloads. The resulting implementations span custom BF16 GEMM, static cuBLASLt Expert-API configuration tables, fused mixture-of-experts backward, shape-gated decoder-layer fusion, native NVFP4 grouped-query attention, and paged prefill attention. In the public SOL-ExecBench leaderboard snapshot recorded on August~20, 2026, KernelArc ranked first on every representative L1, L2, Quantization, and FlashInfer task evaluated. The trajectories support the paper's central motivation: shared multi-agent search can broaden exploration and reach stronger incumbents within a fixed candidate budget, while the value of individual coordination features depends on the kernel and optimization stage.