arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02611cs.DCcs.AIcs.LGcs.SE

KernelBrain:面向智能体的GPU内核优化的粗到细、预算感知搜索

KernelBrain: Coarse-to-Fine, Budget-Aware Search for Agentic GPU Kernel Optimization

Shuai Che, Gang Peng

首次发表
浏览论文内容

中文总结 AI 辅助

KernelBrain是结合LLM引导变异等技术的GPU内核优化智能体,可提升内核质量与搜索效率,在Triton内核任务中获多倍加速且缩短优化时间。

中文摘要 AI 辅助

自动化GPU内核优化在实践中仍具挑战性:生成的变体可能违反正确性约束,运行时测量存在噪声,且搜索常提前停滞。本文提出一种实用优化智能体,结合LLM引导的变异、自适应资源分配、策略门控评估及分析器驱动的诊断。该系统通过低成本评估筛选大量候选,仅向有前景的幸存者分配高保真预算以优化和进化GPU内核。在重要的Triton内核生成任务中,此设计同时提升内核质量与搜索效率,较PyTorch实现0.88倍至6.72倍加速,较最先进的内核智能体实现最高1.4倍加速,优化时间最多降低48%。

英文摘要

Automating GPU kernel optimization remains difficult in practice: generated variants can violate correctness constraints, runtime measurements are noisy, and search often stalls early. We present a practical optimization agent that combines LLM-guided mutation, adaptive resource allocation, policy-gated evaluation, and profiler-informed diagnosis. The system screens many candidates with low-cost evaluation and allocates higher-fidelity budget only to promising survivors to optimize and evolve GPU kernels. On important Triton kernel generation tasks, this design improves both kernel quality and search efficiency, reaching 0.88x-6.72x speedup over PyTorch and up to 1.4x speedup over the state-of-the-art kernel agent, with up to 48% lower optimization time.

发表机构

  • Microsoft(微软公司)

机构由 AI 辅助整理,请以论文原文为准。

↑