KernelBrain:面向智能体的GPU内核优化的粗到细、预算感知搜索
KernelBrain: Coarse-to-Fine, Budget-Aware Search for Agentic GPU Kernel Optimization
浏览论文内容
中文总结 AI 辅助
KernelBrain是结合LLM引导变异等技术的GPU内核优化智能体,可提升内核质量与搜索效率,在Triton内核任务中获多倍加速且缩短优化时间。
中文摘要 AI 辅助
自动化GPU内核优化在实践中仍具挑战性:生成的变体可能违反正确性约束,运行时测量存在噪声,且搜索常提前停滞。本文提出一种实用优化智能体,结合LLM引导的变异、自适应资源分配、策略门控评估及分析器驱动的诊断。该系统通过低成本评估筛选大量候选,仅向有前景的幸存者分配高保真预算以优化和进化GPU内核。在重要的Triton内核生成任务中,此设计同时提升内核质量与搜索效率,较PyTorch实现0.88倍至6.72倍加速,较最先进的内核智能体实现最高1.4倍加速,优化时间最多降低48%。
英文摘要
Automating GPU kernel optimization remains difficult in practice: generated variants can violate correctness constraints, runtime measurements are noisy, and search often stalls early. We present a practical optimization agent that combines LLM-guided mutation, adaptive resource allocation, policy-gated evaluation, and profiler-informed diagnosis. The system screens many candidates with low-cost evaluation and allocates higher-fidelity budget only to promising survivors to optimize and evolve GPU kernels. On important Triton kernel generation tasks, this design improves both kernel quality and search efficiency, reaching 0.88x-6.72x speedup over PyTorch and up to 1.4x speedup over the state-of-the-art kernel agent, with up to 48% lower optimization time.
发表机构
- Microsoft(微软公司)
机构由 AI 辅助整理,请以论文原文为准。