AI 中文总结
针对硬件加速器内核代码生成瓶颈,MKEvolve框架通过迭代共同进化PyTorch模块分解与子模块内核,经拆分融合优化分解并以束搜索改进子内核,实验证明其在正确性、加速及减少LLM令牌使用上有优势。
AI 中文摘要
尽管基于大语言模型(LLM)的代码生成取得了快速进展,但为硬件加速器编写正确且高性能的内核仍然是扩展现代机器学习工作负载的关键瓶颈。我们提出了MKEvolve(模块化内核进化)框架,该框架迭代地共同进化复杂PyTorch模块的模块化分解以及每个子模块的LLM生成内核,通过跨迭代的拆分和融合来优化分解,同时通过LLM驱动的束搜索独立改进每个子内核。实验表明,MKEvolve在提高正确性和加速方面优于端到端直接合成基线,同时将LLM令牌使用量减少了35%。
英文摘要
Despite rapid progress in LLM-based code generation, writing correct and performant kernels for hardware accelerators remains a key bottleneck in scaling modern ML workloads. We present MKEvolve (Modular Kernel Evolve), a framework that iteratively co-evolves a modular decomposition of complex PyTorch modules and the LLM-generated kernel for each submodule, refining the decomposition by splitting and fusing across iterations while independently improving each subkernel via LLM-driven beam search. The resulting kernels are programmatic compositions of independently verified subkernels, making them configurable (subkernel implementations are swappable), interpretable (errors and speedups are traceable to specific subkernels), and readily adaptable to related model architectures. Experiments with Triton on KernelBench L2 and L3, spanning multi-operator sequences and full model architectures, show that MKEvolve improves both correctness and speedup over end-to-end direct synthesis baselines while reducing LLM token usage by up to 35%.