基于编译器的大语言模型Triton内核优化分层诊断
Compiler-Grounded Hierarchical Diagnosis for LLM-Based Triton Kernel Optimization
浏览论文内容
中文总结 AI 辅助
研究针对大语言模型内核优化中编译器优化失败原因难揭示的问题,提出基于编译器的Triton内核分层优化框架,通过跨层诊断实现从模式分类到源级重写的优化,在Ascend NPU上显著加速Triton内核。
中文摘要 AI 辅助
大语言模型的进展实现了内核自动生成与优化,但现有方法多依赖编译反馈等表面信号,难以揭示后端编译器优化失败的原因。因此,本文将内核优化表述为渐进式跨层诊断问题,提出基于编译器的Triton内核分层优化框架。该系统从轻量级模式分类和性能分析诊断逐步升级到中间表示属性和基于编译器的分析,进而提出有依据的源级重写。在Ascend NPU上的实验表明,该系统能显著加速Triton内核,范围从接近基线到大幅提升。
英文摘要
Recent advances in large language models (LLMs) have enabled automated kernel generation and optimization, but most existing approaches rely on surface signals such as compilation feedback and profiling metrics. These signals reveal that a kernel is slow, but not why the backend compiler fails to realize a profitable optimization, especially on emerging accelerators such as NPUs. We therefore formulate kernel optimization as a progressive cross-layer diagnosis problem that links runtime symptoms to IR structure and compiler behavior before rewriting source. Based on this insight, we present our system, a compiler-grounded and hierarchical optimization framework for Triton kernels. the system escalates from lightweight pattern triage and profiling diagnosis to IR attribution and compiler-grounded analysis only when deeper evidence is needed, then proposes evidence-backed source-level rewrites. We implement the system on Triton for Ascend NPUs and evaluate it on 37 successfully converted entries from a standardized NPUKernelBench-derived Ascend 950 benchmark. Across these entries, the system attains a geometric-mean speedup of 4.35$\times$ and a median speedup of 2.73$\times$ from the initial to optimized Triton kernel; 22/37 exceed 2$\times$ and 13/37 exceed 5$\times$. The complete distribution ranges from near-baseline entries to large wins, motivating transparent reporting of the current system's scope and limitations.
发表机构
- Huawei Technologies Co., Ltd(华为技术有限公司)
机构由 AI 辅助整理,请以论文原文为准。