arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AgenticCANN:基于知识增强的智能体演化的自动昇腾C算子生成

AgenticCANN: Automated Ascend C Operator Generation via Knowledge-Augmented Agentic Evolution

Junhao Qiu, Zidong Wang, Yansong Sun, Zhitong Ma, Ping Guo, Qingfu Zhang

arXiv 2607.26661首次发表:更新:

AI 中文总结

本研究针对昇腾C算子生成的独特挑战,提出AgenticCANN知识增强智能体演化框架,在昇腾910B实验中实现高可行性与加速效果,验证了知识注入的通用性优势。

AI 中文摘要

昇腾C算子优化对NPU(神经网络处理单元)推理性能至关重要,但需要深入的硬件知识。尽管大型语言模型(LLM)在自动CUDA内核生成方面展现出潜力,但昇腾C截然不同的编程模型带来了尚未被探索的独特挑战。本文提出AgenticCANN,这是一个专为低语料NPU场景下自动昇腾C算子合成定制的知识增强智能体演化框架。为克服陌生硬件上严重的平台知识缺口,AgenticCANN集成了知识编排生成系统,在开发全生命周期提供结构化、多层次的领域洞见,以解决上游可行性问题。在此基础上,它具备阶段自适应智能体演化策略,可动态调整LLM交互模式以适配特定生成与演化阶段,平衡高探索性候选发现与高收敛性性能。在华为昇腾910B上针对5类模式的6个算子开展的实验表明,本方法在元素级和归一化算子上实现90%至100%的可行性,融合算子上为56%,在1B盘古模型推理内核上实现最高6.65倍加速。进一步分析显示,知识注入使元素级算子的可行性从57%单调提升至86%,证明其具有通用性而非算子特定性的优势。

英文摘要

Ascend C operator optimization is critical for NPU (Neural Processing Unit) inference performance but requires deep hardware expertise. While large language models (LLMs) have shown promise in automated CUDA kernel generation, the fundamentally different programming model of Ascend C introduces unique challenges that remain unexplored. In this paper, we propose AgenticCANN, a knowledge-augmented agentic evolution framework specifically tailored for automated Ascend C operator synthesis in low-corpus NPU environments. To overcome the severe platform knowledge deficit on unfamiliar hardware, AgenticCANN incorporates a knowledge-orchestrated generation system that delivers structured, multi-level domain insights across the development lifecycle to resolve the upstream feasibility bottleneck. Building on this foundation, it features a stage-adaptive agentic evolution strategy that dynamically aligns LLM interaction modes with specific generation and evolution phases, balancing high-exploration candidate discovery with high-convergence performance tuning. Extensive experiments on Huawei Ascend 910B across six operators spanning five pattern categories demonstrate that our method achieves 90 to 100 percent feasibility on elementwise and normalization operators, 56% on fusion operators, and up to 6.65$\times$ speedup on 1B Pangu model inference kernels. Further analysis reveals that knowledge injection monotonically improves feasibility from 57% to 86% on elementwise operators, demonstrating its general rather than operator-specific benefit.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑