arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从经验到专长:面向数据稀缺NPU内核合成的采纳感知记忆学习

From Experience to Expertise: Adoption-Aware Memory Learning for Data-Scarce NPU Kernel Synthesis

Longxiao Fan, Tao Zhang, Han Yan, Jiajun Li, Mingcong Song, Guoping Long, Hongjie Si, Weiwei Sun

arXiv 2609.35568首次发表:更新:

发表机构

Fudan University; University of Science and Technology of China; Huawei Technologies Ltd.(复旦大学; 中国科学技术大学; 华为技术有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出SAGE智能体,通过采纳感知信用分配和选择性整合,在数据稀缺的NPU内核合成中实现95.5%执行率,显著超越基线并加速稀疏注意力43.99倍。

AI 中文摘要

高性能内核是高效加速器执行的基础,但需要专家调优和冗长的手动优化周期。LLM编码智能体有望实现自动化,但其CUDA知识难以迁移到数据稀缺的领域专用架构(DSA),如NPU,这些架构的执行模型和内存层次结构与GPU差异显著。为解决这一迁移差距,后训练方法使LLM适应NPU编程,但依赖稀缺的专家数据和大量训练计算。记忆学习智能体则通过外部记忆进行适应,但其统一的信用分配使得被采纳和未使用的经验具有相同的奖励目标,可能使后续检索排名产生偏差。此外,当学习到的值仅用于指导检索时,跨算子泛化的高价值经验必须被反复检索而非保留在上下文中,从而增加检索开销并削弱跨任务指导。因此,我们提出SAGE,一种用于NPU内核合成的持久自改进智能体。采纳追踪效用估计(ATU)将显式采纳记录与内核评估结果相结合,实现采纳感知的信用分配。效用门控整合(UGC)利用跨算子的正效用和重复采纳,选择并抽象可复用规则至有界常驻上下文。在NPUKernelBench上,SAGE实现了95.5%的执行率,而最强受控基线为84.1%,且86.9%的已解决算子优于torch_npu。使用GLM-5.3,SAGE在稀疏闪存注意力上相比torch_npu参考实现了43.99倍加速。这些结果表明,采纳感知的信用分配和选择性整合使智能体能够跨任务积累和复用硬件特定知识。

英文摘要

High-performance kernels underpin efficient accelerator execution but require expert tuning and lengthy manual optimization cycles. LLM coding agents promise automation, yet their CUDA knowledge transfers poorly to data-scarce domain-specific architectures (DSAs) such as NPUs, whose execution models and memory hierarchies differ substantially from those of GPUs. To address this transfer gap, post-training methods adapt LLMs to NPU programming but depend on scarce expert data and substantial training compute. Memory-learning agents instead adapt through external memory, but their uniform credit assignment gives adopted and unused experiences the same reward target, potentially biasing subsequent retrieval rankings. Moreover, when learned values guide only retrieval, high-value experiences that generalize across operators must be retrieved repeatedly rather than retained in context, thereby increasing retrieval overhead and weakening cross-task guidance. We therefore present SAGE, a persistent self-improving agent for NPU kernel synthesis. Adoption-Traced Utility estimation (ATU) combines explicit adoption records with kernel evaluation outcomes for adoption-aware credit assignment. Utility-Gated Consolidation (UGC) uses positive utility and repeated adoption across operators to select and abstract reusable rules into a bounded resident context. On NPUKernelBench, SAGE achieves a 95.5% execution rate versus 84.1% for the strongest controlled baseline, with 86.9% of solved operators outperforming torch_npu. With GLM-5.3, SAGE achieves a 43.99x speedup over the torch_npu reference on sparse flash attention. These results show that adoption-aware credit assignment and selective consolidation enable agents to accumulate and reuse hardware-specific knowledge across tasks.

Comments30 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑