发表机构
Huawei(华为)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对人工智能代理生成的算子内核在Ascend NPU上缺乏通用评估基线的问题,提出CANN Bench基准测试,涵盖多算子和测试用例,采用三维加权综合评分,为Ascend生态系统提供评估算子编写能力的标准。
AI 中文摘要
人工智能代理现在能够在不同硬件平台上编写、编译和迭代优化低级算子内核。然而,现有基准测试几乎只关注CUDA和Triton,使得编程模型较少暴露的硬件生态系统没有通用评估基线。我们提出了CANN Bench,这是一个针对华为Ascend NPU上人工智能生成的算子代码的开放基准测试。当前版本涵盖53个算子和1060个测试用例,分为四个难度级别。评估采用三维加权综合评分,将编译、功能正确性和性能作为独立轴,为内核生成代理提供有原则的奖励信号。性能根据开箱即用的Ascend上的PyTorch基线和真实NPU硬件上的逐例分析硬件锚定性能(HAP)限制进行分级。评估工具旨在从根本上抵制奖励欺诈。CANN Bench在官方CANN存储库中进行版本控制,旨在实现长期社区共建,为Ascend生态系统提供一个用于人工智能算子编写能力的定量、可重复且可持续维护的标准。
英文摘要
AI agents are now capable of writing, compiling, and iteratively optimizing low-level operator kernels on different hardware platforms. Existing benchmarks, however, focus almost exclusively on CUDA and Triton, leaving hardware ecosystems with less-exposed programming models without a common evaluation baseline. We present CANN Bench, an open benchmark for AI-generated operator code on Huawei's Ascend NPU. The current release covers 53 operators and 1060 test cases organized into four difficulty tiers -- from simple elementwise primitives to MoE dispatch and FlashAttention kernels -- spanning FP16, BF16, FP32, and INT8 precision formats. Evaluation adopts a \textbf{three-dimensional weighted composite score} that treats compilation, functional correctness, and performance as independent axes, providing a principled reward signal for kernel-generation agents. Performance is graded against an out-of-the-box PyTorch-on-Ascend baseline and an analytical per-case Hardware-Anchored Performance (HAP) limit on real NPU hardware, ensuring scores reflect genuine optimization headroom rather than measurement artifacts. The evaluation harness is designed to resist reward hacking from the ground up. CANN Bench is versioned within the official CANN repository and is designed for long-term community co-construction, providing the Ascend ecosystem with a quantitative, reproducible, and sustainably maintained yardstick for AI operator-authoring capability.