发表机构
Covision Lab(Covision实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本白皮书系统评估十款边缘 AI 推理加速器(ASIC NPU、SoC DSP、集成 NPU)对比 NVIDIA RTX A5000 基准,分析吞吐量、延迟、能效等指标,揭示 NPU 在功耗和性能上的优势。
AI 中文摘要
在生产环境中,AI 推理正成为企业 AI 领域的主导成本线。AI 推理市场预计将从 2024 年的 870 亿美元增长至 2032 年的 3490 亿美元(年复合增长率 18.9%)。神经处理单元(NPU)——专为 AI 推理而设计的芯片——正成为纯 GPU 架构的有力替代方案,在相同吞吐量下功耗降低 35-70%。本白皮书系统评估了三个硬件类别中的十款边缘 AI 推理加速器:ASIC NPU(Hailo-8、Hailo-10H、Axelera Metis、Axelera Europa、EdgeCortix Sakura II)、SoC DSP(SiMa MLSoC、Qualcomm QCS6490、QCS8550)以及集成 NPU(Intel Lunar Lake、AMD XDNA2),并以配备 TensorRT 的 NVIDIA RTX A5000 作为生产级基准进行对比。使用涵盖卷积、移动和 Transformer 架构的十二个参考模型作为一致的基准测试套件。分析结果涵盖吞吐量、延迟、模型兼容性、能效、SDK 成熟度和产品生命周期。
英文摘要
AI inference in production settings is becoming the dominant cost line in enterprise AI. The AI inference market is projected to grow from $87B in 2024 to $349B by 2032 (18.9% CAGR). Neural Processing Units (NPUs), chips built specifically for AI inference, are emerging as a compelling alternative to GPU-only architectures, with 35-70% lower power consumption at comparable throughput. This white paper systematically evaluates ten edge AI inference accelerators across three hardware categories: ASIC NPUs (Hailo-8, Hailo-10H, Axelera Metis, Axelera Europa, EdgeCortix Sakura II), SoC DSPs (SiMa MLSoC, Qualcomm QCS6490, QCS8550), and integrated NPUs (Intel Lunar Lake, AMD XDNA2), benchmarked against an NVIDIA RTX A5000 with TensorRT as a production-grade baseline. Twelve reference models spanning convolutional, mobile and transformer architectures are used as a consistent benchmark suite. Results are analysed for throughput, latency, model compatibility, power efficiency, SDK maturity and product lifecycle.