AI 中文总结
研究针对NPU服务LLMs时功耗问题,开发eNPU支持空间细粒度、组件级DVFS,通过重构流水线、引入通信机制、扩展ISA及采用编译器驱动搜索优化,实现能耗降低25.8% - 35.2%且保持SLO保证。
AI 中文摘要
随着神经处理单元(NPUs)快速发展以适应大语言模型(LLMs)不断增长的计算需求,其功耗成为限制因素。研究表明利用动态电压和频率缩放(DVFS)利用服务水平目标(SLO)松弛是提高NPU用于LLM服务能效的有效方法。由于LLMs中的张量运算符在NPU组件间存在不同瓶颈,为使能效最大化,希望为每个组件单独配置频率。本文开发了eNPU,它在硬件和软件上支持NPU上空间细粒度、组件级的DVFS。eNPU重构NPU核心流水线以将组件划分为单独的V/$f$域,引入轻量级跨域通信机制减轻组件间同步开销,并扩展NPU ISA用于亚微秒级DVFS控制。eNPU使用编译器驱动的两级贪婪搜索在SLO约束下共同优化指令调度和每个组件的V/$f$选择。通过在开源NPU核心上实现eNPU的流水线设计并使用生产级NPU模拟器和各种LLMs及生产跟踪评估节能情况。在TPUv4芯片上,eNPU将LLM服务的能耗降低25.8% - 35.2%,面积开销为3.45%,同时保持严格的SLO保证。
英文摘要
As neural processing units (NPUs) evolve rapidly to accommodate the ever-increasing compute demand of large language models (LLMs), their power consumption is becoming a limiting factor. Our study shows that using dynamic voltage and frequency scaling (DVFS) to exploit the service-level objective (SLO) slacks is a promising way to improve NPU energy efficiency for LLM services. And as tensor operators in LLMs exhibit diverse bottlenecks across NPU components, it is desirable to configure the frequency separately for each component to maximize their energy efficiency. In this paper, we develop eNPU that enables hardware and software support for spatially fine-grained, component-level DVFS on NPUs. eNPU refactors the NPU core pipeline to partition components into separate V/$f$ domains. It introduces lightweight cross-domain communication mechanisms to mitigate synchronization overheads across components, and extends the NPU ISA for sub-$μ$s DVFS control. eNPU uses a compiler-driven two-level greedy search to co-optimize instruction scheduling and per-component V/$f$ selection under SLO constraints. We implement eNPU's pipeline design on an open-source NPU core to verify its functionality and evaluate the energy savings with a production-level NPU simulator with various LLMs using production traces. eNPU reduces energy consumption of LLM services by 25.8%--35.2% with 3.45% area overhead on a TPUv4 chip, while preserving strict SLO guarantees.
CommentsAccepted by MICRO'26