用于高效能和低成本的大语言模型服务的异构神经处理单元自动缩放
Auto-Scaling Heterogeneous Neural Processing Units for Energy and Cost-Efficient LLM Serving
AI总结:
研究如何利用异构神经处理单元为大语言模型服务实现能源和成本高效。提出NeuScale自动缩放框架,通过新vPod抽象管理资源,基于屋顶线分析分配,支持动态供应。经模拟器验证,可显著提升成本效率和SLO满意度。
AI中文摘要:
为满足大语言模型(LLM)服务不断增长的计算需求,现代云平台广泛部署了神经处理单元(NPUs)。NPU芯片快速发展导致计算池异构,而云环境缺乏管理NPU异构性的系统和架构支持。本文先对各代真实NPU芯片进行特征研究,证明利用异构NPU芯片的能效、成本效益和性能优势。接着提出NeuScale自动缩放框架,用新vPod抽象管理资源,通过基于直观轻量级屋顶线分析进行最佳vPod分配,支持细粒度动态资源供应。通过生产级NPU模拟器验证,NeuScale能通过充分利用异构NPU资源显著提高成本效率和服务水平目标(SLO)满意度。
英文摘要:
To meet the ever-increasing computing demands of large language model (LLM) services, modern cloud platforms have widely deployed neural processing units (NPUs). These NPU chips have been developed and evolved at an incredibly fast pace, this inevitably produces heterogeneous compute pools backed by different versions of NPU chips. Unfortunately, due to the lack of system and architecture support for managing NPU heterogeneity in the cloud, it is unclear how to best utilize heterogeneous NPUs to maximize the energy and cost efficiency for LLM services. In this paper, we first conduct a characterization study of various generations of real NPU chips to demonstrate the potential benefits on energy/cost efficiency and performance by utilizing heterogeneous NPU chips. To realize these benefits, we present NeuScale, an auto-scaling framework to automatically exploit heterogeneous NPUs for cloud platforms. NeuScale manages heterogeneous NPU resources with a new vPod abstraction, which abstracts the core hardware parameters of different NPU versions and provides compatibility with existing ML frameworks. It makes the best-fit vPod allocations for different LLM inference requests using an intuitive and lightweight roofline-based analysis. It supports fine-grained dynamic NPU resource provisioning by adjusting both the vPod configuration (i.e., scaling up/down) and the number of vPods (e.g., scaling in/out). To validate the benefits of NeuScale at scale, we implement it with a production-level NPU simulator. Our evaluation with popular LLMs shows that NeuScale can significantly improve cost efficiency and service-level objective (SLO) satisfaction rate by best utilizing heterogeneous NPU resources.