面向预填充-解码分离式AI推理的分析型功耗感知配置方法
Analytical Power-Aware Provisioning for Prefill-Decode Disaggregated AI Inference
浏览论文内容
中文总结 AI 辅助
针对PD分离式AI推理,提出分析框架建模服务容量与功耗,推导帕累托前沿,以支持功耗感知的实例配置决策。
中文摘要 AI 辅助
功耗可用性日益制约着AI推理集群的运行,因此需要一种能够同时考虑服务容量和功耗的配置方法。预填充-解码(PD)分离已成为大规模推理服务的主流架构。然而,确定合适的预填充和解码实例数量颇具挑战性,因为服务容量同时取决于工作负载特征、硬件约束、排队和KV缓存预留。现有方法主要依赖性能剖析和模拟,对配置决策如何影响服务容量与功耗之间的权衡提供的分析性洞察有限。本文开发了一个用于PD分离式AI推理的功耗感知配置分析框架。给定推理工作负载和硬件,该框架对已配置部署的服务容量和平均功耗进行建模。服务容量模型基于输入-输出长度的联合分布以及硬件计算和内存限制推导得出。特别地,它明确捕获了由KV缓存预留引起的预填充与解码之间的耦合,以及请求排队的影响。在此基础上,功耗模型将每实例功耗确定为归一化服务吞吐量的函数。综合起来,这些模型确定了候选配置部署之间的服务容量-功耗帕累托前沿,使服务提供商能够随着工作负载或可用功耗的变化选择合适的配置部署。
英文摘要
Power availability increasingly constrains the operation of AI inference fleets, creating a need for provisioning methods that jointly consider serving capacity and power consumption. Prefill--decode (PD) disaggregation has emerged as a prevalent architecture for large-scale inference serving. However, determining the appropriate numbers of prefill and decode instances is challenging because serving capacity depends jointly on workload characteristics, hardware constraints, queueing, and KV-cache reservations. Existing approaches largely rely on profiling and simulation, providing limited analytical insight into how provisioning decisions shape the tradeoff between serving capacity and power consumption. This paper develops an analytical framework for power-aware provisioning of PD-disaggregated AI inference. Given an inference workload and hardware, the framework models the serving capacity and average power consumption of a provisioned deployment. The serving-capacity model is derived from the joint distribution of input--output lengths and hardware compute and memory limits. In particular, it explicitly captures the coupling between prefill and decode induced by KV-cache reservations, as well as the impact of request queueing. On this basis, the power model determines per-instance power consumption as a function of normalized serving throughput. Together, the models determine the serving capacity--power Pareto front among candidate provisioned deployments, enabling the service provider to choose a provisioned deployment as the workload or available power changes.
发表机构
- New York University(纽约大学)
机构由 AI 辅助整理,请以论文原文为准。