面向分离式LLM服务的相位解耦、模型校准功耗控制
Phase-Decoupled, Model-Calibrated Power Control for Disaggregated LLM Serving
浏览论文内容
中文总结 AI 辅助
针对分离式LLM服务,提出相位解耦、模型校准的功耗控制方法,通过延迟门控校准分别优化预填充和解码通道,在MoE模型上实现帕累托改进的能效提升。
中文摘要 AI 辅助
数据中心GPU功耗是限制LLM服务容量的关键约束,生产环境中的服务已转向预填充/解码(PD)分离架构。在分离式B200系统上部署NVIDIA的Max-Q推理配置文件后,我们发现其实际收益有限(+8.6% tokens/J),且依赖于模型,并带来平均端到端延迟成本(+5.2%),这是仅关注吞吐量的评估所无法揭示的;该配置文件还对运行在相反硬件状态下的预填充和解码GPU应用同一设置。我们假设最优功耗设置是所部署的(模型、量化、引擎、硬件)组合的属性,而非GPU类别的属性;每条通道应拥有各自的配置文件;将SLO余量安全地转化为能耗节省需要运行时SLO保护下的延迟门控校准,而非固定配方。我们提出一种相位解耦、模型校准的控制器:预填充通道在SM时钟窗口下运行,其下限通过构造保证延迟;解码通道在自动校准设定的功耗上限下运行,该上限恰好位于实测吞吐量/延迟悬崖之上。由于分离式解码通道消耗平稳的、受内存限制的功耗,功耗上限持续生效,导致POLCA拒绝功耗上限的响应过冲弱点不存在,且GPU自身的功耗管理器在功耗上限下保持吞吐量。在8x B200节点上,在智能体负载下服务Qwen3-Coder-480B(FP8)时,我们的平衡模式在平均端到端延迟+3.5%的情况下实现+20.4% tokens/J,而Max-Q为+8.6% tokens/J和+5.2%延迟,在两个方面均实现帕累托改进。在Qwen3-235B-A22B(NVFP4)上,每种运行模式在每次重复中均满足ITL-p99 SLO;两个供应商配置文件均未满足。解码执行器A/B测试表明,校准的功耗上限优于静态时钟锁定,持续三天的运行节省了通道对32.3%的电能。两个模型均为MoE;稠密模型恢复的收益约少5倍,因此我们将主张限定于MoE服务。
英文摘要
Datacenter GPU power is the binding constraint on LLM serving capacity, and production serving has shifted to prefill/decode (PD) disaggregation. Deploying NVIDIA's Max-Q inference profile on a disaggregated B200 system, we found its realized gain modest (+8.6% tokens/J), model-dependent, and carrying a mean end-to-end latency cost (+5.2%) that throughput-only evaluation does not surface; the profile also applies one setting to prefill and decode GPUs that operate in opposite hardware regimes. We hypothesize that the optimal power setting is a property of the deployed (model, quantization, engine, hardware) combination rather than of the GPU class, that each lane warrants its own profile, and that converting SLO headroom into energy safely requires latency-gated calibration under a runtime SLO guard rather than a fixed recipe. We present a phase-decoupled, model-calibrated controller: the prefill lane runs under an SM-clock window whose floor is a latency guarantee by construction, and the decode lane under a power cap placed by automatic calibration just above a measured throughput/latency cliff. Because a disaggregated decode lane draws flat, memory-bound power, the cap binds continuously, the reactive-overshoot weakness that led POLCA to reject capping is absent, and the GPU's own power manager retains throughput under the cap. On an 8x B200 node serving Qwen3-Coder-480B (FP8) under agentic load, our balanced mode delivers +20.4% tokens/J at +3.5% mean e2e versus +8.6% at +5.2% for Max-Q, a Pareto improvement on both axes. On Qwen3-235B-A22B (NVFP4) every operating mode meets the ITL-p99 SLO in every repetition; both vendor profiles miss it. A decode-actuator A/B shows the calibrated cap beats static clock locks, and a three-day sustained run saves 32.3% of a lane pair's electricity. Both models are MoE; a dense model recovers roughly 5x less, so we scope our claims to MoE serving.
发表机构
- Xenoscube, Inc.(Xenoscube公司)
机构由 AI 辅助整理,请以论文原文为准。