发表机构
National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ECO提出能量导向配置优化,联合搜索部署结构与运行控制,在有限测量预算下用高斯过程与约束贝叶斯优化,显著降低LLM服务GPU能耗并满足SLO。
AI 中文摘要
节能的LLM服务要求在满足延迟和吞吐量服务级别目标(SLO)的同时,最小化服务GPU的能量消耗。注意力-FFN分离(AFD)允许对注意力和专家计算进行独立的资源分配和运行控制,但它们的能量效应在执行流水线中仍然相互耦合。因此,实现其节能潜力需要探索一个分层配置空间,其中部署结构约束了可允许的控制并塑造了它们的端到端效应。找到满足SLO的低能耗配置具有挑战性,因为物理评估成本高昂,且只能测量一小部分候选配置。我们提出了能量导向配置优化(ECO),它在有限的测量预算下联合搜索部署结构及其可允许的运行控制。ECO从校准的阶段行为和流水线依赖中构建结构感知的能量先验,然后使用高斯过程学习残差预测误差。其成本感知的约束贝叶斯优化根据预期能量改进对测量进行优先级排序,同时考虑SLO可行性、执行成功和评估成本,并返回测量到的能耗最低的可行配置。在A6000和A100上使用Qwen和DeepSeek的全部16个场景中,ECO的冻结配置在不重叠的请求上评估,相对于基线平均降低服务能量40.5%,并提高输出令牌率20.7%,同时满足目标SLO。在8个A6000场景中,其选择的可行能量平均比通用约束贝叶斯优化低33.1%,比遗传搜索低25.8%。
英文摘要
Energy-efficient LLM serving requires minimizing serving GPU energy while meeting latency and throughput service-level objectives (SLOs). Attention--FFN disaggregation (AFD) enables separate resource allocation and operating controls for attention and expert computation, but their energy effects remain coupled through the execution pipeline. Realizing its energy-saving potential therefore requires navigating a hierarchical configuration space in which deployment structures constrain admissible controls and shape their end-to-end effects. Finding low-energy configurations that meet SLOs is challenging because physical evaluations are costly and only a small fraction of candidates can be measured. We present Energy-Oriented Configuration Optimization (ECO), which jointly searches deployment structures and their admissible operating controls under a limited measurement budget. ECO constructs a structure-aware energy prior from calibrated stage behavior and pipeline dependencies, then learns residual prediction errors with a Gaussian process. Its cost-aware constrained Bayesian optimization prioritizes measurements according to expected energy improvement while accounting for SLO feasibility, execution success, and evaluation cost, and returns the lowest-energy measured feasible configuration. Across all 16 scenarios on A6000 and A100 with Qwen and DeepSeek, ECO's frozen configurations, evaluated on disjoint requests, reduce serving energy by 40.5\% and increase output token rate by 20.7\% on average relative to baselines while meeting target SLOs. Across the 8 A6000 scenarios, its selected feasible energy averages 33.1\% below generic constrained Bayesian optimization and 25.8\% below genetic search.