发表机构
Department of Electrical Engineering, Tsinghua University; ZJU-UIUC Institute, Zhejiang University; China Datang Technology Innovation Co., Ltd.(清华大学电气工程系; 浙江大学紫金港国际校区; 中国大唐科技创新有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对AI数据中心电力短缺问题,提出模型承诺(MC)框架,联合调度模型部署与跨站点路由,在电网约束下实现100%服务率并降低29.0%运营成本。
AI 中文摘要
AI数据中心在特定时期可能面临电力供应短缺,要求运营商在空间上迁移大语言模型(LLM)推理工作负载以维持服务率。然而,现有的工作负载迁移方法通常假设任何具有足够计算资源的数据中心都能立即服务迁移的请求,这可能导致不可行的转移和未服务的需求。本文提出模型承诺(MC),一种混合整数线性规划框架,在电力约束和电价信号下联合调度模型部署和跨站点请求路由。首先,MC公式化了模型副本加载引入的跨期耦合。其次,它将预填充和解码延迟要求转化为每个副本可服务的需求数量。基于真实数据的案例研究表明,MC使AI数据中心运营商在时变电网条件下实现100%的服务率,并将总运营成本降低29.0%。
英文摘要
AI data centers may face power supply shortages during certain periods, requiring operators to shift large language model (LLM) inference workloads spatially to maintain service rates. However, existing workload-shifting methods typically assume that any data center with sufficient computing resources can immediately serve shifted requests, which may lead to infeasible transfers and unserved demand. This letter proposes model commitment (MC), a mixed-integer linear programming framework that jointly schedules model deployment and cross-site request routing under power constraints and electricity-price signals. First, MC formulates the intertemporal coupling introduced by model replica loading. Second, it translates prefill and decode latency requirements into the amount of demand that each replica can serve. Case studies based on real-world data show that MC enables AI data center operators to achieve a 100% service rate under time-varying grid conditions and reduce total operating cost by 29.0%.
Comments3 pages, 2 figures