arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12915cs.DCcs.AI

InFactPlanner:可持续地理分布式大语言模型(LLM)数据中心规划

InFactPlanner: Planning Sustainable Geo-Distributed LLM Data Centers

  • University of Cyprus(塞浦路斯大学)

机构由 AI 辅助整理,请以论文原文为准。

Nicoletta Tsiopani, Moysis Symeonides, George Pallis, Marios D. Dikaiakos

AI总结:

InFactPlanner是用于LLM推理的可持续地理分布式数据中心部署假设分析的决策支持框架,可通过多维度建模分析得出可持续性与延迟最优选择存在差异等结论。

AI中文摘要:

大语言模型(LLM)推理的快速发展正将可持续性关注点从一次性训练转向持续服务阶段,在此阶段,基础设施决策会影响能源使用、碳排放、水资源消耗和服务质量。然而,运营商通常需要在大规模基础设施建成前比较部署方案,直接测量成本高、速度慢且有时不可行。我们提出InFactPlanner,这是一个由 traces 驱动的决策支持框架,用于对单站点及地理分布式站点的LLM推理进行可持续AI数据中心部署的假设分析。InFactPlanner结合查询 traces、硬件模型配置文件、候选站点配置、PUE(电源使用效率)/WUE(水资源使用效率)参数、可再生能源发电模型以及随时间变化的电网碳强度,以估算功率、能源、碳排放、用水量、延迟和服务器利用率。该框架将底层服务效应抽象为可配置的硬件模型配置文件,支持快速比较站点选择、容量部署、硬件、模型、可再生能源整合及路由选择。我们通过重现参考LLM推理能源估算(偏差小于10%)验证了能源核算流程,评估了其在多个数据中心和服务器数量下的可扩展性,并展示了针对硬件选择、可再生能源部署、地理部署和碳感知路由的场景驱动决策分析。我们的结果表明,可持续性最优选择可能与延迟最优选择不同,且部署的碳价值在很大程度上取决于当地电网结构。

英文摘要:

The rapid growth of LLM inference is shifting sustainability concerns from one-time training to continuous serving, where infrastructure decisions shape energy use, carbon emissions, water consumption, and service quality. Yet operators often need to compare deployment alternatives before large-scale infrastructure is built, making direct measurement costly, slow, and sometimes infeasible. We present InFactPlanner, a trace-driven decision-support framework for what-if analysis of sustainable AI data center deployment for LLM inference across single and geo-distributed sites. InFactPlanner combines query traces, hardware-model profiles, candidate site configurations, PUE/WUE parameters, renewable generation models, and time-varying grid carbon intensity to estimate power, energy, carbon emissions, water use, latency, and server utilization. The framework abstracts low-level serving effects into configurable hardware-model profiles, enabling rapid comparison of site selection, capacity placement, hardware, model, renewable integration, and routing choices. We validate the energy accounting pipeline by reproducing reference LLM inference energy estimates with less than 10% deviation, evaluate scalability across multiple data centers and server counts, and demonstrate scenario-driven decision analyses for hardware selection, renewable placement, geographic deployment, and carbon-aware routing. Our results show that sustainability-optimal choices can differ from latency-optimal ones, and that the carbon value of deployment depends strongly on the local grid mix.

补充信息

↑