从电网到芯片:AI数据中心电力架构、稳定性与灵活性
From Grid to Chip: Power Architecture, Stability, and Flexibility of AI Data Centers
- Aalborg University(奥尔堡大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文从技术视角审视AI数据中心,分析并网瓶颈与政策,梳理电力架构从电网到芯片的演进,提出涵盖机架、设施和系统三层的稳定性框架,并强调电网到芯片协同设计对可扩展AI基础设施的关键作用。
AI中文摘要:
人工智能(AI)计算的快速增长正在将数据中心转变为大规模、动态的电力负荷。其部署主要受能源可用性和并网容量的制约,而电力传输架构、控制系统和计算工作负载在快速电网扰动期间可靠运行的能力进一步加剧了这一制约。本文从技术视角探讨了作为电网交互计算系统的AI数据中心。首先,文章回顾了并网瓶颈、不断演变的并网政策和电网规范要求,这些因素通过工作负载编排、冷却系统、现场资源和储能所提供的时空灵活性,催生了新的技术趋势。其次,文章梳理了电力传输架构从中压电网接口到芯片级的演进过程,讨论了更高电压的直流配电、固态变压器、宽禁带器件、先进的芯片级电力传输以及液冷技术。第三,文章建立了一个涵盖机架级直流母线动态、设施级变流器交互和系统级电网耦合行为的三层稳定性框架。该框架将主导性失稳机制(包括恒功率负荷效应、阻抗交互、强迫振荡和运行模式切换)与相应的建模、评估和缓解方法联系起来。综合这些主题,本文强调电网到芯片的协同设计是构建可扩展AI基础设施的核心要求,将计算工作负载、电力传输系统、能量缓冲和电网运行紧密关联。
英文摘要:
The rapid growth of artificial intelligence (AI) computing is transforming data centers into large, dynamic electrical loads. Their deployment is primarily constrained by energy availability and grid-connection capacity, which is further aggravated by the ability of power-delivery architectures, control systems, and computing workloads to operate reliably during fast grid disturbances. This article presents a technological perspective on AI data centers as grid-interactive computing systems. First, it reviews grid-integration bottlenecks, evolving connection policies, grid-code requirements, which has fostered new technological trends via spatio-temporal flexibility available through workload orchestration, cooling systems, on-site resources, and energy storage. Second, it maps the evolution of power-delivery architectures from medium-voltage grid interfaces to chip-level, discussing higher-voltage DC distribution, solid-state transformers, wide-bandgap devices, advanced chip-level power delivery, and liquid cooling. Third, it establishes a three-level stability framework spanning rack-level DC-bus dynamics, facility-level converter interactions, and system-level grid-coupled behavior. The framework connects dominant instability mechanisms, including constant power load effects, impedance interactions, forced oscillations, and operating-mode transitions, with suitable modeling, assessment, and mitigation approaches. Synthesizing these topics, this article highlights grid-to-chip co-design as a central requirement for scalable AI infrastructure, linking computing workloads, power-delivery systems, energy buffers, and grid operation.