arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PowerSlider:利用相位不对称性实现需求响应下的大语言模型服务

PowerSlider: Exploiting Phase Asymmetry for LLM Serving under Demand Response

Yueying Li, Jiayang Chen, Yuanfan Chen, Leo Han, Haoran Qiu, Esha Choukse, Rodrigo Fonseca, Udit Gupta

arXiv 2608.21719首次发表:更新:

发表机构

Cornell University; Microsoft Azure Research; Cornell Tech(康奈尔大学; 微软Azure研究院; 康奈尔科技学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PowerSlider通过分解大语言模型服务阶段并结合KKT在线求解器,在电网需求响应的时变功率上限下,显著提升了服务吞吐量与延迟性能。

AI 中文摘要

AI推理集群日益受到瞬时功率而非仅能量的约束:电网运营商将新容量与需求响应挂钩,施加随时间变化的功率上限。现有大语言模型服务系统要么优化静态能量目标,要么在负载下丢弃固定优先级层级;无论哪种方式,当功率包络变化时,吞吐量都会崩溃。大语言模型流水线的负载并非均匀:计算密集型的预填充阶段吞吐量随GPU频率几乎线性下降,内存密集型的答案解码阶段在频率降至额定值的0.57倍时仍能维持吞吐量,而推理的思考阶段将KV缓存容量与调度耦合——因此,功率上限应被引导至每瓦性能损失最小的地方。PowerSlider(\textbackslash sys{})通过以下方式实现这一目标:一是新的灵活服务水平目标(Flex SLO)合约,将有界用户松弛转化为优化约束;二是预填充-思考-答案的分解,暴露各阶段的频率和KV控制;三是卡尔曼-库恩-塔克(KKT)在线求解器,每次功率上限变化时在7.7毫秒内重新求解,并有统一的故障安全机制,当动态电压频率调整(DVFS)触及静态功率下限时分段关闭已耗尽的实例。在SGLang框架上使用生产轨迹测试,与五个基线中最佳的47.6%相比,PowerSlider在功率上限降低30%时维持了78.3%的在线吞吐量(提升1.64倍);将延迟关键尾部保持在额定值的1.3倍以内(基线为2.3至6倍,最高达12倍);在重放的加州独立系统运营商(CAISO)电网紧急日中,平均吞吐量达到92%,最低至额定值的0.41倍(低谷时为54%,所有基线均低于7%)。

英文摘要

AI inference clusters are increasingly constrained by instantaneous power, not just energy: grid operators condition new capacity on demand response, imposing time-varying power caps. Existing LLM serving systems optimize a static energy objective or shed fixed priority tiers under load; either way, goodput collapses when the power envelope moves. An LLM pipeline is not a uniform load: compute-bound prefill loses throughput almost linearly with GPU frequency, memory-bound answer decode sustains it down to $0.57\times$ nominal, and reasoning's thinking phase couples KV-cache capacity to scheduling -- so a cap should be steered to where each watt costs the least performance. PowerSlider does so with a new Flex SLO contract that turns bounded user slack into an optimization constraint, prefill--think--answer disaggregation exposing per-stage frequency and KV control, and a Karush--Kuhn--Tucker (KKT) online solver re-solving within 7.7 ms of every cap change, backed by a consolidated fail-safe that power-gates drained instances when DVFS bottoms out on static power. On SGLang with production traces, \sys{} sustains 78.3\% online goodput at a 30\% cap reduction versus 47.6\% for the best of five baselines ($1.64\times$), holds latency-critical tails within $1.3\times$ of nominal (baselines: $2.3$--$6\times$, up to $12\times$), and delivers 92\% mean goodput through a replayed CAISO grid-emergency day bottoming at $0.41\times$ (54\% at the trough; every baseline below 7\%).

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑