arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FluidPD:面向SLO感知的预填充-解码分离LLM服务的原位弹性

FluidPD: In-Place Elasticity for SLO-Aware Prefill-Decode Disaggregated LLM Serving

Kartik Ramesh, Kaidi Fu, Zihan Zheng, Jiahuan Yu, Fabio Oliveira, Carlos Costa, Minjia Zhang

arXiv 2610.06917首次发表:更新:

发表机构

University of Illinois Urbana-Champaign; IBM Research(伊利诺伊大学厄巴纳-香槟分校; IBM研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FluidPD通过原位弹性机制(FluidToken和FluidRole)动态调整预填充与解码工作节点角色,利用压力指数提前感知资源压力,在Azure生产负载上显著提升SLO达成率。

AI 中文摘要

预填充-解码分离正成为LLM服务的常见架构,因为它将具有不同执行模式和SLO目标的两个阶段分开。现有系统通常将固定的预填充/解码工作节点比例与跨工作节点的请求路由相结合。然而,现实工作负载在预填充与解码需求比例上既表现出短时突发,也表现出持续偏移。因此,一个在某一时刻配置良好的系统可能很快变得不匹配,即使在其他地方存在空闲容量,也会导致延迟SLO违规。现有的自动扩展机制可以增加容量,但它们响应缓慢,需要备用GPU,且不能直接解决短时间尺度的阶段不平衡。我们提出FluidPD,一个提供SLO感知原位弹性的P/D分离服务系统。FluidPD引入两种互补机制。FluidToken通过当解码侧有可用余量时将有限部分的预填充计算卸载到解码工作节点来处理瞬时不平衡。FluidRole通过原位重新分配运行中的工作节点在预填充和解码角色之间来处理持续不平衡,避免模型重载和引擎重启。两种机制均由轻量级压力指数引导,这些指数在SLO违规出现之前暴露预填充和解码侧的资源压力。在Azure生产工作负载轨迹上,FluidPD相比静态SGLang将整体SLO达成率提高了最多94.6个百分点,表明SLO感知的原位P/D弹性无需额外配置工作节点即可提升服务质量。

英文摘要

Prefill-decode disaggregation is becoming a common architecture for LLM serving because it separates two phases with distinct execution patterns and SLO objectives. Existing systems typically combine a fixed prefill/decode worker ratio with request routing across workers. However, real-world workloads exhibit both short bursts and sustained shifts in the prefill-to-decode demand ratio. As a result, a configuration that is well provisioned at one time may quickly become mismatched, causing latency SLO violations even when idle capacity exists elsewhere. Existing autoscaling mechanisms can add capacity, but they react slowly, require spare GPUs, and do not directly address short-timescale phase imbalance. We present FluidPD, a P/D-disaggregated serving system that provides SLO-aware in-place elasticity. FluidPD introduces two complementary mechanisms. FluidToken handles transient imbalance by offloading a bounded portion of prefill computation to decode workers when decode-side slack is available. FluidRole handles sustained imbalance by reassigning running workers between prefill and decode roles in place, avoiding model reload and engine restart. Both mechanisms are guided by lightweight pressure indices that expose prefill and decode-side resource pressure before they appear as SLO violations. Across production Azure trace workloads, FluidPD improves overall SLO attainment over static SGLang by up to 94.6 percentage points, demonstrating that SLO-aware in-place P/D elasticity improves service quality without provisioning additional workers.

Comments13 pages, 11 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑