Crossflow:面向智能体LLM服务的预填充-解码弹性机制
Crossflow: Prefill-Decode Elasticity for Agentic LLM Serving
浏览论文内容
中文总结 AI 辅助
针对智能体LLM服务中预填充-解码静态分区导致的容量浪费与排队问题,提出Crossflow弹性机制,通过可撤销租约动态调整边界,提升吞吐量16.2%-17.4%并降低TTFT。
中文摘要 AI 辅助
随着服务容量需求超过训练需求,服务效率变得越来越重要。预填充-解码(P/D)分离通过两个阶段的专业化和隔离来提高服务效率。这些优势依赖于静态分区。然而,阶段需求并非静态。我们观察到,在大型LLM集群中,未缓存输入与输出token的比率在分钟时间尺度上的峰值与均值之比高达4.7倍,而在公开的智能体轨迹中,小时比率在一天内的中位数跨度达24.5倍,而重新分配一个副本需要数十分钟。智能体流量加剧了这种不匹配。将每个池按其第95百分位设置大小,会导致高达17%的集群容量未被使用;而设置低于该值,则会将同样的不平衡转化为排队和未实现的吞吐量。我们提出了Crossflow,它在不改变节点角色的情况下使这一边界具有弹性。每个解码节点发布一个短期、可撤销的租约,该租约限制本地预填充计算、KV容量、传输工作和预计输出。在公开和内部轨迹中,与静态P/D相比,Crossflow在几何均值上将token吞吐量提高了16.2%-17.4%,在高负载下最高提高43.4%,同时在每个评估点降低了平均TTFT。
英文摘要
As serving capacity demand surpasses that of training, serving efficiency becomes increasingly important. Prefill-decode (P/D) disaggregation improves serving efficiency through specialization and isolation of the two phases. These benefits rest on a static partitioning. Phase demand, however, is not static. We observe that in a large LLM fleet the ratio of uncached input to output tokens has peak-to-mean ratios up to 4.7x at minute timescales, and that in a public agentic trace the hourly ratio spans a median 24.5x within a single day, while reassigning a replica takes tens of minutes. Agentic traffic sharpens the mismatch. Sizing each pool at its ninety-fifth percentile leaves up to 17% of cluster capacity unused; sizing below it converts the same imbalance into queueing and unrealized throughput. We present Crossflow, which makes this boundary elastic without changing node roles. Each decode node publishes a short-lived, revocable lease that bounds local-prefill compute, KV capacity, transfer work, and projected output. Across public and internal traces, Crossflow improves token throughput by 16.2-17.4% on geometric mean over static P/D, and by up to 43.4% at high load, while reducing mean TTFT at every evaluated point.
发表机构
- Meta Platforms(元平台公司)
机构由 AI 辅助整理,请以论文原文为准。