arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27085cs.DCcs.AIcs.LG

Crossflow:面向智能体LLM服务的预填充-解码弹性机制

Crossflow: Prefill-Decode Elasticity for Agentic LLM Serving

Yi Xu, Ehsan K. Ardestani, Wenyin Fu, Martin Schatz, Krishna Malladi, Zhan Shu, Adnan Aziz, Shobhit Kanaujia, Ajit Mathews, Chunqiang Tang

首次发表
浏览论文内容

中文总结 AI 辅助

针对智能体LLM服务中预填充-解码静态分区导致的容量浪费与排队问题,提出Crossflow弹性机制,通过可撤销租约动态调整边界,提升吞吐量16.2%-17.4%并降低TTFT。

中文摘要 AI 辅助

随着服务容量需求超过训练需求,服务效率变得越来越重要。预填充-解码(P/D)分离通过两个阶段的专业化和隔离来提高服务效率。这些优势依赖于静态分区。然而,阶段需求并非静态。我们观察到,在大型LLM集群中,未缓存输入与输出token的比率在分钟时间尺度上的峰值与均值之比高达4.7倍,而在公开的智能体轨迹中,小时比率在一天内的中位数跨度达24.5倍,而重新分配一个副本需要数十分钟。智能体流量加剧了这种不匹配。将每个池按其第95百分位设置大小,会导致高达17%的集群容量未被使用;而设置低于该值,则会将同样的不平衡转化为排队和未实现的吞吐量。我们提出了Crossflow,它在不改变节点角色的情况下使这一边界具有弹性。每个解码节点发布一个短期、可撤销的租约,该租约限制本地预填充计算、KV容量、传输工作和预计输出。在公开和内部轨迹中,与静态P/D相比,Crossflow在几何均值上将token吞吐量提高了16.2%-17.4%,在高负载下最高提高43.4%,同时在每个评估点降低了平均TTFT。

英文摘要

As serving capacity demand surpasses that of training, serving efficiency becomes increasingly important. Prefill-decode (P/D) disaggregation improves serving efficiency through specialization and isolation of the two phases. These benefits rest on a static partitioning. Phase demand, however, is not static. We observe that in a large LLM fleet the ratio of uncached input to output tokens has peak-to-mean ratios up to 4.7x at minute timescales, and that in a public agentic trace the hourly ratio spans a median 24.5x within a single day, while reassigning a replica takes tens of minutes. Agentic traffic sharpens the mismatch. Sizing each pool at its ninety-fifth percentile leaves up to 17% of cluster capacity unused; sizing below it converts the same imbalance into queueing and unrealized throughput. We present Crossflow, which makes this boundary elastic without changing node roles. Each decode node publishes a short-lived, revocable lease that bounds local-prefill compute, KV capacity, transfer work, and projected output. Across public and internal traces, Crossflow improves token throughput by 16.2-17.4% on geometric mean over static P/D, and by up to 43.4% at high load, while reducing mean TTFT at every evaluated point.

发表机构

  • Meta Platforms(元平台公司)

机构由 AI 辅助整理,请以论文原文为准。

↑