FlowPrefill: Decoupling Preemption from Prefill Scheduling Granularity to Mitigate Head-of-Line Blocking in LLM Serving
FlowPrefill: 解耦预emption与prefill调度粒度以缓解LLM服务中的头部阻塞
机构 * Tsinghua University(清华大学) ; University of Science and Technology Beijing(北京科技大学)
专题命中 效率与部署 :LLM(title);large language model(abstract);language model(abstract);分类 cs.AI
AI总结 FlowPrefill通过解耦预emption与调度粒度,优化TTFT和吞吐量,缓解LLM服务中的头部阻塞问题。
Comments 13 pages