arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向高效大语言模型服务的截止时间感知自适应预填充分块

Deadline-Aware Adaptive Prefill Chunking for Efficient Large Language Model Serving

Siyu Song, Qi Bai, Jinbo Hao, Kai Li, Chenchen Wang, Jiayu Sun

arXiv 2609.07883首次发表:更新:

发表机构

School of Computer Science and Technology, Beijing Institute of Technology; School of Computer Science and Engineering, Sun Yat-sen University; School of Computer Engineering, Jiangsu Ocean University(北京理工大学计算机科学与技术学院; 中山大学计算机科学与工程学院; 江苏海洋大学计算机工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长提示预填充导致解码延迟超限的问题,提出SLOWeave在线调度方法,动态选择能在最早解码截止时间前完成的最大预填充块,无需调整块大小,在混合与长上下文工作负载下有效吞吐量提升最高3.3倍。

AI 中文摘要

连续批处理提高了大语言模型(LLM)服务的吞吐量,但长提示的预填充会延迟解码迭代,并违反令牌间延迟目标。分块预填充缓解了这种干扰,但其块大小通常是固定的:小块保护解码延迟但反复付出启动开销,而大块提高预填充效率但产生延迟尖峰。我们提出SLOWeave,一种在线调度方法,选择预测能在最早活动解码截止时间之前完成的最大预填充块。该决策无需针对工作负载的块大小调整,并通过在对数时间内在单调迭代成本模型上搜索来计算。我们证明,每当仅解码迭代可行且成本预测器准确时,SLOWeave在保留每个活动请求的下一个令牌截止时间的决策中最大化即时预填充进度。我们在可复现的事件驱动模拟器和迭代级GPU运行时上评估了该方法,涵盖聊天、混合上下文、长上下文和突发工作负载。在25毫秒每输出令牌的目标下,SLOWeave在混合请求上比最强的固定块基线将有效吞吐量提高39%,在长上下文请求上提高38%。在更严格的10毫秒目标下,增益分别升至3.3倍和2.4倍。这些结果将自适应块大小作为有用的服务原语,并为与迭代级LLM运行时集成提供了可实现的控制器。

英文摘要

Continuous batching improves large language model (LLM) serving throughput, but long prompt prefills can delay decode iterations and violate inter-token latency objectives. Chunked prefill mitigates this interference, yet its chunk size is normally fixed: small chunks protect decode latency but repeatedly pay launch overhead, while large chunks improve prefill efficiency but create latency spikes. We introduce SLOWeave, an online scheduling method that selects the largest prefill chunk predicted to finish before the earliest active decode deadline. The decision requires no workload-specific chunk-size tuning and is computed by a logarithmic-time search over a monotone iteration-cost model. We prove that, whenever a decode-only iteration is feasible and the cost predictor is accurate, SLOWeave maximizes immediate prefill progress among decisions that preserve every active request's next-token deadline. We evaluate the method in a reproducible event-driven simulator and an iteration-level GPU runtime across chat, mixed-context, long-context, and bursty workloads. Under a 25ms time-per-output-token objective, SLOWeave improves goodput over the strongest fixed-chunk baseline by 39% on mixed requests and 38% on long-context requests. Under a stricter 10ms objective, the gains rise to 3.3$\times$ and 2.4$\times$, respectively. These results isolate adaptive chunk sizing as a useful serving primitive and provide an implementation-ready controller for integration with iteration-level LLM runtimes.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑