P-PAS:面向长上下文大语言模型服务的预填充压力自适应调度
P-PAS: Prefill-Pressure Adaptive Scheduling for Long-Context LLM Serving
浏览论文内容
中文总结 AI 辅助
针对长上下文LLM服务中静态最大批处理令牌数(MBT)无法适配不同调度压力的问题,提出P-PAS动态调度策略,可在各类负载下维持低延迟,避免固定MBT的局限。
中文摘要 AI 辅助
检索增强生成(RAG)和智能体系统等长上下文大语言模型应用,常需处理数万输入令牌以生成短输出,因此端到端请求延迟是重要的服务指标。研究表明,vLLM中控制令牌调度预算的最大批处理令牌数(MBT)对延迟的影响依赖于调度压力:低调度压力下,更大的令牌预算可降低延迟;高压力下,更小的预算更优。因此,单一静态MBT无法在所有负载场景下表现最佳。本文提出轻量级策略预填充压力自适应调度(P-PAS),该策略根据并发预填充与解码状态动态调整调度预算:低压力下保留大令牌预算,压力升高时限制预填充工作量。在多种模型、工作负载和GPU上的实验显示,P-PAS可在负载变化场景下维持低端到端延迟,避免固定MBT的局限。内核级分析表明,低调度压力下,大预填充块可提升执行效率,但该优势随模型-硬件配置变化;调度压力升高时,小预填充块可减少对活跃解码的干扰,这解释了观测到的负载依赖型MBT敏感性。用于复现结果的代码和制品可在该https URL获取。
英文摘要
Long-context LLM applications such as retrieval-augmented generation (RAG) and agentic systems often process tens of thousands of input tokens to produce short outputs, making end-to-end request latency an important serving objective. We show that the maximum number of batched tokens (MBT), which controls the token scheduling budget in vLLM, has a scheduling-pressure-dependent effect on latency. Larger token budgets can reduce latency under low scheduling pressure, while smaller budgets become preferable under higher pressure. Consequently, no single static MBT performs best across load regimes. We introduce Prefill-Pressure Adaptive Scheduling (P-PAS), a lightweight policy that dynamically adapts the scheduling budget based on concurrent prefill and decode state. P-PAS retains a large token budget under low pressure and constrains prefill work as pressure increases. Across models, workloads, and GPUs, P-PAS maintains low end-to-end latency across changing load regimes, avoiding the limitations of a fixed MBT. Kernel-level profiling shows that large prefill chunks can improve execution efficiency under low scheduling pressure, but that this advantage varies across model--hardware configurations. As scheduling pressure increases, smaller chunks can instead reduce interference with active decoding, explaining the observed load-dependent MBT sensitivity. Code and artifacts for reproducing our results are available at https://github.com/TimoSaemann/ppas-vllm .