arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26523cs.DC

VPP:用于长上下文大语言模型推理中高效分块预填充的虚拟流水线并行技术

VPP: Virtual Pipeline Parallelism for Efficient Chunked Prefill in Long-Context LLM Inference

Yan Shi, Xiaochao Wang, Jingchun Gao, Jintao Luo, Xinyi Zhou, Feng Liu, Kui Luo, Xushi Li, Xinjie Guo, Liangjun Feng

AI总结:

该研究提出VPP虚拟流水线并行技术,通过虚拟阶段布局优化解决长上下文LLM分块预填充的流水线气泡问题,在长序列等场景下显著提升吞吐量并降低气泡率。

AI中文摘要:

分块预填充流水线并行(CPP)是大语言模型(LLM)推理的关键技术。然而,等大小分块会导致延迟不平衡:后续分块需处理更长的前缀键值(KV)缓存,注意力成本更高,进而引发流水线气泡。现有方法通过动态分块调整(Dynamic CPP,DCPP)缓解该不平衡,但实测显示,这种方法以调度开销为代价换取负载均衡,在长序列场景下会变得不利。本研究提出虚拟流水线并行(VPP),其保持分块大小固定,通过虚拟阶段优化流水线布局:V形虚拟阶段遍历方式将每个分块的耗时中间阶段与相邻分块负载较轻的头尾阶段重叠,同时异步通信与流水线打包进一步降低通信停滞和跨请求排空气泡。我们在vLLM-Ascend中实现VPP,并在16颗昇腾910C NPU上,针对三个基于混合专家(MoE)的LLM(序列长度最高达100万token)开展评估。结果显示,在长序列场景下,VPP较DCPP的吞吐量提升最高达13.1%;在混合工作负载下提升6.7%,且在短序列场景下保持原有性能。在51.2万token的DeepSeek-V3.1预填充工作负载上,VPP将流水线气泡率从6.4%降至0.1%,较DCPP降低98.0%。

英文摘要:

Chunked prefill pipeline parallelism (CPP) is a key technique for LLM inference. However, equal-size chunks exhibit imbalanced latency, as later chunks attend longer prefix KV caches and incur higher attention costs, leading to pipeline bubbles. Existing approaches mitigate this imbalance through dynamic chunk resizing (Dynamic CPP, DCPP), but our measurements show that this trades scheduling overhead for load balancing, which becomes unfavorable on long sequences. In this study, we propose Virtual Pipeline Parallelism (VPP), which keeps chunk sizes fixed and optimizes the pipeline layout through virtual stages. A V-shaped virtual-stage traversal overlaps each chunk's expensive middle stages with the lighter head and tail stages of its neighbors, while asynchronous communication and pipelined packing further reduce communication stalls and cross-request drain bubbles. We implement VPP in vLLM-Ascend and evaluate it on three MoE-based LLMs with sequences up to 1M tokens on 16 Ascend 910C NPUs. VPP improves throughput by up to 13.1% over DCPP on long sequences and 6.7% on mixed workloads, while preserving performance on short sequences. On a 512K-token DeepSeek-V3.1 prefill workload, VPP reduces the pipeline bubble ratio from 6.4% to 0.1%, achieving a 98.0% reduction compared with DCPP.

↑