arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型推理中的并行策略表征:基础计算-通信权衡

Characterizing Parallelism Strategies in LLM Inference: Fundamental Compute-Communication Trade-offs

Javad Mirzaei, Jeebak Mitra

arXiv 2610.05305首次发表:更新:

发表机构

Dell Technologies Inc.(戴尔科技公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出统一解析框架,建模TP、PP和HB下分布式LLM推理的延迟分解,揭示计算-通信权衡,指导并行策略选择。

AI 中文摘要

大语言模型(LLM)推理已成为现代AI系统中的主导工作负载,要求服务基础设施在满足严格延迟服务级别目标(SLOs)的同时最大化吞吐量。由于最先进的LLM超出了单个GPU的计算和内存容量,推理通常使用张量并行(TP)、流水线并行(PP)或混合并行(HB)分布在多个GPU上。然而,由于计算、通信、流水线利用率、序列长度、批量大小和模型架构之间的复杂交互,选择最有效的并行策略仍然具有挑战性。现有方法主要依赖经验评估,对这些策略之间的权衡提供的分析性见解有限,特别是在推理的不同预填充和解码阶段。在本文中,我们提出了一个统一的解析框架,用于建模在TP、PP和HB下的分布式LLM推理。该框架将端到端延迟分解为计算、GPU间通信和流水线气泡开销,并推导出解析模型,将TP集体通信、PP点对点通信和流水线利用率捕获为硬件、模型和工作负载特征的函数。该模型进一步表征了预填充和解码的不同执行行为,解释了为什么面向PP的配置有利于计算密集型的预填充,而面向TP的配置通过消除流水线气泡减少解码延迟。在现代LLM的多GPU平台上的实验验证了该模型,并确认了并行策略之间的基础计算-通信权衡。该框架为并行性选择、容量规划和未来LLM服务系统的优化提供了实用指导。

英文摘要

Large Language Model (LLM) inference has become the dominant workload in modern AI systems, requiring serving infrastructures to maximize throughput while meeting strict latency Service-Level Objectives (SLOs). Since state-of-the-art LLMs exceed the compute and memory capacity of a single GPU, inference is commonly distributed across multiple GPUs using tensor parallelism (TP), pipeline parallelism (PP), or hybrid parallelism (HB). However, selecting the most effective parallelism strategy remains challenging due to complex interactions among computation, communication, pipeline utilization, sequence length, batch size, and model architecture. Existing approaches largely rely on empirical evaluation and provide limited analytical insight into the trade-offs among these strategies, particularly across the distinct prefill and decoding phases of inference. In this paper, we present a unified analytical framework for modeling distributed LLM inference under TP, PP, and HB. The framework decomposes end-to-end latency into computation, inter-GPU communication, and pipeline bubble overhead, and derives analytical models that capture TP collective communication, PP point-to-point communication, and pipeline utilization as functions of hardware, model, and workload characteristics. The model further characterizes the differing execution behavior of prefill and decoding, explaining why PP-oriented configurations favor compute-intensive prefill while TP-oriented configurations reduce decoding latency by eliminating pipeline bubbles. Experiments with modern LLMs on multi-GPU platforms validate the model and confirm the fundamental compute-communication trade-off across parallelism strategies. The framework provides practical guidance for parallelism selection, capacity planning, and optimization of future LLM serving systems.

Comments17 Pages, 9 Figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑