EAServe: 面向多模态大语言模型的编码感知分离式服务
EAServe: Encode-Aware Disaggregated Serving for Multimodal Large Language Models
查看机构详情
- University of Georgia(佐治亚大学)
- Western Digital(西部数据)
- University of Texas at Arlington(德克萨斯大学阿灵顿分校)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
EAServe通过将Encode设为EPD流水线控制点,结合运行时管理与混合自动选择,解决多模态LLM服务中GPU利用不平衡问题,相比现有系统提升高达4.3倍有效吞吐量。
中文摘要 AI 辅助
将Prefill和Decode两个阶段分离到独立的GPU池中,现已成为(纯文本)LLM服务的标准优化方法。然而,多模态大语言模型(MLLMs)增加了第三个阶段Encode,给资源分配带来了新的挑战。Encode将图像、视频或音频转换为语言模型可消费的嵌入,形成三阶段的Encode-Prefill-Decode(EPD)流水线。现有框架仅提供部分解决方案:纯文本PD系统缺少Encode阶段,而EPD框架将其作为独立服务暴露,却不调节下游请求流。该流水线还存在结构性资源不平衡:每个请求在进入下游工作之前都必须经过Encode,但即使在高负载下,按请求执行也会使编码GPU严重利用不足,从而饿死下游的Prefill和Decode工作节点。针对此问题,我们将Encode重新定位为EPD流水线的控制点,暴露三个紧密耦合的维度:工作何时进入下游、Prefill在何处执行以及GPU如何共享。我们在EAServe中通过两个协同设计的层实例化了这一方案。其运行时管理负载自适应的微批处理、速率控制的部分卸载到共置的Prefill工作节点,以及动态SM分区以实现可预测的共置。配置层混合自动选择(HAS)通过按阶段容量分析修剪不平衡分配,并通过基于TPE的贝叶斯优化细化剩余部分,从而在GPU分配、编码批大小和卸载比率的联合空间中导航。在覆盖图像、视频和音频的三种MLLM架构上评估,在相同SLO约束下,EAServe相比NVIDIA Dynamo和vLLM分别提供高达4.3倍和1.7倍的更高有效吞吐量,在EPD流水线上维持更均衡和更高的GPU利用率,并且比基线搜索方法更快达到接近最优的配置。
英文摘要
Disaggregating the two stages, Prefill and Decode, onto separate GPU pools is now a standard optimization for (text-only) LLM serving. However, multimodal LLMs (MLLMs), which add a third phase, Encode, pose new challenges for resource allocation. Encode turns images, video, or audio into embeddings that the language model can consume, yielding a three-stage Encode-Prefill-Decode (EPD) pipeline. Existing frameworks offer only partial answers: text-only PD systems lack Encode, while EPD frameworks expose it as a separate service without regulating downstream request flow. The pipeline also carries a structural resource imbalance: every request enters through Encode before downstream work can begin, yet per-request execution leaves the encode GPU severely underutilized even at high loads, starving the downstream Prefill and Decode workers. Addressing this, we reposition Encode as the control point of the EPD pipeline, exposing three tightly coupled dimensions: when work enters downstream, where prefill executes, and how the GPU is shared. We instantiate this in EAServe across two co-designed layers. Its runtime manages load-adaptive micro-batching, rate-controlled partial offload to a co-resident prefill worker, and dynamic SM partitioning for predictable co-location. The configuration layer, Hybrid Auto Selection (HAS), navigates the joint space of GPU allocation, encode batch size, and offload ratio by pruning unbalanced allocations with per-stage capacity profiling and refining the remainder through TPE-based Bayesian optimization. Evaluated on three MLLM architectures spanning image, video, and audio, EAServe delivers up to 4.3x and 1.7x higher goodput than NVIDIA Dynamo and vLLM, respectively, under identical SLO constraints, sustains more balanced and higher GPU utilization across the EPD pipeline, and reaches near-optimal configurations faster than baseline search methods.