arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23130cs.AIcs.PF

从推理引擎到推理控制平面:连接 vLLM、llm-d 与高效分布式 LLM 服务的演进

From Inference Engine to Inference Control Plane: Connecting vLLM, llm-d, and the Evolution of Efficient Distributed LLM Serving

Twinkll Sisodia

首次发表
浏览论文内容

中文总结 AI 辅助

本文综述LLM推理从引擎优化向分布式控制平面演进的趋势,提出推理执行规划器选择可行执行计划,并给出基准图谱、部署指南与评估框架,核心贡献在于综合现有研究而非新基准。

中文摘要 AI 辅助

大语言模型(LLM)推理正从引擎本地优化问题演变为分布式控制问题,涉及可复用状态、阶段放置、异构加速器、网络、自动扩缩容、可靠性以及服务级别目标。本文跨越同行评审的系统研究、开源实现和有记录的生产研究,连接了这一转变。它将 vLLM 和 llm-d 视为互补层:模型服务引擎通过 PagedAttention、连续批处理、内核、量化和并行等机制优化执行,而推理控制平面可以优化执行发生在何处、何时以及何种策略下,跨整个服务器群。贡献在于综合而非新基准;所有报告的性能和部署结果均归因于其原始来源。综合证据表明,现代推理中的稀缺资源正从单纯的原始 FLOPs 转向受管状态、放置、网络移动、可靠性和决策质量。我们提出了一个推理执行规划器,它选择可行的执行计划而不仅仅是端点,包括聚合与分离拓扑、KV 源和传输动作、硬件变体、路由/准入策略以及较慢的扩缩容决策。我们还提供了源本地基准图谱、瓶颈迁移分类法、实用部署指南、基于 SLO-goodput 的评估框架,以及面向智能体、多模态、异构和弹性推理的研究问题。

英文摘要

Large language model (LLM) inference is evolving from an engine-local optimization problem into a distributed control problem involving reusable state, phase placement, heterogeneous accelerators, networking, autoscaling, reliability, and service-level objectives. This paper connects that transition across peer-reviewed systems research, open-source implementations, and documented production studies. It treats vLLM and llm-d as complementary layers: model-serving engines optimize execution through mechanisms such as PagedAttention, continuous batching, kernels, quantization, and parallelism, while an inference control plane can optimize where, when, and under what policy execution occurs across a fleet. The contribution is synthesis rather than a new benchmark; all reported performance and deployment results remain attributed to their original sources. The combined evidence suggests that the scarce resource in modern inference is shifting from raw FLOPs alone toward managed state, placement, network movement, reliability, and decision quality. We propose an Inference Execution Planner that selects feasible execution plans rather than only endpoints, including aggregated versus disaggregated topology, KV source and transfer action, hardware variant, routing/admission policy, and slower scaling decisions. We also provide a source-local benchmark atlas, a bottleneck-migration taxonomy, practical deployment guidance, an evaluation framework based on SLO-goodput, and research questions for agentic, multimodal, heterogeneous, and resilient inference.

补充信息

↑