AI 中文总结
针对全模态生成服务的异构多阶段工作流问题,提出统一服务运行时vLLM-Omni,通过单一编排器管理多阶段流水线,支持文本、语音、图像、视频和动作输出,并在H100/H200上验证了其架构与API。
AI 中文摘要
与智能系统的交互正在超越以文本为中心的聊天机器人和编码智能体。语音原生助手、视觉生成与编辑、世界模型环境以及机器人动作循环,需要能够输出文本、音频、图像、视频和动作的模型。这些模型在执行模式上有所不同:多阶段自回归的全模态和TTS流水线、迭代扩散或流匹配生成器,以及跨步骤携带状态的更长寿命的世界模型或机器人循环。因此,服务不再是一个单一的文本解码循环,而是一个异构的多阶段工作流,涉及跨阶段传输、流式传输和会话形态的交互。现有的推理栈通常针对单一架构家族进行优化。LLM服务器深化了自回归调度和KV管理,而扩散栈则深化了去噪和并行生成。两者都没有为通过独立生成器输出语音、像素或动作的流水线提供共享的控制平面,因此生产部署常常退回到跨不相交运行时的临时组合。我们提出了vLLM-Omni,一个用于全模态生成的统一服务运行时。vLLM-Omni将每个工作负载组织为在单一编排器下的多阶段流水线,该编排器接受请求、跨阶段推进请求,并对流式输出进行解复用。专门的引擎和阶段副本提供计算能力;连接器在数据平面上承载重型负载;面向会话的控制支持长寿命的双向、世界模型和机器人工作负载。本报告涵盖了架构(阶段级KV路径、副本池、多硬件平台和效率栈)以及用于全模态、TTS、图像/视频、世界模型、机器人和双向工作负载的OpenAI兼容和OpenPI API。我们在H100上的多模态夜间CI(TTS和MiniCPM-o在H200上)进行了评估,重点关注Qwen3-Omni。
英文摘要
Interaction with intelligent systems is expanding beyond text-centric chatbots and coding agents. Speech-native assistants, visual generation and editing, world-model environments, and robot action loops require models that emit text, audio, images, video, and actions. These models differ in execution pattern: multi-stage autoregressive omni and TTS pipelines, iterative diffusion or flow-matching generators, and longer-lived world-model or robot loops that carry state across steps. As a result, serving is no longer a single text decode loop, but a heterogeneous multi-stage workflow with cross-stage transfer, streaming, and session-shaped interaction. Existing inference stacks are typically optimized for one architecture family. LLM servers deepen autoregressive scheduling and KV management, while diffusion stacks deepen denoising and parallel generation. Neither provides a shared control plane for pipelines that emit speech, pixels, or actions through separate generators, so production deployments often fall back to ad-hoc composition across disjoint runtimes. We present vLLM-Omni, a unified serving runtime for omni-modality generation. vLLM-Omni organizes each workload as a multi-stage pipeline under a single orchestrator that admits requests, advances them across stages, and demultiplexes streaming outputs. Specialized engines and stage replicas provide compute; a connector carries heavy payloads on the data plane; and session-oriented control supports long-lived duplex, world-model, and robot workloads. This report covers the architecture (stage-level KV paths, replica pools, multi-hardware platforms, and efficiency stack) and OpenAI-compatible and OpenPI APIs for omni, TTS, image/video, world-model, robot, and duplex workloads. We evaluate on the multimodal nightly CI on H100 (TTS and MiniCPM-o on H200), focused on Qwen3-Omni.
Comments34 pages, 11 figures