发表机构
Max Planck Institute for Software Systems(马克斯·普朗克软件系统研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对机器人工厂中VLA模型推理的延迟关键性,提出Robion系统,通过GPU内流分离和智能调度,在边缘服务器上高效服务多机器人多模型,显著提升SLO达成率下的负载能力。
AI 中文摘要
视觉-语言-动作(VLA)模型通过两阶段设计展现出强大的机器人操作能力:先是视觉-语言模型(VLM)阶段,随后是动作扩散变换器(ADiT)阶段。由于机器人必须满足严格的服务水平目标(SLO)以确保安全,VLA推理本质上对延迟极为敏感。满足这些SLO需要高端GPU,然而重量、成本和功耗限制使得无法在机器人上集成此类GPU。先前的工作将VLA推理卸载到边缘服务器,这些服务器在VLA模型上为许多机器人提供服务。然而,当前的VLA系统缺乏在SLO约束下于多GPU服务器上支持多请求、多模型执行的能力,而现有的针对多阶段模型的 serving 系统优化了吞吐量和跨独立GPU的阶段分离,这并不适用于VLA模型毫秒级的阶段。我们设计了Robion,这是首个面向多GPU边缘服务器上多机器人、多模型请求的VLA serving 和管理系统,能够满足SLO。我们的 serving 引擎通过两个流在GPU内分离VLM和ADiT阶段,动态限制VLM流上的流式多处理器(SM),使得ADiT总能找到SM与其并行运行,并通过在这些流上共享来共同部署多个模型,按剩余SLO时间最少优先调度请求。我们的管理引擎支持在多GPU服务器上灵活放置模型,并集成了智能流量控制器,在所选放置下最大化每个模型的批处理,同时限制每个GPU的负载以满足SLO。对于单个模型,在98% SLO达成率下,Robion平均服务的机器人负载比最广泛使用的多阶段 serving 系统vLLM-Omni高6.7倍,比将VLM和ADiT作为单一流水线运行的Monolithic高1.5倍。在一个4-GPU服务器上服务8个不同模型的大规模实验中,Robion在98% SLO达成率下可服务多达64个机器人。
英文摘要
Vision-Language-Action (VLA) models show high robotic manipulation capabilities via a two-stage design: a Vision-Language Model (VLM) stage followed by an Action Diffusion Transformer (ADiT) stage. Since robots must meet strict Service-Level Objectives (SLOs) for safety, VLA inference is inherently latency-critical. Meeting these SLOs requires high-end GPUs, yet weight, cost, and power constraints preclude integrating such GPUs on-robot. Prior works offload VLA inference to edge servers that serve many robots on VLA models. However, current VLA systems lack support for multi-request, multi-model execution on a multi-GPU server under SLOs, while existing serving systems for multi-stage models are optimized for throughput and stage disaggregation across separate GPUs, which are ill-suited for the millisecond-scale stages of VLA models. We design Robion, the first VLA serving and management system for multi-robot, multi-model requests on multi-GPU edge servers that meets SLOs. Our serving engine disaggregates the VLM and ADiT stages within a GPU via two streams, dynamically restricting the SMs on VLM stream so ADiT always finds SMs to run alongside it, and co-locates multiple models by sharing these streams across them, prioritizing requests by least remaining SLO time. Our management engine enables flexible model placements on multi-GPU servers, and integrates an intelligent traffic controller that maximizes per-model batching under the chosen placement while bounding each GPU's load to meet SLOs. For individual models, Robion serves on average 6.7$\times$ and 1.5$\times$ higher robot load within 98% SLO attainment over vLLM-Omni, the most widely used multi-stage serving system, and Monolithic, which runs VLM and ADiT as a single pipeline, respectively. In a large-scale experiment of serving 8 different models on a 4-GPU server, Robion can serve up to 64 robots within 98% SLO attainment.