发表机构
Peking University; Beihang University; Beijing Institute of Technology; Tsinghua University(北京大学; 北航; 北京理工大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对VLA模型训练后处理问题,提出JoyNexus统一服务,解耦相关服务,多租户可通过API调用,引入组批处理提高效率,经评估能减少GPU时间,提升服务利用率。
AI 中文摘要
视觉语言动作(VLA)模型的训练后处理至关重要。现有计算服务通常为单个租户分配专用的GPU和CPU资源,存在基础设施适配负担重、计费模式不合理等问题。为此提出JoyNexus,它解耦了训练模型服务、推理模型服务和环境服务,多租户可通过API调用。还引入组批处理提高效率,通过工作负载模拟和组批处理管道评估,结果表明其能减少GPU时间并提高服务利用率。
英文摘要
The post-training of Vision-Language-Action (VLA) models is essential due to the diversity of simulators, robot embodiments, and task objectives. Existing compute services, whether offered as direct accelerator rental or batch-workload submission, typically allocate an exclusive set of GPU and CPU resources to a single tenant. While this paradigm maximizes client flexibility, it burdens users with infrastructure adaptation, and the fixed card-hour accounting model renders short or bursty workloads both expensive for tenants and inefficient for the service provider. To address these challenges, we present JoyNexus, a unified service for multi-tenant VLA supervised fine-tuning, reinforcement learning, and evaluation. JoyNexus decouples the Training Model Service, Inference Model Service, and Environment Service, each accessed through APIs and backed by resident shared base models with tenant-specific slots. Tenants can directly invoke high-level semantic APIs for training, rollout, and evaluation, or compose custom algorithms using lower-level APIs and their assigned endpoints. Multiple tenants submit workloads concurrently; their action modules, optimizers, rollout records, and policy versions remain isolated, and the service is scheduled by the global Training Queue and Inference Queue. To further improve multi-tenant training efficiency, JoyNexus introduces group batching for heterogeneous VLA data schemas that share a compatible model-facing prefix, enabling a single shared backbone forward pass over grouped samples. Finally, we evaluate JoyNexus through workload simulation and a group-batching pipeline in a realistic embodied scenario. Results show that, compared with isolated single-tenant execution, JoyNexus reduces aggregate GPU time and improves service utilization via cross-tenant scheduling on shared resources.
Comments23 pages, 12 figures