arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

JoyNexus:面向服务的视觉语言动作模型多租户训练后处理

JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models

Haoran Sun, Wentao Zhang, Junyang Hua, Hedan Yang, Yongjian Guo, Yifei Zhang, Xiaolong Xiang, Mingxi Luo, Jing Long, Chen Zhao, Chen Zhou, Wanting Xu, Qiming Yang, Hui Zhang, Song Wang, Xiaodong Bai, Shuai Di, Xu Chu, Xiaotie Deng, Yicheng Gong, Junwu Xiong

arXiv 2607.16074首次发表:更新:

发表机构

Peking University; Beihang University; Beijing Institute of Technology; Tsinghua University(北京大学; 北航; 北京理工大学; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对VLA模型训练后处理问题,提出JoyNexus统一服务,解耦相关服务,多租户可通过API调用,引入组批处理提高效率,经评估能减少GPU时间,提升服务利用率。

AI 中文摘要

视觉语言动作(VLA)模型的训练后处理至关重要。现有计算服务通常为单个租户分配专用的GPU和CPU资源,存在基础设施适配负担重、计费模式不合理等问题。为此提出JoyNexus,它解耦了训练模型服务、推理模型服务和环境服务,多租户可通过API调用。还引入组批处理提高效率,通过工作负载模拟和组批处理管道评估,结果表明其能减少GPU时间并提高服务利用率。

英文摘要

The post-training of Vision-Language-Action (VLA) models is essential due to the diversity of simulators, robot embodiments, and task objectives. Existing compute services, whether offered as direct accelerator rental or batch-workload submission, typically allocate an exclusive set of GPU and CPU resources to a single tenant. While this paradigm maximizes client flexibility, it burdens users with infrastructure adaptation, and the fixed card-hour accounting model renders short or bursty workloads both expensive for tenants and inefficient for the service provider. To address these challenges, we present JoyNexus, a unified service for multi-tenant VLA supervised fine-tuning, reinforcement learning, and evaluation. JoyNexus decouples the Training Model Service, Inference Model Service, and Environment Service, each accessed through APIs and backed by resident shared base models with tenant-specific slots. Tenants can directly invoke high-level semantic APIs for training, rollout, and evaluation, or compose custom algorithms using lower-level APIs and their assigned endpoints. Multiple tenants submit workloads concurrently; their action modules, optimizers, rollout records, and policy versions remain isolated, and the service is scheduled by the global Training Queue and Inference Queue. To further improve multi-tenant training efficiency, JoyNexus introduces group batching for heterogeneous VLA data schemas that share a compatible model-facing prefix, enabling a single shared backbone forward pass over grouped samples. Finally, we evaluate JoyNexus through workload simulation and a group-batching pipeline in a realistic embodied scenario. Results show that, compared with isolated single-tenant execution, JoyNexus reduces aggregate GPU time and improves service utilization via cross-tenant scheduling on shared resources.

Comments23 pages, 12 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑