AI 中文总结
本研究针对MoE增强VLA的专家路由问题,提出KinRT方法,通过运动学聚类监督路由器训练,在两个平台上验证其性能优于现有模型,相关代码与平台将开源。
AI 中文摘要
尽管混合专家(MoE)通过专家专业化增强了视觉-语言智能体(VLA),但由于不同操作任务的动作存在运动学异质性,路由器面临专家路由效率低下的问题,更糟糕的是,推理时无法获取运动学信号。在本研究中,我们首先观察到,大多数语义不同的操作任务可归为多种运动学原型。基于这一发现,我们提出运动学监督显式路由(KinRT),这一范式从隐式的、观测驱动的专家路由转向显式的、运动学引导的专家调度。具体而言,我们对动作轨迹进行运动学聚类,形成多个运动学一致的组,其ID作为监督路由器训练的真值;推理时,路由器仅使用视觉-语言观测来调度专家,无需依赖动作运动学。KinRT引入了一种非对称桥接机制,将训练时动作空间中的任务运动学蒸馏到推理时的观测空间。此外,为评估KinRT的跨平台泛化能力,我们从零开始使用3D打印技术构建了一个经济型DIY机器人(DIYRobot)平台,成本低于2000美元。大量实验表明,KinRT在RoboTwin基准测试中比密集型和MoE型VLA性能提升超过23.26%,在我们引入的DIYRobot平台上提升20.27%。我们的代码和DIYRobot平台将开源。
英文摘要
While MoE augments VLA via expert specialization, router suffers from ineffective expert routing owing to the kinematic heterogeneity of actions across manipulation tasks and, even worse, the unavailability of the kinematic signals at inference time. In this work, we first observe that most semantically distinct manipulation tasks reduce to multiple kinematic archetypes. Motivated by this finding, we propose Kinematics-supervised explicit routing (KinRT), a new paradigm that shifts from implicit, observation-driven expert routing to explicit, kinematics-guided expert dispatching. Specifically, we perform kinematic clustering on action trajectories into multiple kinematically coherent groups, whose IDs serve as ground truth to supervise the training of the router; at inference time, the router dispatches experts only using visual-language observations, without any reliance on action kinematics. KinRT actually introduces an asymmetric bridging mechanism that distills the task kinematics from the action space in training into the observation space at inference. In addition, to assess KinRT's cross-platform generalization, we build an economical, Do-It-Yourself robot (DIYRobot) platform from scratch using 3D-print technology ($<$ 2,000USD). Extensive experiments demonstrate KinRT's superiority over both dense and MoE-featured VLAs by more than 23.26% on RoboTwin benchmark and 20.27% on our introduced DIYRobot platform. Our code and DIYRobot platform will be open-sourced.
Comments9 pages