发表机构
The Hong Kong University of Science and Technology (Guangzhou); COCO Matrix(香港科技大学(广州); COCO Matrix)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对跨具身机器人操作问题,提出DyPES-VLA模型,通过学习共享动态先验与具身特定MoE动作头,在LIBERO等数据集上达到先进操作成功率。
AI 中文摘要
视觉-语言-动作(Vision-Language-Action,VLA)模型已成为机器人操作的强大范式,但为异构机器人具身训练单一通用策略仍是未解决的问题。现有方法存在两个主要局限:其一,它们未充分利用不同视觉与交互数据间共享的动态先验,限制了跨具身迁移能力;其二,它们需要大量手动预处理,将具身特定动作转换为通用格式。为克服这些局限,我们提出DyPES-VLA,一种学习共享动态先验与具身特定控制的跨具身VLA模型。首先,我们通过在跨具身数据上以未来预测目标训练视觉-语言模型(Vision-Language Model,VLM)来学习共享动态先验,驱动共享查询表示捕捉物体运动、接触及交互引发的场景变化。其次,一个具身特定的混合专家(Mixture-of-Experts,MoE)动作头将这些共享动态先验直接转换为每个具身原生动作空间中的可执行控制,无需手动将异构动作预对齐为通用格式。该动作头共享注意力层以捕捉通用时序动作结构,同时其具身特定的前馈专家模块解决不同具身的独特运动学约束与控制语义。作为通用策略,我们的DyPES-VLA在仿真与真实世界评估中达到了先进性能:在LIBERO上的成功率为98.0%,在RoboCasa-GR1上为59.25%,在RoboTwin~2.0上为89.02%。
英文摘要
Vision-Language-Action (VLA) models have become a powerful paradigm for robot manipulation, but training a single generalist policy for heterogeneous robot embodiments remains an open problem. Existing methods have two main limitations. First, they underuse dynamics priors shared across diverse visual and interaction data, limiting cross-embodiment transfer. Second, they require extensive manual preprocessing to convert embodiment-specific actions into a common format. To overcome these limitations, we propose DyPES-VLA, a cross-embodiment VLA that learns shared Dynamics Priors and Embodiment-Specific control. First, we learn shared dynamics priors by training the vision-language model (VLM) with a future-prediction objective on cross-embodiment data, driving the shared query representation to capture object motion, contact, and interaction-induced scene changes. Second, an embodiment-specific Mixture-of-Experts (MoE) action head translates these shared dynamics priors into executable controls directly in each embodiment's native action space, without manually pre-aligning heterogeneous actions into a common format. This head shares attention layers to capture common temporal action structures, while its embodiment-specific feed-forward experts resolve the unique kinematic constraints and control semantics of distinct embodiments. As a generalist policy, our \ourmethod achieves state-of-the-art performance across simulation and real-world evaluations, reaching 98.0% success on LIBERO, 59.25% on RoboCasa-GR1, and 89.02% on RoboTwin~2.0.