arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DyPES-VLA:学习共享动态先验与具身特定控制以实现跨具身操作

DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation

Junfeng Li, Junjie He, Zhide Zhong, Yangyang Zheng, Pingyue Sheng, Jiayu Dong, Ruixin Li, Haodong Yan, Jiaguan Zhu, Tianran Zhang, Runze Yu, Wen Chen, Liuqing Yang, Yuxiang Gao, Haoang Li

arXiv 2608.06374首次发表:更新:

发表机构

The Hong Kong University of Science and Technology (Guangzhou); COCO Matrix(香港科技大学(广州); COCO Matrix)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对跨具身机器人操作问题,提出DyPES-VLA模型,通过学习共享动态先验与具身特定MoE动作头,在LIBERO等数据集上达到先进操作成功率。

AI 中文摘要

视觉-语言-动作(Vision-Language-Action,VLA)模型已成为机器人操作的强大范式,但为异构机器人具身训练单一通用策略仍是未解决的问题。现有方法存在两个主要局限:其一,它们未充分利用不同视觉与交互数据间共享的动态先验,限制了跨具身迁移能力;其二,它们需要大量手动预处理,将具身特定动作转换为通用格式。为克服这些局限,我们提出DyPES-VLA,一种学习共享动态先验与具身特定控制的跨具身VLA模型。首先,我们通过在跨具身数据上以未来预测目标训练视觉-语言模型(Vision-Language Model,VLM)来学习共享动态先验,驱动共享查询表示捕捉物体运动、接触及交互引发的场景变化。其次,一个具身特定的混合专家(Mixture-of-Experts,MoE)动作头将这些共享动态先验直接转换为每个具身原生动作空间中的可执行控制,无需手动将异构动作预对齐为通用格式。该动作头共享注意力层以捕捉通用时序动作结构,同时其具身特定的前馈专家模块解决不同具身的独特运动学约束与控制语义。作为通用策略,我们的DyPES-VLA在仿真与真实世界评估中达到了先进性能:在LIBERO上的成功率为98.0%,在RoboCasa-GR1上为59.25%,在RoboTwin~2.0上为89.02%。

英文摘要

Vision-Language-Action (VLA) models have become a powerful paradigm for robot manipulation, but training a single generalist policy for heterogeneous robot embodiments remains an open problem. Existing methods have two main limitations. First, they underuse dynamics priors shared across diverse visual and interaction data, limiting cross-embodiment transfer. Second, they require extensive manual preprocessing to convert embodiment-specific actions into a common format. To overcome these limitations, we propose DyPES-VLA, a cross-embodiment VLA that learns shared Dynamics Priors and Embodiment-Specific control. First, we learn shared dynamics priors by training the vision-language model (VLM) with a future-prediction objective on cross-embodiment data, driving the shared query representation to capture object motion, contact, and interaction-induced scene changes. Second, an embodiment-specific Mixture-of-Experts (MoE) action head translates these shared dynamics priors into executable controls directly in each embodiment's native action space, without manually pre-aligning heterogeneous actions into a common format. This head shares attention layers to capture common temporal action structures, while its embodiment-specific feed-forward experts resolve the unique kinematic constraints and control semantics of distinct embodiments. As a generalist policy, our \ourmethod achieves state-of-the-art performance across simulation and real-world evaluations, reaching 98.0% success on LIBERO, 59.25% on RoboCasa-GR1, and 89.02% on RoboTwin~2.0.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑