arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24159cs.ROcs.CV

DeVA:用于机器人策略学习的具有物理引导的解耦视频-动作模型

DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning

Mengqi Zhang, Sahil Khose, Simar Kareer, Yuchen Song, Unnat Jain, Judy Hoffman

首次发表
浏览论文内容

中文总结 AI 辅助

研究可泛化机器人操纵策略,提出DeVA模型,通过专门的视频和动作专家、多级特征转移及物理显著引导,实现丰富信息交换,使策略学习更易处理,在模拟和实际实验中展现良好性能。

中文摘要 AI 辅助

可泛化的机器人操纵需要能在执行语言指令时预测视觉场景如何演变的策略。近期视觉-语言-动作模型虽受益于大规模预训练,但其主要静态预训练目标对物理动力学和时间因果关系的监督有限。视频生成模型通过未来预测编码丰富时空先验提供了有前景的基础。现有视频-动作模型存在问题,本文引入DeVA,它有专门的视频和动作专家、多级特征转移及物理显著引导。能在使策略学习更易处理的同时实现丰富信息交换,还通过物理显著引导监督中间视频特征和动作流。在模拟基准和实际部署实验中展现出良好性能。

英文摘要

Generalizable robot manipulation requires policies that can anticipate how visual scenes evolve while executing language instructions. While recent Vision-Language-Action models benefit from large-scale pretraining, their predominantly static pretraining objectives provide limited supervision for physical dynamics and temporal causality, leaving control-relevant knowledge to be learned from downstream robot demonstrations. Video generative models offer a promising foundation by encoding rich spatiotemporal priors through future predictions. However, existing Video-Action Models either couple video and action prediction in a shared backbone, making policy adaptation harder to optimize, or under-utilize video information when guiding the action branch. In this work, we introduce DeVA, a Decoupled Video-Action model with specialized video and action experts, multi-level feature transfer, and physically salient guidance. DeVA transfers representations from multiple video layers to the action expert, enabling rich information exchange while making policy learning more tractable. It further supervises intermediate video features and the action stream with physically salient guidance (affordance/depth). Experiments on both simulation benchmarks and real-world deployment demonstrate strong performance with limited data, faster convergence than a unified architecture, and clear performance gains from physical guidance.

发表机构

  • University of California, Irvine(加利福尼亚大学欧文分校)
  • Georgia Institute of Technology(佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑