arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2604.06168cs.CVcs.RO

动作图像:通过多视图视频生成实现端到端策略学习

Action Images: End-to-End Policy Learning via Multiview Video Generation

Haoyu Zhen, Zixian Gao, Qiao Sun, Yilin Zhao, Yuncong Yang, Yilun Du, Pengsheng Guo, Tsun-Hsuan Wang, Yi-Ling Qiao, Chuang Gan

首次发表 更新
浏览论文内容

中文总结 AI 辅助

本文提出Action Images,通过多视图视频生成实现策略学习,利用像素化的动作表示,使视频模型本身成为零样本策略,提升视频-动作联合生成质量。

中文摘要 AI 辅助

世界动作模型(WAMs)已成为机器人策略学习的有前景方向,因其能利用强大的视频骨干网络来建模未来状态。然而,现有方法往往依赖于单独的动作模块,或使用非像素化的动作表示,难以充分利用预训练视频模型的知识,并限制了跨视角和环境的迁移能力。在本文中,我们提出了Action Images,一种统一的世界动作模型,将策略学习建模为多视图视频生成。不同于将控制编码为低维令牌,我们将7自由度机器人动作转换为可解释的动作图像:基于2D像素的多视角动作视频,明确跟踪机器人臂运动。这种像素化的动作表示使视频骨干网络本身可以作为零样本策略,而无需单独的策略头或动作模块。除了控制外,相同的统一模型还支持视频-动作联合生成、动作条件视频生成和动作标注,基于共享的表示。在RLBench和真实世界评估中,我们的模型实现了最强的零样本成功率,并在视频-动作联合生成质量上优于先前的视频空间世界模型,表明可解释的动作图像是策略学习的有前景途径。

英文摘要

World action models (WAMs) have emerged as a promising direction for robot policy learning, as they can leverage powerful video backbones to model the future states. However, existing approaches often rely on separate action modules, or use action representations that are not pixel-grounded, making it difficult to fully exploit the pretrained knowledge of video models and limiting transfer across viewpoints and environments. In this work, we present Action Images, a unified world action model that formulates policy learning as multiview video generation. Instead of encoding control as low-dimensional tokens, we translate 7-DoF robot actions into interpretable action images: multi-view action videos that are grounded in 2D pixels and explicitly track robot-arm motion. This pixel-grounded action representation allows the video backbone itself to act as a zero-shot policy, without a separate policy head or action module. Beyond control, the same unified model supports video-action joint generation, action-conditioned video generation, and action labeling under a shared representation. On RLBench and real-world evaluations, our model achieves the strongest zero-shot success rates and improves video-action joint generation quality over prior video-space world models, suggesting that interpretable action images are a promising route to policy learning.

发表机构

  • UMass Amherst(马萨诸塞大学阿默斯特分校)
  • NVIDIA(英伟达)
  • Harvard University(哈佛大学)
  • Genesis AI

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑