arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25961cs.RO

一个动作等价于一个补丁:基于PatchWAM的统一世界-动作建模

An Action Is Worth One Patch: Unified World-Action Modeling with PatchWAM

Tianheng Wang, Zhou Xie, Heng Jia, Jianhua Xu, Tong Zhang, Kaicheng Yu

首次发表
浏览论文内容

中文总结 AI 辅助

提出PatchWAM,将连续动作映射为补丁,使单一生成模型统一世界预测与动作生成,无需独立动作头,在LIBERO-Plus和RoboTwin 2.0上分别达到91.8%和96.12%的成功率。

中文摘要 AI 辅助

生成式视觉模型为学习物理动力学表示提供了基础,但其扩展到连续控制引发了一个根本性问题:视觉预测和动作生成是否需要独立的计算路径?现有方法通常引入可训练的动作头或独立的动作专家,以桥接低维状态与高维视觉表示。在本工作中,我们探索当动作以兼容的表示表达时,视觉骨干网络的现有能力是否也能支持控制。为此,我们提出了PatchWAM(补丁世界-动作模型),通过一种称为“动作即补丁”的固定映射,将连续动作视为另一种类型的补丁。这使得单一模型既能预测机器人应如何移动,也能预测场景随后可能呈现的样子。视觉预测和动作生成成为同一生成过程的一部分,无需专门的动作头或独立的动作专家。使用子采样训练窗口的实验显示,该方法优于匹配的双专家控制,而在全数据设置并附加增强演示的情况下,基准评估在LIBERO-Plus上达到91.8%的成功率,在RoboTwin 2.0上达到96.12%。更广泛地,该结果表明,能力无需在可继承之处额外添加:扩展生成骨干网络的约束在于新信号所写入的接口,而非对其建模的能力。

英文摘要

Generative visual models offer a foundation for learning representations of physical dynamics, yet their extension to continuous control raises a fundamental question: do visual prediction and action generation require separate computational pathways? Existing approaches usually introduce trainable action heads or separate action experts to bridge low-dimensional states and high-dimensional visual representations. In this work, we explore whether the visual backbone's existing capacity can also support control when actions are expressed in a compatible representation. Thus, we introduce PatchWAM (Patch World-Action Model), which treats continuous actions as another type of patch through a fixed mapping called Action-as-Patch. This allows a single model to predict both how the robot should move and what the scene may look like afterward. Visual prediction and action generation become parts of the same generative process, without a dedicated action head or separate action expert. Experiments with subsampled training windows show gains over a matched dual-expert control, while benchmark evaluations reach 91.8% success rate on LIBERO-Plus and 96.12% on RoboTwin 2.0 in a full-data setting with additional augmented demonstrations. More broadly, the result suggests that capability need not be added where it can be inherited: the constraint on extending a generative backbone is the interface a new signal is written in, not the capacity to model it.

发表机构

  • Westlake University(西湖大学)
  • Lanzhou University(兰州大学)
  • Zhejiang University(浙江大学)
  • Awomo
  • University of Chinese Academy of Sciences(中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

↑