arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00725cs.RO

SelfWAM:一种用于快速机器人控制的自 grounded 统一世界动作模型

SelfWAM: A Self-Grounded Unified World Action Model for Fast Robot Control

Bikang Pan, Fan Liu, Haotao Lu, Jingya Wang, Ye Shi

首次发表
浏览论文内容

中文总结 AI 辅助

SelfWAM是一种基于MoT架构的统一自 grounded WAM,通过联合预测动作等改进机器人控制,在RoboTwin 2.0等实验中提升了策略性能与推理速度。

中文摘要 AI 辅助

世界动作模型(WAMs)通过联合建模动作与未来观测来改进机器人策略学习,但仅基于任务提示和观测上下文进行未来预测,可能会捕获通用任务进展,而非执行动作的特定后果。我们引入SelfWAM,这是一种基于模态专用混合Transformer(MoT)架构的统一自 grounded WAM,可联合预测动作、动作条件下的未来RGB帧以及机器人自掩码,从而将未来预测建立在机器人可见身体及其动作诱导运动的基础上。在联合训练期间,SelfWAM允许未来视觉查询关注演示动作的干净副本,将视频分支转化为动作特定后果模型,同时保持仅动作的快速推理路径不变。为了让视频学习聚焦于与动作相关的视觉变化,我们使用特定于提示的目标来预测未来机器人自掩码,该目标去除外观细节,提供与条件动作时间演化紧密耦合的目标。干净动作条件和未来自掩码监督共同使未来预测更直接地反映执行动作如何改变机器人的可见运动和周围场景。在RoboTwin 2.0和真实世界操作任务上的实验表明,SelfWAM生成更具动作敏感性的未来,同时保持快速策略推理并提高策略性能。

英文摘要

World Action Models (WAMs) improve robot policy learning by jointly modeling actions and future observations. However, conditioning future prediction only on the task prompt and observation context risks capturing generic task progression rather than the action-specific consequences of the executed action. We introduce SelfWAM, a unified self-grounded WAM built on a modality-specialized Mixture-of-Transformers (MoT) architecture that jointly predicts actions, action-conditioned future RGB frames, and robot self-masks, thereby grounding future prediction in the robot's visible body and its action-induced motion. During joint training, SelfWAM allows future visual queries to attend to a clean copy of the demonstrated action, turning the video branch into an action-specific consequence model while leaving the fast action-only inference path unchanged. To focus video learning on action-relevant visual changes, we use prompt-specific objectives for future robot self-mask prediction, which removes appearance details and provides a target whose temporal evolution is tightly coupled with the conditioning action. Together, clean-action conditioning and future self-mask supervision make future predictions more directly reflect how the executed action changes the robot's visible motion and the surrounding scene. Experiments on RoboTwin 2.0 and real-world manipulation tasks show that SelfWAM produces more action-sensitive futures and preserves fast policy inference, while improving policy performance.

↑