发表机构
Huazhong University of Science & Technology; D-Robotics; Horizon Robotics; Fudan University; Xi’an Jiaotong University(华中科技大学; 地瓜机器人; 地平线机器人; 复旦大学; 西安交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出ReWAM,一种基于DINO特征的表示中心世界-动作模型,通过特征校准、时间瓶颈和动作接地塑造,无需生成式视频预训练即在RoboTwin 2.0和RoboDojo上取得显著成功。
AI 中文摘要
世界-动作模型联合学习机器人策略并预测未来观测,使得表示空间成为控制与预测之间的接口。我们通过受控比较研究该空间的设计,发现重建保真度和预训练感知特征单独使用均不能确保有效的策略学习。这些发现促使我们提出ReWAM,一种基于预训练DINO特征的表示中心的世界-动作模型。特征校准和时间表示瓶颈将这些特征组织成适合动力学建模的紧凑世界状态。动作接地表示塑造仅将动作损失梯度路由到瓶颈,从而使策略塑造表示编码的内容,而世界模型学习其如何演化。无需生成式视频预训练,ReWAM在RoboTwin 2.0上达到93.6%的成功率。在RoboDojo上,使用约600小时的具身预训练数据,其平均得分为12.29,成功率为8.28%。
英文摘要
World-action models jointly learn robot policies and predict future observations, making the representation space an interface between control and prediction. We study the design of this space through controlled comparisons, finding that neither reconstruction fidelity nor pre-trained perceptual features alone ensure effective policy learning. These findings motivate ReWAM, a representation-centric world-action model built on pre-trained DINO features. Feature Calibration and a Temporal Representation Bottleneck organize these features into compact world states suited to dynamics modeling. Action-Grounded Representation Shaping routes only action-loss gradients to the bottleneck, thereby letting the policy shape what the representation encodes while the world model learns how it evolves. Without generative video pre-training, ReWAM achieves 93.6% success on RoboTwin 2.0. On RoboDojo, it achieves an average score of 12.29 and a success rate of 8.28% using approximately 600 hours of embodied pre-training data.
Commentshttps://github.com/hustvl/ReWAM