发表机构
National University of Singapore; A*STAR Institute of Advanced Intelligence and Computing (A*STAR IAIC); Nanyang Technological University(新加坡国立大学; 新加坡科技研究局高级智能与计算研究所; 南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MiniWAM通过PRISM学习紧凑未来表示,联合预测目标与机器人动作,在多仿真基准上性能优于大WAM且训练加速最高达8倍,参数仅0.25B。
AI 中文摘要
世界建模已成为机器人策略的有效协同训练目标,催生了联合预测动作与未来状态的世界-动作模型(World Action Models,WAMs)。然而,大多数WAM在预训练视觉骨干网络的原生表示空间中预测未来状态,导致目标维度高、训练成本大。本文提出MiniWAM,它转而从特权当前-未来转换中学习紧凑未来表示以作为预测目标。为构建这些目标,我们提出了通过逆时空建模的预测表示(Predictive Representations via Inverse Spatiotemporal Modeling,PRISM),该方法结合逆动力学监督与特征重构,在保留有用未来状态信息的同时,突出与控制相关的转换信息。在冻结已学习的PRISM编码器后,MiniWAM被训练为从当前观测中联合预测所得目标与机器人动作。相较于原生未来特征预测,MiniWAM使用的原生未来特征标记减少了65倍,在DINOv3和WAN2.1 VAE特征上均表现更优,同时实现了世界-动作训练最高达8倍的加速。在参数规模为0.25B时,MiniWAM已在LIBERO、LIBERO-Plus和RoboTwin 2.0仿真基准上与大得多的WAM具有竞争力。表示分析进一步表明,PRISM除特征重构外还贡献了行为结构。这些结果证明,有效的世界-动作建模无需预测原生视觉未来,紧凑的预测表示为策略学习提供了强大且高效得多的目标。项目页面可访问:this https URL。
英文摘要
World modeling has emerged as an effective co-training objective for robot policies, giving rise to World Action Models (WAMs) that jointly predict actions and future states. However, most WAMs predict future states in the native representation space of pretrained visual backbones, resulting in high-dimensional targets with substantial training cost. We introduce MiniWAM, which instead predicts compact future representations learned from privileged current-future transitions. To construct these targets, we propose Predictive Representations via Inverse Spatiotemporal Modeling (PRISM), which combines inverse-dynamics supervision with feature reconstruction to emphasize control-relevant transition information while preserving useful future-state information. With the learned PRISM encoder frozen, MiniWAM is trained to jointly predict the resulting targets and robot actions from current observations. With 65$\times$ fewer native future feature tokens, MiniWAM consistently outperforms native future-feature prediction with both DINOv3 and WAN2.1 VAE features, while achieving up to an 8$\times$ speedup in world-action training. At 0.25B parameters, MiniWAM is already competitive with substantially larger WAMs on LIBERO, LIBERO-Plus, and RoboTwin 2.0 simulation benchmarks. Representation analyses further show that PRISM contributes behavioral structure beyond feature reconstruction alone. These results demonstrate that effective world-action modeling does not require predicting native visual futures, and that compact predictive representations provide a strong and substantially more efficient target for policy learning. The project page is available at: https://j1dan.github.io/MiniWAM.