arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17524cs.RO

模态自回归世界-动作模型

Modality-Autoregressive World-Action Models

  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

Adam Hung, Bardienus P. Duisterhof, Deva Ramanan, Jeffrey Ichnowski

AI总结:

本文提出ModAR,一种模态自回归世界-动作模型,通过序列去噪多种未来模态(如点轨迹、DINO特征、深度)来提升动作预测性能,在多个数据规模下取得最高平均成功率,且训练效率高,无需预训练。

AI中文摘要:

世界-动作模型(WAMs)联合建模未来的观测和动作,通常将未来预测为RGB图像。其他视觉模态,如深度、预训练的视觉特征和点轨迹,能够更高效地捕获几何、语义和运动特征。然而,如何在这些WAMs中最佳地组合这些模态仍然是一个开放问题。我们引入了ModAR,这是第一个在预测动作之前以自回归方式去噪多种未来模态的WAM。这使得每个预测都能以先前生成的模态为条件。我们从零开始训练,系统地研究训练数据混合、预测模态和WAM公式如何影响性能。在我们的评估中,WAMs从预测点轨迹、DINO特征和深度图中受益,而额外预测未来RGB并未提供一致的收益。我们还发现,ModAR的序列生成优于现有的WAM公式,在所有评估的数据规模下具有最高的平均成功率。我们还在相同数据上微调了视频模型初始化的WAM Flex-π;ModAR实现了略高的观察平均成功率(75%对72%),同时使用的训练FLOPs约为其1/20,且无需预训练。在三个真实世界的双臂任务上,ModAR优于基线,并随着人类视频而改进。

英文摘要:

World-action models (WAMs) jointly model future observations and actions, typically predicting the future as RGB images. Other visual modalities such as depth, pretrained visual features, and point tracks can more efficiently capture geometric, semantic, and motion features. However, how best to combine these modalities within WAMs remains an open question. We introduce ModAR, the first WAM to autoregressively denoise multiple future modalities before predicting actions. This allows each prediction to condition on previously generated modalities. We train from scratch to systematically study how training-data mixtures, predicted modalities, and WAM formulations affect performance. In our evaluations, WAMs benefit from predicting point tracks, DINO features, and depth maps, while additionally predicting future RGB does not provide a consistent benefit. We also find that ModAR's sequential generation outperforms existing WAM formulations, with the highest average success rate at all evaluated data scales. We also fine-tune the video-model-initialized WAM Flex-$π$ on the same data; ModAR achieves a slightly higher observed average success rate (75% vs. 72%) while using approximately $20\times$ fewer training FLOPs and no pretraining. On three real-world bimanual tasks, ModAR outperforms baselines and improves with human videos.

补充信息

↑