用于统一世界建模的掩码视觉动作
Masked Visual Actions for Unified World Modeling
浏览论文内容
中文总结 AI 辅助
研究如何将动作传达给视频模型用于机器人世界建模,提出掩码视觉动作这一像素空间控制接口,经微调后单个检查点在多场景和实施例中表现出色,在下游操作中能辅助策略评估、改进决策和支持逆建模。
中文摘要 AI 辅助
视频模型吸收了关于视觉世界如何移动、交互以及对接触做出反应的丰富先验知识,使其成为机器人世界建模的有前途的基础。核心挑战在于如何以与它们学习这些交互先验知识的视觉空间对齐的形式将动作传达给此类模型,同时仍基于物理操作。我们引入了掩码视觉动作,这是一种像素空间控制接口,将动作表示为视频中任意实体的部分显示轨迹。揭示机器人运动使模型充当预测场景对低级机器人动作响应的前向动力学模型,而揭示所需对象运动使同一模型恢复与该结果一致的机器人行为。仅用来自真实视频和模拟的15小时掩码示例进行微调,单个检查点就能在不同场景和多个实施例中实现强大的视觉保真度和可控性。在下游操作设置中,该模型生成想象的展开,其结果与用于策略评估的实际世界执行相关,通过在基于模型的规划中对候选未来进行排名来改进决策,并通过从所需对象运动合成机器人运动来支持逆建模。
英文摘要
Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challenge is how to communicate action to such models in a form aligned with the visual space in which they learned these interaction priors, yet still grounded in physical manipulation. We introduce Masked Visual Actions, a pixel-space control interface that expresses action as a partially revealed trajectory of an arbitrary entity in a video. Revealing robot motion makes the model act as a forward dynamics model that predicts the scene's response to low-level robot actions, while revealing desired object motion makes the same model recover robot behavior consistent with that outcome. Finetuned with only 15 hours of masked examples from real videos and simulation, a single checkpoint achieves strong visual fidelity and controllability across diverse scenes and multiple embodiments. In downstream manipulation settings, the model produces imagined rollouts whose outcomes correlate with real-world execution for policy evaluation, improves decision making by ranking candidate futures in model-based planning, and supports inverse modeling by synthesizing robot motion from desired object motion.
发表机构
- Stanford University(斯坦福大学)
- University of Maryland, College Park(马里兰大学帕克分校)
- Harvard University(哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。