AI 中文总结
研究针对现有世界行动模型的局限,提出WorldScape Policy 2.0,用推理增强记忆,构建相关数据集,实现从高级指令自主规划及基于细粒度文本等的可控执行,在模拟和现实平台实验中展现出在多方面的卓越能力。
AI 中文摘要
世界行动模型(WAMs)为机器人操作提供了一个有前景的范例,通过联合建模视觉状态转换和机器人动作。然而,现有WAMs受限于有限的时间上下文、粗略的情节级语言监督和主要基于文本的条件设定,阻碍了任务进展跟踪等。本文引入WorldScape Policy 2.0,一种具有推理增强长短期记忆的可控WAM。其因果短期视觉记忆提供近期观察,长短期事件记忆组织历史VLM输出。还构建了ManipEvent-5M数据集。实验表明该模型在长期自主规划等方面有卓越能力。
英文摘要
World Action Models (WAMs) offer a promising paradigm for robotic manipulation by jointly modeling visual state transitions and robot actions. However, existing WAMs are constrained by limited temporal context, coarse episode-level language supervision, and predominantly text-only conditioning, which hinder task-progress tracking and fine-grained language-video-action grounding while limiting visual-context reasoning and cross-embodiment transfer. In this paper, we introduce WorldScape Policy 2.0, a controllable WAM with reasoning-augmented long short-term memory. Its causal short-term visual memory supplies recent observations as DiT prefill to preserve local interaction dynamics, while its long short-term event memory organizes historical VLM outputs into global-history, local-active, and event-boundary representations for progress-aware retrieval. The retrieved history augments perception and autoregressively generated planning tokens, yielding an implicit subgoal condition for autonomous planning; semantic forcing further transfers event-level instruction semantics into this latent planning pathway. To establish fine-grained multimodal controllability, we construct ManipEvent-5M, an event-grounded embodied pretraining dataset containing nearly 5 million event segments with aligned action trajectories, episode-level task instructions, segment-level subtask captions, goal images, and video demonstrations. These designs provide a unified interface for autonomous planning from high-level instructions and controllable execution from fine-grained text, goal-image, or video-context prompts. Experiments in both simulation and real-world platforms demonstrate superior capabilities in long-horizon autonomous planning, fine-grained instruction following and in-context adaptation.