arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WorldScape Policy 2.0:通过推理增强记忆实现可控的世界行动建模

WorldScape Policy 2.0: Empowering Steerable World Action Modeling with Reasoning-Augmented Memory and In-Context Learning

Haisheng Su, Zongdai Liu, Xin Jin, Haoxuan Dou, Chengming Hu, Baorun Li, Zhanwang Liu, Ruiyan Xu, Jianjie Fang, Xin Zhang, Zhenjie Yang, Xue Yang, Chen Gao, Junchi Yan, Yong Li, Wei Wu

arXiv 2607.18840首次发表:更新:

AI 中文总结

研究针对现有世界行动模型的局限,提出WorldScape Policy 2.0,用推理增强记忆,构建相关数据集,实现从高级指令自主规划及基于细粒度文本等的可控执行,在模拟和现实平台实验中展现出在多方面的卓越能力。

AI 中文摘要

世界行动模型(WAMs)为机器人操作提供了一个有前景的范例,通过联合建模视觉状态转换和机器人动作。然而,现有WAMs受限于有限的时间上下文、粗略的情节级语言监督和主要基于文本的条件设定,阻碍了任务进展跟踪等。本文引入WorldScape Policy 2.0,一种具有推理增强长短期记忆的可控WAM。其因果短期视觉记忆提供近期观察,长短期事件记忆组织历史VLM输出。还构建了ManipEvent-5M数据集。实验表明该模型在长期自主规划等方面有卓越能力。

英文摘要

World Action Models (WAMs) offer a promising paradigm for robotic manipulation by jointly modeling visual state transitions and robot actions. However, existing WAMs are constrained by limited temporal context, coarse episode-level language supervision, and predominantly text-only conditioning, which hinder task-progress tracking and fine-grained language-video-action grounding while limiting visual-context reasoning and cross-embodiment transfer. In this paper, we introduce WorldScape Policy 2.0, a controllable WAM with reasoning-augmented long short-term memory. Its causal short-term visual memory supplies recent observations as DiT prefill to preserve local interaction dynamics, while its long short-term event memory organizes historical VLM outputs into global-history, local-active, and event-boundary representations for progress-aware retrieval. The retrieved history augments perception and autoregressively generated planning tokens, yielding an implicit subgoal condition for autonomous planning; semantic forcing further transfers event-level instruction semantics into this latent planning pathway. To establish fine-grained multimodal controllability, we construct ManipEvent-5M, an event-grounded embodied pretraining dataset containing nearly 5 million event segments with aligned action trajectories, episode-level task instructions, segment-level subtask captions, goal images, and video demonstrations. These designs provide a unified interface for autonomous planning from high-level instructions and controllable execution from fine-grained text, goal-image, or video-context prompts. Experiments in both simulation and real-world platforms demonstrate superior capabilities in long-horizon autonomous planning, fine-grained instruction following and in-context adaptation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑