arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11561cs.RO

记忆即计划:基于记忆锚定规划的世界-动作建模

Memory as Plans: World-Action Modeling with Memory-Grounded Planning

Sizhe Zhao, Haozhe Xie, Weiyu Zhao, Chenchu Zhang, Huan Wang, Chenyang Wang, Qinglin Liu, Shengping Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

针对机器人操作中的非马尔可夫长时程记忆问题,提出MaP-WAM框架,将记忆转化为紧凑计划,通过世界-动作-进度模型执行,在RMBench和真实机器人上取得领先性能。

中文摘要 AI 辅助

主流机器人策略通常采用马尔可夫公式,但许多复杂的现实世界操作任务本质上是非马尔可夫的,需要超出当前观测的长时程记忆。现有的记忆机制通常依赖语言摘要、增长的视觉窗口或其组合,因此可能丢失细粒度的视觉证据,或面临历史覆盖范围与执行效率之间的权衡。我们提出了MaP-WAM,一个“记忆即计划”框架,将依赖记忆的世界-动作建模分解为记忆锚定规划和计划条件执行,并利用长期多模态情节上下文作为规划时的证据,而不是反复用完整历史来调节执行器。MaP-WAM将记忆表示为包含语言指令和稀疏视觉上下文的已完成片段记录,并将这种情节记忆转换为包含下一片段级语言计划和相应视觉引导的紧凑计划。一个世界-动作-进度(WAP)模型在推理时通过联合预测动作块和相应的执行进度来执行每个计划,持续时间未知,通过计划-观测对齐校准预测进度,以实现自适应片段转换和闭环上下文更新。MaP-WAM保持执行器上下文长度固定,而结构化注意力进一步支持规划和执行中的键值缓存。MaP-WAM在RMBench上实现了最先进的性能,成功率达83.3%,在真实机器人任务上达到78.0%的成功率,同时随着任务历史的增长,执行器推理延迟保持近似恒定。

英文摘要

Mainstream robotic policies often adopt a Markovian formulation, but many complex real-world manipulation tasks are inherently non-Markovian, requiring long-horizon memory beyond the current observation. Existing memory mechanisms often rely on language summaries, growing visual windows, or their combinations, and may therefore lose fine-grained visual evidence or face a trade-off between history coverage and execution efficiency. We introduce MaP-WAM, a Memory-as-Plans framework that decomposes memory-dependent world-action modeling into memory-grounded planning and plan-conditioned execution, and uses long-term multimodal episodic context as planning-time evidence rather than repeatedly conditioning the executor on the full history. MaP-WAM represents memory as completed segment records containing language instructions and sparse visual context, and converts this episodic memory into compact plans comprising the next segment-level language plan and corresponding visual guidance. A World-Action-Progress (WAP) model executes each plan over an unknown duration by jointly predicting action chunks and corresponding execution progress at inference time, calibrating predicted progress through plan-observation alignment for adaptive segment transitions and closed-loop context updates. MaP-WAM keeps the executor context length fixed, while structured attention further enables key-value caching in both planning and execution. MaP-WAM achieves state-of-the-art performance on RMBench with an 83.3% success rate and attains 78.0% success on real-robot tasks, while maintaining approximately constant executor inference latency as task history grows.

发表机构

  • Harbin Institute of Technology(哈尔滨工业大学)
  • Nanyang Technological University(南洋理工大学)
  • Shandong University(山东大学)
  • Harbin Institute of Technology (Weihai) Qingdao Research Institute(哈尔滨工业大学(威海)青岛研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑