arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11308cs.ROcs.AI

2AM:将智能体侧记忆作为长时程操作中可引导动作模型的引导

2AM: Grounding Agent-Side Memory as Guidance for Steerable Action Models in Long-Horizon Manipulation

Yutong Hu, Fengjiao Chen, Xuezhi Cao, Renaud Detry

首次发表
浏览论文内容

中文总结 AI 辅助

2AM通过让多模态智能体持有任务记忆并以结构化提示引导无状态动作模型,在LIBERO-Mem上以76.3%完成率大幅超越基线,证明记忆可留在智能体侧且引导精度至关重要。

中文摘要 AI 辅助

长时程机器人操作需要记忆,但记忆不一定非要存储在动作策略内部。为了应对此类任务,当前的智能体系统通常将视觉-语言-动作模型(VLA)与规划器和几何工具相结合,有时还会使用额外的深度信息或经过标定的几何信息。这些系统混淆了归因:性能提升可能来自更丰富的观测或替代的运动工具,而失败可能源于策略本身或语言接口描述不充分。我们通过一种刻意受限的设计来隔离这一问题:减少工具种类,但增加接口带宽。2AM让多模态智能体成为任务记忆的唯一持有者,并让一个基于RGB的、情节性无状态的动作模型成为任务相关运动的唯一执行者。智能体将交互历史编译为子任务语言以及可选的2D抓取、放置和移动提示,这些提示在不同时间尺度上绑定其物理意图。为了教会视觉-语言-动作模型(VLA)这种可引导性,我们为演示数据增加了结构化的提示标签,并在条件丢弃、空间噪声和时间抖动下进行训练,以容忍不完美的智能体输出。在LIBERO-Mem基准上,在没有深度信息、在线几何或基于规划器的物体运动的情况下,2AM达到了76.3%的平均完成率,比最强报告基线14.8%提高了61.5个百分点,同时实现了63.0%的宽松成功率和11.8%的严格成功率。这些结果表明,任务记忆可以保留在智能体侧。它们进一步表明,动作模型的能力不仅取决于策略学到了什么,还取决于智能体引导它的精确程度。

英文摘要

Long-horizon robot manipulation requires memory, but not necessarily inside the action policy. To address such tasks, current agentic systems often combine VLAs with planners and geometric tools, sometimes using additional depth or calibrated geometry. These systems confound attribution: gains may come from richer observations or alternative motor tools, while failures may stem from either the policy or an under-specified language interface. We isolate this question through a deliberately constrained design: less tool breadth, but greater interface bandwidth. 2AM makes a multimodal Agent the sole holder of task memory and a single RGB-based, episodically stateless Action Model the sole executor of task-relevant motion. The Agent compiles interaction history into subtask language and optional 2D grasp, place, and move hints that bind its physical intention at different time scales. To teach this steerability to the VLA, we augment demonstrations with structured hint labels and train under condition dropout, spatial noise, and temporal jitter to tolerate imperfect Agent outputs. On LIBERO-Mem, without depth, online geometry, or planner-based object motion, 2AM reaches 76.3% average completion, a 61.5-point improvement over the strongest reported baseline of 14.8%, together with 63.0% relaxed and 11.8% strict success. These results show that task memory can remain Agent-side. They further show that Action Model capability depends not only on what the policy has learned, but on how precisely the Agent can steer it.

发表机构

  • KU Leuven(鲁汶大学)
  • Meituan Inc.(美团公司)

机构由 AI 辅助整理,请以论文原文为准。

↑