发表机构
Nanyang Technological University; Griffin Labs; École Centrale de Lyon(南洋理工大学; 格里芬实验室; 里昂中央理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MemBodied提出固定大小的循环联想记忆,通过关联状态和回合锚点保留历史信息,在RMBench和LIBERO-Long任务上显著提升成功率,且参数开销小。
AI 中文摘要
视觉-语言-动作模型为通用机器人控制提供了坚实基础,然而绝大多数策略并不保留和利用当前观测之外的回合级信息。这一限制在依赖于仅存在于过去观测中的信息的历史依赖操作任务中尤为关键。在上下文中保留过去的观测有助于恢复这些信息,但代价是上下文不断膨胀、推理延迟显著增加。为此,我们提出了MemBodied,一种固定大小的回合记忆,包含两个互补组件:一个关联状态,记录跨策略调用的交互;以及一个回合锚点,保留初始场景的紧凑表示作为参考。在每次策略调用时,模型基于当前输入和记忆组件来生成动作,而非直接使用过去的观测。在需要记忆的五个RMBench任务中,MemBodied的平均成功率是无状态策略的7.81倍,是普通循环记忆的2.98倍,同时以少10倍的额外参数超越了最强记忆增强基线1.3倍。在完全可观测的LIBERO-Long套件上,它达到了90.6%,比无状态π0策略提高了5.4%。这些发现支持MemBodied作为扩展策略上下文用于历史依赖操作的一种实用替代方案。
英文摘要
Vision-Language-Action models provide a strong foundation for general-purpose robot control, yet a vast majority of policies do not preserve and leverage episode-level information beyond the current observation. This limitation is consequential in history-dependent manipulation tasks that depend on information available only in past observations. Retaining past observations in context can aid in recovering this information, but at the significant cost of ever-growing, bloated context and inference latency. We thus introduce MemBodied, a fixed-size episodic memory with two complementary components: an associative state that records interactions across policy calls and an episode anchor that preserves a compact representation of the initial scene as a reference. At each policy call, the model conditions action generation on the current input and the memory components, rather than directly using past observations. Across five evaluated RMBench tasks requiring memory, MemBodied achieves $7.81\times$ the mean success rate of a stateless policy and $2.98\times$ of vanilla recurrent memory, while outperforming the strongest memory-augmented baseline by $1.3\times$ with $10\times$ fewer added parameters. On the fully observable LIBERO-Long suite, it reached 90.6%, a 5.4% improvement over the stateless $π_0$ policy. These findings support MemBodied as a practical alternative to expanding the policy context for history-dependent manipulation.