OnEvoMemory:通过机器人在线试跑演化预训练机器人策略的记忆机制
OnEvoMemory: Evolving Memory through Online Robot Rollouts for Pretrained Robot Policies
浏览论文内容
中文总结 AI 辅助
针对现有机器人记忆机制依赖外部模型的问题,提出OnEvoMemory价值引导记忆模块,结合离线初始化与在线试跑优化记忆选择,提升了基础VLA策略在长时程机器人操作中的性能。
中文摘要 AI 辅助
长时程机器人操作要求策略能跟踪已完成的子任务和关键交互事件,但现有记忆机制严重依赖外部模型或预定义更新规则。为解决该问题,我们提出OnEvoMemory,一种面向预训练机器人策略的价值引导记忆模块。该模块维护近期上下文、高价值经验和显著转移,同时从轨迹结果中学习应保留的经验。离线演示初始化记忆先验,而成功与失败的在线试跑则优化记忆选择,帮助策略识别任务阶段转移并避免重复已完成的子任务。在长时程操作基准上的实验表明,OnEvoMemory通过离线初始化和在线记忆演化提升了基础VLA策略的性能。
英文摘要
Long-horizon robot manipulation requires policies to track completed subtasks and critical interaction events. However, existing memory mechanisms heavily rely on external models or predefined update rules. To address this, we propose OnEvoMemory, a value-guided memory module for pretrained robot policies. It maintains recent context, high-value experiences, and salient transitions, while learning which experiences should be retained from trajectory outcomes. Offline demonstrations initialize the memory prior, whereas successful and unsuccessful online rollouts refine memory selection, helping the policy recognize task-stage transitions and avoid repeating completed subtasks. Experiments on long-horizon manipulation benchmarks show that OnEvoMemory improves the performance of the base VLA policy through both offline initialization and online memory evolution.
发表机构
- Shanghai Jiao Tong University(上海交通大学)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。