发表机构
Xidian University(西安电子科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出即插即用模块 RS,通过双分支结构压缩视觉历史并组织经验,适配 pi0 后显著提升机器人任务成功率,在真实机器人实验中表现优异。
AI 中文摘要
长 horizon 机器人策略需要在不扩展视觉-语言-动作(VLA)上下文的情况下,紧凑访问近期观测结果和可复用经验。我们提出 Remember Smarter(RS),即一种即插即用模块,包含互补的视觉历史和双曲经验记忆分支。其视觉分支使用双向空间 Mamba 和因果时间 Mamba 压缩多视图补丁历史,再通过残差交叉注意力将生成的记忆暴露给面向动作的隐藏状态,同时保持 VLM 视觉令牌流不变。其经验分支将成功的 VLM 最终层状态存储在 Poincare VAE 空间中,按层次组织,并异步将检索到的经验转换为测地线提示令牌,且不阻塞动作推理。当适配到 pi0 时,RS 将 LIBERO-Plus 的总成功率从 53.6% 提升至 70.6%,并在旨在评估记忆保留和经验利用的真实机器人实验中取得显著性能提升。
英文摘要
Long-horizon robot policies require compact access to recent observations and reusable experience without expanding the vision-language-action (VLA) context. We introduce Remember Smarter (RS), a plug-and-play module with complementary visual-history and hyperbolic experience-memory branches. Its visual branch compresses multi-view patch histories using bidirectional spatial Mamba and causal temporal Mamba, then exposes the resulting memory to action-facing hidden states through residual cross-attention while leaving the VLM visual-token stream unchanged. Its experience branch stores successful final-layer VLM states in a Poincare VAE space, organizes them hierarchically, and asynchronously converts retrieved experience into geodesic prompt tokens without blocking action inference. When adapted to pi0, RS increases total success on LIBERO-Plus from 53.6% to 70.6% and achieves substantial performance gains in real-robot experiments designed to evaluate memory retention and experience utilization.
Comments19 pages, 7 pages