发表机构
Southeast University(东南大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
StateMem提出单状态残差记忆框架,通过预测误差更新记忆标记并自适应路由缓存前缀,在LIBERO、RoboMemArena及真实任务中显著提升成功率并降低刷新率。
AI 中文摘要
依赖记忆的机器人操作通常需要后续动作利用早期交互的信息。现有的视觉-语言-动作(VLA)策略主要依赖当前观测,限制了历史信息的保留。记忆增强的VLA,如MemoryVLA,通过外部记忆库解决了这一限制,但需要显式的存储和检索。为解决这些限制,我们提出了StateMem,一种用于VLA策略的单状态残差记忆框架,它利用预测误差通过低秩残差更新持久记忆标记,并自适应地路由缓存前缀。一个无需训练的控制器在线调整路由阈值,而快速校正则在缓存重用期间补偿过时的前缀特征。我们在LIBERO、RoboMemArena和真实世界操作任务上评估了StateMem。在LIBERO上,StateMem实现了97.6%的平均成功率,并将平均VLM前缀刷新率相对于完全刷新降低了20.25%。在RoboMemArena的遮挡类别中,StateMem在单VLA方法中取得了最佳性能,达到了21.8%的任务成功率(TSR)和44.3%的累积成功率(CSR)。在六个真实世界操作任务中,它实现了平均成功率+21%的提升。
英文摘要
Memory-dependent robotic manipulation often requires later actions to use information from earlier interactions. Existing vision-language-action (VLA) policies primarily rely on current observations, limiting historical information retention. Memory-augmented VLAs, such as MemoryVLA, address this limitation with external memory banks but require explicit storage and retrieval. To address these limitations, we propose StateMem, a single-state residual memory framework for VLA policies that uses prediction error to update a persistent memory token through low-rank residuals and to adaptively route cached prefixes. A training-free controller adjusts the routing threshold online, while fast correction compensates for stale prefix features during cache reuse. We evaluate StateMem on LIBERO, RoboMemArena, and real-world manipulation tasks. On LIBERO, StateMem achieves an average success rate of 97.6% and reduces the average VLM prefix refresh rate by 20.25% relative to full refresh. In the Occlusion category of RoboMemArena, StateMem achieves the best performance among single-VLA methods, reaching 21.8% Task Success Rate (TSR) and 44.3% Cumulative Success Rate (CSR). Across six real-world manipulation tasks, it achieves +21% in average success rate.
Comments9 pages, 5 figures, conference