arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

StateMem:用于视觉-语言-动作策略的单状态残差记忆与自适应推理

StateMem: Single-State Residual Memory with Adaptive Inference for Vision-Language-Action Policies

Wenzhuo Li, Qiongfeng Shi, Yi Zhou

arXiv 2609.22684首次发表:更新:

发表机构

Southeast University(东南大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

StateMem提出单状态残差记忆框架,通过预测误差更新记忆标记并自适应路由缓存前缀,在LIBERO、RoboMemArena及真实任务中显著提升成功率并降低刷新率。

AI 中文摘要

依赖记忆的机器人操作通常需要后续动作利用早期交互的信息。现有的视觉-语言-动作(VLA)策略主要依赖当前观测,限制了历史信息的保留。记忆增强的VLA,如MemoryVLA,通过外部记忆库解决了这一限制,但需要显式的存储和检索。为解决这些限制,我们提出了StateMem,一种用于VLA策略的单状态残差记忆框架,它利用预测误差通过低秩残差更新持久记忆标记,并自适应地路由缓存前缀。一个无需训练的控制器在线调整路由阈值,而快速校正则在缓存重用期间补偿过时的前缀特征。我们在LIBERO、RoboMemArena和真实世界操作任务上评估了StateMem。在LIBERO上,StateMem实现了97.6%的平均成功率,并将平均VLM前缀刷新率相对于完全刷新降低了20.25%。在RoboMemArena的遮挡类别中,StateMem在单VLA方法中取得了最佳性能,达到了21.8%的任务成功率(TSR)和44.3%的累积成功率(CSR)。在六个真实世界操作任务中,它实现了平均成功率+21%的提升。

英文摘要

Memory-dependent robotic manipulation often requires later actions to use information from earlier interactions. Existing vision-language-action (VLA) policies primarily rely on current observations, limiting historical information retention. Memory-augmented VLAs, such as MemoryVLA, address this limitation with external memory banks but require explicit storage and retrieval. To address these limitations, we propose StateMem, a single-state residual memory framework for VLA policies that uses prediction error to update a persistent memory token through low-rank residuals and to adaptively route cached prefixes. A training-free controller adjusts the routing threshold online, while fast correction compensates for stale prefix features during cache reuse. We evaluate StateMem on LIBERO, RoboMemArena, and real-world manipulation tasks. On LIBERO, StateMem achieves an average success rate of 97.6% and reduces the average VLM prefix refresh rate by 20.25% relative to full refresh. In the Occlusion category of RoboMemArena, StateMem achieves the best performance among single-VLA methods, reaching 21.8% Task Success Rate (TSR) and 44.3% Cumulative Success Rate (CSR). Across six real-world manipulation tasks, it achieves +21% in average success rate.

Comments9 pages, 5 figures, conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑