EvoMem-VLA:用于长时程机器人操作的状态演化记忆
EvoMem-VLA: State-Evolution Memory for Long-Horizon Robot Manipulation
浏览论文内容
中文总结 AI 辅助
针对VLA模型在长时程任务中丢失历史信息的问题,提出EvoMem-VLA,通过状态演化记忆编码历史变化,在仿真和真实任务上显著提升成功率。
中文摘要 AI 辅助
大多数视觉-语言-动作(VLA)模型依赖当前观测,一旦任务相关信息离开视野便会丢失,从而限制了在长时程、依赖记忆的任务上的性能。现有工作尝试融入压缩的历史特征或稀疏的视觉关键帧。然而,孤立的快照可能使策略对过去交互中发生了什么变化以及接下来应执行哪个动作感到不确定。为克服这一局限,我们提出EvoMem-VLA,通过显式编码并保留历史状态之间观测到的变化来构建状态演化记忆。这些变化表示保留了交互结果的证据,使策略能够超越孤立快照追踪任务进展。具体而言,我们引入条件增量标记化,将有序帧对编码为有方向的、源条件化的增量标记,每个标记关联其对应的状态证据。共享的VLM骨干支持任务自适应路由:普通长时程任务遵循直接动作路径,而多阶段任务使用子任务路径,生成可执行的子任务作为动作生成的额外输入。凭借为每个仿真基准单独联合训练的策略,EvoMem-VLA在RMBench上达到80.7%的成功率,在RoboMME上达到82.0%,并在涵盖两种机器人本体的四项真实世界任务中达到83.8%。这些结果在全部三个评估设置中均显著优于先前最先进水平。
英文摘要
Most vision-language-action (VLA) models rely on current observations and lose task-relevant evidence once it leaves view, limiting performance on long-horizon, memory-dependent tasks. Existing efforts incorporate compressed historical features or sparse visual keyframes. However, isolated snapshots can leave the policy uncertain about what changed during past interactions and which action should follow. To overcome this limitation, we propose EvoMem-VLA, which constructs state-evolution memory by explicitly encoding and retaining observed changes between historical states. These change representations preserve evidence of interaction outcomes, allowing the policy to track task progress beyond isolated snapshots. Specifically, we introduce conditional delta tokenization to encode ordered frame pairs into directional, source-conditioned delta tokens, each associated with its corresponding state evidence. A shared VLM backbone supports task-adaptive routing: normal long-horizon tasks follow a direct action route, whereas multi-stage tasks use a subtask route that generates an executable subtask as an additional input for action generation. With a single jointly trained policy for each simulation benchmark, EvoMem-VLA achieves success rates of 80.7\% on RMBench, 82.0\% on RoboMME and 83.8\% across four real-world tasks spanning two robot embodiments. These results represent substantial improvements over the previous state of the art in all three evaluation settings.
发表机构
- Xi’an Jiaotong University(西安交通大学)
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
- The Chinese University of Hong Kong(香港中文大学)
- Xbotics Embodied AI Community(Xbotics具身智能社区)
机构由 AI 辅助整理,请以论文原文为准。