arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2606.23589cs.RO

KEMO: 面向长程机器人操作的VLA策略事件驱动关键帧记忆

KEMO: Event-Driven Keyframe Memory for Long-Horizon Robot Manipulation with VLA Policies

  • Hong Kong Embodied AI Lab(香港具身智能实验室)
  • The Chinese University of Hong Kong(香港中文大学)
  • University of Electronic Science and Technology of China(电子科技大学)
  • Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

Yihan Zeng, Minghao Ye, Yiyuan Chen, Yide Shentu, Philipp Wu, Zike Yan, Zhongyu Li

AI总结:

提出KEMO,一种轻量级插件式记忆框架,通过事件驱动选择关键帧并编码为紧凑记忆令牌,结合交叉注意力和门控残差融合,提升VLA策略在长程操作中的任务成功率23.6%。

AI中文摘要:

长程机器人操作仍然具有挑战性,因为相似的观察可能出现在不同的执行阶段,而适当的动作取决于先前完成的操作。记忆可以通过使策略从执行历史中推断任务进度来解决这种歧义。然而,现有的记忆增强方法通常要么保留需要压缩的密集历史,要么主要依赖可能丢弃早期任务相关事件的近期上下文。在这项工作中,我们提出了KEMO,一个轻量级插件式记忆框架,自动选择性保留与任务相关状态变化的关键帧,用于VLA策略。KEMO结合机器人运动学和视觉过滤来检测事件,将选定的关键帧编码为紧凑的时间排序记忆令牌,并通过交叉注意力和门控残差融合与当前视觉特征集成,用于VLA训练。检测到的事件还定义了关键转换附近的高权重训练样本。我们在各种真实世界双臂操作任务上评估KEMO,这些任务包含2到6个计分子任务,轨迹长度从830步到2846执行步(持续时间从28到95秒)。与无记忆基线(例如,$\pi_{0.5}$)相比,KEMO将总任务成功率提高了23.6%,阶段完成率提高了34.1%。消融实验表明,事件驱动的关键帧选择优于均匀采样和最近帧保留,而提出的门控融合和关键帧对齐损失加权提供了互补的增益。

英文摘要:

Long-horizon robot manipulation remains challenging because similar observations may occur at different execution stages, while the appropriate action depends on previously completed operations. Memory can address this ambiguity by enabling policies to infer task progress from execution history. However, existing memory-augmented approaches often either retain dense histories that require compression or rely primarily on recent context that may discard earlier task-relevant events. In this work, we propose propose KEMO, a lightweight plug-in memory framework that automatically selectively preserves keyframes associated with task-relevant state changes for VLA policies. KEMO combines robot kinematics with visual filtering to detect events, encodes the selected keyframes as compact temporally ordered memory tokens, and integrates them with current visual features through cross-attention and gated residual fusion for VLA training. The detected events also define higher-weight training samples near critical transitions. We evaluate KEMO on various real-world dual-arm manipulation tasks spanning 2 to 6 scored subtasks, and trajectory length ranging from 830 steps to 2846 execution steps (durations from 28 to 95 seconds). Compared with the memory-free baseline (e.g., $π_{0.5}$), KEMO improves aggregate Task Success Rate by 23.6\% and Stage Completion Rate by 34.1\%. Ablations show that event-driven keyframe selection outperforms uniform sampling and recent-frame retention, while the proposed gated fusion and keyframe-aligned loss weighting provide complementary gains.

↑