发表机构
Harbin Institute of Technology; The University of Sydney; StellarEdge AI(哈尔滨工业大学; 悉尼大学; 星边人工智能)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对冻结VLA策略无法跟踪任务进度的问题,提出AGM框架,通过物理证据验证子目标实现闭环,在RoboMME基准和物理机器人上均取得显著性能提升。
AI 中文摘要
冻结的视觉-语言-动作(VLA)策略可提供广泛的操作技能,但执行开环动作块时不跟踪任务进度,导致智能体无法可靠决定是继续、重试还是终止。外部记忆是自然的解决方案,但如果将尝试的动作视为已完成进度,会将局部执行错误转化为持续的任务状态错误。我们提出基于成就的记忆(Achievement-Grounded Memory,AGM),这是一种适用于冻结VLA策略的轻量级闭环框架,它将任务表示为带有进度指针的子目标序列,仅在当前子目标经物理证据验证后才更新此记忆。本体感觉交互线索决定何时进行验证,而连贯点跟踪和语言条件跨视角比较(通过一个仅243万参数的验证头从冻结基础模型获取)决定已完成的内容。AGM由此将开环执行转化为执行、验证和进度的闭环,在测试时无需大模型推理即可保持策略冻结。在RoboMME计数基准上,AGM在PickXTimes和BinFill任务上的表现平均比最强的记忆增强基线高出多个百分点,该框架在物理机器人上也取得了同样显著的增益。因此,可靠的具身记忆更多依赖于严格的状态更新,而非记忆容量。
英文摘要
Frozen vision-language-action (VLA) policies offer broad manipulation skills but execute open-loop action chunks without tracking task progress, so the agent cannot reliably decide whether to continue, retry, or terminate. External memory is a natural remedy, yet we find it can be harmful when attempted actions are recorded as completed progress: transient execution failures become persistent task-state errors, and such a memory can underperform no progress memory at all. We propose Achievement-Grounded Memory (AGM), a lightweight closed-loop framework for frozen VLA policies. AGM represents a task as a static subgoal sequence with a dynamic progress pointer and advances the pointer on physically verified achievement rather than on attempts. Proprioceptive gripper-load cues decide when to verify; coherent point tracking verifies grasps, and language-conditioned cross-view comparison, read by a single trained 2.43M-parameter verification head, verifies placements. The policy, tracker, and encoder remain frozen, the head is the only trained component, and deployment needs no auxiliary vision-language model. On the RoboMME Counting benchmark, AGM reaches 100.0% on PickXTimes and 84.0% on BinFill, surpassing the strongest memory-augmented baseline by 7.7 points on the four-task average, and the gains carry over to a physical robot, where AGM reaches 100.0% and 82.0%. These results suggest that reliable embodied memory depends more on disciplined state updates than on memory capacity.
Comments26 pages, 9 figures