发表机构
College of Control Science and Engineering, Zhejiang University(浙江大学控制科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长时程机器人操作中的阶段混淆问题,提出PACE方法,利用进度对齐上下文和因果执行记忆,在多个基准上将成功率显著提升。
AI 中文摘要
基于演示的条件策略为指定机器人行为提供了自然接口,然而当视觉相似状态在不同阶段重复出现,或演示与执行以不同速度进行时,长时程操作仍然困难。我们将由此产生的失败模式识别为阶段混淆,并引入进度对齐执行上下文(PACE),这是一种有状态方法,可根据已实现的执行进度持续重新解释完整演示。PACE将演示压缩为有序的多模态提示词元,并使用仅训练时的双边缘注意力监督来暴露其潜在阶段结构。在执行过程中,一个情节局部的快速权重记忆因果地编码已实现的动作-观察转换,并调制提示交叉注意力,为统一的扩散动作专家生成进度对齐上下文,无需测试时阶段标签或阶段特定策略。PACE在LIBERO-Gen Goal Chain上将成功率从88.9%提升至94.0%,在Spatial Combination上从79.1%提升至83.3%,在两步Block Routing任务上从33.3%提升至73.3%。失败分析进一步表明,结构化演示对齐和因果执行记忆共同缓解了阶段混淆。
英文摘要
Demonstration-conditioned policies provide a natural interface for specifying robot behavior, yet long-horizon manipulation remains difficult when visually similar states recur across different stages or when demonstrations and executions proceed at different speeds. We identify the resulting failure mode as stage confusion and introduce Progress-Aligned Context for Execution (PACE), a stateful method that continually reinterprets a complete demonstration according to realized execution progress. PACE compresses the demonstration into ordered multimodal prompt tokens and uses training-only dual-edge attention supervision to expose its latent stage structure. During execution, an episode-local fast-weight memory causally encodes realized action-observation transitions and modulates prompt cross-attention, producing a progress-aligned context for a unified diffusion action expert without test-time stage labels or stage-specific policies. PACE improves success from 88.9% to 94.0% on LIBERO-Gen Goal Chain, from 79.1% to 83.3% on Spatial Combination, and from 33.3% to 73.3% on the two-step Block Routing tasks. Failure analysis further indicates that structured demonstration alignment and causal execution memory jointly mitigate stage confusion.