发表机构
The University of Sydney(悉尼大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出寄存器路由延迟融合(RRDF),通过屏蔽直接紧凑-视觉注意力并利用寄存器工作空间分阶段跨模态交互,控制视觉运动模仿中的信息流,在模拟和真实任务上匹配或优于密集ACT,并提升外观偏移鲁棒性。
AI 中文摘要
视觉运动模仿策略将高维视觉观测与本体感觉等紧凑信号相结合,其融合拓扑结构决定了这些信息流在何时以及通过哪些令牌进行交互。在密集令牌融合中,视觉令牌可能从第一个编码器层直接关注紧凑令牌,使得用于动作预测的紧凑线索能够在视觉空间表征形成的早期阶段影响其形成。我们探究控制这一路由是否能改善视觉响应性和策略行为。我们提出了寄存器路由延迟融合(RRDF),该方法屏蔽了紧凑-视觉直接注意力,并通过一个学习到的寄存器工作空间分阶段进行跨模态交互。其隔离-收集-路由调度保护了早期流分离的前缀,并随后仅允许通过寄存器介导的交换,同时紧凑条件信息仍可供原生动作生成器使用。在五个模拟任务和三个真实机器人任务中,RRDF在标称条件下匹配或优于密集ACT。四个模拟任务上的外观偏移评估和三个真实机器人任务上的保持位置评估也倾向于RRDF。相位匹配输入探针显示测得的态到图像敏感性较低,而消融实验表明仅添加寄存器并不能完全复现全部性能提升。这些结果支持在保留紧凑动作条件信息的同时控制跨模态传播。
英文摘要
Visuomotor imitation policies combine high-dimensional visual observations with compact signals such as proprioception, and their fusion topology determines when and through which tokens these streams interact. In dense token fusion, visual tokens may attend directly to compact tokens from the first encoder layer, allowing action-predictive compact cues to influence spatial visual representations early in their formation. We ask whether controlling this route improves visual responsiveness and policy behavior. We introduce Register-Routed Delayed Fusion (RRDF), which masks direct compact-visual attention and stages cross-modal interaction through a learned register workspace. Its isolate-collect-route schedule protects an early stream-separated prefix and later permits only register-mediated exchange, while compact conditioning remains available to the native action generator. Across five simulation tasks and three real-robot tasks, RRDF matches or improves dense ACT under nominal conditions. Appearance-shift evaluations on four simulation tasks and held-out-position evaluations on three real-robot tasks also favor RRDF. Phase-matched input probes show lower measured state-to-image sensitivity, while ablations indicate that adding registers alone does not reproduce the full performance gain. These results support controlling cross-modal propagation while retaining compact action conditioning.