arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27314cs.RO

CoRe-WAM:用于世界动作模型的对应对齐时间残差

CoRe-WAM: Correspondence-Aligned Temporal Residuals for World Action Models

Bin Zhou, Jialong Liu, Jianan Wang, Changhao Chen, Kani Chen

首次发表
浏览论文内容

中文总结 AI 辅助

CoRe-WAM通过对应对齐的时间残差模块,在冻结预训练骨干上仅优化159万参数,使策略利用历史视觉变化,在50个RoboTwin 2.0任务上以5,000次更新达到92.22%的干净成功率,比基线Motus高3.56个百分点。

中文摘要 AI 辅助

比较当前和过去的观测有助于机器人理解场景变化并在操作过程中选择后续动作。然而,当物体或相机移动时,比较同一图像位置的视觉特征可能会混合不同的场景内容。我们引入了CoRe-WAM,一种世界动作模型,通过参数高效的时间接口整合对应对齐的视觉变化。其TraceDelta模块利用冻结跟踪模型的对应关系,将历史视觉特征传输到当前位置,然后在共享的预训练特征空间中计算符号差异。因此,对应关系决定了哪些历史内容与当前进行比较,而不是作为单独的轨迹表示进入策略。一个轻量级适配器将这些差异转换为有效性门控残差,补充当前的视觉条件,使策略能够利用近期变化以及当前场景信息。基于Motus构建,CoRe-WAM保持预训练骨干权重冻结,并优化159万个参数。在5,000次更新的适应预算下,CoRe-WAM在50个RoboTwin 2.0任务中实现了92.22%的干净成功率,比Motus高出3.56个百分点;在随机化评估中,它实现了89.60%的成功率,提升了2.58个百分点。将TraceDelta集成到基于StarVLA的策略中,将干净成功率从58.10%提高到67.62%,支持了时间接口在Motus之外的迁移。

英文摘要

Comparing current and past observations helps robots understand scene changes and select subsequent actions during manipulation. However, comparing visual features at the same image location can mix different scene content when objects or the camera move. We introduce CoRe-WAM, a world-action model that incorporates correspondence-aligned visual changes through a parameter-efficient temporal interface. Its TraceDelta module uses correspondences from a frozen tracking model to transport historical visual features to current locations before computing signed differences in a shared pretrained feature space. Correspondence thus determines which historical content is compared with the present, rather than entering the policy as a separate trajectory representation. A lightweight adapter converts these differences into validity-gated residuals that supplement current visual conditioning, allowing the policy to use recent changes alongside current-scene information. Built on Motus, CoRe-WAM keeps the pretrained backbone weights frozen and optimizes 1.59 million parameters. With a 5,000-update adaptation budget, CoRe-WAM achieves 92.22% clean success across 50 RoboTwin 2.0 tasks, 3.56 percentage points above Motus; on randomized evaluation, it achieves 89.60% success, a 2.58-point gain. Integrating TraceDelta into a StarVLA-based policy improves clean success from 58.10% to 67.62%, supporting transfer of the temporal interface beyond Motus.

发表机构

  • The Hong Kong University of Science and Technology(香港科技大学)
  • Astribot
  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

↑