发表机构
Shenyang Institute of Automation, Chinese Academy of Sciences; Mohamed bin Zayed University of Artificial Intelligence; Anhui University; Xiaomi Corporation; Fudan University; University of Trento(中国科学院沈阳自动化研究所; 穆罕默德·本·扎耶德人工智能大学; 安徽大学; 小米公司; 复旦大学; 特伦托大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出动作经验字典(AED),通过编码历史动作轨迹为共享嵌入,支持技能复用和跨任务关系建模,并引入运动感知过渡损失,在仿真和真实世界实验中验证了其有效性。
AI 中文摘要
世界动作模型(WAMs)将视觉动态预测与动作生成相结合,然而它们并未明确支持跨操作任务的动作经验复用。此外,现有的WAMs难以捕捉潜在的跨任务语义关系,这些关系本可指导目标动作预测,因为冗余的背景元素干扰了关键视觉信息的提取。为解决这些挑战,我们开发了一种新颖的动作经验字典(AED),它将历史物理动作轨迹编码为共享的动作嵌入,以支持技能复用和跨任务关系建模。具体而言,我们首先聚合历史动作以与视觉观察对齐,并使用预训练的动作分词器从AED中检索动作嵌入。随后,我们通过交叉注意力对池化后的嵌入进行视觉条件化,并将其前置到带噪声的动作标记中,从而为预测提供交互上下文和动作意图。为建模与动作相关的运动并减少对无关背景线索的依赖,我们引入了一种运动感知的过渡损失,该损失在随机时间间隔上监督视觉特征变化的预测。在仿真基准和真实世界跨具身设置上的实验验证了我们AED的有效性。匿名项目网站可在\ref{此https URL}{AED}获取。
英文摘要
World Action Models (WAMs) couple visual dynamics prediction with action generation, yet they do not explicitly support the reuse of action experience across manipulation tasks. Furthermore, existing WAMs struggle to capture underlying cross-task semantic relationships that could guide target action prediction, as redundant background elements interfere with the extraction of key visual information. To address these challenges, we develop a novel Action Experience Dictionary (AED) that encodes historical physical action trajectories into shared action embeddings to support skill reuse and model cross-task relationships. Specifically, we first aggregate historical actions to align with visual observations and retrieve action embeddings from the AED using a pretrained action tokenizer. Subsequently, we visually condition the pooled embeddings through cross-attention and prepend them to noisy action tokens, providing interaction context and action intent for prediction. To model action-related motion and reduce reliance on irrelevant background cues, we introduce a motion-aware transition loss that supervises visual feature change prediction over random temporal intervals. Experiments on simulation benchmarks and in real-world cross-embodiment settings verify the effectiveness of our AED. The project code is available at https://github.com/JiahuaDong/AED .