arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18462cs.CVcs.AI

CSWAM:世界动作模型中用于分布外泛化的更好因果语义表示

CSWAM: Better Causal Semantic Representations for Out-of-Distribution Generalization in World Action Models

发表机构美的集团AIRC · 同济大学
查看机构详情
  • AIRC, Midea Group(美的集团AIRC)
  • Tongji University(同济大学)

机构由 AI 辅助整理,请以论文原文为准。

Tianbin Liu, Jian Zhu, Taiyi Su, Jianjun Zhang, Chong Ma, Zitai Huang, Weiyi Lu, Yi Xu

首次发表
浏览论文内容

中文总结 AI 辅助

CSWAM通过引入基于V-JEPA的因果语义专家增强FastWAM,利用观测历史提升世界动作模型在分布偏移下的泛化能力,在仿真和真实机器人实验中显著提高成功率。

中文摘要 AI 辅助

FastWAM风格的世界动作模型能够实现高效的动作仅推理,但在视觉分布偏移下泛化能力较差。其面向重建的表示强调外观特定细节,限制了对未见场景和物体的泛化。在没有观测历史的情况下,模型也缺乏时间证据来在陌生的视觉条件下稳健地识别与任务相关的状态变化和运动。为解决这些局限,我们提出了因果语义世界动作模型(CSWAM),它通过基于V-JEPA 2.1构建的因果语义专家来增强FastWAM。V-JEPA提供对语义状态变化和运动的时间接地表示,且较少依赖外观特定细节。该专家从当前和过去观测的稀疏历史中学习其未来演化,并通过因果注意力将历史派生上下文共享给视频和动作流。在推理时,CSWAM基于当前视频状态和观测到的语义历史来调节动作去噪,保留了高效的动作仅推理。我们进行了仿真和真实机器人实验,以评估分布偏移下的泛化能力。通过具身预训练,CSWAM在RoboTwin 2.0从干净到随机化的迁移中将随机化成功率从10.16%提升至45.18%,比FastWAM提高了35.02个百分点。在两个真实机器人任务和三个OOD难度级别上,CSWAM将平均成功率从27.5%提升至70.0%,比FastWAM提高了42.5个百分点。

英文摘要

FastWAM-style world action models enable efficient action-only inference, but generalize poorly under visual distribution shifts. Their reconstruction-oriented representations emphasize appearance-specific details, limiting generalization to unseen scenes and objects. Without observation history, the model also lacks temporal evidence for robustly identifying task-relevant state changes and motion in unfamiliar visual conditions. To address these limitations, we present the Causal Semantic World Action Model (CSWAM), which augments FastWAM with a causal semantic expert built on V-JEPA 2.1. V-JEPA provides temporally grounded representations of semantic state changes and motion with less dependence on appearance-specific details. The expert learns their future evolution from a sparse history of current and past observations and shares the history-derived context with both the video and action streams through causal attention. At inference, CSWAM conditions action denoising on the current video state and observed semantic history, retaining efficient action-only inference. We conduct simulation and real-robot experiments to evaluate generalization under distribution shifts. With embodied pretraining, CSWAM raises Randomized success on RoboTwin 2.0 Clean-to-Randomized transfer from 10.16% to 45.18%, a gain of 35.02 percentage points over FastWAM. Across two real-robot tasks and three OOD difficulty levels, CSWAM improves average success over FastWAM by 42.5 percentage points, from 27.5% to 70.0%.

补充信息

↑