arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ST-WAM:面向视觉分布偏移下鲁棒操作的语义-时间世界动作模型

ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts

Mingxin Wang, Bin Hu, Bin Qian, Kaitao Jiang, Haoning Wu, Feng Yan, Bowen Jing, Ruiyang Hao, Enyi Wang, Kangning Niu, Yandan Yang, Mu Xu, Yan Wang, Houde Liu, Tianlun Li

arXiv 2607.28993首次发表:更新:

AI 中文总结

该研究针对视觉分布偏移下机器人操作的鲁棒性问题,提出ST-WAM模型,采用DINOv3等技术,在多个基准测试中取得优异性能,大幅提升了偏移场景下的零样本操作表现。

AI 中文摘要

世界动作模型(WAMs)已成为一种有前景的范式,它联合建模机器人动作与未来视觉动态。然而,其对基于像素生成的未来监督的依赖,会将与动作相关的状态转换和与任务无关的视觉内容混淆,从而限制了视觉分布偏移下的鲁棒性。我们发现了训练分布幻觉,这是一种反复出现的现象:以视觉偏移观测为条件的未来会生成训练域内容,而非忠实于当前场景。受控帧三元组诊断进一步表明,DINOv3特征在视觉偏移下比Wan-VAE隐变量更稳定,同时能更好地保留任务状态区分度。我们不修正预测的未来,而是提出语义-时间WAM(ST-WAM),通过将DINOv3用作未来预测和历史检索的共享语义表示,同时保留细粒度VAE动态来提升动作鲁棒性。其双空间未来专家(DSFE)联合预测未来VAE隐变量和DINO特征,而当前锚定意图检索(CAIR)则在当前视觉-语言上下文下,从近期DINO历史中检索与任务相关的证据。ST-WAM以端到端方式训练,无需额外的具身预训练或任务特定标注,且推理时无需显式未来生成。它在LIBERO上达到98.7%,在RoboTwin 2.0上达到92.8%;更重要的是,与Fast-WAM相比,它将零样本LIBERO-Plus性能提升了21.3个百分点,且视觉偏移下的真实世界成功率从25.8%提高到61.5%,增幅超一倍。这些结果表明,语义-时间建模能有效补充像素生成动态,实现鲁棒操作。

英文摘要

World Action Models (WAMs) have emerged as a promising paradigm by jointly modeling robot actions and future visual dynamics. However, their reliance on pixel-generative future supervision can entangle action-relevant state transitions with task-irrelevant visual content, limiting robustness under visual distribution shifts. We identify Training-Distribution Hallucination, a recurring phenomenon in which futures conditioned on visually shifted observations hallucinate training-domain content rather than remain faithful to the current scene. A controlled frame-triplet diagnosis further shows that DINOv3 features remain more stable across visual shifts while better preserving task-state distinctions than Wan-VAE latents. Rather than correcting the predicted futures, we propose Semantic-Temporal WAM (ST-WAM) to improve action robustness by using DINOv3 as a shared semantic representation for future prediction and history retrieval while retaining fine-grained VAE dynamics. Its Dual-Space Future Experts (DSFE) jointly predict future VAE latents and DINO features, while Current-Anchored Intent Retrieval (CAIR) retrieves task-relevant evidence from recent DINO history under the current visual-language context. ST-WAM is trained end-to-end without additional embodied pretraining or task-specific annotations, and requires no explicit future generation at inference. It achieves 98.7% on LIBERO and 92.8% on RoboTwin 2.0; more importantly, compared with Fast-WAM, it improves zero-shot LIBERO-Plus performance by 21.3 percentage points and more than doubles real-world success under visual shifts from 25.8% to 61.5%. These results demonstrate that semantic-temporal modeling effectively complements pixel-generative dynamics for robust manipulation.

Comments9 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑