AI 中文总结
该研究针对视觉分布偏移下机器人操作的鲁棒性问题,提出ST-WAM模型,采用DINOv3等技术,在多个基准测试中取得优异性能,大幅提升了偏移场景下的零样本操作表现。
AI 中文摘要
世界动作模型(WAMs)已成为一种有前景的范式,它联合建模机器人动作与未来视觉动态。然而,其对基于像素生成的未来监督的依赖,会将与动作相关的状态转换和与任务无关的视觉内容混淆,从而限制了视觉分布偏移下的鲁棒性。我们发现了训练分布幻觉,这是一种反复出现的现象:以视觉偏移观测为条件的未来会生成训练域内容,而非忠实于当前场景。受控帧三元组诊断进一步表明,DINOv3特征在视觉偏移下比Wan-VAE隐变量更稳定,同时能更好地保留任务状态区分度。我们不修正预测的未来,而是提出语义-时间WAM(ST-WAM),通过将DINOv3用作未来预测和历史检索的共享语义表示,同时保留细粒度VAE动态来提升动作鲁棒性。其双空间未来专家(DSFE)联合预测未来VAE隐变量和DINO特征,而当前锚定意图检索(CAIR)则在当前视觉-语言上下文下,从近期DINO历史中检索与任务相关的证据。ST-WAM以端到端方式训练,无需额外的具身预训练或任务特定标注,且推理时无需显式未来生成。它在LIBERO上达到98.7%,在RoboTwin 2.0上达到92.8%;更重要的是,与Fast-WAM相比,它将零样本LIBERO-Plus性能提升了21.3个百分点,且视觉偏移下的真实世界成功率从25.8%提高到61.5%,增幅超一倍。这些结果表明,语义-时间建模能有效补充像素生成动态,实现鲁棒操作。
英文摘要
World Action Models (WAMs) have emerged as a promising paradigm by jointly modeling robot actions and future visual dynamics. However, their reliance on pixel-generative future supervision can entangle action-relevant state transitions with task-irrelevant visual content, limiting robustness under visual distribution shifts. We identify Training-Distribution Hallucination, a recurring phenomenon in which futures conditioned on visually shifted observations hallucinate training-domain content rather than remain faithful to the current scene. A controlled frame-triplet diagnosis further shows that DINOv3 features remain more stable across visual shifts while better preserving task-state distinctions than Wan-VAE latents. Rather than correcting the predicted futures, we propose Semantic-Temporal WAM (ST-WAM) to improve action robustness by using DINOv3 as a shared semantic representation for future prediction and history retrieval while retaining fine-grained VAE dynamics. Its Dual-Space Future Experts (DSFE) jointly predict future VAE latents and DINO features, while Current-Anchored Intent Retrieval (CAIR) retrieves task-relevant evidence from recent DINO history under the current visual-language context. ST-WAM is trained end-to-end without additional embodied pretraining or task-specific annotations, and requires no explicit future generation at inference. It achieves 98.7% on LIBERO and 92.8% on RoboTwin 2.0; more importantly, compared with Fast-WAM, it improves zero-shot LIBERO-Plus performance by 21.3 percentage points and more than doubles real-world success under visual shifts from 25.8% to 61.5%. These results demonstrate that semantic-temporal modeling effectively complements pixel-generative dynamics for robust manipulation.
Comments9 pages, 5 figures