arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.04996cs.RO

DreamWAM:超越RGB未来预测的世界动作模型

DreamWAM: Beyond RGB Future Prediction for World Action Models

Shanglin Yuan, Weiheng Zhao, Xin Shi, Haoyi Jiang, Xianda Guo, Liu Liu, Wenyu Liu, Wei Sui, Xinggang Wang

首次发表
浏览论文内容

中文总结 AI 辅助

DreamWAM通过结构化超越RGB的世界建模,在LIBERO及真实操作任务中提升了世界动作模型的鲁棒性与成功率,相关代码模型已公开。

中文摘要 AI 辅助

世界动作模型(World Action Models, WAMs)通过预测观测世界的演化来学习与动作相关的表征。现有大多数WAMs将未来定义在RGB空间中,此时与任务相关的状态转移会与纹理、光照、背景和视角等无关变化纠缠在一起。本文认为WAMs应显式预测与动作相关的未来状态,而非仅依赖RGB预测。我们提出DreamWAM,其将未来预测重新定义为超越RGB的结构化世界建模,通过外观、运动、几何和语义的互补视图来表征未来状态。训练期间,DreamWAM结合了RGB与运动的联合潜在去噪,以及用于几何和语义的轻量门控残差分支;VideoDiT与ActionDiT之间的共享注意力使动作分支能从这些未来状态预测中学习,而所有超越RGB的监督分支在推理阶段被禁用,部署时仍仅使用RGB。在无回滚和联合视频-动作推理设置下,DreamWAM在LIBERO数据集上分别将匹配的仅RGB基线从97.30%提升至98.40%、从98.00%提升至98.90%;在未见过的LIBERO-Plus扰动下,提升幅度更大,从51.36%增至63.44%、从69.16%增至75.47%。该鲁棒性也延伸至真实世界操作:在光照、背景和物体布局的未见过变化下,DreamWAM的平均成功率为74.4%,而Fast-WAM-Joint为55.6%。这些结果表明,鲁棒的世界-动作学习不仅依赖于预测未来,还依赖于以对动作重要的形式表征未来。代码和模型已公开于该httpsURL。

英文摘要

World Action Models (WAMs) learn action-relevant representations by predicting how the observed world will evolve. Most existing WAMs define this future in RGB space, where task-relevant state transitions are entangled with nuisance variations in texture, illumination, background, and viewpoint. We argue that WAMs should explicitly predict action-relevant future state rather than relying on RGB prediction alone. We introduce DreamWAM, which reformulates future prediction as structured world modeling beyond RGB, representing future states through complementary views of appearance, motion, geometry, and semantics. During training, DreamWAM combines joint latent denoising of RGB and motion with lightweight gated residual branches for geometry and semantics. Shared attention between VideoDiT and ActionDiT allows the action branch to learn from these future-state predictions, while all beyond-RGB supervision branches are disabled at inference and deployment remains RGB-only. Across both no-rollout and joint video-action inference, DreamWAM consistently improves the matched RGB-only baselines on LIBERO, from 97.30\% to 98.40\% and from 98.00\% to 98.90\%, respectively. The gains become larger under unseen LIBERO-Plus perturbations, from 51.36\% to 63.44\% and from 69.16\% to 75.47\%. The same robustness extends to real-world manipulation, where DreamWAM attains an average success rate of 74.4\% across unseen changes in lighting, background, and object layout, compared with 55.6\% for Fast-WAM-Joint. These results show that robust world-action learning depends not only on predicting the future, but on representing it in a form that matters for action. The code and models are publicly released at https://github.com/hustvl/DreamWAM.

↑