arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DreamTrue:具备反事实后训练的动作忠实机器人世界模型

DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-Training

Junyan Li, Ruizhi Li, Yu Liu, Xiangshuo Liu, Mingchao Sun, Hongyu Pan, Mu Xu, Lue Fan, Zhaoxiang Zhang

arXiv 2610.12468首次发表:更新:

发表机构

Institute of Automation, Chinese Academy of Sciences (CASIA); Alibaba Group; Amap(中国科学院自动化研究所; 阿里巴巴集团; 高德地图)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出DreamTrue机器人世界模型,通过离线几何校准和反事实后训练解决动作跟随与交互覆盖问题,在AgiBot上实现最优动作跟随,将交互缺陷率降至6.25%,获2026年AgiBot世界挑战赛世界模型赛道第一。

AI 中文摘要

我们提出DreamTrue,这是一种多视角、跨 embodiment( embodiment 指具身形态)的机器人世界模型,用于实现动作忠实且物理合理的视频预测。在现有机器人数据集上训练此类模型面临两个障碍:校准不精确会损害动作跟随能力,而不成功交互的覆盖有限会使预测偏向成功结果。为提升跨具身形态的动作跟随能力,我们将动作轨迹渲染为图像空间条件,并引入离线几何校准以对齐这些条件与目标视频。为扩大交互覆盖范围,我们引入反事实后训练,修改记录的动作轨迹并在更广泛的动作和接触配置下生成未来视频。为在无配对真实未来的情况下为这些预测提供反馈,我们构建了人工标注的视频数据集,涵盖机器人、物体及交互缺陷,并利用其训练具身视频奖励模型。该模型的分数指导强化学习后训练,以实现更物理合理的交互结果。在AgiBot平台上,DreamTrue实现了最先进的动作跟随性能,同时将人工评估的交互缺陷率从48.12%降至6.25%。值得注意的是,我们的模型在2026年AgiBot世界挑战赛的世界模型赛道中排名第一。项目页面可通过此URL获取。

英文摘要

We present DreamTrue, a multi-view, cross-embodiment robot world model for action-faithful and physically plausible video prediction. Training such a model on existing robot datasets faces two obstacles: imprecise calibration can impair action following, while limited coverage of unsuccessful interactions can bias predictions toward successful outcomes. To improve action following across embodiments, we render action trajectories into image-space conditions and introduce offline geometric calibration to align these conditions with the target videos. To broaden interaction coverage, we introduce counterfactual post-training, modifying recorded action trajectories and generating future videos under a wider range of actions and contact configurations. To provide feedback on these predictions without paired ground-truth futures, we construct a human-annotated video dataset covering robot, object, and interaction defects and use it to train an embodied video reward model. Its scores guide reinforcement-learning post-training toward more physically plausible interaction outcomes. On AgiBot, DreamTrue attains state-of-the-art action following, while reducing the human-assessed interaction defect rate from from 48.12% to 6.25%. Notably, our model ranks first in the world model track of the AgiBot World Challenge 2026. The project page can be found at https://brave-eai.github.io/DreamTrue.

Commentsproject page: https://brave-eai.github.io/DreamTrue; code: https://github.com/brave-eai/DreamTrue

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑