arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FutureWorlds:从替代未来中学习机器人世界模型

FutureWorlds: Learning Robotic World Models from Alternative Futures

Hao Wu, Shengju Qian, Weiyan Wang, Fan Xu, Fan Zhang, Yuanpeng He, Qingsong Wen, Yuxuan Liang

arXiv 2610.01019首次发表:更新:

发表机构

HKUST (GZ); CUHK; Tencent; USTC; PKU; Squirrel Ai Learning(香港科技大学(广州); 香港中文大学; 腾讯; 中国科学技术大学; 北京大学; 松鼠Ai学习)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FutureWorlds提出统一框架,结合多样束搜索和记忆条件策略优化,从替代未来中学习机器人世界模型,在多个数据集上显著提升预测质量与运动准确性。

AI 中文摘要

机器人世界模型预测以动作条件化的未来场景,为理解动作结果提供了基础。然而,将替代性预测转化为有用的学习信号仍然具有挑战性:相似的候选方案限制了信息丰富的质量比较,而分叉的轨迹则需要持续维护各自的历程。我们引入了FutureWorlds,一个统一了候选构建、历程维护和基于相对质量学习的框架。基于多模态离散自回归模型,FutureWorlds在强化学习过程中使用多样束搜索来构建在置信度和多样性之间取得平衡的候选未来。候选特定的有界记忆保留了场景状态,并确保生成和策略评分使用匹配的历程。我们进一步提出了MemSPO(记忆条件搜索引导策略优化),它将视频轨迹奖励转换为组相对优势以优化世界模型。在RT-1、BridgeV2和RoboCasa上,相对于每个数据集上的最强基线,FutureWorlds将32帧预测的LPIPS分别降低了14.78%、20.84%和9.12%。在固定评估配置下,仅200次MemSPO更新就能进一步提高生成质量,并支持超出训练范围的持续预测。记忆消融、解码敏感性分析和光流评估表明,这些收益不仅限于视觉质量,还扩展到更准确的运动预测和更一致的对象状态。项目页面和代码:此https URL。

英文摘要

Robotic world models predict action-conditioned future scenes, providing a foundation for understanding action outcomes. However, turning alternative predictions into useful learning signals remains challenging: similar candidates limit informative quality comparisons, while diverging trajectories require persistent maintenance of their individual histories. We introduce FutureWorlds, a framework that unifies candidate construction, history maintenance, and learning from relative quality. Built on a multimodal discrete autoregressive model, FutureWorlds uses diverse beam search during reinforcement learning to construct candidate futures that balance confidence and diversity. Candidate-specific bounded memory preserves scene states and ensures that generation and policy scoring use matching histories. We further propose MemSPO (Memory-Conditioned Search-Guided Policy Optimization), which converts video trajectory rewards into group-relative advantages to optimize the world model. On RT-1, BridgeV2, and RoboCasa, FutureWorlds reduces LPIPS for 32-frame predictions by 14.78%, 20.84%, and 9.12%, respectively, relative to the strongest baseline on each dataset. Under fixed evaluation configurations, only 200 MemSPO updates further improve generation quality and support continued prediction beyond the training horizon. Memory ablations, decoding sensitivity analysis, and optical-flow evaluation show that these gains extend beyond visual quality to more accurate motion prediction and more consistent object states. Project page and code: https://github.com/Alexander-wu/FutureWorlds.

Comments32 pages, including references and appendix. Code: https://github.com/Alexander-wu/FutureWorlds

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑