AI 中文总结
本文提出面向机器人操纵的开源世界模拟器BWM,通过动作条件化自回归预测实现高保真,兼具数据引擎与策略评估器功能,在WorldArena挑战赛中表现最优并开源相关资源。
AI 中文摘要
可靠的机器人学习需要一种世界模拟器,能在物理硬件执行前预测动作后果,包括高风险和易失败的结果。现有物理模拟器需要大量资产构建与校准,仍存在现实差距;视频生成器往往无法精确控制对机器人细粒度动作的响应。本文提出Boundless World Model(BWM),一种面向机器人操纵的开源、低成本、高保真世界模拟器。BWM是一种动作条件化的世界模型,结合初始环境引导、动态视觉历史以及时间对齐的机器人动作条件,对未来观测进行有状态自回归预测。我们通过轨迹回放、重叠片段采样和初始观测增强构建动作对齐的训练片段。BWM兼具数据引擎和策略评估器的功能:作为数据引擎,它用动作对齐的rollout扩充模仿学习数据;作为策略评估器,它用于闭环评估、风险预判和策略排序。在WorldArena基准测试和物理机器人上的实验表明,BWM在数据引擎和策略评估器设置下,模拟器保真度和功能效用均有所提升,在WorldArena挑战赛的Track 1及两个Track 2应用中均排名第一。我们发布了BWM开源生态系统,包括模型检查点、训练与推理代码,以及数据生成和策略评估接口。
英文摘要
Reliable robot learning requires a world simulator that can predict action consequences before execution on physical hardware, including risky and failure-prone outcomes. Existing physics simulators require substantial asset construction and calibration and still face a sim-to-real gap, while video generators often lack precise control over their responses to fine-grained robot actions. In this paper, we present the Boundless World Model (BWM), an open-source, low-cost, high-fidelity world simulator for robot manipulation. BWM is an action-conditioned world model that combines initial-environment guidance, dynamic visual history, and temporally aligned robot-action conditioning for stateful autoregressive prediction of future observations. We construct action-aligned training clips through trajectory replay, overlapping clip sampling, and initial-observation enhancement. BWM serves as a data engine that augments imitation-learning data with action-aligned rollouts, and as a policy evaluator for closed-loop assessment, risk anticipation, and policy ranking. Experiments on the WorldArena benchmark and physical robots demonstrate improved simulator fidelity and functional utility across the data-engine and policy-evaluator settings. BWM ranks first overall in the WorldArena Challenge across Track 1 and its two Track 2 applications. We release the BWM open-source ecosystem, including model checkpoints, training and inference code, and interfaces for data generation and policy evaluation.