发表机构
Saint Petersburg State University; Shenzhen Kaihong Digital Industry Development Co., Ltd.; Chinese Academy of Sciences (CAS) – Shenzhen Institute of Advanced Technology; Harbin Institute of Technology; Chongqing Research Institute of HIT(圣彼得堡国立大学; 深圳开鸿数字产业发展有限公司; 中国科学院深圳先进技术研究院; 哈尔滨工业大学; 哈尔滨工业大学重庆研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FIRM-WM通过状态分解和公共重置干预分支,解决无奖励视觉规划中目标可比性与历史依赖动态的冲突,在多个基准上超越LeWM并降低规划时间。
AI 中文摘要
无奖励潜在世界模型可以从离线视频中学习,并通过预测的潜在未来优化动作来解决新的图像-目标任务。这一设定对规划状态提出了两个要求:其坐标必须能与目标图像进行比较;此外,其动态必须保留速度、运动趋势、接触以及其他超出目标坐标的历史依赖信息。离线训练造成了第二个不匹配:每条记录轨迹仅揭示一个事实未来,而基于采样的规划器需要比较从同一状态出发但未实际执行的多种动作。我们提出了FIRM-WM(事实-干预循环世界模型),一个围绕这两个差距设计的紧凑像素世界模型。其循环状态将类型化的、可与目标比较的配置与一个128维动态纤维分离,后者用于预测但被排除在最终目标成本之外。广泛的事实轨迹提供了状态覆盖,而公共重置干预分支为替代动作序列提供了观测结果。在执行每个分支之前,我们重置环境并恢复环境状态设置接口所暴露的相同记录值。在匹配的CEM规划和三个独立的完整流水线随机种子下,FIRM-WM在TwoRoom上达到99.0±1.0%,在Reacher上达到92.7±2.1%,在OGBench-Cube上达到88.0±3.0%,而LeWM分别为89.0%、88.0%和70.0%。部署的模型使用2.98-3.42M参数,并在这些任务上记录了2.13-11.60倍更低的规划时间。
英文摘要
Reward-free latent world models can learn from offline videos and solve new image--goal tasks by optimizing actions through predicted latent futures. This setting places two demands on the planning state: its coordinates must be comparable with a goal image. Moreover, its dynamics must retain velocity, motion trend, contact, and other history--dependent information beyond those goal coordinates. Offline training creates a second mismatch: each recorded trajectory reveals one factual future, whereas a sampling--based planner compares many actions that were not taken from the same state. We introduce FIRM-WM (Factual--Interventional Recurrent World Model), a compact pixel world model designed around these two gaps. Its recurrent state separates a typed, goal--comparable configuration from a 128-dimensional dynamic fiber used for prediction but excluded from the terminal goal cost. Broad factual trajectories provide state coverage, while common--reset intervention branches provide observed outcomes for alternative action sequences. Before executing each branch, we reset the environment and restore the same recorded values exposed by the environment's state--setting interface. Under matched CEM planning and three independent full-pipeline seeds, FIRM-WM reaches 99.0$\pm$1.0% on TwoRoom, 92.7$\pm$2.1% on Reacher, and 88.0$\pm$3.0% on OGBench-Cube, compared with 89.0%, 88.0%, and 70.0% for LeWM. The deployed model uses 2.98--3.42M parameters and records 2.13--11.60$\times$ lower planning time on these tasks.