发表机构
University of California, Davis(加州大学戴维斯分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对世界模型训练中局部与长视目标不匹配的问题,提出DPWM非递归架构,采用端到端长视终点目标训练,显著提升了连续控制等基准的长时预测性能。
AI 中文摘要
世界模型应支持在扩展时间范围内进行想象,但大多数世界模型仍通过局部少步预测目标进行训练,并通过递归展开自身预测来部署。这造成了根本性不匹配:少步损失优化局部转移保真度,而长时预测依赖于误差和梯度如何在整个轨迹中传播。结果,训练过程中对终点具有不同下游影响的转移被同等对待,且小的局部误差会通过递归推理被放大。我们认为,通过端到端终点预测目标直接优化能更好地实现长时精度。为实例化这一范式,我们引入直接预测世界模型(Direct Prediction World Model,DPWM),这是一种非递归架构,可将任意长度的动作序列压缩为单个嵌入,并在单次前向传播中预测终点观测。该设计在预测和梯度传播中均避免了递归展开,使得在递归自回归训练变得不稳定的视域下,长时端到端训练成为可能。实验表明,在连续控制和基于像素的基准测试中,DPWM较递归世界模型基线显著提升了长时终点预测性能,且随着预测视域增大,提升幅度更大。我们进一步表明,当用相同的长时终点目标重新训练时,递归基线也会获得类似收益,支持了我们的核心主张:训练目标而非特定骨干选择是长时预测精度的主要驱动因素。我们的结果表明,世界模型可受益于在其最终使用的时间尺度上进行训练和评估,将重点从局部转移建模转向长时预测精度。
英文摘要
World models are expected to support imagination over extended temporal horizons, yet most are still trained through local few-step prediction objectives and deployed by recursively rolling out their own predictions. This creates a fundamental mismatch: few-step losses optimize local transition fidelity, while long-horizon prediction depends on how errors and gradients propagate through the entire trajectory. As a result, transitions with different downstream influence on the endpoint are treated uniformly during training, and small local errors are amplified through recursive inference. We argue that long-horizon accuracy is better achieved by optimizing directly, through an end-to-end endpoint prediction objective. To instantiate this paradigm, we introduce the Direct Prediction World Model (DPWM), a non-recursive architecture that compresses an action sequence of arbitrary length into a single embedding and predicts the endpoint observation in a single forward pass. This design avoids recurrent rollout in both prediction and gradient propagation, making long-horizon end-to-end training practical at horizons where unrolled autoregressive training becomes unstable. Empirically, DPWM substantially improves long-horizon endpoint prediction over recursive world-model baselines on continuous-control and pixel-based benchmarks, with larger gains as the prediction horizon increases. We further show that recurrent baselines benefit similarly when retrained with the same long-horizon endpoint objective, supporting our central claim that the training objective, rather than the particular backbone choice, is the main driver of long-horizon prediction accuracy. Our results suggest that world models can benefit from being trained and evaluated at the temporal scales where they are ultimately used, shifting the focus from local transition modeling toward long-horizon predictive accuracy.