基于认知不确定性的自适应展开截断用于高效离线世界模型训练
Adaptive Rollout Truncation Based on Epistemic Uncertainty for Efficient Offline World Model Training
- Technical University of Munich(慕尼黑工业大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出基于认知不确定性的自适应展开截断策略,用于离线世界模型训练,通过阈值终止展开以减少计算量,实验表明在保持精度的同时减少约72%的展开计算。
AI中文摘要:
准确的神经世界模型是基于模型的机器人技术的核心,它使机器人能够根据先前观测到的轨迹预测未来状态。多步自回归训练提高了长时域预测能力,但固定的展开时域也增加了计算成本,并且当模型仍不准确时可能放大早期训练误差。现有的训练方案通常在优化过程中使用相同的展开长度,与模型当前的预测可靠性无关。我们提出了一种认知不确定性驱动的自适应展开策略,用于遵循自动课程训练方案的离线世界模型训练。模型并非总是展开到固定时域,而是当认知不确定性超过从预热阶段校准的阈值时终止自回归展开。我们研究了两种不确定性估计器:具有共享循环主干和蒙特卡洛丢弃的五头集成。一个两阶段预热程序在启用自适应截断之前稳定了不确定性估计。在ANYmal-D和ANT上的实验表明,基于集成的自适应截断在匹配或改进固定时域训练和RWM-U基线的预测精度的同时,需要显著更少的累积展开步骤。按照所提出的方法在ANYmal-D上训练世界模型,达到与基线相当的性能,但展开计算量减少了约72%。这些结果表明,认知不确定性不仅对下游策略正则化有用,而且对使世界模型训练本身更加计算高效也有用。
英文摘要:
Accurate neural world models are central to model-based robotics, where they enable robots to predict future states from previously observed trajectories. Multi-step autoregressive training improves long-horizon prediction, but fixed rollout horizons also increase computational cost and can amplify early training errors when the model is still inaccurate. Existing training schemes typically use the same rollout length throughout optimization, independent of the model's current predictive reliability. We propose an epistemic uncertainty-driven adaptive rollout strategy for offline world model training following an auto-curriculum training scheme. Instead of always unrolling to a fixed horizon, the model terminates autoregressive rollouts once epistemic uncertainty exceeds a threshold calibrated from a warm-up phase. We study two uncertainty estimators: a five-head ensemble with a shared recurrent backbone and Monte Carlo Dropout. A two-stage warm-up procedure stabilizes uncertainty estimates before we enable adaptive truncation. Experiments on ANYmal-D and ANT show that ensemble-based adaptive truncation matches or improves the prediction accuracy of fixed-horizon training and the RWM-U baseline while requiring substantially fewer cumulative rollout steps. Training a world model on ANYmal-D following the presented approach reaches comparable final performance with the baselines with roughly 72% less rollout computation. These results indicate that epistemic uncertainty is useful not only for downstream policy regularization, but also for making world model training itself more compute-efficient.