预测潜在世界模型的闭环性能:用于MPC和基于模型的强化学习的离线检查点选择,在非马尔可夫奖励的LunarLander中
Reward Observability and the Limits of Offline Checkpoint Selection in RSSM World Models
浏览论文内容
中文总结 AI 辅助
提出一套基于最优控制理论的结构化验证时诊断指标,用于从离线检查点中预测潜在世界模型的闭环性能,其中复合奖励可观测性分数(CROF)在LunarLander任务上有效选择检查点,显著提升MPC和基于模型的强化学习性能。
中文摘要 AI 辅助
我们研究如何仅从验证时诊断指标预测学习到的潜在世界模型的下游闭环性能。从世界模型训练运行中选择正确的检查点很困难:在闭环性能崩溃后很久,验证损失和多步预测RMSE仍在改善。我们提出了一套源自最优控制理论的结构化验证时诊断指标,并将其应用于具有成形奖励的Gymnasium LunarLander v3。我们在其上训练RSSM [5, 4]世界模型,并将每个检查点的CEM-MPC回报视为闭环质量的标准。通过评估40个指标与该标准,我们发现最强的单一预测因子是奖励可观测性分数(ROF),它衡量奖励预测器对可观测子空间的依赖程度。我们将ROF与三个结构正则化器组合成一个单数离线检查点选择分数,即复合奖励可观测性分数(CROF)。CROF选择的世界模型训练了一个基于模型的A2C策略,该策略在真实环境交互次数少约65倍的情况下,比公平评估的无模型A2C基线高出约24.5回报点,并且同一世界模型还驱动了一个强大的零样本CEM-MPC策略。代码和数据:此 https URL。
英文摘要
We study the closed-loop properties of a recurrent state-space model (RSSM) world model trained on human demonstrations in Gymnasium's LunarLander-v3. We use the trained world model for zero-shot CEM model-predictive control (MPC) and for actor-critic (A2C) training in imagination. Scored on 100 held-out episodes, the selected model-based A2C policy (trained on world-model checkpoint 280) reaches a mean return of +189.5, matching the best model-free A2C checkpoint (+183.7; 600- and 1000-step episode caps respectively) with ~65x fewer real training transitions. We also compare world-model MPC with a behaviour-cloning (BC) policy trained on the successful demonstrations. The BC policy matches MPC's mean return only under stochastic action selection, and on the same 20 episodes it has one catastrophic episode where MPC has none. We then introduce the Reward Observability Fraction (ROF), the Euclidean fraction of the reward gradient in the observable subspace of the linearized latent dynamics, and show that the next H observations carry Fisher information about every direction in this subspace and none about directions orthogonal to it. ROF itself however depends on how the latent is scaled: rescaling it changes ROF but not the model, so raw levels are not comparable across models. Forcing the reward head onto the posterior-corrected latent z raises ROF, and the rise survives a coordinate-invariant check on the pair of runs we tested. Finally, we test whether ROF or other offline metrics can predict the closed-loop collapse of MPC. Collapse varies between training runs with identical data and configuration. ROF does not predict it, none of the 108 offline summaries we screened passes a permutation test, and the best candidate fails on new runs. Predicting collapse offline from the model and logged data alone remains open.