AI 中文总结
针对基于模型的离线强化学习中模型误差导致轨迹偏离数据分布的问题,提出分布内想象(IDI)框架,通过轨迹级支持估计与截断控制,结合轨迹正则化强化学习,在有限数据下提升性能。
AI 中文摘要
基于模型的离线强化学习(MBORL)通过模型生成的轨迹提高样本效率。然而,累积模型误差可能使想象轨迹偏离离线数据分布,导致不真实的合成数据和不稳定的策略优化。许多现有方法主要使用转移级不确定性来控制轨迹展开。我们提出分布内想象(IDI),一种轨迹展开控制框架,它在学习到的表示空间中估计轨迹支持,并自适应截断离开离线轨迹流形的展开。结合轨迹正则化强化学习(熵正则化强化学习的扩展),IDI在数据有限的环境中持续提升性能。实验表明,轨迹支持预测展开失败的能力显著优于转移级不确定性,强调了在MBORL中轨迹级展开控制的重要性。
英文摘要
Model-based offline reinforcement learning (MBORL) improves sample efficiency through model-generated trajectories. However, accumulative model error can drive imagined trajectories outside the offline data distribution, leading to unrealistic synthetic data and unstable policy optimization. Many existing methods primarily control rollouts using transition-level uncertainty. We propose \emph{in-distribution imagination} (IDI), a rollout control framework that estimates trajectory support in a learned representation space and adaptively truncates rollouts that leave the offline trajectory manifold. Combined with trajectory-regularized RL, an extension of entropy-regularized RL, IDI consistently improves performance in limited-data settings. Experiments show that trajectory support predicts rollout failure substantially better than transition-level uncertainty, highlighting the importance of trajectory-level rollout control in MBORL.
Comments11 pages, 3 figures, RLC 2026 MBRL Workshop