基于不确定性引导混合动力学的稳定多步回滚
Stable Multi-Step Rollouts via Uncertainty-Guided Hybrid Dynamics
AI总结:
本文提出一种与模型无关的不确定性引导混合动力学框架,用于解决基于模型的强化学习中多步回滚不稳定的问题,在杜芬振子实验中实现了稳定长 horizon 预测并改善了成本-努力权衡。
AI中文摘要:
多步回滚对于基于模型的强化学习(RL)和预测控制至关重要,但学习到的动力学模型在递归应用时往往变得不稳定,导致发散和不可靠的策略更新。本文提出一种与模型无关的混合动力学框架,该框架通过不确定性引导切换律将可证收缩的标称模型与灵活的偏移模型融合。切换信号由校准的认知不确定性导出,仅在系统离开标称区域时激活,确保每个模型在其可靠性范围内运行。在明确的平滑性和有界性假设下,我们证明所得混合预测器产生全局有界的递归多步回滚:轨迹在标称区域内保持李雅普诺夫稳定,在偏移期间表现出至多仿射增长。为在实践中说明该理论,我们在基于模型的RL方案中实例化混合动力学框架,该方案使用真实单步转换进行价值学习,使用混合回滚进行策略改进。在非线性杜芬振子上的实验表明,与稳定基线相比,该方法实现了稳定的长 horizon 预测并改善了成本-努力权衡。
英文摘要:
Multi-step rollouts are essential for model-based reinforcement learning (RL) and predictive control, yet learned dynamics models often become unstable when recursively applied, leading to divergence and unreliable policy updates. This paper proposes a model-agnostic hybrid dynamics framework that blends a provably contracting nominal model with a flexible excursion model through an uncertainty-guided switching law. The switching signal is derived from calibrated epistemic uncertainty and activates only when the system leaves the nominal region, ensuring that each model operates within its reliability regime. Under clearly stated smoothness and boundedness assumptions, we show that the resulting hybrid predictor yields globally bounded recursive multi-step rollouts: trajectories remain Lyapunov-stable in the nominal region and exhibit at most affine growth during excursions. To illustrate the theory in practice, we instantiate the hybrid dynamics framework within a model-based RL scheme that uses real one-step transitions for value learning and hybrid rollouts for policy improvement. Experiments on a nonlinear Duffing oscillator demonstrate stable long-horizon prediction and improved cost-effort trade-offs relative to a stabilizing baseline.