发表机构
McGill University; Mila – Quebec Artificial Intelligence Institute(麦吉尔大学; 米拉-魁北克人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种基于矩闭合的不确定性下解析规划方法,通过二次动作值参数化实现闭式贝尔曼备份,在连续控制中降低目标方差,为感知分布的规划提供了合理框架。
AI 中文摘要
随机环境中有效的基于模型的强化学习需要规划时考虑预测不确定性,解析传播完整状态分布是实现这一目标的合理方法,但传统上需采用受限的策略或奖励结构才能保持可处理性。因此,现代深度强化学习大多退而采用随机采样(会引入显著的目标方差)或完全忽略预测协方差的确定性点估计。本文研究是否能在无这些约束的情况下实现感知分布的规划。采用二次动作值参数化,我们首先将贝尔曼备份简化为仅对状态值函数的期望;核心思路是预测转移分布与值函数类之间的兼容性原则,在此原则下该期望在分布的矩中是解析的。我们将该原则实例化为高斯转移模型与径向基值函数配对,得到闭式备份,可同时传播预测均值和协方差。实验表明,我们的方法在连续控制的随机观测下可降低目标方差,产生校准良好的预测不确定性,为基于学习的分布模型规划提供了合理框架。
英文摘要
Effective model-based reinforcement learning in stochastic environments requires planning that accounts for predictive uncertainty. Propagating full state distributions analytically offers a principled way to do this, but has traditionally required restrictive policy or reward structures to remain tractable. Consequently, modern deep reinforcement learning has largely retreated to either stochastic sampling, which introduces significant target variance, or deterministic point estimates that ignore predictive covariance entirely. We investigate whether distribution-aware planning is possible without these constraints. Using a quadratic action-value parameterization, we first reduce the Bellman backup to an expectation over the state-value function alone; the key idea is then a compatibility principle between the predictive transition distribution and the value function class, under which this expectation is analytic in the distribution's moments. We instantiate this principle with a Gaussian transition model paired with a radial-basis value function, yielding a closed-form backup that propagates both predictive mean and covariance. Empirically, our approach reduces target variance and yields well-calibrated predictive uncertainty under stochastic observations in continuous control, providing a principled framework for planning with learned distribution models.
CommentsTo appear in Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence (UAI 2026), PMLR