状态与结果不确定性下的强化学习:一个基础性分布视角
Reinforcement Learning under State and Outcome Uncertainty: A Foundational Distributional Perspective
浏览论文内容
中文总结 AI 辅助
本文提出分布点基值迭代(DPBVI),将分布强化学习扩展到部分可观测马尔可夫决策过程,通过新的分布贝尔曼算子和psi-向量表示回报分布,为风险敏感控制奠定基础。
中文摘要 AI 辅助
在许多现实世界的规划任务中,智能体必须应对环境状态的不确定性以及所选策略结果的可变性。作为迈向部分可观测环境下更安全算法的第一步,我们同时处理这两种形式的不确定性。具体而言,我们将分布强化学习(DistRL)——该范式为完全可观测域建模整个回报分布——扩展到部分可观测马尔可夫决策过程(POMDPs),使智能体能够学习每个条件计划的回报分布。具体地,我们为部分可观测性引入了新的分布贝尔曼算子,并证明了它们在 supremum p-Wasserstein 度量下的收敛性。我们还通过 psi-向量提出了这些回报分布的有限表示,推广了 POMDP 求解器中的经典 alpha-向量。在此基础上,我们开发了基于分布的点基值迭代(DPBVI),它将 psi-向量整合到标准的点基备份过程中,从而弥合了 DistRL 与 POMDP 规划之间的鸿沟。通过跟踪回报分布,DPBVI 为未来在需要谨慎管理罕见高影响事件的领域中进行风险敏感控制奠定了基础。我们提供源代码以促进在部分可观测性下鲁棒决策的进一步研究。
英文摘要
In many real-world planning tasks, agents must tackle uncertainty about the environment's state and variability in the outcomes of any chosen policy. We address both forms of uncertainty as a first step toward safer algorithms in partially observable settings. Specifically, we extend Distributional Reinforcement Learning (DistRL)-which models the entire return distribution for fully observable domains-to Partially Observable Markov Decision Processes (POMDPs), allowing an agent to learn the distribution of returns for each conditional plan. Concretely, we introduce new distributional Bellman operators for partial observability and prove their convergence under the supremum p-Wasserstein metric. We also propose a finite representation of these return distributions via psi-vectors, generalizing the classical alpha-vectors in POMDP solvers. Building on this, we develop Distributional Point-Based Value Iteration (DPBVI), which integrates psi-vectors into a standard point-based backup procedure-bridging DistRL and POMDP planning. By tracking return distributions, DPBVI lays the foundation for future risk-sensitive control in domains where rare, high-impact events must be carefully managed. We provide source code to foster further research in robust decision-making under partial observability.