arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30856cs.LG

学习机会约束马尔可夫决策过程与贝尔曼分布证书

Learning Chance-Constrained MDPs with Bellman Distributional Certificates

Chenbei Lu, Hongyu Yi

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出贝尔曼分布证书方法,用于解决机会约束马尔可夫决策过程,通过构建约束违反概率的贝尔曼递归,实现模型基上下界匹配,并提供无模型策略梯度算法,在合成和IEEE 14节点基准上验证安全性。

中文摘要 AI 辅助

安全强化学习(RL)通常强制执行期望成本约束,但此类期望安全性可能无法控制罕见高成本轨迹的发生概率。机会约束马尔可夫决策过程(CCMDPs)施加了更强的概率级别要求,但被广泛认为更难处理,因为机会约束是非凸的,且依赖于完整轨迹而非贝尔曼线性期望。在本文中,我们揭示了这种计算难度并不必然意味着更高的统计代价。对于具有固定有界后继支持且可访问认证规划预言机的表格型折扣CCMDPs,我们建立了一个基于模型的上界,并匹配了对数项以内的下界。在技术上,我们的关键思想是“贝尔曼分布证书”,它在策略选择之前为约束违反概率构建贝尔曼递归。该证书可在候选策略之间重用;结合共享的行式反向KL置信集,它提供了策略均匀的轨迹KL转移,而无需对策略或时间预算贝尔曼表进行联合界。对于随机策略,我们给出了一种无模型方差缩减的策略梯度算法,具有有限样本期望KKT残差保证,并对每个被接受的策略进行独立验证。在合成CCMDPs和IEEE 14节点储能控制基准上的数值实验展示了所提算法的安全性和机制行为。

英文摘要

Safe reinforcement learning (RL) commonly enforces expected-cost constraints, but such expectation safety may fail to control the probability of rare high-cost trajectories. Chance-constrained MDPs (CCMDPs) impose a stronger probability-level requirement, but are widely viewed as harder because the chance constraint is nonconvex and depends on the full trajectory rather than a Bellman-linear expectation. In this paper, we reveal that this computational difficulty does not necessarily imply a higher statistical price. For tabular discounted CCMDPs with fixed bounded successor support and access to a certified planning oracle, we establish a model-based upper bound, with a matching lower bound up to logarithmic terms. Technically, our key idea is the \emph{Bellman distributional certificate}, which constructs a Bellman recursion for constraint violation probabilities before policy selection. The certificate can be reused across candidate policies; combined with shared row-wise reverse-KL confidence sets, it gives a policy-uniform trajectory-KL transfer without a union bound over policies or time--budget Bellman tables. For stochastic policies, we give a model-free variance-reduced policy-gradient algorithm with a finite-sample expected KKT-residual guarantee and independent validation of every accepted policy. Numerical experiments on synthetic CCMDPs and an IEEE 14-bus energy storage control benchmark illustrate the safety and mechanism behavior of the proposed algorithms.

发表机构

  • Cornell University AI for Science Institute(康奈尔大学人工智能科学研究所)
  • Cornell University(康奈尔大学)
  • University of Washington(华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑