发表机构
University of Edinburgh; Politecnico di Milano; University College London(爱丁堡大学; 米兰理工大学; 伦敦大学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对逐步约束的表格型CMDP,提出SVAE算法,利用安全子图结构进行方差自适应探索,实现实例相关的最优遗憾界和约束违反界,并证明其必要性。
AI 中文摘要
我们研究了具有逐步安全约束的回合制表格型约束马尔可夫决策过程中的在线学习问题。在此类设置中,约束条件诱导出一个安全子图,该子图决定了可行策略下累积奖励的方差,从而影响学习的难度。然而,利用这一结构需要在控制约束违反的同时学习哪些动作是安全的。我们提出了安全方差自适应探索(SVAE)算法,这是一种高效算法,它学习候选安全子图并在其中进行方差自适应乐观规划。以高概率,SVAE在K个回合中实现累积遗憾阶为$\widetilde{\mathcal{O}}(\sqrt{SAH\min\{\mathbb{V}_\Sigma,K\mathrm{Var}^{\star}\}}+S\sqrt{AH^3\min\{K,\mathcal{C}\}}+S^2AH^2)$,其中H是单个回合的视界,S和A分别是状态和动作的数量。这里,$\mathrm{Var}^{\star}$是安全策略中的最大回报方差,$\mathbb{V}_\Sigma$是在遇到第一个不安全动作之前累积的方差,$\mathcal{C}$捕捉了消除被错误视为潜在安全的动作的统计复杂度。SVAE还实现了$\widetilde{\mathcal{O}}(H\sqrt{SAK}+S^2AH^2)$的逐步约束违反以及一个关于K呈多对数增长的间隙相关违反界。最后,我们建立了一个下界,表明对这些实例特定量的依赖是不可避免的。
英文摘要
We study online learning in episodic tabular constrained Markov decision processes with step-wise safety constraints. In such a setting, the constraints induce a safe subgraph that shapes the variance of cumulative rewards under feasible policies and, consequently, the difficulty of learning. Exploiting this structure, however, requires learning which actions are safe while controlling constraint violations. We propose Safe Variance-Adaptive Exploration (SVAE), an efficient algorithm that learns candidate safe subgraphs and performs variance-adaptive optimistic planning within them. With high probability, SVAE achieves cumulative regret of order $\widetilde{\mathcal{O}}(\sqrt{SAH\min\{\mathbb{V}_Σ,K\mathrm{Var}^{\star}\}}+S\sqrt{AH^3\min\{K,\mathcal{C}\}}+S^2AH^2)$ over $K$ episodes, where $H$ is the horizon of a single episode, while $S$ and $A$ are the numbers of states and actions, respectively. Here, $\mathrm{Var}^{\star}$ is the maximum return variance among safe policies, $\mathbb{V}_Σ$ is the variance accumulated before the first unsafe action is encountered, and $\mathcal{C}$ captures the statistical complexity of eliminating actions incorrectly considered potentially safe. SVAE additionally attains $\widetilde{\mathcal{O}}(H\sqrt{SAK}+S^2AH^2)$ step-wise constraint violation and a gap-dependent violation bound that is polylogarithmic in $K$. Finally, we establish a lower bound showing that dependence on these instance-specific quantities is unavoidable.