发表机构
University of Edinburgh; Politecnico di Milano(爱丁堡大学; 米兰理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对带随机硬约束的对抗马尔可夫决策过程,提出MA-OPS算法,结合Slater裕度的乐观搜索与策略悲观评估,获得最优悔值界并证明其最优性。
AI 中文摘要
我们研究带随机硬约束的回合式约束马尔可夫决策过程,该过程具有对抗性损失。具体而言,从一个已知的、具有裕度d的严格可行策略出发,我们力求在每回合都满足期望成本约束的同时获得最优悔值。在该设定下,Stradi等人(2025)表明,精心设计的混合规则可达到阶为$\tilde{\text{O}}(\frac{\text{√}T}{\text{min}\{d,d^2\}})$的悔值;他们还为同一设定提供了阶为$\text{Ω}(\frac{\text{√}T}{ρ})$的下界,其中ρ是离线问题的Slater裕度,可远大于d。在本研究中,我们基于他们的方法,以获得悔值对这些裕度的最优依赖关系。具体而言,我们提出MA-OPS算法,该算法结合了对Slater裕度的乐观搜索与对所选策略的悲观评估,以安全地学习具有大可行性裕度的策略;随后利用该策略在每回合满足约束的同时最小化悔值。特别地,我们证明MA-OPS可达到悔值$\tilde{\text{O}}(\frac{\text{√}T}{ρ}+\frac{1}{dρ})$;最后,我们提供匹配的下界,表明悔值界中对T、d、ρ的依赖关系在对数因子范围内是最优的。
英文摘要
We study episodic constrained Markov decision processes with adversarial losses under stochastic hard constraints. Specifically, starting from a known strictly feasible policy with margin $d$, we seek to obtain optimal regret while satisfying the expected cost constraints in every episode. In this setting, Stradi et al. (2025) show that a carefully designed mixing rule attains regret of order $\widetilde{\mathcal{O}}(\sqrt{T}/\min\{d,d^2\})$. Interestingly, they also provide a lower bound of order $Ω(\sqrt{T}/ρ)$ for the same setting, where $ρ$ is the Slater margin of the offline problem and can be much larger than $d$. In this work, we build on their approach to obtain optimal regret dependence on these margins. Specifically, we propose MA-OPS, an algorithm that combines an optimistic search for the Slater margin with a pessimistic evaluation of the selected policies to safely learn a policy with a large feasibility margin. This policy is then used to minimize regret while satisfying the constraints at every episode. In particular, we show that MA-OPS attains regret $\widetilde{\mathcal{O}}(\sqrt{T}/ρ+ 1/(dρ))$. Finally, we provide a matching lower bound, showing that the dependence on $T$, $d$, $ρ$ in the regret bound is optimal up to logarithmic factors.