发表机构
KAIST; Washington State University; Seoul National University(韩国科学技术院; 华盛顿州立大学; 首尔国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对未知转移的对抗性线性约束MDP,提出结合自适应FTRL、收缩值估计和指数Lyapunov函数的新原始-对偶算法,实现最优$\tilde{O}(\sqrt{K})$遗憾与约束违反,无需Slater条件。
AI 中文摘要
我们研究了具有未知转移的回合制对抗性线性约束马尔可夫决策过程(CMDPs),其中损失函数和约束函数均可能随回合而对抗性变化。此前最优算法实现了$\tilde{\mathcal{O}}(K^{3/4})$的遗憾和累积约束违反,与关于回合数$K$的最优$\tilde{\mathcal{O}}(\sqrt{K})$依赖之间存在差距。我们通过提出一种新的原始-对偶算法来弥合这一差距,该算法在无需假设Slater条件的情况下实现了$\tilde{\mathcal{O}}(\sqrt{K})$的遗憾和累积约束违反。主要挑战在于,学习线性CMDPs需要对具有受控覆盖数的值函数类进行均匀集中,而约束在线学习中的标准技术(如策略混合)可能使该函数类更加复杂。我们的算法结合了自适应跟随正则化领导者(FTRL)、收缩值估计和指数Lyapunov函数。自适应对偶正则化器抵消了原始遗憾界中对偶权重的依赖,从而无需策略混合。我们进一步表明,FTRL更新中的归一化将策略参数限制为独立于对偶权重的大小,这解释了所得策略类为何仍与均匀集中兼容。在特征访问下,计算复杂度与状态空间的大小无关。
英文摘要
We study episodic adversarial linear constrained Markov decision processes (CMDPs) with unknown transitions, where both the loss and constraint functions may vary adversarially across episodes. The best previous algorithm achieves $\widetilde{\mathcal{O}}(K^{3/4})$ regret and cumulative constraint violation, leaving a gap to the optimal $\widetilde{\mathcal{O}}(\sqrt{K})$ dependence on the number of episodes $K$. We close this gap by proposing a new primal dual algorithm that achieves $\widetilde{\mathcal{O}}(\sqrt{K})$ regret and cumulative constraint violation without assuming Slater's condition. We further extend the algorithm to achieve the same $\widetilde{\mathcal{O}}(\sqrt{K})$ guarantees for regret and hard constraint violation, which does not allow constraint violations to cancel across episodes. The main challenge is that learning linear CMDPs requires uniform concentration over a value function class with a controlled covering number, whereas standard techniques in constrained online learning, such as policy mixing, can make this class more complex. Our algorithm combines adaptive Follow the Regularized Leader (FTRL), contracted value estimation, and an exponential Lyapunov function. An adaptive dual regularizer offsets the dependence on the dual weights in the primal regret bound, removing the need for policy mixing. We further show that the normalization in the FTRL update bounds the policy parameters independently of the magnitudes of the dual weights, which explains why the resulting policy class remains compatible with uniform concentration. Under feature access, the computational complexity is independent of the size of the state space.