arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39093cs.LG

通过状态增强学习无限时域平均奖励约束马尔可夫决策过程

Learning Infinite-Horizon Average-Reward CMDPs via State Augmentation

Kihyun Yu, Seoungbin Bae, Dabeen Lee

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出首个计算高效算法,通过状态增强和Huber势重塑奖励,在弱连通假设下实现无限时域平均奖励CMDP的$\sqrt{T}$最优遗憾与约束违反。

中文摘要 AI 辅助

我们研究弱连通假设下的无限时域平均奖励约束马尔可夫决策过程(CMDPs)。现有针对该设置的高概率保证要么需要计算效率低下的算法,要么对交互次数$T$的依赖次优。我们提出了据我们所知第一个在表格设置下以高概率实现$\tilde{\mathcal{O}}(\sqrt{T})$遗憾和累积约束违反的计算高效算法。$\sqrt{T}$的依赖在忽略对数因子的情况下是最优的。我们的方法将累积约束违反纳入状态,并通过Huber势的差分定义重塑奖励。新增状态决定了对进一步违反的惩罚,而奖励函数在增强状态空间上保持不变。由于新增状态具有已知的确定性动态,只需估计原始转移核。Huber势的有界斜率保持每步奖励有界,且势差分可望远镜式求和,将重塑回报与原始累积奖励和终端势联系起来。这些性质使我们能够应用有限时域近似和带裁剪的乐观值迭代,如无约束平均奖励MDP中所用,而不恶化关于$T$的遗憾率。

英文摘要

We study infinite-horizon average-reward constrained Markov decision processes (CMDPs) under the weakly communicating assumption. Existing high-probability guarantees for this setting either require computationally inefficient algorithms or have suboptimal dependence on the number of interactions $T$. We propose, to the best of our knowledge, the first computationally efficient algorithm that achieves $\widetilde{\mathcal{O}}(\sqrt{T})$ regret and cumulative constraint violation with high probability in the tabular setting. The $\sqrt{T}$ dependence is optimal up to logarithmic factors. Our approach incorporates cumulative constraint violation into the state and defines a reshaped reward through differences of a Huber potential. The added state determines the penalty on further violations while the reward function remains fixed on the augmented state space. Since the added state has known deterministic dynamics, only the original transition kernel needs to be estimated. The bounded slope of the Huber potential keeps the per-step reward bounded, and the potential differences telescope to relate the reshaped return to the original cumulative reward and the terminal potential. These properties allow us to apply finite-horizon approximation and optimistic value iteration with clipping, as used in unconstrained average-reward MDPs, without worsening the regret rate in $T$.

发表机构

  • KAIST(韩国科学技术院)
  • Seoul National University(首尔大学)

机构由 AI 辅助整理,请以论文原文为准。

↑