平均奖励马尔可夫决策过程中策略镜像下降的尖锐收敛性与样本复杂度
Sharp Convergence and Sample Complexity of Policy Mirror Descent for Average-Reward MDPs
浏览论文内容
中文总结 AI 辅助
本文针对平均奖励MDPs中的策略镜像下降,提出统一主递归分析,证明精确、非精确表格及线性函数逼近下的线性收敛,并将LFA样本复杂度从$t_{\mathrm{mix}}^5$改进至$t_{\mathrm{mix}}^3/\varepsilon^2$,且证明该复杂度不可改进。
中文摘要 AI 辅助
策略镜像下降(PMD)在折扣马尔可夫决策过程(MDPs)中已有成熟的有限时间理论,但在平均奖励设定下知之甚少,而平均奖励是许多控制应用更自然的目标。我们给出了遍历平均奖励MDPs中PMD的有限时间、有限样本分析,该分析围绕一个单一的主递归展开,该递归控制任何评论家(critic)下的收敛性,无需外部正则化。其特化形式在精确、非精确表格和线性函数逼近(LFA)更新中产生线性速率,其中精确PMD具有超线性区域。我们以端到端样本复杂度阶$t_{\mathrm{mix}}^3/\varepsilon^2$补充这些收敛结果,在表格(依赖于$|S||A|$)和LFA(依赖于$d$)设定中均如此。我们的LFA样本复杂度将先前最佳的$t_{\mathrm{mix}}^5$混合依赖改进为$t_{\mathrm{mix}}^3$,且匹配的信息论下界确立了评论家在两种设定下$t_{\mathrm{mix}}^3/\varepsilon^2$的样本复杂度不可改进。
英文摘要
Policy mirror descent (PMD) has a mature finite-time theory in discounted Markov decision processes (MDPs), but less is known in the average-reward setting, a more natural objective for many control applications. We give a finite-time, finite-sample analysis of PMD in ergodic average-reward MDPs built around a single master recursion that governs convergence for any critic, without external regularization. Its specializations yield linear rates for exact, inexact-tabular, and linear function approximation (LFA) updates, with a superlinear regime for exact PMD. We complement these convergence results with end-to-end sample complexities of order $t_{\mathrm{mix}}^3/\varepsilon^2$ in both tabular ($|S||A|$-dependent) and LFA ($d$-dependent) settings. Our LFA sample complexity sharpens the prior best $t_{\mathrm{mix}}^5$ mixing dependence to $t_{\mathrm{mix}}^3$, and matching information-theoretic lower bounds establish that the critic's $t_{\mathrm{mix}}^3/\varepsilon^2$ sample complexity is unimprovable in both settings.
发表机构
- The Ohio State University(俄亥俄州立大学)
机构由 AI 辅助整理,请以论文原文为准。