分布鲁棒平均奖励强化学习:弱通信条件下的有限样本保证
Distributionally Robust Average-Reward Reinforcement Learning: Finite-Sample Guarantees under Weak Communication
浏览论文内容
中文总结 AI 辅助
本文研究弱通信平均奖励分布鲁棒强化学习,提出无需先验知识的算法,给出估计鲁棒最优平均奖励和学习的样本复杂度保证,并验证了$n^{-1/2}$收敛速率。
中文摘要 AI 辅助
我们研究了弱通信条件下平均奖励设置中的分布鲁棒强化学习(DR-RL)。我们的主要结果为估计鲁棒最优平均奖励和学习近最优策略提供了有限样本保证,涵盖了SA-矩形和S-矩形结构,以及基于散度和基于距离的不确定集。具体而言,对于Kullback--Leibler和$f_k$-散度球,我们建立了显式半径条件,在这些条件下鲁棒平均奖励贝尔曼方程允许常数增益解;而对于全变差和Wasserstein球,任何正半径都足够,且不需要名义MDP是弱通信的。我们的算法无需先验知识,对于估计鲁棒最优平均奖励,样本复杂度为$\widetilde O(|\mathcal{S}||\mathcal{A}|p_{\wedge}^{-1}\operatorname{Span}^{2}(u_{\delta}^{\ast})\epsilon^{-2})$;对于学习$\epsilon$-最优策略,样本复杂度为$\widetilde O(|\mathcal{S}||\mathcal{A}|p_{\wedge}^{-2}\operatorname{Span}^{2}(u_{\delta}^{\ast})\epsilon^{-2})$。这里,$p_{\wedge}$是最小的正名义转移概率,$u_{\delta}^{\ast}$是鲁棒最优偏差函数。我们进一步提供了$\operatorname{Span}(u_{\delta}^{\ast})$的几乎紧的显式上界。最后,我们通过数值实验验证了预测的$n^{-1/2}$收敛速率。
英文摘要
We study distributionally robust reinforcement learning (DR-RL) in the average-reward setting under weak communication. Our main result provides finite-sample guarantees for estimating the robust optimal average reward and learning a near-optimal policy, covering both SA-rectangular and S-rectangular structures with divergence-based and distance-based uncertainty sets. Specifically, for Kullback--Leibler and $f_k$-divergence balls, we establish explicit radius conditions under which the robust average-reward Bellman equation admits a constant-gain solution, while for total variation and Wasserstein balls, any positive radius suffices without requiring the nominal MDP to be weakly communicating. Our algorithm is prior-knowledge-free and achieves sample complexities of $\widetilde O(|\mathcal{S}||\mathcal{A}|p_{\wedge}^{-1}\operatorname{Span}^{2}(u_δ^{\ast})ε^{-2})$ for estimating the robust optimal average reward and $\widetilde O(|\mathcal{S}||\mathcal{A}|p_{\wedge}^{-2}\operatorname{Span}^{2}(u_δ^{\ast})ε^{-2})$ for learning an $ε$-optimal policy. Here, $p_{\wedge}$ is the smallest positive nominal transition probability and $u_δ^{\ast}$ is a robust optimal bias function. We further provide an almost-tight explicit upper bound on $\operatorname{Span}(u_δ^{\ast})$. Finally, we validate the predicted $n^{-1/2}$ convergence rate through numerical experiments.