发表机构
Yale University; University of Pennsylvania(耶鲁大学; 宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对平均报酬准则下的分布鲁棒马尔可夫决策过程,推导了极小极大最优样本复杂度界,并提出两种插件约简学习程序以实现该复杂度。
AI 中文摘要
分布鲁棒马尔可夫决策过程为模型不确定性下的序列决策提供了原则性框架。我们研究在平均报酬准则下,学习一个ε-最优鲁棒策略所需的样本数量。生成模型提供来自名义转移核的样本,而策略性能在半径最多为σ的(s,a)矩形总变差不确定性集上进行评估。令H₀和H_σ分别表示名义和鲁棒最优偏差跨度。我们确定σH₀为区分高容忍和低容忍 regime 的扰动尺度。我们的匹配上界和下界表明,忽略对数因子,极小极大总样本复杂度为NSA ≍ (SA/ε²)乘以分段函数,当ε≳σH₀时,取min{H₀,H_σ};当ε≲σH₀时,取min{H₀,H_σ}+σH_σ²。其中S和A分别为状态和动作数量,N为每个状态-动作对的样本数。样本复杂度由类似名义AMDP结果的线性跨度项和仅在低容忍 regime 出现的鲁棒性特定项组成。我们使用基于约简的插件程序达到这些速率,该程序选择约简(名义或鲁棒)及其折扣因子:一种使用已知跨度参数做出这些选择的跨度知情程序,以及一种从数据中校准这两个选择的跨度无关程序。
英文摘要
Distributionally robust Markov decision processes provide a principled framework for sequential decision making under model uncertainty. We study how many samples are necessary and sufficient to learn an $\varepsilon$-optimal robust policy under the average-reward criterion. A generative model provides samples from the nominal transition kernel, whereas policy performance is evaluated over $(s,a)$-rectangular total-variation uncertainty sets of radius at most $σ$. Let $H_0$ and $H_σ$ denote the nominal and robust optimal bias spans, respectively. We identify $σH_0$ as the perturbation scale separating high- and low-tolerance regimes. Our matching upper and lower bounds show that, up to logarithmic factors, the minimax total sample complexity is $$ NSA \asymp \frac{SA}{\varepsilon^2}\begin{cases} \min\{H_0,H_σ\}, & \varepsilon\gtrsimσH_0,\\ \min\{H_0,H_σ\}+σH_σ^2, & \varepsilon\lesssimσH_0. \end{cases} $$ Here $S$ and $A$ are the numbers of states and actions, and $N$ is the number of samples per state-action pair. The sample complexity consists of a linear-span term that resembles the nominal AMDP results and a robustness-specific term that appears only in the low-tolerance regime. We attain these rates using reduction-based plug-in procedures that select the reduction---nominal or robust---and its discount factor: a span-informed procedure that makes these choices using known span parameters, and a span-agnostic procedure that calibrates both choices from data.