arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

递归熵风险强化学习在生成模型下的近最优样本复杂度

Near-Optimal Sample Complexity for Recursive Entropic Risk Reinforcement Learning with a Generative Model

Amirparsa Bahrami, Oliver Mortensen, Mohammad Sadegh Talebi

arXiv 2610.06931首次发表:更新:

发表机构

Sharif University of Technology; University of Copenhagen(谢里夫理工大学; 哥本哈根大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对递归熵风险强化学习,在生成模型下提出精细的MB-RS-QVI分析,获得近最优样本复杂度,消除上下界指数差距,仅留多项式差距。

AI 中文摘要

本文研究了在风险参数 \\(\beta\neq 0\\) 的递归熵风险偏好下,有限折扣马尔可夫决策过程(MDPs)中价值学习和策略学习的样本复杂度,假设可以访问 MDP 的生成模型。我们对基于模型的风险敏感 Q 值迭代(MB-RS-QVI)方法进行了精细分析,该方法是一种先前工作中引入的插件式基于模型的方法,并推导了学习最优 \\(Q\\) 值函数和 \\(\varepsilon\\)-最优策略的 \\((\varepsilon,\delta)\\)-PAC 保证。我们的界在有效视界 \\(1/(1-\gamma)\\) 的指数依赖方面优于该设置下现有的最佳保证。特别是,它们在 \\(|\beta|/(1-\gamma)\\) 的指数依赖以及 \\(S\\)、\\(A\\)、\\(\varepsilon\\) 和 \\(|\beta|\\) 方面与现有下界匹配,直至对数因子。因此,我们的分析消除了先前已知上下界之间的指数差距,仅留下有效视界上的多项式差距。

英文摘要

In this paper, we study the sample complexities of value and policy learning in finite discounted Markov decision processes (MDPs) under recursive entropic risk preferences with risk parameter \(β\neq 0\), assuming access to a generative model of the MDP. We provide a refined analysis of model-based risk-sensitive Q-value iteration (MB-RS-QVI), a plug-in model-based method introduced in prior work, and derive \((\varepsilon,δ)\)-PAC guarantees for both learning the optimal \(Q\)-value function and an \(\varepsilon\)-optimal policy. Our bounds improve the exponential dependence on the effective horizon \(1/(1-γ)\) compared with the best existing guarantees for this setting. In particular, they match the existing lower bounds in their exponential dependence on \(|β|/(1-γ)\), as well as in \(S\), \(A\), \(\varepsilon\), and \(|β|\), up to logarithmic factors. Consequently, our analysis removes the exponential gap between the previously known upper and lower bounds, leaving only a polynomial gap in the effective horizon.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑