无连续性假设的CVaR-UCBVI的近极小极大主导阶遗憾
Continuity-Free Near-Minimax Leading-Order Regret for CVaR-UCBVI
浏览论文内容
中文总结 AI 辅助
本文针对有限时间表格型CVaR强化学习,证明Bernstein CVaR-UCBVI算法无需连续性假设即可达到近极小极大主导阶遗憾,其主导项匹配极小极大下界,实现了该问题的近最优算法设计。
中文摘要 AI 辅助
针对有限时间框架下的表格型条件风险价值(CVaR)强化学习,现有研究证明,对于任意归一化回报律,其主导阶遗憾界为\ud835\udf51(τ⁻¹√(SAK));在密度下界条件下,可得到更精确的\ud835\udf51(√(SAK/τ))速率。本文证明,同一Bernstein CVaR-UCBVI算法无需连续性假设即可达到该更精确速率,关键在于选定预算自界:回合缺口的条件方差至多为τ加价值估计宽度。将其代入原始Bernstein分解,以高概率得到任意归一化回报律(包括原子、混合及连续律)下的遗憾界\ud835\udf51(√(SAK/τ)+(SAHK^(1/4)+S²AH)/τ),其中τ⁻¹/²主导项与预期遗憾极小极大下界仅差对数因子。因此,Bernstein CVaR-UCBVI在主导阶 regime 下对全回报律类达到极小极大最优,低阶项仍保留τ⁻¹依赖关系。
英文摘要
For finite-horizon tabular CVaR reinforcement learning, prior work proves a $\widetilde{O}(τ^{-1}\sqrt{SAK})$ leading regret bound for arbitrary normalized return laws and the sharper $\widetilde{O}(\sqrt{SAK/τ})$ rate under a density lower bound. We show that the same Bernstein CVaR-UCBVI algorithm attains the sharper rate without continuity assumptions. The key is a selected-budget self-bound: the conditional variance of the episode shortfall is at most $τ$ plus the value-estimation width. Substitution into the original Bernstein decomposition yields, with high probability, $\widetilde{O}(\sqrt{SAK/τ}+(SAHK^{1/4}+S^2AH)/τ)$ regret for arbitrary normalized return laws, including atomic, mixed, and continuous laws. The $τ^{-1/2}$ leading term matches the expected-regret minimax lower bound up to logarithmic factors. Thus Bernstein CVaR-UCBVI is minimax-optimal over the full return-law class in the leading-order regime; the lower-order terms retain their $τ^{-1}$ dependence.