arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28960cs.LG

无连续性假设的CVaR-UCBVI的近极小极大主导阶遗憾

Continuity-Free Near-Minimax Leading-Order Regret for CVaR-UCBVI

Yuanlong Chen

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对有限时间表格型CVaR强化学习,证明Bernstein CVaR-UCBVI算法无需连续性假设即可达到近极小极大主导阶遗憾,其主导项匹配极小极大下界,实现了该问题的近最优算法设计。

中文摘要 AI 辅助

针对有限时间框架下的表格型条件风险价值(CVaR)强化学习,现有研究证明,对于任意归一化回报律,其主导阶遗憾界为\ud835\udf51(τ⁻¹√(SAK));在密度下界条件下,可得到更精确的\ud835\udf51(√(SAK/τ))速率。本文证明,同一Bernstein CVaR-UCBVI算法无需连续性假设即可达到该更精确速率,关键在于选定预算自界:回合缺口的条件方差至多为τ加价值估计宽度。将其代入原始Bernstein分解,以高概率得到任意归一化回报律(包括原子、混合及连续律)下的遗憾界\ud835\udf51(√(SAK/τ)+(SAHK^(1/4)+S²AH)/τ),其中τ⁻¹/²主导项与预期遗憾极小极大下界仅差对数因子。因此,Bernstein CVaR-UCBVI在主导阶 regime 下对全回报律类达到极小极大最优,低阶项仍保留τ⁻¹依赖关系。

英文摘要

For finite-horizon tabular CVaR reinforcement learning, prior work proves a $\widetilde{O}(τ^{-1}\sqrt{SAK})$ leading regret bound for arbitrary normalized return laws and the sharper $\widetilde{O}(\sqrt{SAK/τ})$ rate under a density lower bound. We show that the same Bernstein CVaR-UCBVI algorithm attains the sharper rate without continuity assumptions. The key is a selected-budget self-bound: the conditional variance of the episode shortfall is at most $τ$ plus the value-estimation width. Substitution into the original Bernstein decomposition yields, with high probability, $\widetilde{O}(\sqrt{SAK/τ}+(SAHK^{1/4}+S^2AH)/τ)$ regret for arbitrary normalized return laws, including atomic, mixed, and continuous laws. The $τ^{-1/2}$ leading term matches the expected-regret minimax lower bound up to logarithmic factors. Thus Bernstein CVaR-UCBVI is minimax-optimal over the full return-law class in the leading-order regime; the lower-order terms retain their $τ^{-1}$ dependence.

补充信息

↑