arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22227cs.LGmath.OC

基于平滑分位数目标的风险敏感强化学习

Risk-Sensitive Reinforcement Learning with Smoothed Quantile Objectives

Mohammad Alipour-Vaezi, Huaiyang Zhong, Sajad Khodadadian

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对强化学习分位数目标的不稳定性问题,提出UCB-BQRL算法,引入缓冲分位数准则与EVI-BQ过程,建立了其regret界与信息论下界,并证明相关评估问题为PP难。

中文摘要 AI 辅助

强化学习(RL)近年来取得了巨大成功,但其经典基础未考虑目标函数的风险敏感性,而风险敏感性在医疗、金融等多个领域至关重要。整合风险敏感性的常用方法是优化累积奖励分布的特定分位数,但精确分位数目标是非平滑的,在回报分布受微小扰动时会突变,导致必须从数据中学习转移模型时难以可靠优化。受这种不稳定性的启发,本文提出UCB-BQRL,这是一种基于模型的乐观学习算法,它维护转移核的置信集,并使用下缓冲分位数准则进行规划。该缓冲准则通过对附近的低分位数取平均来平滑精确分位数目标,从而提高在转移估计误差下的稳定性。为了在每一轮计算缓冲分位数策略,本文引入EVI-BQ,这是一种精确动态规划过程。本文为UCB-BQRL建立了高概率 regret 界,其在对数因子下的缩放为$\u24e2(\u2113^{\tau/\rho_\tau}+H^2\sqrt{SAT})$,其中$\u03c1_\tau$是根级左平台阈值,是依赖于问题的常数。此外,本文为任何处理分位数目标函数的算法的 regret 建立了信息论下界$\u03a9(H/\rho_\tau\sqrt{AT})$。最后,本文证明,即使在两状态、单动作的有限时间MDP中,对于固定策略,精确点分位数评估和精确下缓冲分位数评估在多项式时间图灵归约下是PP难的。

英文摘要

Reinforcement Learning (RL) has achieved tremendous success in recent years. However, the classical foundations of RL do not account for the risk sensitivity of the objective function, which is critical in various fields, including healthcare, finance, etc. A popular approach to incorporate risk sensitivity is to optimize a specific quantile of the cumulative reward distribution. However, exact quantile objectives are non-smooth and can change abruptly under small perturbations of the return distribution, making them difficult to optimize reliably when the transition model must be learned from data. Motivated by this instability, we develop UCB-BQRL, a model-based optimistic learning algorithm that maintains confidence sets for the transition kernel and plans using a lower-buffered quantile criterion. The buffered criterion smooths the exact quantile objective by averaging nearby lower quantiles, thereby improving stability under transition-estimation error. To compute the buffered-quantile policy at each episode, we introduce EVI-BQ, an exact dynamic-programming procedure. We establish a high-probability regret bound for UCB-BQRL, which up to logarithmic factors scales as $\mathcal{O}(\mathrm{e}^{τ/ρ_τ}+H^2\sqrt{SAT})$, where $ρ_τ$ is denoted as the root-level left-plateau threshold, which is a problem-dependent constant. Further, we establish an information-theoretic lower bound of $Ω(H/ρ_τ\sqrt{AT})$ for the regret of any algorithm dealing with a quantile objective function. Finally, we prove that the exact point-quantile evaluation and exact lower-buffered quantile evaluation are PP-hard under polynomial-time Turing reductions, even for a fixed policy in a two-state, one-action finite-horizon MDP.

发表机构

  • Virginia Tech(弗吉尼亚理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑