arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21327cs.LGcs.AImath.OC

带缓冲分位数目标的深度强化学习

Deep Reinforcement Learning with Buffered Quantile Objectives

  • Virginia Tech(弗吉尼亚理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Mohammad Alipour-vaezi, Sajad Khodadadian

AI总结:

提出无模型深度强化学习框架Deep-BQRL,通过缓冲分位数目标和集成探索实现风险敏感决策,在资产出售等任务中优于PPO和TRPO。

AI中文摘要:

基于分位数的强化学习通过优化累积收益分布的指定分位数,为风险敏感决策提供了一种可解释的方法。尽管具有吸引力,但在点分位数目标下进行学习具有挑战性:收益分布的微小扰动可能导致分位数发生突变,而精确的分位数敏感规划需要计算量大的分布优化。下缓冲分位数通过平均目标水平正下方的相邻分位数来缓解前一个困难,在保留底层点分位数目标的同时提供更平滑的替代目标。然而,基于这一原理的现有方法仍然是基于模型的,并依赖于显式的收益律规划,限制了其在小型表格问题之外的应用。我们开发了Deep-BQRL,一个无模型的分布强化学习框架,将缓冲分位数学习扩展到神经函数逼近。该方法直接从采样转移中学习条件收益分位数,从学习到的分位数函数的相关区域构建缓冲动作得分,并使用集成分歧来引导探索。增强的输入表示允许学习到的策略响应轨迹信息,而无需显式重现精确规划所需的分位数状态递归。在资产出售最优停止问题和滑动的FrozenLake上的实验将Deep-BQRL与基于模型的UCB-BQRL以及表格PPO和TRPO实现进行了比较。在资产出售中,Deep-BQRL在报告的目标水平下获得了比PPO和TRPO更小的平均累积点分位数策略差距,而UCB-BQRL保持了最小的差距。学习到的停止决策也随目标分位数变化,为该方法的风险敏感行为提供了可解释的说明。

英文摘要:

Quantile-based reinforcement learning provides an interpretable approach to risk-sensitive decision-making by optimizing a prescribed quantile of the cumulative-return distribution. Despite this appeal, learning under a point quantile objective is challenging: quantiles can change abruptly under small perturbations of the return distribution, and exact quantile-sensitive planning requires computationally demanding distributional optimization. Lower-buffered quantiles alleviate the former difficulty by averaging neighboring quantiles immediately below the target level, providing a smoother surrogate while preserving the underlying point-quantile objective. Existing methods based on this principle, however, remain model-based and rely on explicit return-law planning, limiting their applicability beyond small tabular problems. We develop Deep-BQRL, a model-free distributional reinforcement-learning framework that extends buffered-quantile learning to neural function approximation. The method learns conditional return quantiles directly from sampled transitions, constructs buffered action scores from the relevant region of the learned quantile function, and uses ensemble disagreement to guide exploration. An augmented input representation allows the learned policy to respond to trajectory information without explicitly reproducing the quantile-state recursion required by exact planning. Experiments on an asset-selling optimal-stopping problem and slippery FrozenLake compare Deep-BQRL with model-based UCB-BQRL and tabular PPO and TRPO implementations. In asset selling, Deep-BQRL attains smaller mean cumulative point-quantile policy gaps than PPO and TRPO at the reported target levels, while UCB-BQRL retains the smallest gaps. The learned stopping decisions also vary with the target quantile, providing an interpretable illustration of the method's risk-sensitive behavior.

↑