arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

具有动态UBSR测度的马尔可夫决策过程的在线策略评估

Online Policy Evaluation for MDPs with Dynamic UBSR Measures

Weikai Wang, Erick Delage

arXiv 2607.23030首次发表:更新:

发表机构

GERAD & Department of Decision Sciences, HEC Montréal(GERAD与蒙特利尔高等商学院决策科学系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究在具有动态UBSR测度的MDPs中进行策略评估,提出UBSR-TD算法及变体,通过纳入损失函数使现有算法可适应动态UBSR设置,经实验验证了方法有效性。

AI 中文摘要

开发用于策略评估的高效函数近似方法是风险感知强化学习中的一项基本挑战。现有方法要么专注于受限的风险测度类别,要么依赖模拟器,限制了它们在完全在线设置中的适用性。在这项工作中,我们提出了计算高效的在线学习算法,用于在具有动态基于效用的短缺风险(UBSR)测度的马尔可夫决策过程(MDPs)中进行策略评估,采用线性函数近似。具体而言,我们引入了UBSR-TD算法,建立了其几乎必然收敛的条件,并开发了几种旨在加速收敛的变体。我们的公式表明,通过将损失函数纳入时间差分误差,现有的风险中性MDPs的策略评估算法可以很容易地适应动态UBSR设置。数值实验支持了我们的理论发现,并且在具有保质期不确定性的易腐库存管理问题中的应用证明了所提出方法的实际有效性。

英文摘要

Developing efficient function-approximation methods for policy evaluation is a fundamental challenge in risk-aware reinforcement learning. Existing approaches either focus on restrictive classes of risk measures or rely on access to a simulator, limiting their applicability in fully online settings. In this work, we propose computationally efficient online learning algorithms for policy evaluation in Markov decision processes (MDPs) with dynamic utility-based shortfall risk (UBSR) measures under linear function approximation. Specifically, we introduce the UBSR-TD algorithm, establish conditions under which it converges almost surely, and develop several variants designed to accelerate convergence. Our formulation shows that existing policy evaluation algorithms for risk-neutral MDPs can be readily adapted to dynamic UBSR settings by incorporating a loss function into the temporal-difference error. Numerical experiments support our theoretical findings, and an application to a perishable inventory management problem with shelf-life uncertainty demonstrates the practical effectiveness of the proposed methods.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑