风险感知通用效用马尔可夫决策过程
Risk-Aware General-Utility Markov Decision Processes
浏览论文内容
中文总结 AI 辅助
研究风险感知通用效用马尔可夫决策过程,聚焦熵风险度量,提出基于蒙特卡洛树搜索的方法,可解决该过程并达任意精度,实验表明此方法在多种任务的GUMDPs中优化风险感知行为时成功。
中文摘要 AI 辅助
我们研究具有风险感知目标的通用效用马尔可夫决策过程(GUMDPs)。在此框架中,智能体旨在优化目标值分布的风险度量,目标函数取决于智能体策略引发的状态访问频率。首先,我们激发、提出并形式化了风险感知GUMDPs,使智能体和决策者能通过风险厌恶权衡预期性能,并受益于GUMDPs框架下丰富的目标集。我们聚焦于熵风险度量(ERM)。其次,展示了如何借助在线规划技术解决带ERM目标的风险感知GUMDPs,提出基于蒙特卡洛树搜索(MCTS)的方法,能以任意期望精度解决风险感知GUMDPs。最后,给出一组实验结果,表明该方法在多种任务(标准MDPs、最大状态熵探索、模仿学习和多目标MDPs)的GUMDPs中优化一系列风险感知行为时是成功的。
英文摘要
We study general-utility Markov decision processes (GUMDPs) with risk-aware objectives. In this framework, an agent aims to optimize a risk measure of the distribution of objective values, where the objective function depends on the frequency of visitation of states induced by the agent's policy. First, we motivate, propose, and formalize risk-aware GUMDPs, which enable agents and decision makers to trade off expected performance by risk aversion while benefiting from the rich set of objectives that can be cast under the framework of GUMDPs. We focus our attention on the entropic risk measure (ERM). Second, we show how we can solve risk-aware GUMDPs with ERM objectives by resorting to online planning techniques. In particular, we propose an approach based on Monte Carlo Tree Search (MCTS) to provably solve risk-aware GUMDPs up to any desired accuracy. Third, we provide a set of experimental results showcasing that our approach is successful when optimizing for a spectrum of risk-aware behaviors in the context of GUMDPs under diverse tasks (standard MDPs, maximum state entropy exploration, imitation learning, and multi-objective MDPs).