发表机构
Technion – Israel Institute of Technology(以色列理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对在线POMDP规划中期望成本掩盖高风险状态的问题,提出对每步信念上的即时成本应用CVaR,保留期望累积回报为目标,使任何基于期望的规划器仅通过改变成本计算即可实现风险敏感,并给出有限时间性能保证。
AI 中文摘要
在线POMDP规划器优化期望累积成本,当信念在高成本状态上分配显著质量时,这可能会掩盖危险状态。现有的风险规避方法对价值函数应用静态或动态条件风险价值(CVaR),捕捉轨迹级风险,但存在两个不足:(i)通过将即时成本保留为信念上状态依赖成本的期望,信念内部的风险未被处理;(ii)通过修改价值函数,它们需要新的定制算法,而不是重用现有的基于期望的规划器。我们转而将CVaR应用于每一步信念上的即时成本,直接针对当前状态每步的不确定性。标准期望累积回报被保留为目标,因此所得问题具有标准MDP结构:任何基于期望的POMDP规划器都可以通过仅改变成本计算而变得风险敏感。我们继承了策略评估和稀疏采样的有限时间保证——估计误差与风险水平无关——并且作为我们的核心理论结果,证明了粒子信念MDP替代与原始POMDP之间差距的有限时间界,这共同产生了从真实POMDP价值到算法估计的端到端保证。在风险中性极限下,该公式恢复了标准的基于期望的规划。
英文摘要
Online POMDP planners optimize the expected cumulative cost, which can mask dangerous states when the belief places significant mass on high-cost states. Existing risk-averse methods apply static or dynamic Conditional Value at Risk (CVaR) to the value function, capturing trajectory-level risk, but share two gaps: (i) by retaining the immediate cost as an expectation of a state-dependent cost over the belief, the risk \emph{within} the belief is left unaddressed; and (ii) by modifying the value function, they require new tailored algorithms rather than reusing existing expectation-based planners. We instead apply CVaR to the immediate cost over the belief at each step, directly targeting per-step uncertainty about the current state. The standard expected cumulative return is retained as the objective, so the resulting problem has a standard MDP structure: any expectation-based POMDP planner can be made risk-sensitive by changing only the cost computation. We inherit finite-time guarantees for policy evaluation and sparse sampling---with estimation error independent of the risk level---and, as our central theoretical result, prove a finite-time bound on the gap between the particle belief MDP surrogate and the original POMDP, which together yield an end-to-end guarantee from the true POMDP value to the algorithmic estimate. In the risk-neutral limit, the formulation recovers standard expectation-based planning.