发表机构
Inria; University of Liverpool; University of Birmingham(法国国家信息与自动化研究所; 利物浦大学; 伯明翰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究一般Borel状态空间上的Restart POMDP,通过充分统计量表示将其简化为完全可观察MDP,证明最优策略在不同成本准则下具有阈值结构,还分析了状态空间偏序时最优阈值的变化规律及平均成本准则下的相关结果。
AI 中文摘要
我们研究一般Borel状态空间上的Restart POMDP(部分可观察马尔可夫决策过程),控制器要么让隐藏状态未被观察地演化,要么重启系统并观察新状态。利用由最后观察到的状态和自重启以来经过的时间构成的充分统计量表示,我们将该问题简化为完全可观察的马尔可夫决策过程(MDP)。在自然的单步成本恶化条件下,我们证明对于折现成本和总非折现成本准则,最优策略在经过时间上具有阈值结构。当状态空间是偏序的且核是随机单调的时,我们进一步表明最优阈值随状态非递增。对于平均成本准则,在几何遍历性和瞬态增益占优的额外假设下,我们在证明最优阈值和相对值函数的一致有界性后,通过消失折现方法建立了类似的阈值结果。
英文摘要
We study a Restart POMDP (Partially Observable Markov Decision Process) on a general Borel state space, where the controller either lets the hidden state evolve unobserved or restarts the system and observes the new state. Exploiting a sufficient-statistic representation consisting of the last observed state and the elapsed time since restart, we reduce the problem to a fully observed MDP. Under a natural one-step cost deterioration condition, we prove that optimal policies have a threshold structure in the elapsed time for both the discounted and total undiscounted cost criteria. When the state space is partially ordered and the kernel is stochastically monotone, we further show that the optimal threshold is nonincreasing in the state. For the average cost criterion, under additional assumptions of geometric ergodicity and domination of the transient gain, we establish analogous threshold results via the vanishing discount approach, after showing the uniform boundedness of the optimal thresholds and relative value functions.
Comments8 pages, the paper has been accepted to IEEE CDC 2026