在具有小策略集的不确定马尔可夫决策过程中优化极小极大遗憾
Optimizing Minimax Regret in Uncertain MDPs with Small Sets of Policies
浏览论文内容
中文总结 AI 辅助
针对不确定马尔可夫决策过程的策略数量约束问题,提出k-自适应策略合成算法KAPS,可优化小策略集的极小极大遗憾,实验显示从1个策略增至2个时遗憾降幅最大。
中文摘要 AI 辅助
现实应用中的序贯决策往往涉及环境模型的不确定性,不确定马尔可夫决策过程(UMDPs)将可能的环境表示为一组马尔可夫决策过程(MDPs),这些MDPs共享状态和动作,但转移概率和奖励可能不同。针对所有可能的MDP优化单个策略可能会牺牲性能,而为每个MDP单独优化策略则可能违反可准备和部署的策略数量的操作、监管或可解释性约束。我们考虑模型不确定性在执行前不久得到解决的场景,允许从预先准备的有限策略集中选择最合适的策略。我们引入k-自适应策略合成,它在极小极大遗憾目标下优化这样的k个策略集合。我们证明该问题是NP难的,并开发了KAPS,这是一种具有特定问题边界和启发式方法的精确嵌套分支定界算法,KAPS共同优化哪些MDP共享一个策略以及策略本身。在各种UMDP基准上的实验表明,当从1个策略增加到2个策略时,遗憾始终会出现最大幅度的降低;在单策略场景中,KAPS在解质量上与现有方法具有竞争力,且能更频繁地证明最优性。
英文摘要
Sequential decision-making in real-world applications often involves uncertainty about the environment's model. Uncertain Markov decision processes (UMDPs) represent the possible environments as a set of MDPs with shared states and actions but potentially different transition probabilities and rewards. Optimizing a single policy across all possible MDPs may sacrifice performance, while preparing an individually optimized policy for every MDP may violate operational, regulatory, or interpretability constraints on the number of policies that can be prepared and deployed. We consider settings in which model uncertainty is resolved shortly before execution, allowing the most suitable policy to be selected from a limited set prepared in advance. We introduce $k$-adaptable policy synthesis, which optimizes such a set of $k$ policies under a minimax-regret objective. We prove that the problem is NP-hard and develop KAPS, an exact nested branch-and-bound algorithm with problem-specific bounds and heuristics. KAPS jointly optimizes which MDPs share a policy and the policies themselves. Experiments across various UMDP benchmarks show that the largest reduction in regret consistently occurs when increasing from one to two policies. In the single-policy setting, KAPS is competitive with existing methods in solution quality and proves optimality substantially more often.