爆炸式DecPOMDPs的奇特案例:通过策略计数控制这一“火势”
The Curious Case of Exploding DecPOMDPs: Containing the Fire through Policy Counting
浏览论文内容
中文总结 AI 辅助
针对DecPOMDPs策略空间爆炸问题,本文提出策略计数DecPOMDPs及对应的策略计数动态规划方法,实现智能体数量维度的可处理性。
中文摘要 AI 辅助
分散式部分可观测马尔可夫决策过程(DecPOMDPs)是对不确定性下多智能体决策进行建模的通用框架。然而,DecPOMDPs存在智能体数量相关的指数级复杂度。应对智能体数量导致的难解性的一种方法是研究智能体间表现出某种对称性的智能体划分,通过计数实现紧凑编码。但即便模型复杂度和评估成本降至多项式依赖,策略空间仍会爆炸。本文将焦点从计数智能体重定向为计数策略,这使所谓的策略计数DecPOMDPs在智能体数量维度上具备可处理性(可处理性)。此外,我们提出基于紧凑表示的策略计数动态规划,以高效求解策略计数DecPOMDPs。
英文摘要
Decentralised partially observable Markov decision processes (DecPOMDPs) provide a general framework for modelling multi-agent decision making under uncertainty. However, DecPOMDPs are known to suffer from exponential complexity in the number of agents. One way to combat this intractability in agent numbers is to look at partitions of agents that exhibit a form of symmetry among agents, allowing for a compact encoding by counting. However, a challenge arises as the policy space explodes, even though the model complexity and evaluation cost reduce to a polynomial dependence. In this paper, we redirect our focus from counting agents to counting policies, which actually enables tractability in agent numbers for so called policy-counted DecPOMDPs. Further, we present policy-counted dynamic programming using the compact representation to solve policy-counted DecPOMDPs efficiently.
发表机构
- University of Münster(明斯特大学)
机构由 AI 辅助整理,请以论文原文为准。