发表机构
Ruhr-University Bochum; Eindhoven University of Technology; Oxford University; Radboud University(波鸿鲁尔大学; 埃因霍温理工大学; 牛津大学; 拉德堡德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对马尔可夫决策过程(MDPs)提出属性驱动因果抽象技术,通过保留模型特性缩减规模,经理论与实验验证,该技术可生成小型抽象模型以计算近最优策略,且能推广至大规模MDP模型。
AI 中文摘要
马尔可夫决策过程(MDPs)是广泛应用的决策模型,通常基于状态变量及其取值在因子化状态空间上进行定义。状态数量的指数级增长使得MDPs中的许多推理任务极具挑战性。抽象是缩减MDPs规模、缓解可扩展性问题的有效技术。本研究针对因子化MDPs引入因果关系概念,提出一种新颖的属性驱动因果抽象技术,该技术可保留原始MDP模型的诸多特性。为此,我们依赖状态变量谓词间的因果关系,识别出那些在满足或违反给定抽象属性方面具有相同原因的状态。我们在理论和实验层面,针对MDPs、区间MDPs或随机博弈等不同模型类型,对多种因果MDP抽象进行了比较。评估结果证明了我们方法的潜力:在多个标准基准测试中,我们获得了小型抽象模型,能够据此计算原始MDP的近最优策略;此外,我们的因果抽象通常可推广至相关的大规模MDP模型。
英文摘要
Markov Decision Processes (MDPs) are widely used as decision-making models, commonly specified over factored state spaces through state variables and their valuations. The exponential blowup in the number of states renders many reasoning tasks in MDPs challenging. Abstractions are promising techniques to reduce MDPs and thus mitigate scalability issues. In this work, we introduce a notion of causality on factored MDPs and a novel property-driven causal abstraction technique that retains many characteristics of the original MDP model. For this, we rely on causal relations over state variable predicates and identify those states that share the same reasons for fulfilling or violating a given abstraction property. We theoretically and empirically compare various causal MDP abstractions using different model types such as MDPs, interval MDPs, or stochastic games. Our evaluation demonstrates the potential of our approach: For several standard benchmarks, we obtain small abstractions that allow us to compute near-optimal policies for the original MDP. Furthermore, our causal abstractions often generalize to related large-scale MDP models.