发表机构
Univ Rennes; Inria; CNRS; IRISA; University of Sydney; Université Libre de Bruxelles; Rice University(雷恩大学; 法国国家信息与自动化研究所; 法国国家科学研究中心; 雷恩信息与自动化研究所; 悉尼大学; 布鲁塞尔自由大学; 莱斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出用于折扣和收益的分类器自动机,将其应用于马尔可夫决策过程,可综合出针对不连续效用函数的最优或ε-最优策略,且证明该综合问题是PSPACE难的。
AI 中文摘要
将连续数据分类到离散区间是人工智能中的基础操作。我们提出分类器自动机,这是一种确定性自动机,可读取无限奖励序列并确定其折扣和属于有限多个区间中的哪一个。分类器自动机是比较器自动机的推广,比较器自动机是仅含两个区间的特殊情况,已在定量综合中被证明有用。我们的主要技术贡献是构建了一种分类器自动机,其状态空间大小与区间数量呈线性关系,而非通过比较器自动机的叉积得到的指数级大小。随后我们将分类器自动机应用于马尔可夫决策过程,利用它可综合策略,以最大化效用函数可能不连续时折扣和收益的期望效用。对于分段常数效用函数,所得算法是精确的,且运行时间为伪多项式;对于分段Lipschitz效用函数(这类函数包括所有在有限个跳跃点间斜率有界的效用函数),算法仍为伪多项式时间,并能生成ε-最优策略。我们还证明,所研究的综合问题即使对于分段常数效用函数也是PSPACE难的。
英文摘要
Categorizing continuous data into discrete bins is a fundamental operation in artificial intelligence. We introduce the categorizer automaton, a deterministic automaton that reads an infinite sequence of rewards and identifies which of finitely many bins contains its discounted sum. Categorizer automata generalize comparator automata, the special case of two bins, which have already proven useful in quantitative synthesis. Our main technical contribution is the construction of a categorizer automaton whose state space is linear in the number of bins, rather than exponential as obtained by a cross-product of comparator automata. We then apply categorizer automata to Markov decision processes, where they allow one to synthesize policies that maximize the expected utility of a discounted-sum payoff for utility functions that may be discontinuous. For piecewise-constant utility functions, the resulting algorithm is exact and runs in pseudo-polynomial time. For piecewise-Lipschitz utility functions, a class that includes any utility with bounded slope between finitely many jumps, it again runs in pseudo-polynomial time and yields an $\varepsilon$-optimal policy. We also show that the synthesis problem considered is PSPACE-hard already for piecewise-constant utilities.