基于LP的子模定向问题的强且紧凑的马尔可夫决策过程策略
Strong and Compact Policies for Submodular Markov Decision Processes via LP-Based Submodular Orienteering
浏览论文内容
中文总结 AI 辅助
针对子模马尔可夫决策过程,提出基于LP的子模定向算法,利用Sherali-Adams层次和Round-or-Cut思想,实现多项式时间$O(n^{\varepsilon})$近似,并揭示近似比与历史依赖的权衡。
中文摘要 AI 辅助
为马尔可夫决策过程(MDPs)寻找策略是强化学习和运筹学等领域中的核心问题。在此问题中,我们需要反复选择智能体应执行的动作。根据动作和智能体的当前状态,智能体获得奖励并随机转移到新状态。目标是在长度为$H$的有限时间范围内最大化期望奖励。我们考虑最近引入的一个变体,它将模型中传统的可加奖励函数推广为单调子模函数,从而能够捕捉一系列有趣的应用。在没有随机成分的情况下,该问题等价于子模定向问题,其目标是在有向图中找到一条$s$-$t$路径,在长度约束下最大化单调子模函数。我们提出了一种新颖的基于LP的子模定向算法,利用了Sherali-Adams层次和Round-or-Cut的思想。我们的保证与已知的子模定向拟多项式时间对数近似相当,但也扩展到子模马尔可夫决策过程的设置。在多项式时间范围内,对于每个$\varepsilon >0$,我们给出了$O(n^{\varepsilon})$-近似(对于子模MDP为$O(H^{\varepsilon})$),其中$n$是顶点数,这即使对于子模定向问题也是未知的。在我们工作之前,子模MDP的最佳已知近似保证的近似比与$H$成线性关系。除了这些算法结果之外,我们的方法还揭示了近似保证与智能体决策所依赖的先前访问顶点数量之间的权衡。
英文摘要
Finding policies for Markov Decision Processes (MDPs) is a central problem in areas such as Reinforcement Learning and Operations Research. Here, we have to repeatedly choose an action that should be performed by an agent. Depending on the action and the current state of the agent, the agent collects a reward and randomly transitions into a new state. The goal is to maximize the reward in expectation over a finite time horizon of length $H$. We consider a recently introduced variant that generalizes the traditionally additive reward function in the model to a monotone submodular one, which allows for capturing a range of interesting applications. Without the stochastic component, this problem is equivalent to the Submodular Orienteering problem, where the goal is to find an $s$-$t$ walk in a directed graph maximizing a monotone submodular function under a length constraint. We present a novel LP-based algorithm for Submodular Orienteering using ideas from the Sherali-Adams hierarchy and Round-or-Cut. Our guarantees are comparable to the known quasi-polynomial time logarithmic approximation for Submodular Orienteering, but also extend to the setting of Submodular Markov Decision Processes. In the polynomial time regime, we present an $O(n^{\varepsilon})$-approximation (and $O(H^{\varepsilon})$ for Submodular MDPs) for every $\varepsilon >0$, where $n$ is the number of vertices, which was unknown even for Submodular Orienteering. Prior to our work, the best known approximation guarantee for Submodular MDPs had an approximation ratio linear in $H$. Beyond these algorithmic results, our methods reveal a trade-off between the approximation guarantee and the number of previously visited vertices on which an agent conditions its decision.
发表机构
- University of Southern Denmark(南丹麦大学)
- ETH Zurich(苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。