arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05021cs.ROcs.SYeess.SY

马尔可夫决策过程中基于切换策略的最优约束sc-LTL规划

Optimal Constrained sc-LTL Planning in MDPs via Switching Policies

Zetong Xuan, Yu Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对MDP上带sc-LTL目标与安全约束的规划问题,提出归约为约束可达性问题的新方法,构造切换策略并通过线性规划求最优策略,经网格世界案例验证有效。

中文摘要 AI 辅助

我们研究马尔可夫决策过程(MDPs)上规划问题的最优策略综合,其中目标和安全约束均由共安全线性时序逻辑(sc-LTL)指定。由于sc-LTL规范的复杂性,所研究的问题本质上是非马尔可夫的,可能需要策略随机化来平衡目标与约束。我们提出一种新方法,将约束sc-LTL规划问题归约为扩展模型上的约束可达性问题。随后证明,由各sc-LTL规范的平稳策略构造的一类切换策略,足以满足约束可达性问题的最优性。该发现使我们能通过易处理的线性规划计算最优策略。网格世界案例研究表明,所提切换策略可实现目标与安全约束间的最优权衡,验证了最优性与易处理性。

英文摘要

We study the synthesis of optimal policies for planning problems on Markov decision processes with both objectives and safety constraints specified in co-safe linear temporal logic (sc-LTL). Our problems are inherently non-Markovian due to the complexity of the sc-LTL specification and may require policy randomization to balance the objective and constraint. We propose a novel approach that reduces the constrained sc-LTL planning problem to a constrained reachability problem on an extended model. We then show that a class of switching policies constructed from stationary policies for the individual sc-LTL specifications is sufficient for optimality for the constrained reachability problem. Our finding enables a tractable linear program to compute the optimal policy. A grid world case study demonstrates that our switching policies can achieve the optimal trade-off between the objective and the safety constraint and validates both optimality and tractability.

发表机构

  • University of Florida(佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑