arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向协同安全线性时序逻辑规划的无模型精确策略迭代

Exact Model-Free Policy Iteration for Co-safe LTL Planning

Zetong Xuan, Yu Wang

arXiv 2608.05047首次发表:更新:

发表机构

University of Florida(佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对有限马尔可夫决策过程中sc-LTL目标的无模型强化学习问题,提出两步法无模型策略迭代算法,经理论证明与随机网格世界实验验证可收敛至最优策略。

AI 中文摘要

本研究针对有限马尔可夫决策过程中协同安全线性时序逻辑(sc-LTL)目标,开展无模型强化学习研究,该问题可通过标准乘积构造简化为最大可达性目标。针对该问题,直接基于样本的自举方法(如TD或Q学习)可能因Bellman方程解的非压缩性与非唯一性而无法收敛至最优策略。我们提出一种新的两步无模型强化学习方法:首先利用折扣代理识别用于解决该非唯一性的钳位集,随后应用无折扣策略评估与贪心策略改进,且保证能找到最优解。我们证明了策略评估步骤的几乎必然收敛性,以及策略迭代算法在有限步内终止于最优策略的特性。这些理论结果通过随机网格世界上的数值实验得到验证。

英文摘要

This work studies model-free reinforcement learning for co-safe linear temporal logic (sc-LTL) objectives in finite Markov decision processes, which can be reduced to maximal reachability objectives via the standard product construction. For this problem, direct sample-based bootstrap methods (e.g., TD or Q-learning) may fail to converge to optimal policies due to the noncontractive nature and nonuniqueness of solutions to the Bellman equation. We develop a new two-step model-free reinforcement learning method that first uses a discounted surrogate to identify a clamp set that resolves this nonuniqueness, and then applies undiscounted policy evaluation and greedy policy improvement with guarantees of finding an optimal solution. We prove almost-sure convergence of the policy evaluation step and finite termination of the policy iteration algorithm at an optimal policy. These theoretical results are validated through numerical experiments on a stochastic grid world.

Comments7 pages, 1 figure. Accepted for publication in IEEE Control Systems Letters. To be presented at the 2026 IEEE Conference on Decision and Control

DOI:10.1109/LCSYS.2026.3702192

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑