发表机构
University of Florida(佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对有限马尔可夫决策过程中sc-LTL目标的无模型强化学习问题,提出两步法无模型策略迭代算法,经理论证明与随机网格世界实验验证可收敛至最优策略。
AI 中文摘要
本研究针对有限马尔可夫决策过程中协同安全线性时序逻辑(sc-LTL)目标,开展无模型强化学习研究,该问题可通过标准乘积构造简化为最大可达性目标。针对该问题,直接基于样本的自举方法(如TD或Q学习)可能因Bellman方程解的非压缩性与非唯一性而无法收敛至最优策略。我们提出一种新的两步无模型强化学习方法:首先利用折扣代理识别用于解决该非唯一性的钳位集,随后应用无折扣策略评估与贪心策略改进,且保证能找到最优解。我们证明了策略评估步骤的几乎必然收敛性,以及策略迭代算法在有限步内终止于最优策略的特性。这些理论结果通过随机网格世界上的数值实验得到验证。
英文摘要
This work studies model-free reinforcement learning for co-safe linear temporal logic (sc-LTL) objectives in finite Markov decision processes, which can be reduced to maximal reachability objectives via the standard product construction. For this problem, direct sample-based bootstrap methods (e.g., TD or Q-learning) may fail to converge to optimal policies due to the noncontractive nature and nonuniqueness of solutions to the Bellman equation. We develop a new two-step model-free reinforcement learning method that first uses a discounted surrogate to identify a clamp set that resolves this nonuniqueness, and then applies undiscounted policy evaluation and greedy policy improvement with guarantees of finding an optimal solution. We prove almost-sure convergence of the policy evaluation step and finite termination of the policy iteration algorithm at an optimal policy. These theoretical results are validated through numerical experiments on a stochastic grid world.
Comments7 pages, 1 figure. Accepted for publication in IEEE Control Systems Letters. To be presented at the 2026 IEEE Conference on Decision and Control
DOI:10.1109/LCSYS.2026.3702192