发表机构
Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出Quasar,首个无模型Q学习算法,用于无最大端分量MDP的可达性,通过时序差分更新实现渐近最优,内存从O(|S|^2|A|)降至O(|S||A|),样本效率显著优于基于模型方法。
AI 中文摘要
强化学习(RL)用于可达性规范是序贯决策的基础。先前的工作确立了向最优策略的渐近收敛性,但仅通过基于模型的方法实现,这些方法必须显式估计底层马尔可夫决策过程(MDP)的转移概率。我们提出Quasar,这是首个在无非终止最大端分量(MECs)的MDP片段上具有可达性渐近保证的无模型算法,而该片段是每个MDP通过标准MEC商归约所得到的基本构建块。我们的算法遵循经典Q学习方法,使用时序差分更新收敛到最优策略,而无需学习转移概率。由此产生的学习器将内存占用从基于模型方法所需的O(|S|^2|A|)减少到O(|S||A|)。在标准化的定量验证基准集上,我们的算法以比先前基于模型的最先进方法少数个数量级的样本收敛到最优策略。这些结果共同构成了向可达性学习实际部署迈出的具体一步,并随之推动了规范引导的RL的实际应用。
英文摘要
Reinforcement learning (RL) for reachability specifications is fundamental to sequential decision-making. Prior work establishes asymptotic convergence to optimal policies, but only through model-based methods that must explicitly estimate the transition probabilities of the underlying Markov Decision Process (MDP). We present Quasar, the first model-free algorithm with asymptotic guarantees for reachability on the fragment of MDPs free of non-terminal maximal end components (MECs), a building block to which every MDP reduces by the standard MEC quotient. Our algorithm follows the classical Q-learning approach, using temporal-difference updates to converge to an optimal policy without ever learning the transition probabilities. The resulting learner reduces the memory footprint from the O(|S|^2|A|) that model-based methods require to O(|S||A|). On the standardized Quantitative Verification Benchmark Set, our algorithm converges to the optimal policy with orders of magnitude fewer samples than the previous model-based state-of-the-art. Together these results are a concrete step toward the practical deployment of reachability learning and, with it, of specification-guided RL.
Comments15 pages, 4 figures