arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于参数化动作马尔可夫决策过程的知识与梯度引导强化学习

Knowledge- and Gradient-Guided Reinforcement Learning for Parametrized Action Markov Decision Processes

Jonas Ehrhardt, René Heesch, Oliver Niggemann

arXiv 2607.12924首次发表:更新:

发表机构

HSU-AI Institute for Artificial Intelligence(HSU人工智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究参数化动作马尔可夫决策过程中的强化学习,提出KGRL算法,利用领域知识修剪动作、约束参数空间,结合梯度细化参数,能提供局部解释并提高样本效率,优于现有强化学习基线。

AI 中文摘要

本文研究参数化动作马尔可夫决策过程(PAMDP)中的强化学习,其中每个决策由符号动作和数值参数组成。在这种设置下,强化学习算法通常用一次性估计器确定参数,导致训练样本效率低下。虽然在大多数PAMDP环境中有明确但不完整的知识,却很少直接用于提高强化学习智能体的训练样本效率。我们提出了新颖的神经符号知识与梯度引导强化学习(KGRL)算法。KGRL利用Datalog知识库中的领域知识为给定状态推导适用动作和可行参数集,修剪决策空间中的非适用动作并约束剩余动作的参数空间。然后使用基于梯度的参数细化循环在智能体训练和部署期间估计最优参数。通过记录轨迹上激活的规则,KGRL还提供关于动作修剪和参数约束的局部过程解释。总体而言,KGRL引导智能体的探索和部署朝着可行且有约束意识的决策,同时提高训练期间的样本效率。在样本效率和情节回报方面,KGRL均优于PAMDP的现有强化学习基线。

英文摘要

In this paper, we study Reinforcement Learning in Parametrized Action Markov Decision Processes (PAMDP), where each decision consists of a symbolic action and numerical parameters. In such settings Reinforcement Learning algorithms typically determine parameters with one-shot estimators, which makes their training sample inefficient. Though in most PAMDP environments explicit but incomplete knowledge (e.g., rules, safety constraints, or expert heuristics) is available, it is rarely directly used to increase the sample-efficiency of training Reinforcement Learning agents. We step into this gap and propose our novel Neuro-Symbolic Knowledge- and Gradient-Guided Reinforcement Learning (KGRL) algorithm. KGRL uses domain knowledge in a Datalog knowledge base to derive the set of applicable actions and feasible parameters for a given state. This allows it to prune non-applicable actions from the decision-space and constrain the parameter spaces of the remaining actions. We then use a gradient-based parameter refinement loop to estimate the optimal parameters during training and deployment of the agent. By recording activated rules along the trajectory, KGRL additionally provides local procedural explanations on the pruning of actions and constraining of parameters. Overall, KGRL guides the agent's exploration and deployment toward feasible and constraint-aware decisions, while increasing sample efficiency during training. KGRL outperforms state-of-the-art RL baselines for PAMDPs in both, sample efficiency and episodic return.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑