arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从黑箱到可执行逻辑:通过Prolog专家系统实现可解释强化学习

From Black Box to Executable Logic: Explainable Reinforcement Learning through Prolog Expert Systems

Eduardo C. Garrido-Merchán

arXiv 2607.15459首次发表:更新:

发表机构

Institute of Research in Technology (IIT), Universidad Pontificia Comillas(技术研究所(IIT),庞培法布拉大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究如何让深度强化学习策略从黑箱变为可解释的可执行逻辑程序,提出三阶段转换方法,经理论分析和实验验证,该方法能实现策略转换,在特定任务中表现良好,如在两室任务和部分连续控制任务中达到相应效果。

AI 中文摘要

训练好的深度强化学习策略是一个黑箱,我们探讨能否将其重写为可执行逻辑程序,使其具有可解释性,该程序能再现其行为,可供人阅读、逻辑引擎运行及优化器编辑。我们提出一个三阶段事后转换方法,先提取冻结近端策略优化教师,以经典关系学习方式从其决策中诱导有序规则列表,输出Prolog程序,由现成逻辑引擎执行每个决策;后续扩展阶段编辑规则库,仅当策略评估证明回报增加时才接受编辑。我们证明了四个保证。回报损失边界使提炼程序成为有限马尔可夫决策过程中的机器可检查证书,扩展循环单调改进并终止。对于连续观察设置,我们回答了转换是否可行:命题阈值实例化随着分辨率B增长将网络转换到任意保真度,分歧为O(1/B),回报差距以相同速率缩小,匹配的下界表明对于倾斜决策边界,成本在观察维度上是指数级的。实证上,在有16944个可达状态的两室钥匙与门任务中,扩展后的Prolog程序在每个种子中都获得精确最优回报,在预算受限情况下,十个种子中有十个在精确回报上超过随机教师。在三个连续控制任务中,发出的程序替代网络,在Acrobot上用十一个子句与神经教师在噪声范围内匹配,在CartPole上恢复约97%的回报,而在更精细控制的LunarLander上仅部分恢复,恰好是指数下界预测的上限。

英文摘要

A trained deep reinforcement learning policy is a black box, and we ask whether it can be made explainable by rewriting it as an executable logic program that reproduces its behaviour and that a person can read, a logic engine can run, and an optimizer can edit. We present a three-stage post-hoc transformation that extracts a frozen proximal policy optimization teacher, induces an ordered rule list from its decisions in the manner of classical relational learning, and emits the result as a Prolog program whose every decision is executed by an off-the-shelf logic engine; a subsequent expansion stage edits the rule base and accepts an edit only when policy evaluation certifies a return increase. We prove four guarantees. A return-loss bound makes the distilled program a machine-checkable certificate in a finite Markov decision process, and the expansion loop improves monotonically and terminates. For the continuous-observation setting we answer whether the conversion is possible at all: the propositional threshold instantiation converts the network to arbitrary fidelity as the resolution B grows, with disagreement O(1/B) and a return gap that closes at the same rate, and a matching lower bound shows the cost is exponential in the observation dimension for an oblique decision boundary. Empirically, on a two-room key-and-door task with 16,944 reachable states the expanded Prolog program attains exact optimal return in every seed and, in a budget-capped regime, exceeds the stochastic teacher on exact return in ten of ten seeds. On three continuous-control tasks the emitted program substitutes the network, matching the neural teacher within noise on Acrobot with eleven clauses and recovering about 97% of its return on CartPole, while on the finer-control LunarLander it recovers only partially, exactly the ceiling the exponential lower bound predicts.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑