arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.13655cs.AI

通过归纳逻辑编程解释强化学习智能体

Explaining Reinforcement Learning Agents via Inductive Logic Programming

Celeste Veronese, Edoardo Zorzi, Daniele Meli, Alessandro Farinelli

首次发表
浏览论文内容

中文总结 AI 辅助

该研究致力于推进可解释强化学习,引入客观指标量化策略可解释性。采用归纳逻辑编程提取符号表示并定义新指标,实验表明这些指标能突出特定动作学习动态,提供细粒度洞察,揭示多智能体模式,助力策略转移与泛化。

中文摘要 AI 辅助

可解释强化学习(XRL)旨在使强化学习(RL)策略更透明和可解释,这在安全关键和以人为本的场景中是关键要求。然而,它大多基于用户研究,缺乏共享评估指标。基于逻辑的可解释人工智能(XAI)方法提供了紧凑、可读的决策抽象,但逻辑表示的可解释性程度的系统量化仍是开放问题。本文旨在通过引入RL设置中策略可解释性的客观和面向规划的指标来推进XRL的技术水平。同时,通过提供一种原则性方法来量化逻辑规则的可解释性,为XAI的逻辑领域做出贡献。我们采用归纳逻辑编程(ILP)提取RL策略的符号表示并定义一组新的可解释性指标,包括激活率、特征覆盖率、句法距离和语义距离。不同RL领域的实验表明,所提出的指标突出了超出全局回报的特定动作学习动态,提供了超越经典全局特征重要性估计方法的对领域特征的细粒度洞察,并揭示了多智能体RL中的协调、专业化和适应模式。此外,它们为特定动作策略的转移和泛化提供了关键见解。

英文摘要

Explainable Reinforcement Learning (XRL) seeks to make Reinforcement Learning (RL) policies more transparent and interpretable, a key requirement in safety-critical and human-centric scenarios. However, it is mostly based on user studies, thus targeting the needs of a specific audience and lacking shared evaluation metrics. On the other hand, logic-based approaches within eXplainable Artificial Intelligence (XAI) provide compact, human-readable abstractions of decision-making. However, the systematic quantification of the explainability degree of logical representations remains an open problem. This work aims to advance the state of the art in XRL by introducing objective and planning-oriented metrics for policy explainability in RL settings. At the same time, it contributes to the field of logic for XAI by providing a principled way to quantify the explainability of logical rules, moving beyond common-sense assessments and simple propositional fragments. We employ Inductive Logic Programming (ILP) to extract symbolic representations of RL policies and define a novel set of explainability metrics, including activation rate, feature coverage, syntactic distance and semantic distance. These metrics quantify alignment between symbolic rules and agent behavior, the role of features in decision-making, and the evolution of policies during training and across agents in single and multi-agent RL. Experiments across different RL domains show that the proposed metrics highlight action-specific learning dynamics beyond global return, provide fine-grained insights into domain features beyond classical approaches for global feature importance estimation, and uncover coordination, specialization, and adaptation patterns in MARL. Moreover, they provide crucial insights for the transfer and generalization of action-specific policies.

发表机构

  • University of Verona(维罗纳大学)
  • Sapienza University of Rome(罗马第一大学)

机构由 AI 辅助整理,请以论文原文为准。

↑