arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21209cs.LOcs.AI

策略感知自主智能体的可解释性框架

Explainability Framework for Policy-Aware Autonomous Agents

  • Miami University(迈阿密大学)

机构由 AI 辅助整理,请以论文原文为准。

Heather Merhout, Daniela Inclezan

AI总结:

针对策略感知自主智能体,借鉴社会科学见解,用回答集编程语言结合Python实现可解释性框架,利用违反策略的惩罚创建对比性解释,并通过调查评估框架。

AI中文摘要:

在人工智能领域,智能体是能自主决策以达成目标的系统。随着此类系统在日常生活中愈发普遍,对添加可解释性特征的需求增加。我们提出一个框架,概述如何为策略感知智能体(即在决策框架中纳入规则执行策略的智能体)生成可理解的解释。该框架借鉴社会科学见解设计,用回答集编程语言实现,借助Python辅助信息提取和自然语言翻译。利用智能体违反策略时的惩罚来检测与原行动相反场景中的不良事件,从而创建对比性解释,这是可解释性框架的核心组件。通过人类参与者对程序生成解释提供反馈的调查来评估该框架。

英文摘要:

In the field of Artificial Intelligence, an agent is a system which is able to autonomously make decisions in order to reach a desired goal. As these systems grow more prevalent in our day-to-day lives, there has been an increased need to add explainability features which can provide an account for an agent's behavior. We therefore propose a framework that outlines how to produce comprehensible explanations for policy-aware agents, or agents which have rule-enforcing policies incorporated in their decision-making framework. This framework is designed using insights from the social sciences on how to produce good explanations. It is implemented in the Answer Set Programming language while using Python to assist with information extraction and natural-language translation. Because these agents incur penalties when violating policies, we are able to leverage these penalties to detect undesirable events in scenarios that are counterfactual to the agents' original actions. This lends itself to creating contrastive explanations (e.g., "the agent performed this action because, had it not, undesirable event X would have occurred."), which formulate the core component for our explainability framework. The framework is evaluated using a survey wherein human participants provide feedback on our program-generated explanations.

补充信息

↑