arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06063cs.AI

通过执行轨迹解释AI智能体

Explaining AI Agents Through Execution Traces

Vittoria Vineis, Fabiano Veglianti, Lorenzo Antonelli, Claudia Di Carlo, Matteo Silvestri, Gabriele Tolomei

首次发表
浏览论文内容

中文总结 AI 辅助

提出一种基于执行轨迹的事后XAI框架,将智能体行为转化为结构化报告和忠实解释,跨架构适用,优于朴素LLM解释。

中文摘要 AI 辅助

AI智能体越来越多地被部署在现实世界环境中,它们与外部工具交互并做出顺序决策,而人类监督有限。这产生了对智能体做了什么以及为什么做的可靠且可审计的解释的迫切需求。然而,传统的可解释人工智能(XAI)方法无法为此类交互式、多步骤系统提供所需的过程级透明度,从而推动了针对AI智能体专门设计的方法的范式转变。为弥补这一差距,我们提出了一个事后XAI框架,该框架将智能体的冗长执行轨迹转换为结构化报告和忠实于其可观察行为的自然语言解释。由于该框架仅依赖执行轨迹,因此适用于不同的智能体架构、环境和任务。跨多个基准和架构的人类与自动化评估表明,我们的框架能生成高质量、轨迹忠实的解释,同时可靠地识别无依据的声明、不合理的行为和证据缺口,优于朴素的LLM生成解释。

英文摘要

AI Agents are increasingly deployed in real-world settings, where they interact with external tools and make sequential decisions with limited human oversight. This creates a pressing need for reliable and auditable explanations of what an agent did and why. However, traditional Explainable AI (XAI) methods fall short of providing the process-level transparency required for such interactive, multi-step systems, motivating a paradigm shift toward approaches specifically designed for AI Agents. To address this gap, we present a post-hoc XAI framework that transforms a lengthy agent's execution trace into a structured report and a faithful natural-language explanation explicitly grounded in its observable behavior. Because it relies solely on execution traces, the framework applies across different agent architectures, environments, and tasks. Human and automated evaluations across multiple benchmarks and architectures show that our framework produces high-quality, trace-faithful explanations while reliably identifying unsupported claims, unjustified actions, and evidence gaps, outperforming naive LLM-generated explanations.

发表机构

  • Sapienza University of Rome(罗马第一大学)

机构由 AI 辅助整理,请以论文原文为准。

↑