发表机构
Tsinghua University; Microsoft Research; Microsoft; University of Illinois Urbana-Champaign(清华大学; 微软研究院; 微软公司; 伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出神经符号方法AGENTSCOPE,通过将智能体轨迹行为抽象为结构化表示并结合LLM引导推理,在Who&When、AgentErrata数据集上显著提升智能体失败的故障定位与归因准确率。
AI 中文摘要
随着大语言模型(LLM)智能体的普及,理解和诊断智能体失败的能力对于实现卓越效能与可信度至关重要。由于智能体失败常表现为漫长且复杂的轨迹,手动在海量数据中排查问题是不可行的。然而,传统软件漏洞诊断技术难以应对LLM智能体失败,而完全依赖LLM作为判断者会产生不可靠的诊断结果。为克服这些挑战,本文提出AGENTSCOPE,一种用于智能体失败模式诊断的新型神经符号方法。AGENTSCOPE的核心原理是基于智能体轨迹将其行为抽象为结构化表示;此外,AGENTSCOPE引入神经不变量概念以指定智能体行为属性。AGENTSCOPE利用LLM引导的推理,在结构化表示与神经不变量的基础上,精准定位轨迹中的失败步骤及其类型。我们在公开可用的智能体失败数据集Who&When,以及我们创建的更全面的数据集AgentErrata上验证了AGENTSCOPE的有效性,结果显示AGENTSCOPE在故障定位与归因准确率上显著优于当前最优方法。本研究表明,将结构化抽象与LLM引导的推理相结合,可实现对智能体失败的有效、可靠且可解释的诊断。
英文摘要
With the proliferation of LLM agents, the ability to understand and diagnose failures in agents is essential to achieving superior effectiveness and trustworthiness. As agent failures often manifest via long and complex trajectories, manually finding the needles in the haystack is untenable. However, traditional diagnosis techniques for software bugs can hardly address LLM agent failures, while completely relying on LLMs as the judge yields unreliable diagnosis results. To overcome these challenges, this paper presents AGENTSCOPE, a new neuro-symbolic approach for agent failure mode diagnosis. The key principle of AGENTSCOPE is to abstract agent behavior, based on its trajectories, into structured representations. Furthermore, AGENTSCOPE introduces the concept of neural invariants to specify agent behavior properties. AGENTSCOPE leverages LLM-guided reasoning atop the structured representation against neural invariants to pinpoint both the failure step and its type in the trajectory. We show the effectiveness of AGENTSCOPE on publicly available agent failure datasets (Who&When) and a more comprehensive dataset created by us (AgentErrata), where AGENTSCOPE significantly outperforms the current state of the art in fault localization and attribution accuracy. Our work shows that integrating structured abstractions with LLM-guided reasoning enables effective, reliable, and interpretable diagnosis for agent failures.