发表机构
University of Wisconsin-Madison; Microsoft Research(威斯康星大学麦迪逊分校; 微软研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究基于大语言模型的智能体系统失败归因问题,提出OAT方法将其转化为单类学习,用神经控制微分方程建模成功轨迹动态模式,实验表明该方法比基线快且F1分数更高,是诊断智能体系统失败的有效方向。
AI 中文摘要
基于大语言模型的智能体系统的失败归因,即识别失败轨迹中的哪些步骤导致任务失败,对于调试和改进这些系统至关重要。现有方法要么依赖计算成本高的基于提示的管道,要么需要对带有步骤级错误注释的失败轨迹进行训练后处理,而这些注释收集成本高且难以扩展。我们认为实用的失败归因模型应轻量级且无需对失败数据进行步骤级监督即可训练。为此,我们解决无监督失败归因问题,即仅在成功轨迹上训练,并在推理时给定失败轨迹识别错误步骤。我们提出了OAT,将此问题转化为使用神经控制微分方程的单类学习,在潜在空间中对成功轨迹的动态模式进行建模。在推理时,根据失败轨迹中每个步骤与在成功轨迹上学习的动态的偏差为其分配异常分数,然后用于形成一组错误步骤集合。仅在100条成功轨迹上进行训练的实验表明,OAT比基于提示的基线快200至5000倍,并且在域内和分布外数据集中分别以高出20%和7%的F1分数持续优于它们,这表明OAT是诊断智能体系统失败的一个有前途且高效的方向。
英文摘要
Failure attribution for LLM-based agentic systems, i.e., identifying which steps in a failure trajectory caused the task to fail, is critical for debugging and improving these systems. Existing approaches either rely on prompting-based pipelines, which are computationally expensive, or require post-training on failure trajectories with step-level error annotations, which are costly to collect and difficult to scale. We argue that a practical failure attribution model should be lightweight and trainable without step-level supervision on failure data. To this end, we address unsupervised failure attribution, i.e., training exclusively on successful trajectories and identifying error steps at inference time given a failure trajectory. We propose OAT, which casts this problem as one-class learning with neural controlled differential equations, modeling the dynamical pattern of successful trajectories in latent space. At inference time, each step in a failure trajectory is assigned an anomaly score based on its deviation from the dynamics learned on successful trajectories, which is then used to form a set of error steps. With training on only 100 successful trajectories, experiments show that OAT is 200--5000 $\times$ faster than prompting-based baselines, and, at the same time, consistently outperforms them in both in-domain and out-of-distribution datasets with +20% and +7% F1 scores, respectively, demonstrating that OAT is a promising and efficient direction for diagnosing agentic system failures.