arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当智能体执行失败时:从遥测数据中检测和定位运行时故障

When Agentic Executions Fail: Detecting and Localizing Runtime Faults from Telemetry

Chenkai Zhang, Yiran Li, Yifang Tian, Michalis Bachras, Hans-Arno Jacobsen

arXiv 2608.14680首次发表:更新:

发表机构

University of Toronto(多伦多大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出 AGENTCHAOSBENCH 基准,在智能体系统遥测数据中检测定位运行时故障,实验显示现有 LLM 基线对该任务的解决效果仍远未达标。

AI 中文摘要

基于大语言模型(LLM)的智能体系统的可靠性是整个执行过程的属性(包括其工具调用、模型调用、安全护栏(guardrails)以及智能体间消息),而非仅取决于最终答案,然而仅评估任务结果几乎无法揭示运行失败的方式或原因。我们提出 AGENTCHAOSBENCH,这是一个用于从智能体系统的执行遥测数据中检测和定位运行时故障的基准。我们运行了五个异构应用程序,这些应用程序通过智能体到智能体协议(Agent-to-Agent protocol)协调智能体,并通过模型上下文协议(Model Context Protocol)调用工具,在其工具、模型、安全护栏以及智能体间边界处注入了十种操作故障(包括不可用或缓慢的工具、损坏或过大的响应、延迟、循环或路由错误的委派,以及绕过安全护栏),同时设置了无故障的对照组。由此产生的数据集包含 275 条经过清理的轨迹:250 条故障执行轨迹,涵盖十种故障类型,以及 25 条无故障对照组轨迹。每条故障轨迹都与相同输入的无故障执行轨迹对齐;故障类型标签以及适用时的位置标签会被从诊断中排除。对于结构化的单轨迹输入,第一组零样本 LLM 基线表明该任务远未解决:参数规模达 140 亿的局部检测器仅达到 13.6%-19.2% 的 top-1 故障类型准确率,而前沿模型 DeepSeek-v4-pro 仅达到 24.8%;同时识别故障类型及其位置的最高准确率为 22%;依赖参考的故障(尤其是绕过安全护栏的故障)在单轨迹场景下仍几乎无法解决。对齐的参考可改善选定的相对故障,但无法解决安全护栏绕过问题。被排除的标签和紧凑的预测格式支持基于 LLM 的方法和非 LLM 诊断方法的可复现比较。

英文摘要

Reliability in LLM-based agentic systems is a property of the whole execution (its tool calls, model calls, guardrails, and inter-agent messages), not of the final answer alone, yet evaluating only task outcomes reveals little about how or why a run fails. We present AGENTCHAOSBENCH, a benchmark for detecting and localizing runtime faults in agentic systems from their execution telemetry. We run five heterogeneous applications that coordinate agents over the Agent-to-Agent protocol and call tools through the Model Context Protocol, and inject ten types of operational fault (unavailable or slow tools, corrupted or oversized responses, and delayed, looped, or misrouted delegations and bypassed guardrails) at their tool, model, guardrail, and inter-agent boundaries, alongside a no-fault control. The resulting dataset contains 275 sanitized traces: 250 faulty executions spanning ten fault types and 25 no-fault controls. Each faulty trace is aligned with the no-fault execution of the same input; fault-type labels and, where applicable, location labels are held out from diagnosis. On structured single-trace inputs, a first set of zero-shot LLM baselines shows the task is far from solved: local detectors up to 14B parameters reach only 13.6-19.2% top-1 fault-type accuracy and the frontier DeepSeek-v4-pro only 24.8%, while jointly identifying the fault type and its location tops out at 22%; reference-dependent faults (above all a bypassed guardrail) stay near-unsolved from a single trace. An aligned reference improves selected relative faults but does not resolve guardrail bypass. The held-out labels and compact prediction format support reproducible comparison of LLM-based and non-LLM diagnosis methods.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑