发表机构
University of Illinois Urbana-Champaign; University of Toronto; Google; Stanford University(伊利诺伊大学厄巴纳-香槟分校; 多伦多大学; 谷歌公司; 斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型代理故障调试难题,提出开源框架AgentDebugX,其核心DeepDebug通过多轮诊断定位根源,在基准测试中表现出色,能修复更多失败任务,还通过多种方式公开工作流程并提供错误中心。
AI 中文摘要
大语言模型代理故障难以调试,因为错误出现的步骤往往并非其根源。现有可观测性工具重放执行跟踪,但在识别根本原因或将诊断转化为恢复方面支持不足。我们提出AgentDebugX,一个开源调试框架,将调试组织成检测、归因、恢复和重新运行的闭环。其核心DeepDebug通过全局轨迹理解、结构引导调查和交叉检查进行多轮根本原因诊断。在Who and When基准测试中,DeepDebug在两个测试的开放权重主干上的评估方法中达到最佳严格归因准确率。在GAIA上,DeepDebug在单次重新运行中修复了73个失败任务中的13个。AgentDebugX通过多种方式公开此工作流程,并提供错误中心用于共享和重用故障诊断修复包。
英文摘要
LLM agent failures are difficult to debug because the step where an error surfaces is often not the one that caused it. Existing observability tools replay execution traces but provide little support for identifying the root cause or translating diagnosis into recovery. We present AgentDebugX, an open-source debugging framework that organizes debugging as a closed loop of Detect, Attribute, Recover, and Rerun. At its core, DeepDebug performs multi-turn root-cause diagnosis through global trajectory understanding, structure-guided investigation, and cross-examination. On the Who and When benchmark, DeepDebug achieves the best strict attribution accuracy among the evaluated methods on both tested open-weight backbones, reaching 28.8 percent exact agent-and-step accuracy on qwen3.5-9b versus 21.7 percent for the strongest single-pass baseline. On GAIA, DeepDebug repairs 13 of 73 failed tasks in a single rerun, compared with 4 to 6 for three decoupled self-correction baselines, improving overall accuracy from 55.8 percent to 63.6 percent. AgentDebugX exposes this workflow through a Python library, CLI, web console, and installable agentic skill, and provides an opt-in Error Hub for sharing scrubbed failure-diagnosis-repair bundles and reusing them as debugging memory.