发表机构
Tsinghua University; Shenzhen International Graduate School, Tsinghua University; Tencent Hunyuan(清华大学; 清华大学深圳国际研究生院; 腾讯混元)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长视野智能体轨迹调试的级联错误与关键错误定位难题,提出TrajDebug框架,构建含486条轨迹的TrajErrBench基准,实验显示其性能优于现有基线,可提供改进智能体的可操作反馈。
AI 中文摘要
基于大语言模型(LLM)的智能体系统在复杂领域展现出卓越能力,但存在级联错误和调试困难的问题。关键错误检测旨在定位失败轨迹中导致最终失败的最早错误步骤,然而该领域进展面临两大挑战:其一,长轨迹难以识别单个错误,因为判断某一步骤的证据可能分散在遥远的指令、观测和先前上下文中;其二,失败轨迹通常包含多个具有不同下游影响的局部错误,其中仅部分错误需为最终失败负责。本研究提出TrajDebug,这是一种错误生命周期追踪框架,通过多粒度历史压缩和基于证据的错误识别解决长轨迹错误发现问题,并通过追踪每个错误的解决状态和终端影响支持关键归因。我们进一步构建TrajErrBench,这是一个包含486条手动标注的失败轨迹的基准数据集,源自Tau2Bench和SWE-Bench Pro,涵盖现实工具使用和编码场景。在不同智能体基准上的实验表明,TrajDebug相较于现有基线实现了最佳整体性能,应用研究进一步证明其诊断结果可为提升下游智能体成功率提供可操作的反馈。我们将发布代码和数据以促进进一步研究。
英文摘要
LLM-based agentic systems have shown remarkable capabilities in complex domains, while suffering from cascading errors and difficulty in debugging. Critical error detection aims to locate the earliest error step in a failed trajectory that is responsible for the final failure. However, progress faces two main challenges. First, long trajectories make it difficult to identify individual errors, since the evidence for judging a step may be scattered across distant instructions, observations, and prior context. Second, failed trajectories often contain multiple local errors with different downstream effects, only some of which remain responsible for the final failure. In this work, we propose TrajDebug, an error-lifecycle tracing framework that addresses long-trajectory error discovery with multi-granularity history compression and evidence-based error identification, and supports critical attribution by tracing each error's resolution status and terminal impact. We further construct TrajErrBench, a benchmark of 486 manually annotated failed trajectories from Tau2Bench and SWE-Bench Pro, covering realistic tool-use and coding scenarios. Experiments across diverse agent benchmarks show that TrajDebug achieves the best overall performance over existing baselines, and application studies further demonstrate that its diagnoses provide actionable feedback for improving downstream agent success. We will release the codes and data to facilitate further research.