arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

利用轨迹图实现智能体大语言模型系统的预执行错误诊断

Leveraging Trajectory Graphs for Pre-Execution Error Diagnosis in Agentic LLM Systems

Xu Zheng, Zhuomin Chen, Chaohao Lin, Hua Wei, Haifeng Chen, Wei Cheng, Dongsheng Luo

arXiv 2607.27443首次发表:更新:

发表机构

Florida International University; Arizona State University; NEC Laboratories America; Singapore Management University(佛罗里达国际大学; 亚利桑那州立大学; 美国 NEC 实验室; 新加坡管理大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出Trajectory Graph Copilot框架,以Graph Debugger为核心,通过预执行错误诊断提升LLM智能体完成长周期任务的能力,在多基准测试中平均提升14.69%通过率。

AI 中文摘要

基于大语言模型(LLM)的智能体在各类复杂交互任务中展现出卓越性能,但在具身AI等领域常见的长周期交互任务中仍存在不足。这类场景的复杂性与庞大动作空间会引发复合错误,单个次优动作即可破坏完整轨迹,导致智能体在低效或不可恢复路径上耗尽有限的步骤预算。为在无需高成本微调的情况下解决该问题,研究借鉴软件调试中分析执行日志以提前发现错误的思路,提出名为“Trajectory Graph Copilot”的新型框架,作为LLM智能体的“副驾驶”在动作执行前诊断潜在错误。该框架核心为“Graph Debugger”,将历史轨迹建模为概率图,采用图神经网络识别易引发失败的连续动作模式;作为主动诊断沙盒,该方法对潜在缺陷动作发出早期预警,促使智能体自我修正,通过预执行错误诊断避免高成本失误,显著提升智能体完成长周期任务的能力。在四个基准测试集上针对三个LLM智能体开展的大量实验显示,该方法平均提升了14.69%的通过率。

英文摘要

Large Language Model~(LLM)-based agents have demonstrated exceptional performance across a wide range of complex interactive tasks. However, they often struggle with long-horizon interactive tasks common in domains, such as embodied AI. The complexity and vast action spaces in these settings lead to compounding errors, where a single suboptimal action can derail an entire trajectory, causing the agent to exhaust its limited step budget on inefficient or unrecoverable paths. To overcome this without costly fine-tuning, we draw inspiration from software debugging, where execution logs are analyzed to preemptively catch errors. We propose \textit{Trajectory Graph Copilot}, a novel framework that acts as a ``copilot'' for LLM agents by diagnosing potential action errors before they are executed. At its core,\textit{Graph Debugger} models historical trajectories as a probabilistic graph and uses a Graph Neural Network to identify sequential action patterns that frequently lead to failure. Functioning as a proactive diagnostic sandbox, our method provides early warnings on potentially flawed actions, prompting the agent to self-correct. This pre-action error diagnosis prevents costly mistakes, significantly enhancing the agent's ability to complete long-horizon tasks successfully. The extensive experiments on four benchmarks with three LLM agents demonstrate a $14.69\%$ pass ratio improvement on average.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑