发表机构
Dalian University of Technology(大连理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ParaRecover是一个过程级基准,通过14种错误类型和10,626个实例评估多轮并行工具使用智能体的错误定位与恢复能力,并提出SDE评分标准以提升其反思性恢复能力。
AI 中文摘要
现有的智能体基准主要评估最终任务成功率或工具调用正确性,对智能体能否可靠地诊断并从中级执行失败中恢复提供的洞察有限。这一局限性在多轮并行工具使用场景中尤为关键,因为错误可能跨依赖分支传播并引发级联故障。我们引入了ParaRecover,一个用于评估多轮并行工具使用智能体中错误定位与恢复的过程级基准。该基准基于一个包含14种错误类型的细粒度分类法,涵盖规划依赖、工具选择和参数匹配,包含10,626个实例,跨越两个难度级别。为了实现细粒度的、面向过程的评估,我们进一步提出了SDE评分标准,该标准在智能体执行过程中衡量结构完整性、诊断推理和演化策略。对十多个主流大语言模型的评估显示,即使是最先进的模型在多轮错误传播、隐式工具使用失败和精确重新规划方面仍然存在困难。此外,我们证明了SDE评分标准为改进智能体的反思性恢复能力提供了有效的监督信号。我们的数据和代码可在该https URL获取。
英文摘要
Existing agent benchmarks mainly evaluate final task success or tool-call correctness, providing limited insight into whether agents can reliably diagnose and recover from intermediate execution failures. This limitation becomes particularly critical in multi-turn parallel tool-use scenarios, where errors may propagate across dependent branches and trigger cascading failures. We introduce ParaRecover, a process-level benchmark for evaluating error localization and recovery in multi-turn parallel tool-use agents. Built upon a fine-grained taxonomy of 14 error types covering planning dependencies, tool selection, and argument matching, the benchmark comprises 10,626 instances spanning two difficulty levels. To enable finegrained, process-oriented evaluation, we further propose the SDE rubric, which measures structural integrity, diagnostic reasoning, and evolutionary strategy during agent execution.Experiments across more than ten mainstream LLMs reveal that even state-of-the-art models still struggle with multi-turn error propagation,implicit tool-use failures, and precise replanning. Moreover, we demonstrate that the SDE rubric provides effective supervision signals for improving agents' reflective recovery capabilities. Our data and code are available at https://github.com/gbw206/ParaRecover.
CommentsAccepted to EMNLP 2026 Main Conference