arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ToolRobustBench:工具调用智能体的分阶段扰动评估与故障诊断

ToolRobustBench: Stage-Wise Perturbation Evaluation and Failure Diagnosis for Tool-Calling Agents

YiShan Zheng, Yuan Wu, Yi Chang

arXiv 2608.23635首次发表:更新:

AI 中文总结

本研究推出ToolRobustBench,这一工具调用智能体的分阶段诊断基准,通过四类扰动族评估7个模型等的15456个实例,发现工具输出/观测扰动是主要瓶颈,可诊断工具调用故障来源与传播。

AI 中文摘要

大型语言模型(LLMs)将工具调用作为智能体的基础能力,使其能够调用外部系统完成文本生成之外的任务。然而,干净的端到端(E2E)成功无法确定工具使用故障的来源或其在调用过程中的传播方式。我们推出ToolRobustBench,这是一个用于工具调用智能体的分阶段诊断基准,其中工具调用智能体是一种LLM系统,负责选择工具、提供结构化参数并解释其返回的反馈。ToolRobustBench将四类扰动族与工具使用流程对应:工具接口、用户意图、工具输出/观测以及运行时环境扰动。它将故障归因于工具选择、模式适配、参数绑定、工具输出/运行时反馈处理以及端到端任务成功。对7个模型、16个采样本地工具、4类扰动族及14个子类型的15456个单实例的实验显示,干净性能较高但不均匀,且稳健性大幅下降,其中工具输出/观测扰动是主要瓶颈。混合族实验揭示了非加性故障模式,无法通过孤立的单族结果解释。因此,ToolRobustBench提供了一个确定性且级联感知的基准,用于诊断超越干净工具调用准确率的稳健性;

英文摘要

Large language models (LLMs) rely on tool calling as a fundamental agent capability, enabling them to invoke external systems and complete tasks beyond text generation. However, clean end-to-end (E2E) success cannot identify where a tool-use failure originates or how it propagates through a call. We introduce ToolRobustBench, a stage-wise diagnostic benchmark for tool-calling agents, where a tool-calling agent is an LLM system that selects a tool, supplies structured arguments, and interprets its returned feedback. ToolRobustBench aligns four perturbation families with the tool-use pipeline: tool-interface, user-intent, tool-output/observation, and runtime-environment perturbations. It attributes failures to tool selection, schema grounding, argument binding, tool-output/runtime-feedback handling, and E2E task success. Experiments on 15,456 single-family instances across 7 models, 16 sampled local tools, 4 perturbation families, and 14 subtypes show high but non-uniform clean performance and substantial robustness degradation, with tool-output/observation perturbation the dominant bottleneck. Mixed-family experiments reveal non-additive failure patterns that are not explained by isolated single-family results. Thus, ToolRobustBench provides a deterministic and cascade-aware benchmark for diagnosing robustness beyond clean tool-calling accuracy;

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑