arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PluginEval:用于函数调用细粒度错误归因的诊断基准

PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling

Dongjie Xu, Julius, Hanchi Dong, Minghua Tang, Yuxuan Sun, Ziwei Nie, Zicheng Liu, Dujun Qing, Jiajie Xu

arXiv 2608.08700首次发表:更新:

AI 中文总结

PluginEval是缓解现有函数调用基准局限的两阶段诊断基准,可细粒度归因错误,评估显示不同模型在难度级别和错误类别中存在性能差异,且其LLM判断器与人类标注一致性良好。

AI 中文摘要

随着大语言模型越来越多地作为自主智能体运行,工具路由的可靠评估至关重要。当前基准存在三个结构性局限:遵循幂律的数据分布导致罕见场景代表性不足;缺乏对抗性硬负样本掩盖了模型间的性能差异;标注流程依赖未通过执行验证的LLM判断。本文引入PluginEval,这一基准通过两阶段框架系统性缓解上述局限。首先,我们将工具路由形式化为三个决策序列,将生成与验证分离:LLM提出候选调用,而确定性验证与真实API执行提供可靠质量信号。其次,我们按能力、意图和边界分解每个插件,以识别触发与排除场景。随后,我们生成不同难度级别的查询以填补覆盖缺口,包括针对三种失败模式的对抗性负样本,并将其返回至第一阶段进行标注。该过程形成闭环迭代直至覆盖收敛。评估时,我们超越整体准确率:锚定黄金标注的LLM判断将错误分类为漏调用、误调用或参数错误,为每个模型生成详细错误画像。我们评估五个模型系列,包括专有模型与开放权重模型,分析其在不同难度级别和错误类别中的性能,并通过与人类标注的一致性验证该判断器。

英文摘要

Reliable evaluation of tool routing is critical as Large Language Models increasingly operate as autonomous agents. Current benchmarks face three structural limitations: data distributions that follow a power law leave rare scenarios underrepresented; the absence of adversarial hard negatives obscures performance differences across models; and annotation pipelines depend on LLM judgments that have not been validated through execution. In this paper, we introduce PluginEval, a benchmark constructed through a two-stage framework that systematically mitigates these limitations. First, we formulate tool routing as a sequence of three decisions and separate generation from verification. LLMs propose candidate calls, while deterministic validation and real API execution provide reliable quality signals. Second, we decompose each plugin by capability, intent, and boundary to identify trigger and exclusion scenarios. We then generate queries at different difficulty levels to fill coverage gaps, including adversarial negatives targeting three failure modes, and return them to the first stage for annotation. This process creates a closed loop that iterates until coverage converges. For evaluation, we move beyond aggregate accuracy. An LLM judge anchored to gold annotations classifies failures as missed calls, spurious calls, or parameter errors, producing a detailed error profile for each model. We evaluate five model families, including proprietary models and models with open weights, analyze their performance across difficulty levels and error categories, and validate the judge through agreement with human annotations.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑