arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26623cs.AI

AgentJudgeBench:用于评估智能体工具调用的多难度基准

AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling

发表机构ServiceNow人工智能部门
查看机构详情
  • ServiceNow AI(ServiceNow人工智能部门)

机构由 AI 辅助整理,请以论文原文为准。

Abhigya Verma, Amit Kumar Saha, Seganrasan Subramanian, Sai Harshitha Aluru

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出AgentJudgeBench基准,评估LLM评判在智能体工具调用工作流上的可靠性,发现其对齐度随难度下降,真实值并非总有益,还给出了实用评估指南。

中文摘要 AI 辅助

大语言模型(LLM)评判被广泛用于评估智能体工具调用系统,但其在结构化、依赖驱动的工作流上的可靠性仍未得到充分研究。我们提出AgentJudgeBench,这是首个针对工作流有向无环图(DAG)上的智能体工具调用的LLM评判可靠性进行系统研究的基准,区别于更广泛的开放式文本或偏好评估类LLM评判任务。该基准包含3808个实例,涵盖6种DAG拓扑结构和3个难度层级,在有真实值和无真实值两种条件下,搭配5种生成器(3B至70B的开源权重模型及GPT-5.4)和6种评判器(20B到前沿规模模型)进行评估。评判器的对齐度随任务难度单调下降,无真实值时下降速度快1.5倍;在无真实值的困难查询中,6种评判器无论规模如何,均收敛到77%-82%的狭窄区间,揭示出结构上限主要由任务难度驱动,尽管其高度部分取决于较弱生成器的提示词,仅靠模型容量无法克服。真实值的作用并非一概有益:它使GPT-5.4的对齐度降低1.5个百分点,使Gemini-2.5-Pro的对齐度降低3.9个百分点,与过度锚定的情况一致。在缓解策略中,思维链推理和评判器温度的影响可忽略不计,而结构化评估 rubric 可使对齐度提升最多6.5个百分点,但在评判器-生成器对之间的通用性不一致。在有真实值时,QwQ-32B与程序化参考的匹配度最佳;人工验证研究显示,GPT-OSS-120B是最符合人类对齐的评判器;无真实值时,前沿评判器在共享上限内仅表现出微弱优势。这些结果揭示了当前LLM评判的基本局限,并为智能体系统的可靠评估提供了实用指南。

英文摘要

LLM judges are widely used to evaluate agentic tool-calling systems, yet their reliability on structured, dependency-driven workflows remains largely unexamined. We present AgentJudgeBench, the first benchmark to systematically study LLM-as-a-judge reliability for agentic tool-calling over workflow DAGs, as distinct from the broader LLM-as-a-judge task of open-ended text or preference evaluation. The benchmark comprises 3,808 instances spanning six DAG topologies and three difficulty tiers, evaluated with five generators (3B-70B open-weight models and GPT-5.4) and six judges (20B to frontier scale) under paired with- and without-ground-truth conditions. Judge alignment degrades monotonically with task difficulty, 1.5x faster without ground truth, and on hard queries without ground truth all six judges converge to a narrow 77-82% band regardless of scale, revealing a structural ceiling driven primarily by task difficulty, though its height is partly prompt-dependent for weaker generators, that model capacity alone cannot overcome. Ground-truth exposure is not uniformly beneficial: it reduces alignment for GPT-5.4 (1.5 pp) and Gemini-2.5-Pro (3.9 pp), consistent with over-anchoring. Among mitigation strategies, chain-of-thought reasoning and judge temperature both have negligible effect, while structured evaluation rubrics improve alignment by up to 6.5 pp but do not generalize uniformly across judge-generator pairs. With ground truth, QwQ-32B best matches the programmatic reference, while a human validation study identifies GPT-OSS-120B as the most human-aligned judge; without it, frontier judges lead only marginally within the shared ceiling. These results expose fundamental limitations of current LLM judges and yield practical guidelines for reliable evaluation in agentic systems.

补充信息

↑