AgentJudgeBench:用于评估智能体工具调用的多难度基准
AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling
查看机构详情
- ServiceNow AI(ServiceNow人工智能部门)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究提出AgentJudgeBench基准,评估LLM评判在智能体工具调用工作流上的可靠性,发现其对齐度随难度下降,真实值并非总有益,还给出了实用评估指南。
中文摘要 AI 辅助
大语言模型(LLM)评判被广泛用于评估智能体工具调用系统,但其在结构化、依赖驱动的工作流上的可靠性仍未得到充分研究。我们提出AgentJudgeBench,这是首个针对工作流有向无环图(DAG)上的智能体工具调用的LLM评判可靠性进行系统研究的基准,区别于更广泛的开放式文本或偏好评估类LLM评判任务。该基准包含3808个实例,涵盖6种DAG拓扑结构和3个难度层级,在有真实值和无真实值两种条件下,搭配5种生成器(3B至70B的开源权重模型及GPT-5.4)和6种评判器(20B到前沿规模模型)进行评估。评判器的对齐度随任务难度单调下降,无真实值时下降速度快1.5倍;在无真实值的困难查询中,6种评判器无论规模如何,均收敛到77%-82%的狭窄区间,揭示出结构上限主要由任务难度驱动,尽管其高度部分取决于较弱生成器的提示词,仅靠模型容量无法克服。真实值的作用并非一概有益:它使GPT-5.4的对齐度降低1.5个百分点,使Gemini-2.5-Pro的对齐度降低3.9个百分点,与过度锚定的情况一致。在缓解策略中,思维链推理和评判器温度的影响可忽略不计,而结构化评估 rubric 可使对齐度提升最多6.5个百分点,但在评判器-生成器对之间的通用性不一致。在有真实值时,QwQ-32B与程序化参考的匹配度最佳;人工验证研究显示,GPT-OSS-120B是最符合人类对齐的评判器;无真实值时,前沿评判器在共享上限内仅表现出微弱优势。这些结果揭示了当前LLM评判的基本局限,并为智能体系统的可靠评估提供了实用指南。
英文摘要
LLM judges are widely used to evaluate agentic tool-calling systems, yet their reliability on structured, dependency-driven workflows remains largely unexamined. We present AgentJudgeBench, the first benchmark to systematically study LLM-as-a-judge reliability for agentic tool-calling over workflow DAGs, as distinct from the broader LLM-as-a-judge task of open-ended text or preference evaluation. The benchmark comprises 3,808 instances spanning six DAG topologies and three difficulty tiers, evaluated with five generators (3B-70B open-weight models and GPT-5.4) and six judges (20B to frontier scale) under paired with- and without-ground-truth conditions. Judge alignment degrades monotonically with task difficulty, 1.5x faster without ground truth, and on hard queries without ground truth all six judges converge to a narrow 77-82% band regardless of scale, revealing a structural ceiling driven primarily by task difficulty, though its height is partly prompt-dependent for weaker generators, that model capacity alone cannot overcome. Ground-truth exposure is not uniformly beneficial: it reduces alignment for GPT-5.4 (1.5 pp) and Gemini-2.5-Pro (3.9 pp), consistent with over-anchoring. Among mitigation strategies, chain-of-thought reasoning and judge temperature both have negligible effect, while structured evaluation rubrics improve alignment by up to 6.5 pp but do not generalize uniformly across judge-generator pairs. With ground truth, QwQ-32B best matches the programmatic reference, while a human validation study identifies GPT-OSS-120B as the most human-aligned judge; without it, frontier judges lead only marginally within the shared ceiling. These results expose fundamental limitations of current LLM judges and yield practical guidelines for reliable evaluation in agentic systems.