arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

智能体评估可靠性:更多任务(并不总是)能修复智能体排行榜

Agent Evaluation Reliability: More Tasks Won't (Always) Fix An Agent Leaderboard

Michael Hardy, Ruhana Azam, Anka Reuel, Mykel Kochenderfer, Sanmi Koyejo

arXiv 2610.00651首次发表:更新:

发表机构

Stanford University; University of Illinois Urbana-Champaign(斯坦福大学; 伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出贝叶斯方差分解框架分析智能体排行榜可靠性,发现可靠性依赖测量目标,脚手架影响显著,更多任务收益有限,汇集多样化基准可低成本提升跨任务排名。

AI 中文摘要

智能体评估越来越多地被用于比较大型语言模型(LLM)并为部署决策提供信息,然而排名不仅可能反映模型本身,还可能反映评估条件(如脚手架或任务)的影响。这使得可靠性依赖于声明:一个能可靠地对已部署系统进行排名的评估,可能无法可靠地对底层模型进行排名。我们探讨当前智能体评估能可靠支持哪些结论,以及哪些额外的评估能改进这些结论。我们开发了一个用于稀疏、不平衡智能体排行榜的贝叶斯方差分解框架,并将其应用于来自整体智能体排行榜(Holistic Agent Leaderboard)和港湾指数(Harbor Index)的22个基准。该框架将信号(与预期声明相关的性能差异)与噪声(仍可能改变排名的无关变异)分离开来。我们发现:(1)可靠性取决于测量目标。固定的模型-脚手架系统排名可靠(0.935-0.994),而底层模型的可靠性显著较低(0.148-0.841)。(2)脚手架的选择可以改变结论。脚手架间可靠性衡量脚手架是否保持模型排名,表明脚手架效应在不同评估中差异显著。(3)更多任务无法解决所有不确定性。即使有无限多个类似构造的任务,当不确定性由有限的脚手架覆盖主导时,基准的模型排名可靠性最多提高0.097。(4)汇集多样化的基准可以在较低成本下改善跨任务排名。对于跨多样化智能体任务的排名,汇集基准在相同的任务预算下将预期可靠性从0.44提高到0.75,并可将预期成本降低高达83%。评估设计应遵循预期声明:确定分数或排名应意味着什么,诊断限制其可靠性的因素,并将评估预算花在重要的不确定性来源上。

英文摘要

Agent evaluations are increasingly used to compare LLMs and inform deployment decisions, yet ranks can reflect not only the model but also the effects of the evaluation conditions such as the scaffolds or tasks. This makes reliability claim-dependent: an evaluation that reliably ranks deployed systems may not reliably rank underlying models. We ask which conclusions current agent evaluations reliably support and what additional evaluation would improve them. We develop a Bayesian variance-decomposition framework for sparse, imbalanced agent leaderboards and apply it to 22 benchmarks from the Holistic Agent Leaderboard and Harbor Index. The framework separates signal, performance differences relevant to the intended claim, from noise, irrelevant variation that can still change rankings. We find: (1) Reliability depends on the measurement goal. Fixed model-scaffold systems are ranked reliably (0.935-0.994), while underlying-model reliability is substantially lower (0.148-0.841). (2) Scaffold choice can change conclusions. Inter-scaffold reliability measures whether scaffolds preserve model rankings, showing that scaffold effects vary substantially across evaluations. (3) More tasks cannot resolve all uncertainty. Even infinitely many similarly constructed tasks improve model-ranking reliability of a benchmark by at most 0.097 when uncertainty is dominated by limited scaffold coverage. (4) Pooling diverse benchmarks can improve cross-task rankings at lower cost. For rankings across diverse agentic tasks, pooling benchmarks raises projected reliability from 0.44 to 0.75 at the same task budget and can reduce projected cost by up to 83\%. Evaluation design should follow the intended claim: identify what a score or ranking should mean, diagnose what limits its reliability, and spend evaluation budget on the sources of uncertainty that matter.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑