arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.07785cs.AI

LLM智能体排行榜排名究竟在比较什么?

What Does an LLM-Agent Leaderboard Rank Actually Compare?

Wei-Jung Huang

AI总结:

本文研究LLM智能体排行榜分数的估计含义,提出基于估计目标的成对比较程序,并指出在任务差异和不确定性下,排名差异常无法证明优劣性结论。

AI中文摘要:

LLM智能体排行榜会引发一种常见的推断:排名靠前的智能体就是更好的智能体。当系统在任务组合、标签来源、发布细节或成本规则上存在差异时,公开的评估日志可能并不支持这一结论。我们研究排行榜分数估计的是什么,以及何时它们能证明成对优劣性结论的合理性。我们提出的基于估计目标的成对比较程序明确了比较目标和测量来源,检查共同支持,并使用明确的置信规则和实际差异幅度来评估受支持的差异。受控检查在已知有限样本条件下评估决策标签,并表明在判断对目标重新加权的敏感性时,为何必须包含不确定性。在SWE-bench、AgentRewardBench和tau2-bench上,接近的排名差异往往无法解决;代理标签和效用规则也可能改变被选中的系统。DataAgentBench和Open Agent展示了从较粗略的公开记录中仍可估计的内容。排行榜分数总结了已发布的评估,而细粒度的优劣性主张还额外依赖于用于解释差异的估计目标和置信规则。

英文摘要:

An LLM-agent leaderboard invites a familiar inference: an agent ranked above another is the better agent. Public evaluation logs may not support that conclusion when systems differ in task mixture, label source, release detail, or cost rule. We study what leaderboard scores estimate and when they justify pairwise superiority conclusions. Our estimand-aware pairwise procedure states the comparison target and measurement source, checks common support, and evaluates the supported difference using a stated uncertainty rule and practical margin. Controlled checks evaluate the decision labels under known finite-sample conditions and show why uncertainty must be included when judging sensitivity to target reweighting. Across SWE-bench, AgentRewardBench, and tau2-bench, close rank differences are often unresolved; proxy labels and utility rules can also change which system is selected. DataAgentBench and Open Agent show what remains estimable from coarser public records. A leaderboard score summarizes a released evaluation, whereas a fine-grained superiority claim additionally depends on the estimand and uncertainty rule used to interpret the difference.

补充信息

↑