多少任务足以做出智能体基准决策?对公共大语言模型智能体基准的重放分析
How Many Tasks Are Enough for Agent Benchmark Decisions? A Replay Analysis of Public LLM Agent Benchmarks
浏览论文内容
中文总结 AI 辅助
研究智能体基准测试中多少任务足以做决策,通过重放公共任务级记录,提出部分预算需满足支持完整基准决策、覆盖任务组及控制未解决比较比例等条件,还指出部分评估报告应包含的内容。
中文摘要 AI 辅助
智能体基准测试通常在所有任务运行后比较两个智能体,但昂贵的评估使得部分运行很诱人。仅任务比例并不能表明部分运行是否支持与完整基准测试相同的成对结论。我们通过重放来自SWE-bench、AppWorld和tau-bench的已完成公共任务级记录来研究这个问题。只有当部分预算支持完整基准测试的决策、覆盖所需任务组且未解决的比较不超过目标比例时,才被视为足够。所需任务比例差异很大。在5个百分点预算网格上严格的0个百分点阈值下,AppWorld在15%时首次满足所有目标,tau-bench在25%时满足,SWE-bench Verified在90%时满足;SWE-bench Lite在主要覆盖规则下95%时未满足所有目标。部分评估报告应说明一个智能体必须比另一个智能体表现好多少、任务如何选择、需要什么覆盖规则、使用什么决策规则以及可能有多少比较未解决。
英文摘要
Agent benchmarks often compare two agents after all tasks have run, but costly evaluations make partial runs tempting. A task fraction alone does not show whether a partial run supports the same pairwise conclusion as the completed benchmark. We study this question by replaying completed public task-level records from SWE-bench, AppWorld, and tau-bench. A partial budget counts as enough only when it supports the completed benchmark's decision, covers required task groups, and leaves no more than a target fraction of comparisons unresolved. The required task fraction varies sharply. At the strict 0 percentage point threshold on a 5 percentage point budget grid, AppWorld first meets all targets at 15 percent, tau-bench at 25 percent, and SWE-bench Verified at 90 percent; SWE-bench Lite does not meet all targets by 95 percent under the primary coverage rule. Partial-evaluation reports should state how much one agent must outperform another, how tasks are selected, what coverage rule is required, what decision rule is used, and how many comparisons may remain unresolved.