arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

梅西耶:用于跨基准智能体评估的高分辨率语料库

Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation

Stefan Krsteski, Charlotte Meyer, Guillaume Allegre, Tony O'Halloran, Alexandre Sallinen

arXiv 2607.25891首次发表:更新:

发表机构

Andromede AI(仙女座人工智能公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对智能体评估中任务等碎片化问题,引入梅西耶语料库,整合多基准、智能体、任务和验证器等信息,标准化记录并分类。通过该语料库发现基准进展不均衡,指出多验证器任务评分问题,得出能力量表,为智能体评估提供基础可复用设施。

AI 中文摘要

在交互式环境中评估人工智能智能体受到任务、支架、验证器和评分规则碎片化的阻碍。现有工作聚焦狭窄场景,规模有限或需昂贵重运行,导致许多经验记录无法比较。我们引入梅西耶,一个包含957,253条记录的统一语料库,涵盖30个基准、714个智能体、11,891个任务和74,205个验证器。梅西耶整合公共基准分数,并通过在六个代表性不足的专业和科学领域(包括最近的法律基准)的五智能体运行进行补充。每条记录按模型、支架、环境、任务、验证器和聚合规则标准化,还有用于职业和行业分析的SOC/NAICS分类。利用该语料库,我们发现前沿进展在基准类型间不均衡,“函数调用”已饱和,“编程”进展最快,“企业工作流程”最具挑战性。此外,反事实重新评分表明多验证器任务中严格的全通过聚合会掩盖进展并人为改变智能体排名。从这些标准化记录中,我们得出与Epoch的评估能力指数排名在Spearman相关系数为0.81时相符的能力量表,且可按领域、职业、动作空间或验证器类型进行细化。梅西耶为智能体能力扩展、基准审核和评估失败的细粒度分析提供了基础的、可重复使用的基础设施。

英文摘要

Comprehensively evaluating AI agents across interactive environments is difficult due to fragmented tasks, scaffolds, verifiers, and scoring rules. Unfortunately, existing efforts to unify these evaluations are limited in scale and domain, making costly reruns necessary and leaving available data incomparable. We introduce MESSIER, a unified corpus of 957,611 records spanning 30 benchmarks, 745 agents, 11,891 tasks, and 74,263 verifiers. MESSIER combines public evaluation results with new runs on six underrepresented professional and scientific benchmarks, standardizing their heterogeneous components into a common schema. Using this corpus, we show that frontier progress is uneven across benchmark groups, with function-calling evaluations largely saturated, programming improving fastest, and enterprise workflows remaining most challenging. Counterfactual rescoring further shows that strict all-pass scoring in multi-verifier tasks can alter agent rankings. Finally, we derive capability scores from our corpus that correlate with Epoch's Evaluation Capability Index rankings at Spearman \r{ho} = 0.84. The scores can also be estimated for subsets defined by domain, occupation, action space, or verifier type. In essence, MESSIER is a reusable resource for studying agent performance at scale, and a basis for designing better evaluations.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑