Harbor适配器与Harbor-Index:面向大规模智能体评估的基础设施及精选元数据集
Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation
浏览论文内容
中文总结 AI 辅助
该研究推出Harbor适配器与Harbor-Index,构建统一智能体评估基础设施,适配80余个基准,评估8类模型,精选82个高难度任务,为语言模型智能体提供可靠全面的评估支撑。
中文摘要 AI 辅助
对数量不断增长的智能体基准进行评估颇具挑战性,因为这些基准往往需要复杂的环境和智能体集成。本文提出Harbor适配器,这是一种用于智能体基准评估的统一基础设施,主要贡献有三点:其一,开发了基准适配器,可将80余个基准适配为可评估任意智能体的形式,并通过严格的代码审查和一致性实验对其进行验证;其二,在54个基准上对涵盖不同能力层级的8个模型开展大规模评估,每个模型均使用Terminus-2及3种原生 harness(评估框架)中的一种运行,从而能比以往更广泛地分析智能体的能力与失败模式;其三,推出Harbor-Index,这是从适配后的基准套件中经难度筛选、AI与人工审核及审核-修复循环优化得到的精选集合,包含29个基准中的82个难度高、多样性强、质量优的任务。Harbor-Index在保留大规模智能体评估的挑战性与广度的同时,运行成本可控,所有评估的模型-harness配置的通过率均未超过30%,最强模型(带Codex的GPT-5.5)的通过率为28.0%。本文将适配器、评估结果、深入分析及Harbor-Index作为开源成果发布,以支持对语言模型智能体开展更可靠、全面的评估。
英文摘要
Evaluating agents on the growing number of agentic benchmarks is challenging because they often require complex environments and agent integrations. We introduce Harbor Adapters, a unified evaluation infrastructure for agentic benchmarks. Our work makes three contributions. First, we develop benchmark adapters that port more than 80 benchmarks to evaluate arbitrary agents, and validate them through rigorous code review and parity experiments. Second, we conduct a large-scale evaluation of 8 models spanning capability tiers across 54 benchmarks; every model is run with Terminus-2 and with one of 3 native harnesses. This enables a broader analysis of agent capabilities and failure modes than was previously possible. Third, we introduce Harbor-Index, a curated set of 82 difficult, diverse, and high-quality tasks spanning 29 benchmarks, refined from the adapted suite through difficulty filtering, AI and human audit, and an audit-and-fix loop. Harbor-Index preserves the challenge and breadth of large-scale agentic evaluations while being affordable to run; no evaluated model-harness configuration exceeds 30% pass rate, and the strongest (GPT-5.5 with Codex) reaches 28.0%. We release the adapters, evaluation results, in-depth analysis, and Harbor-Index as open-source artifacts to support more reliable and comprehensive evaluation of language-model agents.