发表机构
UK AI Security Institute; Poznan University of Technology; Generality Labs(英国人工智能安全研究所; 波兹南理工大学; Generality实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究开发AI扫描器检测智能体基准中四类有效性问题,在Inspect Evals基准测试中识别出五个常用基准的多项质量问题,为自动审计基准质量提供了概念验证。
AI 中文摘要
前沿模型的能力常通过智能体基准评估,要信任结果,基准需准确测量其声称的内容且无失效缺陷。此前对 SWE-Bench-Verified 等基准的人工审计已发现 transcript 中的多项有效性问题,但人工审查难以规模化,且自动方法能否可靠暴露损害基准有效性的缺陷尚不明确。本文中,我们开发了 AI 扫描器以检测四类有效性问题:访问 ground truth、工具故障、猜测漏洞及答案格式歧义。我们为每类问题制定了评分规则以指导人工标注,并在 Inspect Evals 基准的保留测试集上,基于人工标签评估扫描器。我们的扫描器在五个广泛使用的基准中识别出多项已验证的质量问题,包括随机人工检查难以发现的案例。并非所有案例都被识别,且扫描器性能在不同基准、标准和模型间差异显著。我们强调了需解决的多项开放挑战,以改进扫描器从而提出更有力的质量保证主张,包括评估领域中削弱扫描器性能的更广泛标准化缺口。总体而言,这些结果为使用自动 transcript 分析更广泛地审计基准质量提供了概念验证。
英文摘要
Capabilities of frontier models are often assessed using agentic benchmarks. To trust these results, benchmarks must accurately measure what they claim to and be free from invalidating flaws. Previous manual audits of benchmarks such as SWE-Bench-Verified have uncovered several validity issues in transcripts. However, manual review is difficult to scale, and it is unclear whether automated methods can reliably surface flaws that compromise benchmark validity. In this paper, we developed AI scanners to detect four types of validity issues: ground truth access, tool failure, guessing vulnerability, and answer format ambiguity. We produced grading rubrics for each to instruct human labeling, and evaluated the scanners against human labels on a held-out test set of Inspect Evals benchmarks. Our scanners identified several verified quality issues in five widely used benchmarks, including cases unlikely to be caught by random manual inspection. Not all cases were identified, and scanner performance varied substantially across benchmarks, criteria and models. We highlight several open challenges to be addressed to improve scanners for stronger quality assurance claims, including broader standardization gaps in the evaluation field that degrade scanner performance. Together, these results serve as a proof of concept for using automated transcript analysis to audit benchmark quality more broadly.
Comments49 pages, 19 figures, Preprint