EvalResearchBench:AI代理能否自行设计评估?
EvalResearchBench: Can AI Agents Design Their Own Evaluations?
浏览论文内容
中文总结 AI 辅助
本研究提出EvalResearchBench基准,探究AI代理能否自主设计评估器,实验表明最佳评估器与目标基准的排序一致性约75%,但仍有改进空间。
中文摘要 AI 辅助
递归自我改进(RSI)依赖于评估反馈来评估进展并指导进一步研究,然而反复运行复杂基准测试成本高昂且减缓迭代速度。人类专家通过选择基准子集或设计紧凑的套件来降低这一成本。我们探究AI代理能否自动化这一设计过程,并引入EvalResearchBench(ERB),一个用于自主评估研究的基准。给定目标材料、开发参考、候选API以及固定的时间和API预算,一个被称为研究者的代理选择或合成任务,实现评分器,并在试点测试中对其进行修订,然后冻结一个可执行的评估器,用于编码、协作和推理。我们研究了9个研究者和13个候选模型,并将每个冻结的评估器与14个目标基准在分数一致性和成对一致性上进行比较。最好的评估器对约75%的候选对进行排序,与目标一致,低于目标间分歧所设定的91%上限。没有哪个研究者在所有指标上领先,且在开发目标上表现最佳的评估器并非在隐藏于研究者的密封目标上表现最佳。一个由人类设计的公开任务样本仍然是一个强大的基线,且执行成本最低的评估器达到了最高的成对一致性。代理通过试点反馈修复任务和评分器,但它们的评估器仍可能截断答案、耗尽评估预算,或让少数问题主导一个领域的得分。
英文摘要
Recursive self-improvement (RSI) relies on evaluation feedback to assess progress and guide further research, yet repeatedly running complex benchmarks is costly and slows iteration. Human experts reduce this cost by selecting benchmark subsets or designing compact suites. We ask whether AI agents can automate this design process and introduce EvalResearchBench (ERB), a benchmark for autonomous evaluation research. Given target materials, development references, candidate APIs, and fixed time and API budgets, an agent called the researcher selects or synthesizes tasks, implements graders, and revises them in pilot tests before freezing an executable evaluator for coding, co-work, and reasoning. We study 9 researchers and 13 candidate models and compare each frozen evaluator with 14 target benchmarks on score concordance and pairwise agreement. The best evaluators order about 75\% of candidate pairs as the targets do, below the 91\% ceiling set by disagreements among the targets. No researcher leads on every metric, and the best evaluator on development targets is not the best on sealed targets hidden from the researcher. A human-designed sample of public tasks remains a strong baseline, and the evaluator with the lowest recorded execution cost attains the highest pairwise agreement. Agents repair tasks and graders through pilot feedback, yet their evaluators can still truncate answers, exhaust the evaluation budget, or let a few questions dominate a domain score.
发表机构
- Oregon State University(俄勒冈州立大学)
- AG2 AI
- University of California, San Diego(加州大学圣迭戈分校)
- Pennsylvania State University(宾夕法尼亚州立大学)
机构由 AI 辅助整理,请以论文原文为准。