arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18744cs.AIcs.CLcs.SE

自动生成的评估指标:从自身盲点进化评估器

Metrics That Write Themselves: Evolving an Evaluator from Its Own Blind Spots

发表机构AWS生成式AI创新中心
查看机构详情
  • AWS Generative AI Innovation Center(AWS生成式AI创新中心)

机构由 AI 辅助整理,请以论文原文为准。

Xing Zhang, Yanwei Cui, Guanghui Wang, Zhihao Lin, Peiyang He

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出EvalCEGAR方法,借鉴程序验证的反例引导抽象精化进化自动评估指标,在MBPP+等基准上缩小评分差距,效果优于人工算子且无需模型调用费用。

中文摘要 AI 辅助

智能体在可靠自动指标下提升迅速,无指标时则停滞不前,而最需要智能体的应用(如报告生成)恰恰是无人知晓如何评分的领域。指标能否自动生成?判断答案优劣难度大,但指出答案存在的错误相对容易,因此本研究进化出的指标是一组小型Python算子,每个算子标记候选答案存在的一种命名缺陷或弃权(不执行),最终通过投票得出结果。直接让模型生成算子不可行:183个候选仅实现96种不同行为,且仅来自庞大空间中的一个狭窄区域。EvalCEGAR则借鉴程序验证中的反例引导抽象精化方法,将算子池视为抽象并搜索碰撞——即两个答案被算子评分完全相同,但一个正确、一个错误。该碰撞对(而非提示)才是生成需求,当碰撞无法被循环中的所有尝试解决时,循环会扩大算子可读取的范围,而非重新采样。在MBPP+和HumanEval+(其隐藏单元测试提供精确真值)上,该循环生成了一个55行的算子,在428个未见过的任务上,该算子将“无标记”与“完美过滤器”之间的差距缩小了15.4%(+0.0065,p=0.0010),且标记量仅为人工编写的最佳算子的四分之一;在未见过的基准上,该算子在三分之一的标记上与人工算子效果完全匹配。8次运行中有6次生成了此类算子,且全部6次均在样本外有效;而将15个人工编写的算子合并为一个过滤器时,准确率反而下降。在几乎不相交的候选集上,相同信息下的LLM评判器与该算子效果相当,但LLM评判器需对每个候选永久收取一次模型调用费用,而该算子无需任何费用。

英文摘要

Agents improve quickly against a reliable automatic metric and stall without one, and the applications that need them most, report generation among them, are the ones nobody knows how to score. Can the metric write itself? Saying what makes an answer good is hard; pointing at something wrong with one is easier, so the metric we evolve is a pool of small Python operators that each flag a candidate for one named defect, or abstain, and vote. Asking a model for operators directly does not work: 183 candidates realise only 96 distinct behaviours, from one narrow region of an enormous space. EvalCEGAR instead borrows counterexample-guided abstraction refinement from program verification. It reads the pool as an abstraction and searches for a collision, two answers the operators score identically, one correct and one not. That pair, not a prompt, is the authoring request, and when a collision defeats every attempt the loop widens what an operator may read rather than resampling. On MBPP+ and HumanEval+, a sandbox whose hidden unit tests give exact ground truth, the loop writes a 55-line operator that closes 15.4% of the gap between flagging nothing and a perfect filter on 428 unseen tasks (+0.0065, p=0.0010) at a quarter of our best hand-written operator's flags. On the benchmark it never saw it matches that operator's effect exactly on a third of the flags. Six of eight runs admit such an operator and all six help out of sample; our 15 hand-written operators applied together as one filter lose accuracy. An LLM judge on the same information ties that delta on a nearly disjoint set of candidates, and charges a model call per candidate forever where the operator charges none.

补充信息

↑