arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.12790cs.AIcs.CLcs.MA

谁来评估评估者?用于自我改进的语言模型代理的协同进化评估指标和技能

Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents

  • Amazon(亚马逊)

机构由 AI 辅助整理,请以论文原文为准。

Xing Zhang, Guanghui Wang, Yanwei Cui, Ziyuan Li, Wei Qiu, Bing Zhu, Peiyang He

AI总结:

研究自我进化智能体系统中评估指标缺失问题,提出指标可进化,通过“双棘轮”协同进化指标与技能循环,在多任务中保留提升效果,还阐述安全性源于锚定规则与外部审核,为无可靠自动验证器场景提供架构。

AI中文摘要:

自我进化的智能体系统通过创建、修正和淘汰自身技能来提升,但每个这样的循环都基于一个隐藏假设:已有可靠的评估指标。在许多实际应用中并非如此。我们提出三点主张。其一,指标可以进化:我们的指标循环在完整的进化生命周期下搜索小缺点检测器的组合,训练使其与十个项目的锚定参考集一致,通过对未标记输出的共识进行正则化,并根据一个从未读取过的留出锚进行审核,产生一个透明、可检查的指标而非不透明的评判。其二,由于不存在要超越的指标,标准是恢复准确指标本应实现的结果,我们的指标与生命周期管理的技能循环协同进化的“双棘轮”做到了这一点:在代码生成(MBPP+)、企业文本到SQL(Spider~2.0-Snow)和无参考报告生成中,它保留了由真实情况或最佳可用评分标准驱动的相同技能循环所实现的留出提升的88%-110%。其三,安全性来自锚定规则加上外部审核:移除锚定保护会使指标沦为空洞的检测器,而移除生命周期则不会;当进化后的技能在报告评分标准上耍手段时,一个独立评判者发现了问题,一个检测器修复了问题,并且一个任务感知评判者在77%的已决对中更喜欢进化后的输出而非进化前的基线。我们认为这种预期失败的架构在不存在可靠自动验证器的任何地方都是正确的默认选择。

英文摘要:

Self-evolving agent systems create, revise, and retire their own skills, but every such loop assumes a reliable evaluation metric already exists. In many real applications none does. We show the metric itself can be the evolving object: our loop searches compositions of small typed drawback detectors under a full evolutionary lifecycle, selecting for agreement with a ten-item anchored reference set and regularizing by consensus over unlabeled outputs. What evolves is the function that grades one output, never the fixed task sets it is scored on, and what comes out is an inspectable expression rather than an opaque judge. It is also valid: on code generation it gains 0.21 agreement with hidden ground truth on a locked set that metric selection never reads (paired $p=0.014$), beating the bare LLM judge it contains. Validity is where safety lives: removing the anchor guards collapses the metric into a vacuous always-pass detector while removing the detector lifecycle does not, inverting the lesson from skill evolution. That collapse warns this line of work that downstream task score cannot validate a self-evolved evaluator, since the collapsed metric trains skills just as well. Task score answers only sufficiency, and an evolved metric suffices: \emph{Double Ratchet}, co-evolving the metric with a lifecycle-managed skill loop, retains 88--110\% of the lift ground truth or a hand-written rubric buys, across MBPP+, Spider~2.0-Snow, and report generation. When evolved skills gamed the report rubric, an independent judge caught it and one added detector repaired it.

补充信息

↑