谁来验证验证器?可检查评分器与自我改进智能体的协同演化
Who Verifies the Verifier? Co-Evolving Inspectable Graders with Self-Improving Agents
浏览论文内容
中文总结 AI 辅助
该研究将可检查的缺陷检测器构成的协同演化验证器作为核心,在MBPP+等任务中优于裸LLM评判器,可替代真实值或评分规则,且下游任务得分无法证明其有效性。
中文摘要 AI 辅助
我们改变了智能体:它真的变得更好了吗?每个自我改进智能体循环都会成百上千次地回答这个问题,而每次回答都来自一个验证器。对于开放式任务,不存在这样的验证器,因此该循环会被赋予人工编写的评分规则,或来自类似自身模型的裸大语言模型(LLM)评判输出,这会引发奖励黑客行为和共同的盲点。我们将验证器作为演化对象:它是一种可检查的表达式,基于聚类失败案例合成,由小型、大多为确定性的缺陷检测器构成,在生成时受门控约束,且选择依据是其与包含10个项目的锚定参考集的一致性,以及对未标记输出的共识,而非智能体的得分。在MBPP+数据集上,它在每个种子上比人工编写的种子组合获得了+0.21的保留一致性,最终优于其包含的裸LLM评判器。一项发现应改变协同演化验证器的验证方式:移除锚定保护会使验证器坍缩为空洞的、始终通过的评分器,但这种坍缩后的验证器训练技能的效果同样好。下游任务得分无法证明自我演化验证器的有效性,得分确实能回答充分性问题,在此处演化验证器可替代:Double Ratchet将验证器与生命周期管理的技能循环配对,在代码生成、企业文本转SQL、无参考报告生成任务中,保留了真实值或评分规则为同一循环带来的提升的88%-110%。当演化出的技能操纵报告评分规则时,外部评判器会捕获该问题,一个新增的检测器修复了它;而评判器本身在获得任务契约前是错误的。
英文摘要
We changed the agent: did it actually get better? Every self-improving agent loop answers this hundreds of times, and every answer comes from a verifier. On open-ended tasks none exists, so the loop is handed a hand-written rubric or a bare LLM judge grading output from a model like itself, inviting reward hacking and shared blind spots. We make the verifier the evolving object: an inspectable expression over small, mostly deterministic drawback detectors, synthesized from clustered failures, gated at birth, and selected for agreement with a ten-item anchored reference set plus consensus over unlabeled outputs, never for the agent's score. On MBPP+ it gains +0.21 held-out agreement over the hand-authored seed composition, on every seed, and ends ahead of the bare LLM judge it contains. One finding should change how co-evolved verifiers are validated: removing the anchor guards collapses the verifier into a vacuous always-pass grader, yet that collapsed verifier trains skills just as well. Downstream task score cannot certify a self-evolved verifier. Score does answer sufficiency, and there an evolved verifier can substitute: Double Ratchet, pairing the verifier with a lifecycle-managed skill loop, retains 88-110% of the lift that ground truth or a rubric buys the same loop, across code generation, enterprise text-to-SQL, and reference-free report generation. When evolved skills gamed the report rubric, an outer judge caught it and one added detector repaired it; the judge itself was wrong until given the task contract.
发表机构
- AWS Forward Deployed Engineering(AWS 前端部署工程部门)
- HSBC Holdings Plc.(汇丰控股有限公司)
机构由 AI 辅助整理,请以论文原文为准。