学习破解评判者的策略
Learning Strategies to Break Judges
AI总结:
提出智能体引导方法,通过对抗性变异和策略提炼,发现并分析智能体评判者在数学推理中的可解释失败模式,揭示前沿可靠性下降。
AI中文摘要:
随着AI智能体超越人类表现,系统设计者直接评估它们并理解其失败模式变得极其困难。因此,智能体本身正被广泛部署用于评估、评判和提供模型轨迹的反馈。但这引发了一个重要问题:我们如何信任评判者?在这项工作中,我们提出了一种智能体引导的方法来发现智能体评判者的弱点,从而揭示可解释的失败机制。我们的方法聚焦于数学推理,分两个阶段进行:首先,我们部署对抗性智能体通过引入错误来变异一组正确的证明,试图误导评判者——换言之,注入评判者无法捕捉的错误。然后,我们将这些尝试提炼成一小部分变异策略,使我们能够分析评判者的失败模式。为确保这些策略不过度拟合初始证明集,我们通过将变异策略应用于一组保留的证明并查询同一评判者来评估它们。我们将方法部署在GPT-5.6-sol和Claude Opus 5上,分别配对其智能体编排器Codex和Claude Code。这些既用作引入错误的变异器,也用作评估数学推理正确性的评判者。我们发现,在所有智能体评判者中,我们能够提炼出持续绕过其评估的变异策略,从而使我们能够确定可操作的失败模式。我们的分析还揭示,评判者的可靠性在前沿有所下降:奥林匹克级证明或研究生级数学文本中的错误被更一致地检测到,而研究级手稿中的缺陷更可能逃脱检测。
英文摘要:
As AI agents surpass human performance, it becomes exceedingly hard for system designers to evaluate them directly and understand their failure modes. Consequently, agents themselves are being deployed extensively to evaluate, judge, and provide feedback on model traces. But this raises an important question: how can we trust the judge? In this work, we propose an agent-guided method to find weaknesses of agentic judges that expose interpretable failure mechanisms. Our method focuses on mathematical reasoning and proceeds in two stages: first, we deploy adversarial agents to mutate a set of sound proofs by introducing errors, attempting to misguide judges---in other words, injecting errors that judges are unable to catch. Then, we distill these attempts into a small set of mutation strategies which allow us to analyze the failure modes of the judges. To ensure that these strategies are not overfit to the initial set of proofs, we evaluate them by applying the mutation strategies to a held-out set of proofs and querying the same judge. We deploy our method on GPT-5.6-sol and Claude Opus 5, paired with their agent orchestrators, Codex and Claude Code, respectively. These are used both as mutators to introduce errors and as judges to evaluate correctness of mathematical reasoning. We find that across all the agentic judges, we are able to distill mutation strategies that consistently bypass their evaluations, thereby enabling us to ascertain actionable failure modes. Our analysis also reveals that judge reliability degrades at the frontier: errors in Olympiad-level proofs or graduate-level mathematical texts are detected more consistently, whereas flaws in research-level manuscripts are more likely to escape detection.