发表机构
Evaluator Integrity(评估者诚信机构)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究发现主流LLM评估框架未实现优先提交式评判,其变体易被操纵,优先提交式评判未消除锚点且继承评判者错误,小模型可抵抗操纵,还验证了自身工具的缺陷。
AI 中文摘要
大语言模型(LLM)评判者是对另一系统输出进行评分的模型,可能会被其所评分的系统操纵。近期研究发现一种有效的防御措施:评判者先自行解决任务并提交答案,仅当候选答案与该答案匹配时才接受。我们将此称为优先提交式评判,并探究已发布软件是否实现了该机制,以及其代价如何。我们审计了8种广泛使用的评估框架的默认评判者配置,在涉及的24种配置中,无一种实现了该机制;9种实现了一种被文献判定为无效的变体,且共享一个可追溯至复制排版错误的祖先提示词。在受控实验中,无法获取正确答案的普通N-best搜索算法,会针对其中一种配置(严格按文档使用)优化代码:在区间合并任务中,某一随机种子下评判者接受了96个候选中的90个,另一随机种子下接受了93个,所有被接受的候选均通过了搜索可见的所有测试,但未通过其无法访问的保留测试集;评判者识别出缺陷代码行并将其作为满分依据。优先提交式评判消除了该操纵效果:两个随机种子下96个候选均被拒绝。在第二项任务中,该机制在两个随机种子下均使情况恶化:评判者提交的答案错误,其中一个随机种子下的种群收敛至该错误答案,这是我们的主要发现。优先提交式评判并未消除可被操纵的锚点,只是将其从候选答案转移至评判者自身的答案,因此评估效果仅与评判者在该任务上的表现相当。该前提条件可预先低成本测量,且具有任务局部性而非规模依赖性:规模较小的评判者解决了前沿评判者失败的任务,并在前沿评判者无法抵抗操纵的情况下抵抗了操纵。我们还验证了自身工具:我们标准中的15项主张有5项与原文不符,2项保留检查不符合其规范。
英文摘要
LLM judges, models that score another system's output, can be gamed by the systems they score. Recent work identifies one defence that works: the judge solves the task itself first and commits to that answer, then accepts a candidate only if the two match. We call this commit-first judging, and ask whether shipped software implements it, and what it costs. We audit the default judge configurations of eight widely used evaluation frameworks. Of the 24 configurations in scope, none implement it. Nine implement a variant the literature measures as ineffective, and share one ancestor prompt, traceable through a copied typographical error. In a controlled experiment, an ordinary best-of-N search with no access to correct answers optimises code against one of these configurations, used exactly as documented. On an interval merging task the judge accepted 90 of 96 candidates in one seed and 93 of 96 in the other; every accepted candidate passed every test the search could see and failed a held-out suite it could not. The judge identified the defective line and cited it as grounds for a perfect score. Commit-first judging removed the effect: 0 of 96 in both seeds. On a second task it made matters worse in both seeds: the judge's committed answer was wrong, and in one seed the population converged on it. This is our main finding. Commit-first judging does not remove the anchor that gets gamed, it moves it from the candidate to the judge's own answer, so evaluation is only as good as the judge is at the task. That precondition is cheap to measure in advance, and is task local rather than scale dependent: a smaller judge solved a task the frontier judge failed and resisted gaming where it did not. We also validate our own instruments: five of fifteen claims in our criteria were wrong against verbatim sources, and two held-out checks were unjustified by their specifications.
Comments11 pages, 4 figures