多智能体代码评判何时真正有依据?两种无标签测量,以及一个拒绝猜测的评判者
When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess
- University of Windsor(温莎大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出无标签测量方法,揭示多智能体代码评判在缺乏独立且差异化的证据时会盲目裁决,通过门控拒绝无法判定的比较,将准确率从20.7%提升至36.9%。
AI中文摘要:
当一个语言模型评判另一个模型的代码是否正确时,它不会报告证据的缺失。它会返回一个带有推理过程的自信裁决,这与它有依据时给出的裁决无法区分。多智能体验证将评判分解为可核查的声明,并针对证据逐一验证,是一种有前景的应对方法,并且在证据是一组检索到的文档时效果良好。我们认为此类方法对其证据有两个要求:证据必须独立于被审查的答案,并且证据必须在被比较的两个候选方案之间存在差异。第二个条件在检索文档时自动成立,但在代码评判中不再成立。我们运行了MARCH这一已发表的框架,在两个代码评判基准上进行了超过80项按条件逐单元的测量,发现它在78%至95%的比较中判定两个解决方案同样好,准确率仅为4.4%,而同一模型直接提问时的准确率为43.7%。更简单的问题或更大的评判模型都无法改变这一结果。从流水线自身日志中提取的两项测量无需标签即可解释这一现象。基于其中一项进行门控,流水线会拒绝其无法做出的比较,并将准确率从20.7%提升至36.9%,同时仍能回答所有比较中的一半。其贡献并非一个更准确的评判者,而是一种无标签的方法,用以判断评判者何时缺乏做出回答的依据。
英文摘要:
When one language model judges whether another's code is correct, it does not report the absence of evidence. It returns a confident verdict with reasoning attached, indistinguishable from a verdict it had grounds for. Multi-agent verification, which decomposes a judgment into checkable claims and verifies each against evidence, is a promising response and works well when the evidence is a set of retrieved documents. We argue such methods require two things of their evidence: it must be independent of the answer under review, and it must differ between the two candidates being compared. The second condition holds automatically with retrieved documents and stops holding in code judging. Running MARCH, a published framework unmodified over 80 condition-by-cell measurements on two code judging benchmarks, we find it declares both solutions equally good on 78 to 95% of comparisons, reaching 4.4% accuracy where the same model asked directly reaches 43.7%. Neither easier problems nor a larger judge changes this. Two measurements taken from the pipeline's own logs explain it without needing labels. Gating on one of them, the pipeline declines the comparisons it cannot make and raises its accuracy from 20.7 to 36.9% while still answering half of all comparisons. The contribution is not a more accurate judge, but a label-free way to tell when a judge has no basis for its answer.