发表机构
Stony Brook University(石溪大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示从LLM评判者首个token读取裁决会高估位置偏差,在未承诺配对中强制读取导致89.7%翻转,建议报告裁决token开头比率以修正偏差。
AI 中文摘要
从LLM评判者首个生成token的logits中读取其裁决是廉价的,无需进行生成,而这正是约束解码和似然评分评估工具所产生的结果。我们表明,这种读取方式会单向扭曲位置偏差:在我们测试的所有条件下,它都会高估位置偏差,因此以这种方式获得的数值表现为上界。其机制在于,评判者并不总是以裁决token开头,对于三个Qwen3评判者,在12%至49%的配对中如此,而对于Llama-3.1-8B和Phi-3.5-mini,这一比例低于3%;在这些配对中强制读取会返回首先展示的响应,而非真正的判断。在评判者未做出承诺的924个配对中,当响应被交换时,强制读取在89.7%的配对中翻转,而生成后读取的比例为47.5%(配对差异+0.422,95%置信区间[+0.365, +0.467])。这种扭曲对所测量的内容具有特异性:它使位置偏差移动42个百分点,而在十种条件中的七种中,评判者准确率的变化不到一个百分点,因此它误导的是审计评判者的人,而非使用评判者的人。第二种较小的失败甚至发生在评判者确实以裁决token开头时,因为它有时以一个字母开头,然后推理到另一个字母,在0至5.5%的配对中出现,其比率与合规性不相关。我们建议在报告任何位置偏差数值时,同时报告评判者以裁决token开头的比率,这只需一次前向传播,无需标签。
英文摘要
Reading an LLM judge's verdict from the logits of its first generated token is cheap, requires no generation, and is exactly what constrained decoding and likelihood-scoring evaluation harnesses produce. We show that this readout distorts position bias in one direction: it overstates it in every condition we test, so figures obtained this way behave as upper bounds. The mechanism is that judges do not always lead with a verdict token, on 12% to 49% of pairs for three Qwen3 judges and under 3% for Llama-3.1-8B and Phi-3.5-mini, and forcing a read on those pairs returns whichever response was shown first rather than a judgment. Pooled over the 924 pairs where a judge did not commit, the forced read flips on 89.7% of them when the responses are swapped, against 47.5% read after generation (paired difference +0.422, 95% CI [+0.365, +0.467]). The distortion is specific to what is measured: it moves position bias by 42 points while moving judge accuracy by under one point in seven of ten conditions, so it misleads whoever audits a judge rather than whoever uses one. A second, smaller failure occurs even when the judge does lead with a verdict token, since it sometimes opens with one letter and reasons its way to the other, on 0 to 5.5% of pairs at a rate uncorrelated with compliance. We recommend reporting the rate at which a judge leads with a verdict token, which costs one forward pass and no labels, alongside any position-bias figure.