发表机构
Stony Brook University(石溪大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究利用残差流激活预测LLM评判对候选顺序的敏感性,线性探针在JudgeBench和MT-Bench上均优于非激活基线,表明判决前激活可有效预测位置翻转。
AI 中文摘要
候选回答的呈现顺序可能会改变LLM评判的最终判决。通常,检测这种位置翻转需要对每对候选回答以两种顺序分别进行评判,这会使评判次数翻倍。我们研究了在初始判决前记录的残差流激活是否能够预测翻转的发生。我们使用嵌套分组交叉验证,在534个JudgeBench对和三个Qwen3评判模型及Llama-3.1-8B上评估了正则化线性探针。线性探针的AUROC达到0.621-0.850,比使用口头化置信度、判决标签logits、回答长度和评判初始选择的组合基线高出0.062-0.113的AUROC。在JudgeBench上训练后冻结的线性探针,在未进行MT-Bench拟合或重新校准的情况下,在1802个MT-Bench比较上实现了0.685-0.853的AUROC。这些结果表明,判决前的激活能够支持对候选顺序敏感性的预测,并且优于本文评估的非激活预测器。
英文摘要
The order in which candidate responses are presented can change an LLM judge's verdict. Detecting such a position flip ordinarily requires judging each pair in both orders, which doubles the number of judgments. We investigate whether residual stream activations recorded immediately before the initial verdict can predict a flip. We use nested grouped cross-validation to evaluate regularized linear probes on 534 JudgeBench pairs for three Qwen3 judges and Llama-3.1-8B. The linear probes achieve AUROCs of .621-.850 and outperform a combined baseline that uses verbalized confidence, verdict-label logits, response lengths, and the judge's initial choice by .062-.113 AUROC. Linear probes trained on JudgeBench and then frozen achieve AUROCs of .685-.853 on 1,802 MT-Bench comparisons without MT-Bench fitting or recalibration. These results show that pre-verdict activations support prediction of susceptibility to candidate order and outperform the non-activation predictors evaluated here.
CommentsIt was submitted to neurips workshop (JUDGE workshop) and received a review score 6.5 combined with one strong accept and one above threshold accept