发表机构
Meta Superintelligence Labs; FAIR at Meta(元超级智能实验室; 元公司FAIR研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出Wiggle框架测试LLM评判者的认知稳定性,发现9个前沿模型在压力下裁决波动显著,且多数压力会损害真实值,基线陪审团多数强度可预测波动项目。
AI 中文摘要
大型语言模型(LLM)评判者已成为模型评估、在线评分和奖励建模的核心基础设施。评判者通常通过在黄金数据上的准确率进行验证,但准确率几乎无法说明它们在重新提示、挑战或持续反驳下是否稳定。我们提出了Wiggle框架,这是一种用于LLM评判者认知稳定性的统一压力测试。该框架将评判者的稳健性分解为三个维度:机械一致性(在重新提示和重构下的稳定性)、单轮置信度(在单次挑战下的稳定性)以及多轮持续性(在持续或自适应压力下的稳定性)。我们使用该框架研究了9个前沿模型,涵盖安全、毒性、AI写作检测和政治回应评估的14项评判任务。所有模型作为评判者都表现出显著的波动——在静态反驳下翻转裁决的概率为25%至71%,在对抗性LLM说服者作用下则为62%至91%。关键的是,我们发现成功改变评判者裁决的压力几乎总是相对于真实值产生净损害。除了该框架本身,我们还确定基线陪审团多数强度是预测哪些项目会波动的最有效的单次信号。总体而言,这是首个在评判语境下对机械性、从众性和可说服性测试进行的跨数据集公平比较。
英文摘要
LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling. Judges are typically validated by accuracy on golden data, but accuracy says little about whether they are stable under re-prompting, challenge, or sustained pushback. We introduce the \emph{Wiggle Framework}, a unified stress test for epistemic stability in LLM judges. The framework decomposes judge robustness along three dimensions: Mechanical Consistency (stability under re-prompting and reframing), Single-turn Conviction (stability under a single challenge), and Multi-turn Persistence (stability under sustained or adaptive pressure). We use the framework to study 9 frontier models across 14 judging tasks spanning safety, toxicity, AI writing detection, and political-response evaluation. Every model exhibits substantial wiggle as a judge --- flipping verdicts 25--71\% of the time under static pushback, and 62--91\% with an adversarial LLM persuader. Critically, we find that pressure that succeeds in changing a judge's verdict is almost always net-corrupting with respect to ground truth. Beyond the framework itself, we identify baseline jury majority strength as the most effective single-shot signal for anticipating which items wiggle. Taken together, this is the first apples-to-apples cross-dataset comparison of mechanical, conformity, and persuadability tests in a judging context.