LLM 评判者的二阶响应定律:提示不稳定性的无偏估计
Second-Order Response Laws for LLM Judges: Debiased Estimation of Prompt Instability
浏览论文内容
中文总结 AI 辅助
该研究针对 LLM 评判者,推导了提示不稳定性的无偏估计量,分离提示内与跨提示差异,经实验验证校正估计更准确,可单独估计提示鲁棒性。
中文摘要 AI 辅助
LLM 评判者(LLM judges)通常通过单个提示(prompt)和仅少数重复调用进行评估。当它们的判决结果存在差异时,目前尚不清楚这种差异是来自单个提示内的采样噪声,还是不同提示间的系统性差异。我们使用二阶响应定律对这种差异进行形式化:该定律描述了由既定提示策略诱导的、以提示为条件的判决分布的分布。对于提示不稳定性的二次度量,我们证明,在有限重复预算下,常用的插件估计量存在向上偏差,因为它混淆了提示内噪声与提示间差异。我们通过提示内与跨提示一致性的差异,为采样提示和既定固定提示普查分别推导了无偏估计量。在交叉提示-答案顺序设计下,同一框架可分离提示、顺序、交互及残差调用差异,同时将无效完成输出保留为结果。已知定律的模拟和字节级相同的实时零测试均恢复了预测的有限重复次数(finite-$R$)膨胀。在匹配的 Qwen 研究中,校正后的低重复估计值比插件估计值更接近独立获取的 $R=16$ 参考值,且在小重复预算下增益最大。对四个冻结评判者配置的匹配面板显示出配置特定的膨胀幅度和分量轮廓。因此,提示鲁棒性可与有限调用噪声分开估计。
英文摘要
LLM judges are often evaluated with a single prompt and only a few repeated calls. When their verdicts vary, it remains unclear whether the variation comes from sampling noise within a prompt or systematic differences across prompts. We formalize this distinction using a second-order response law: the distribution of prompt-conditioned verdict distributions induced by a declared prompt policy. For a quadratic measure of prompt instability, we show that the usual plug-in estimator is biased upward at finite repeat budgets because it confounds within-prompt noise with between-prompt variation. We derive unbiased estimators for both sampled prompts and declared fixed prompt censuses from the difference between within- and across-prompt agreement. Under a crossed prompt-by-answer-order design, the same framework separates prompt, order, interaction, and residual call variation, while retaining invalid completed outputs as outcomes. Known-law simulations and a byte-identical live null recover the predicted finite-$R$ inflation. In a matched Qwen study, corrected low-repeat estimates are closer to an independently acquired $R=16$ reference than plug-in estimates, with the largest gains at small repeat budgets. A matched panel across four frozen judge configurations exhibits configuration-specific inflation magnitudes and component profiles. Prompt robustness can therefore be estimated separately from finite-call noise.