分解LLM裁判的不确定性以定位专家标注
Decomposing LLM-Judge Uncertainty to Target Expert Labels
浏览论文内容
中文总结 AI 辅助
本文提出用贝叶斯模型分解LLM裁判的偶然与认知不确定性,以指导专家标注,在ChaosNLI上认知排序比总不确定性多消除83%错误。
中文摘要 AI 辅助
LLM裁判在大规模上评估输出。专家应仅在其最不确定的地方进行标注。其自然的升级信号混淆了两种不确定性:偶然不确定性,即专家群体中的真实分歧,这种分歧无法通过标注减少;以及认知不确定性,即裁判的无知,这种无知可以通过标注减少。一个小的贝叶斯模型将两者分离:对已收集标签进行回归,学习在多大程度上信任黑盒裁判的预测。两个分量均以简单公式形式得出,无需采样或额外的裁判调用。这些分量在真实LLM裁判上针对确切已知真相进行隔离,且所声称的置信度并不能指导其实际错误。在真实人类分歧(ChaosNLI)上,对于相同的专家标签,认知排序比总不确定性多消除83%的错误,尽管简单地升级最少标注的项目在那里同样有效。我们证明了可以估计裁判无知之处而非专家真正分歧之处,并提议利用这一点来指导专家标注。
英文摘要
An LLM judge evaluates outputs at scale. Experts should label only where it is least sure. Its natural escalation signal conflates two uncertainties: aleatoric, real disagreement in the expert pool, which labels cannot reduce, and epistemic, the judge's ignorance, which labels do reduce. A small Bayesian model separates them: a regression on labels already collected learns how far to trust a black-box judge's prediction. Both components follow as simple formulas, with no sampling or further judge calls. The components isolate on a real LLM judge against exactly known truth, and stated confidence is no guide to its actual error. On real human disagreement (ChaosNLI) the epistemic ranking removes 83% more error than total uncertainty for the same expert labels, though simply escalating the least-labelled items does as well there. We demonstrate we can estimate where a judge is ignorant rather than where experts genuinely disagree, and propose using this to direct expert labelling. Code and data are available at https://github.com/composo-ai/judge-uncertainty-decomposition.
发表机构
- Composo AI
机构由 AI 辅助整理,请以论文原文为准。