发表机构
South China Normal University; University of Southern California; Fudan University; East China Normal University(华南师范大学; 南加州大学; 复旦大学; 华东师范大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
BayesJudge是面向人类与LLM判断的不确定性感知贝叶斯元评估方法,可分离响应质量与呈现效应,在SummEval数据集上成功检测LLM的顺序偏差并推断评分者行为特征。
AI 中文摘要
AI评估流程常产生冲突判断而非明确标签,在LLM成对评估中这一冲突尤为明显:分歧可能源于模糊样本、未明确的评估准则、异质或不稳定的人类评分者,或是LLM判断因响应顺序调换而改变。我们提出BayesJudge,这是一个针对人类-LLM冲突判断流的在线贝叶斯元评估层。对于每一次比较,BayesJudge会估计相对于评分组的两个响应的后验判断分布,当存在平局或模糊标签时会对应状态。同时,它会估计评分者特有的人类混淆矩阵和LLM的呈现顺序偏差。该方法利用开放平局标签保持模糊性可观测,并通过成对调换顺序的判断调用将响应质量与呈现效应分离。我们构建了精确的在线后验递归,并采用可扩展的Rao-Blackwell化假设密度SMC近似用于流式推理。受控合成实验证明了对预设评估者参数的恢复,并展示了两种协议层面的可识别性机制:开放平局标签暴露模糊性质量,调换顺序的成对判断将样本偏好与位置偏差分离。在真实世界的SummEval数据集上,BayesJudge成功检测到LLM判断输出中系统性的呈现顺序效应,在无评分者元数据的情况下推断出不同的专家和众包工作者行为特征,并生成与人类分歧相关的后验不确定性估计。我们的代码可在该httpsURL获取。
英文摘要
AI evaluation pipelines often produce conflicting judgments rather than clean labels. In pairwise LLM evaluation, this conflict is especially visible: disagreement can arise from ambiguous items, underspecified rubrics, heterogeneous or unstable human raters, or an LLM judge whose verdict changes when the response order is swapped. We propose BayesJudge, an online Bayesian meta-evaluation layer for conflicting human-LLM judgment streams. For each comparison, BayesJudge estimates a panel-relative posterior verdict distribution over the two responses, with a tie or ambiguity state when such labels are available. At the same time, it estimates rater-specific human confusion matrices and LLM presentation-order bias. The method uses tie-open labels to keep ambiguity observable and paired order-swapped judge calls to separate response quality from presentation effects. We formulate the exact online posterior recursion and use a scalable Rao-Blackwellized assumed-density SMC approximation for streaming inference. Controlled synthetic experiments demonstrate recovery of prespecified evaluator parameters and illustrate two protocol-level identifiability mechanisms: tie-open labels expose ambiguity mass, and order-swapped paired judgments separate item preference from position bias. On a real-world SummEval dataset, BayesJudge successfully detects systematic presentation-order effects in LLM judge outputs, infers distinct expert and crowdworker behavior signatures without rater metadata, and produces posterior uncertainty estimates that correlate with human disagreement. Our code is available at https://anonymous.4open.science/r/BayesJudge-0879.
CommentsAccepted at NeurIPS 2026