arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

去偏作为测量干预:LLM-as-a-Judge评估中的校准平局与分辨率损失

Debiasing as a Measurement Intervention: Calibrated Ties and Resolution Loss in LLM-as-a-Judge Evaluation

Liang Zhao, Yong Wang, Jiangzhe Chen

arXiv 2609.12439首次发表:更新:

发表机构

Shanghai University of International Business and Economics(上海对外经贸大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究将LLM-as-a-judge去偏视为测量干预,提出TraceJudgeBench基准,发现反引文提示虽抑制偏差但损害分辨率,需联合报告偏差抑制、分辨率保持、平局成本与协议成本。

AI 中文摘要

LLM-as-a-judge协议通常通过指示评判者忽略诸如引文格式、来源标签和证据展示风格等呈现线索来进行去偏。我们表明,这种干预可以抑制偏差,同时损害测量仪器的分辨率。我们引入TraceJudgeBench,一个用于审计RAG和智能体工作流评估中类引文伪影的诊断基准,涵盖内容等价对、引文消融、正确性冲突、人工验证的软性和中度质量差距、提示强度阶梯、解耦评判以及受控工作流排序探针。在GPT-5.5、Claude Sonnet 4.6和DeepSeek V4-Flash上,更强的反引文提示将较差引文的胜率从高达50.5%降至0%;然而,在严格的压力测试端点之前,某些操作点已将验证过的中度差距决策转换为平局,而正确性冲突的准确率保持在93.0%或以上。第二个包含50对的FinQA中度差距构建重现了定性前沿,开放权重模型Qwen2.5-14B-Instruct-AWQ和Gemma-3-12B-IT的运行重现了中心HotpotQA前沿。TRACE式解耦在报告设置中恢复了96.5-100.0%的更好普通分辨率。人工验证区分了平局的三种含义:正确的等价平局、校准的软边界平局以及在验证的质量差距上破坏分辨率的平局。我们将去偏视为一种测量干预,其偏差抑制、分辨率保持、平局成本和协议成本必须联合报告。补充工件包含基准划分、提示、原始评判输出、验证摘要和分析。

英文摘要

LLM-as-a-judge protocols are commonly debiased by instructing judges to ignore presentation cues such as citation formatting, source labels, and evidence-display style. We show that this intervention can suppress bias while damaging the resolution of the measurement instrument. We introduce TraceJudgeBench, a diagnostic benchmark for auditing citation-like artifacts in RAG and agent-workflow evaluation, covering content-equivalent pairs, citation ablations, correctness conflicts, human-validated soft and moderate quality gaps, prompt-strength ladders, decoupled judging, and a controlled workflow-ranking probe. Across GPT-5.5, Claude Sonnet 4.6, and DeepSeek V4-Flash, stronger anti-citation prompts reduce worse-cited wins from up to 50.5% to 0%; yet some operating points already convert validated moderate-gap decisions into Tie before the strict stress-test endpoint, while correctness-conflict accuracy remains at or above 93.0%. A second, 50-pair FinQA moderate-gap construction reproduces the qualitative frontier, and open-weight Qwen2.5-14B-Instruct-AWQ and Gemma-3-12B-IT runs reproduce the central HotpotQA frontier. TRACE-style decoupling recovers 96.5-100.0% better-plain resolution across the reported settings. Human validation separates three meanings of Tie: correct equivalence Tie, calibrated soft-boundary Tie, and resolution-destroying Tie on validated quality gaps. We frame debiasing as a measurement intervention whose bias suppression, resolution retention, Tie cost, and protocol cost must be reported jointly. The supplementary artifact contains benchmark splits, prompts, raw judge outputs, validation summaries, and analysis.

CommentsAccepted at Emnlp 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑