LLM作为评判者系统中的锚定偏差:先验分数损害评估独立性
Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation Independence
浏览论文内容
中文总结 AI 辅助
该研究发现LLM-as-a-Judge系统存在锚定偏差,先验分数作为元数据会系统性偏移LLM的评估结果,且现有缓解措施效果有限,需针对性设计评估方案。
中文摘要 AI 辅助
大型语言模型(LLM)越来越多地对生成内容进行评估,催生了LLM-as-a-Judge(LLM作为评判者)范式。这些系统如今会对输出打分、筛选内容,并在生产流程中控制迭代优化,通常假设每次判断独立于早期评估。我们通过三种提示条件验证该假设:无元数据、修订框架,以及包含修订、尝试次数和先验分数字段的锚定元数据。结果显示,即使仅将先验分数作为上下文元数据,也会锚定判断并系统性地将评分向其值偏移。在192000次尝试评估(185271次成功)中,8个被评估模型中有7个,其对20个固定文本的锚定元数据总效应的95%任务分层自举置信区间低于0。衡量评分分布差异的标准化指标Cohen's d的绝对值达到0.71。对选定模型-任务探针的词元级分析显示出阈值式响应模式:引入锚定元数据会导致输出评分概率发生显著重新分布,而在测试的阈值以下范围内改变锚定值则产生相对较小的额外变化。在带有人工标注真值的分类行业数据中,锚定元数据会阻止48%的错误修正,并将10.18%的正确判断翻转为分配的错误标签,表明该偏差不仅限于数值评分,还延伸至分类决策。无论是思维链(Chain-of-Thought)还是忽略元数据的警告,均未降低总效应,尽管在行业实验中,该警告相对于基线提升了配对准确率效应。可靠的LLM评估需要精心的上下文工程,而非对公正性的假设,有效的缓解措施必须针对预期模型和任务或领域进行验证。
英文摘要
Large language models (LLMs) increasingly assess generated content, giving rise to the LLM-as-a-Judge paradigm. These systems now score outputs, filter content, and gate iterative refinement in production pipelines, where each judgment is often assumed to be independent of earlier evaluations. We test this assumption using three prompt conditions: no metadata, revision framing, and anchored metadata containing revision, attempt, and prior-score fields. We show that prior scores, even when included only as context metadata, anchor judgments and systematically shift ratings toward their values. Across 192,000 attempted evaluations (185,271 successful), seven out of the eight evaluated models have 95% task-stratified bootstrap intervals below zero for the total anchored-metadata effect on 20 fixed texts. Cohen's $d$, a standardized measure of the difference between score distributions, reaches an absolute value of 0.71. Token-level analysis of selected model-task probes suggests a threshold-like response pattern: introducing anchored metadata produces a marked redistribution of output-score probabilities, while changing the anchor value within the tested below-threshold range produces comparatively little additional variation. On categorical industry data with human-labeled ground truth, anchored metadata blocks 48% of error corrections and flips 10.18% of correct judgments toward an assigned wrong label, demonstrating the bias extends beyond numerical scoring to categorical decisions. Neither Chain-of-Thought nor a metadata-disregard warning reduces the total effect, although the warning improves the paired accuracy effect relative to baseline in the industry experiment. Reliable LLM evaluation demands careful context engineering rather than an assumption of impartiality. Effective mitigation must be validated for the intended model and task or domain.
发表机构
- Infobip(英福比普(Infobip))
机构由 AI 辅助整理,请以论文原文为准。