评估大语言模型作为评判者:关于基于大语言模型的自动文本生成评估中的评分标准伪影
Judging LLM-as-a-Judge: Concerning Rubric Artifacts in LLM-based Automated Text Generation Evaluation
- Centre for Responsible AI (CeRAI)(负责任人工智能中心(CeRAI))
- Wadhwani School of Data Science and AI (WSAI)(瓦德瓦尼数据科学与人工智能学院(WSAI))
- IIT Madras(印度理工学院马德拉斯分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究发现LLM-as-a-Judge评估中,仅评分标准就能让分类器达到可观预测性能,且评判者无法可靠更新决策,凸显此类评估的可靠性问题及方法学研究的必要性。
AI中文摘要:
大语言模型作为评判者(LLM-as-a-Judge)的流程正被越来越多地用于评估人工智能生成的文本,其核心假设是:评判结果源于基于评分标准对候选响应进行推理。我们表明,这一假设值得进一步审视。仅在评分标准文本上训练、不访问任何被评估响应的分类器,在评判输出上能达到可观的预测性能。这表明评分标准的表述编码了可恢复的评估信号,使得评分可部分独立于模型输出被预测。最后,反事实扰动实验显示,当候选响应或评分标准准则被反转时,评判者往往无法可靠地更新其决策。我们的发现引发了对基于评分标准的大语言模型评估可靠性的担忧,并强调需要进一步研究通过大语言模型进行自动评估的方法。
英文摘要:
LLM-as-a-Judge pipelines are increasingly used to evaluate AI-generated text, based on the assumption that judgments arise from reasoning over candidate responses with respect to a rubric. We show that this assumption warrants further scrutiny. Classifiers trained only on rubric text, without access to any evaluated response, achieve nontrivial predictive performance on judge outputs. This suggests that rubric formulations encode recoverable evaluative signals, allowing scores to be partially anticipated independently of model outputs. Finally, counterfactual perturbations reveal that judges often fail to reliably update their decisions when either the candidate response or the rubric criterion is reversed. Our findings raise concerns about the reliability of rubric-based LLM evaluation and highlight the need for further methodological study of automated evaluation via LLMs.