当一致性不代表可靠性:评估本地LLM裁判与人类评分的对比
When Consistency Does Not Mean Reliability: Evaluating Local LLM Judges Against Human Ratings
- Indian Institute of Technology Kharagpur(印度理工学院卡拉格普尔分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究评估两个本地LLM裁判(LLaMA-3-8B和Qwen2.5-7B)与人类评分的相关性,发现其自一致性高但人类一致性低,强调需同时评估一致性与人类对齐。
AI中文摘要:
大型语言模型(LLMs)越来越多地被用于评估其他语言模型的响应。这种方法被称为“LLM作为裁判”(LLM-as-a-Judge),比人工评估更快、更便宜。然而,裁判可能产生一致的评分,但不一定与人类评估者一致。在本研究中,我们使用两个本地开源权重LLM裁判,即LLaMA-3-8B和Qwen2.5-7B,来研究这个问题。我们评估了由指令微调的GPT-2(124M)模型生成的300个响应,这些响应针对100个问题,涵盖五个类别:事实知识、指令遵循、数学、推理和写作。每个响应由九名人类标注者评分,并由每个LLM裁判使用相同的评分标准评估三次。我们使用皮尔逊相关系数、斯皮尔曼相关系数、平均绝对误差(MAE)、有符号偏差和自一致性,将裁判评分与人类平均评分进行比较。LLaMA-3-8B与人类评分的皮尔逊相关系数为0.275,而Qwen2.5-7B达到0.340。它们的MAE分别为27.71和18.64。尽管一致性有限,但两个裁判都表现出较高的自一致性,LLaMA-3-8B的精确一致性率为97.3%,Qwen2.5-7B为92.3%。这些结果表明,高自一致性并不一定意味着与人类判断的高一致性。我们的发现强调了在使用本地LLM作为自动裁判时,需要同时评估一致性和人类对齐。
英文摘要:
Large language models (LLMs) are increasingly used to evaluate the responses of other language models. This approach, known as LLM-as-a-Judge, is faster and cheaper than human evaluation. However, a judge may produce consistent scores without necessarily agreeing with human evaluators. In this work, we study this issue using two local open-weight LLM judges, LLaMA-3-8B and Qwen2.5-7B. We evaluate 300 responses generated by an instruction-tuned GPT-2 (124M) model for 100 questions covering five categories: factual knowledge, instruction following, mathematics, reasoning, and writing. Each response is scored by nine human annotators and is evaluated three times by each LLM judge using the same rubric. We compare the judge scores with the average human scores using Pearson correlation, Spearman correlation, mean absolute error (MAE), signed bias, and self-consistency. LLaMA-3-8B shows a Pearson correlation of 0.275 with human scores, while Qwen2.5-7B achieves 0.340. Their MAEs are 27.71 and 18.64, respectively. Despite this limited agreement, both judges show high self-consistency, with exact consistency rates of 97.3\% for LLaMA-3-8B and 92.3\% for Qwen2.5-7B. These results show that high self-consistency does not necessarily indicate high agreement with human judgments. Our findings highlight the need to evaluate both consistency and human alignment when using local LLMs as automatic judges.