arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估对矩阵:基于事实的检索增强生成中大型语言模型评判器的答案配对元评估

Eval-Pair Matrix: Answer-Paired Meta-Evaluation of LLM Judges for Grounded RAG

Sriram Selvam, Anneswa Ghosh

arXiv 2607.10626首次发表:更新:

AI 中文总结

研究针对基于事实的检索增强生成中大型语言模型评判器自我宽容难识别问题,引入Eval-Pair Matrix协议,通过特定流程生成答案并评估,经实验得出同模型效应等结果,强调RAG评判器研究应报告多方面内容。

AI 中文摘要

大型语言模型作为评判器的评估广泛用于检索增强生成(RAG),但重用同一模型家族作为生成器和评判器会使自我宽容难以识别。我们引入了Eval-Pair Matrix,这是一种用于基于事实的RAG的可控元评估协议。从GaRAGe问题和基础段落开始,我们为每个记录引入一个隐藏的答案因果矛盾,使用GPT、Grok和Gemini模型从扰动段落生成答案,然后使用相同模型作为盲评判器根据原始段落评估每个答案。实验包含300个核心记录、897个标记的生成器输出和2683个评判裁决,形成一个交叉的3x3矩阵;主要分析使用275个完全验证的记录。我们通过在完全相同的候选答案上配对评判器来估计同模型效应,而不是比较不同答案的对角线和非对角线单元格。这改变了解释:对角线和非对角线F1相似,配对的同模型召回效应接近零。唯一稳健的配对差距是对于避免诱导主张的答案,匹配评判器的标记较低。有针对性的人工评估发现,审查的明显误报是替代源错误检测、标记诱导主张是否被采用的错误或不明确的情况;没有一个被判定为真正的误报。经验教训是方法上的:RAG评判器研究应报告完整矩阵、答案配对效应、行为层次和标签任务对齐。

英文摘要

LLM-as-a-judge evaluation is widely used for retrieval-augmented generation (RAG), but reusing the same model family as both generator and judge makes self-leniency difficult to identify. We introduce Eval-Pair Matrix, a controlled meta evaluation protocol for source-grounded RAG. Starting from GaRAGe questions and grounding passages, we induce one hidden answer-causal contradiction per record, generate answers from perturbed passages with GPT, Grok, and Gemini models, and then use the same models as blind judges to evaluate each answer against the original passages. The experiment contains 300 core records, 897 labeled generator outputs, and 2,683 judge verdicts in a crossed 3 x 3 matrix; the primary analysis uses 275 fully validated records. Instead of comparing diagonal and off-diagonal cells across different answers, we estimate same-model effects by pairing judges on the exact same candidate answer. This changes the interpretation: diagonal and off diagonal F1 are similar, and the paired same-model recall effect is near zero (-0.5 pp; 95% cluster bootstrap CI [-2.7, +1.7]). The only robust paired gap is lower matching-judge flagging for answers that avoided the induced claim (-4.3 pp). A targeted human evaluation finds that reviewed apparent false positives are alternate source-error detections, mistakes in labeling whether the induced claim was adopted, or unclear cases; none were adjudicated as genuine false alarms. The lesson is methodological: RAG judge studies should report full matrices, answer-paired effects, behavior strata, and label-task alignment.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑