arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型能写出可靠的评分标准吗?实验复现的元评估

Can LLMs Write Reliable Rubrics? A Meta-Evaluation for Experiment Reproduction

Hanhua Hong, Yizhi Li, Jiaoyan Chen, Luu Gia Huy, Sophia Ananiadou, Jung-jae Kim, Chenghua Lin

arXiv 2607.12835首次发表:更新:

发表机构

The University of Manchester; Institute for Infocomm Research (I²R), A*STAR; IQuest Research; ELLIS Manchester; University of Information Technology, VNU(曼彻斯特大学; 资讯通信研究院(I²R),新加坡科技研究局; IQuest研究公司; ELLIS曼彻斯特; 越南国家大学信息技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究大语言模型生成的论文复现评分标准,将其 reformulate 为清单样式,在两个骨干模型上评估四种生成设置,通过语义相似性和分数对齐进行元评估,结果显示增强设置改善下游评估对齐,同时指出其存在过于细化、偏向高分和适应性差等问题。

AI 中文摘要

基于评分标准的评估是评估基于大语言模型的研究代理的开放式输出的一种有前景的方法,特别是在论文复现中,直接的论文与存储库比较容易产生幻觉。然而,构建特定于论文的评分标准需要大量专家工作,限制了如PaperBench等基准的可扩展性。在这项工作中,我们进行了首次针对论文复现的大语言模型生成评分标准的系统元评估。我们将评分标准重新制定为清单样式格式,并在两个骨干模型上评估四种生成设置。我们通过语义相似性对生成的评分标准进行内在元评估,并通过与真实评分标准的分数对齐进行外在元评估。我们的结果表明,增强设置大幅改善了下游评估对齐,最强设置接近人类基线,而内在收益较为适度。进一步分析表明,大语言模型生成的评分标准往往过于细化,偏向高分,且对论文领域适应性较差,凸显了其优势和局限性。

英文摘要

Rubric-based evaluation is a promising approach for assessing open-ended outputs from LLM-based research agents, particularly in paper reproduction, where direct paper-to-repository comparison is prone to hallucination. However, constructing paper-specific rubrics requires substantial expert effort, limiting the scalability of benchmarks such as PaperBench. In this work, we present, to our knowledge, the first systematic meta-evaluation of LLM-generated rubrics for paper reproduction. We reformulate rubrics into a checklist-style format and evaluate four generation settings across two backbone models. We meta-evaluate generated rubrics intrinsically by semantic similarity and extrinsically by score alignment with ground-truth rubrics. Our results show that the augmented settings substantially improves downstream evaluation alignment, with the strongest setting approaching the human baseline, while intrinsic gains are more modest. Further analyses reveal that LLM-generated rubrics are often overly fine-grained, biased toward high scores, and less adaptive to paper domains, highlighting both the affordances and limitations.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑