arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28626cs.CLcs.AI

大型语言模型是否会仔细审查其评审内容?一项关于评分校准、错误检测及作者身份影响的多模态审计研究

Do large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects

Emad Alharbi

首次发表
浏览论文内容

中文总结 AI 辅助

本研究以两款多模态LLM为评审者评估2026 ICLR投稿,发现其评分高于人类、错误检测率低,提供图表会降低错误检测率,作者身份对评审无影响,编辑决策与简单评分平均一致。

中文摘要 AI 辅助

大型语言模型(LLMs)越来越多地被用于生成同行评审意见,这促使人们审视它们的批判性评估能力。本研究将Qwen2.5-VL-72B和Pixtral-Large-124B这两款多模态大型语言模型作为评审者,对2026年国际学习表征会议的165篇投稿进行评估,该会议的举办时间晚于这两款模型的训练截止时间。手稿以作者身份被遮蔽、替换为高声望机构或低声望机构的形式呈现给模型,且分为仅文本或带图表的文本两种格式。此外,研究人员在55篇手稿中插入了145个可验证的明显错误,以评估在自然提示和面向验证的提示下的错误识别能力。在所有手稿组(包括被拒投稿)中,LLM的评分范围为7.0至8.1,而人类的平均评分范围为3.4至6.8。在自然提示下,模型检测到12.1%的已验证错误;而一句验证指令将检测率提升至22.2%,但仍有78%的错误未被检测到。提供图表会降低错误检测率,同时提高评审评分。没有任何视觉错误能根据对应图表得到可靠验证,且有一半仅文本评审描述了未提供的图表。作者身份对评审评分和错误检测均无影响。LLM的编辑决策与简单的评分平均结果完全匹配。

英文摘要

Large language models (LLMs) are increasingly used to generate peer reviews, prompting examination of their capacity for critical evaluation. This study evaluates two multimodal LLMs, Qwen2.5-VL-72B and Pixtral-Large-124B, as reviewers across 165 submissions to the 2026 International Conference on Learning Representations, a venue that postdates both models' training cutoffs. Manuscripts were presented to both models with author identities blinded, replaced with high-prestige affiliations, or replaced with low-prestige affiliations, and in either text-only or text-with-figure format. Additionally, 145 verifiably detectable errors were inserted into 55 manuscripts to assess error identification under natural and verification-oriented prompts. Across all manuscript groups, including rejected submissions, LLM scores ranged from 7.0 to 8.1, whereas human mean scores ranged from 3.4 to 6.8. The models detected 12.1\% of the verified errors under natural prompting, and a one-sentence verification instruction increased detection to 22.2\%; however, 78\% of the errors remained undetected. Providing figures reduced error detection while increasing review scores. No visual error was reliably verified against its corresponding figure, and half of the text-only reviews described figures that were not provided. Author identity did not influence either review scores or error detection. LLM editorial decisions exactly matched those produced by simple score averaging.

发表机构

  • University of Tabuk(塔布克大学)

机构由 AI 辅助整理,请以论文原文为准。

↑