大语言模型(LLM)评审与人类同行评审的契合度有多高?
How Closely Do LLM Reviews Align with Human Peer Review?
AI总结:
该研究对比OpenAI GPT-5.4、Google Gemini 3.1 Pro Preview等三款LLM与人类对300篇ICLR 2026投稿的评审,发现LLM能区分论文是否被接收,但无法复现人类对口头报告与海报论文的区分,且评审侧重点存在差异。
AI中文摘要:
大语言模型(LLM)正越来越多地被用于生成科学评审,但现有评估很少在同一受控环境中考察不同提供商的LLM是否同时契合会议决策与人类评审优先级。我们将OpenAI GPT-5.4、Google Gemini 3.1 Pro Preview、Anthropic Claude Opus 4.6的评审,与300篇主题匹配的ICLR 2026投稿的人类评审及最终决策进行对比,这些投稿平均分为口头报告、海报和被拒论文三类。在移除决策信息后,每个模型均使用相同的指令和评分标准评审所有论文。本研究在三个互补维度开展了跨提供商分析:与宽泛及细粒度决策类别的契合度、推荐量表使用的差异、所识别缺陷的主题一致性。所有三款LLM均能区分被接收与被拒论文,但均未复现人类评分中存在的口头报告与海报论文的区分。评分模式具有提供商特异性:Gemini给出的评分系统性偏高,而OpenAI和Claude对被拒及海报论文的评分更接近人类,但对口头报告论文更为严苛。人类与LLM评审的侧重点也存在差异,LLM更常指出缺少基线对比,人类则更常提出计算效率方面的担忧。这些结果表明,宽泛的决策契合并不意味着与更精细的人类判断或评审优先级达成一致。
英文摘要:
Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether different providers align with both conference decisions and human reviewing priorities within the same controlled setting. We compare reviews from OpenAI GPT-5.4, Google Gemini 3.1 Pro Preview, and Anthropic Claude Opus 4.6 with human reviews and final decisions for 300 topic-matched ICLR 2026 submissions, equally divided among oral, poster, and rejected papers. Each model reviewed every paper using identical instructions and rating scales after decision information was removed. Our study contributes a cross-provider analysis of three complementary dimensions: alignment with broad and fine-grained decision categories, differences in recommendation-scale usage, and thematic agreement in identified weaknesses. All three LLMs distinguished accepted from rejected papers, but none reproduced the oral versus poster distinction present in human ratings. Scoring patterns were provider-specific: Gemini assigned systematically higher ratings, while OpenAI and Claude were closer to humans for rejected and poster papers but more critical of oral papers. Human and LLM reviews also differed in emphasis, with LLMs more frequently identifying missing baseline comparisons and humans more often raising computational-efficiency concerns. These results show that broad decision alignment does not imply agreement with finer human judgments or reviewing priorities.