arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越类人性:映射大语言模型(LLMs)与人类评审者的科学评审特征

Beyond Human-Likeness: Mapping the Scientific Critique Profiles of LLMs and Human Reviewers

Yunhan Yang, Mike Thelwall, Guoxiu He

arXiv 2609.01895首次发表:更新:

发表机构

School of Information, Journalism and Communication, The University of Sheffield; School of Economics and Management, East China Normal University(谢菲尔德大学信息、新闻与传播学院; 华东师范大学经济与管理学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究以ICLR 2025同行评审数据为基础,对比人类与不同提示下的LLM评审特征,发现二者评审侧重不同,LLM辅助评审改变了评审功能构成,需区分LLM放大内容与人类负责判断领域。

AI 中文摘要

大语言模型(LLMs)作为同行评审工具的价值日益受到讨论,但人们常通过类人性、感知有用性或与评审意见的文本重叠度来评估其价值。本研究将关注点从LLMs是否类似人类评审者转向它们执行何种科学评审功能。利用ICLR 2025的同行评审数据,我们将人类评审与基线提示和专家提示下生成的LLM评审进行比较。我们通过两种评审行为(弱点评审和科学质疑)来操作化科学评审,并使用五个理论导向框架对逐点评审文本进行标注:安德森的知识类型、图尔敏的论证模型、格雷瑟的问题深度、SOLO认知复杂度以及哈蒂的反馈功能。结果揭示了差异化的评审特征:人类评审更强调科学框架和修改指导,更常识别高阶弱点并提出以改进为导向的问题;LLM评审在解释深度、整合推理和明确论证结构方面表现出更高的比率。专家提示并未使LLM评审整体更类人,它部分缩小了部分差距,但主要放大了LLM特有的整合与形式论证倾向。这些发现表明,LLM辅助的同行评审改变了评审文本的功能构成,因此区分LLM放大的评审内容与需要人类优先处理和负责任判断的领域至关重要。

英文摘要

Large language models (LLMs) are increasingly discussed as tools for peer review, but their value is often assessed through human-likeness, perceived usefulness, or textual overlap with reviewer comments. This study shifts attention from whether LLMs resemble human reviewers to what functions of scientific critique they perform. Using ICLR 2025 peer-review data, we compare human reviews with LLM reviews generated under baseline and expert prompts. We operationalize scientific critique through two review acts, weakness critique and scientific questioning, and annotate point-level review text using five theory-guided frameworks: Anderson's knowledge types, Toulmin's argumentation model, Graesser's question depth, SOLO cognitive complexity, and Hattie's feedback functions. The results reveal a differentiated critique profile. Human reviews placed greater emphasis on scientific framing and revision guidance, more often identifying higher-order weaknesses and asking questions oriented toward improvement. LLM reviews showed higher rates of explanatory depth, integrative reasoning, and explicit argument structuring. Expert prompting did not make LLM critique uniformly more human-like; it partially narrowed some gaps but mainly amplified LLM-specific tendencies toward integration and formal argumentation. These findings show that LLM-assisted peer review changes the functional composition of review text, making it important to distinguish LLM-amplified critique from areas requiring human prioritization and accountable judgement.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑