发表机构
Technical University of Munich(慕尼黑工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究调研111个顶尖AI/NLP会议及医学期刊的AI评审政策,评估ICLR 2026和《自然·通讯》的AI评审质量,发现AI评审存在系统性缺陷,主张采用多维度评估。
AI 中文摘要
AI辅助同行评审作为支持科学出版流程的工具正受到越来越多的讨论和采用,但目前人们对出版 venues 如何规范其使用、以及当前AI评审系统的能力如何缺乏系统性理解。针对这些问题,我们首先对111个顶尖AI/NLP会议及医学期刊中面向评审人的AI政策展开调研,发现两个研究社区之间存在显著的监管差异。其次,我们利用包含原始手稿提交内容及数百份人工与机器生成评审意见的新型数据集,对ICLR 2026和《自然·通讯》(Nature Communications)的AI生成同行评审进行评估。我们采用大语言模型作为评审人(LLM-as-a-Judge)、评分一致性、评审细致度以及与人工评审关注点的重叠度等互补评估指标,对比开源模型与专有模型生成的评审意见。研究结果显示,当前大语言模型(LLM)能够生成详细且流畅的评审意见,但存在系统性缺陷,例如建议过于积极、批评内容泛化、证据支撑不均衡。我们证明仅靠 aggregate 质量评分会高估评审质量,并主张对AI生成的同行评审进行多维度评估。
英文摘要
AI-assisted peer review is increasingly discussed and adopted as a tool to support the scientific publishing process, yet there is little systematic understanding of how publication venues regulate its use or of how capable current AI review systems are. We address these questions by first surveying reviewer-facing AI policies across 111 leading AI/NLP conferences and medical journals, revealing substantial regulation differences between the two communities. Second, we evaluate AI-generated peer reviews at ICLR 2026 and Nature Communications using a novel dataset comprising original manuscript submissions and several hundred human- and machine-generated reviews. We compare reviews produced by open-source and proprietary models using complementary evaluation metrics, including LLM-as-a-Judge, score alignment, granularity, and overlap with human reviewers' concerns. Our results show that current LLMs can generate detailed and fluent reviews but exhibit systematic weaknesses, such as overly positive recommendations, generic criticism, and uneven evidence grounding. We demonstrate that aggregate quality scores alone can overestimate review quality and argue for multi-dimensional evaluation of AI-generated peer reviews.
Comments30 pages, 19 figures, 10 tables. Currently under peer review. GitHub link for code and data is given in the paper