发表机构
Keio University; NEC Corporation(庆应义塾大学; 日本电气公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究分析不同类型审稿人指南对基于大语言模型的自动同行评审的影响,通过实验发现官方会议指南效果最佳,模仿审稿人指南较差,且严格评分标准会降低性能,强调主观和整体评分的重要性。
AI 中文摘要
同行评审是科学研究中的重要环节,工作量的增加使其自动化变得愈发必要。本研究分析了不同类型的审稿人指南,如官方会议指南以及通过大语言模型从高质量人工评审中生成的模仿审稿人指南,对自动同行评审的影响。实验表明官方会议指南产生的评审结果与人类判断最一致,而模仿审稿人指南通常效果较差。此外,严格执行评分标准会降低性能,凸显了允许主观和整体评分的重要性。
英文摘要
Peer review is an essential process in scientific research, yet the growing workload has made its automation increasingly necessary. In this study, we analyze how different types of reviewer guidelines, such as official conference guidelines and reviewer-imitating ones generated from high-quality human reviews using LLMs, affect automated peer review. Our experiments show that official conference guidelines produce review results most consistent with human judgments, suggesting that evaluation criteria refined through conference practice serve as effective guidance for automated reviewing as well. In contrast, reviewer-imitating guidelines were generally less effective than official conference guidelines. Furthermore, enforcing strict rubric-style scoring consistently degraded performance, highlighting the importance of allowing subjective and holistic scoring.
Comments18 pages, 2 figures, ACL 2026 Findings