发表机构
Sakana AI; National University of Singapore(Sakana AI; 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM辅助同行评审仅聚焦模仿人类评审的问题,提出以错误检测为核心的多层评审(MLR)框架及含人工插入错误的可扩展基准,该方法与人类评审高度一致且错误检测性能优异,但存在对抗性操纵脆弱性,需强化鲁棒性。
AI 中文摘要
科学出版的快速发展给同行评审带来了压力,尤其是在机器学习领域,引发了评审质量下降和评审人工作量增加的担忧。大语言模型(LLM)已被提议作为自动化评审助手,但对其的评估主要集中在模仿人类撰写的评审,而非支持同行评审的核心功能。在此,我们引入以验证为中心的视角看待LLM辅助同行评审,强调错误检测是一项关键且资源密集型任务。我们提出了一个可扩展的基准,用于评估评审系统识别逻辑矛盾的能力,该基准通过向会议论文中人工插入错误构建而成,可产生明确的评估目标并支持系统比较。我们进一步提出了多层评审(MLR)框架,该框架在生成评审前优先详细理解稿件,更贴合人类评审实践,同时提高了令牌效率。在评估中,我们的方法表现出与人类评审分数的高度一致性,实现了高错误检测性能,并为评审人关注点提供了补充视角。这些改进可归因于底层LLM的选择和我们系统的设计。同时,我们证实了其对对抗性操纵的持续脆弱性,强调了自动化评审系统鲁棒性的必要性。我们的研究结果凸显了严格、以错误为中心的评估对指导基于LLM的工具在同行评审及其他关键科学工作流中负责任部署的重要性。
英文摘要
The rapid growth of scientific publishing has strained peer review, particularly in machine learning, raising concerns about declining review quality and increasing reviewer workload. Large language models (LLMs) have been proposed as automated review assistants, yet their evaluation has focused largely on imitating human-written reviews rather than supporting the core functions of peer review. Here, we introduce a verification-centric perspective on LLM-assisted peer review, emphasizing error detection as a critical and resource-intensive task. We present a scalable benchmark that evaluates review systems' ability to identify logical contradictions, constructed through synthetic insertion of errors into conference papers, yielding unambiguous evaluation targets and enabling systematic comparison. We further propose a Multi-Layered Review (MLR) framework that prioritizes detailed manuscript comprehension before review generation, aligning more closely with human reviewing practices while improving token efficiency. Across evaluations, our approach demonstrates strong alignment with human review scores, achieves high error detection performance, and provides complementary perspectives on reviewer focus. These improvements can be attributed to both the choice of the underlying LLM and the design of our system. At the same time, we corroborate persistent vulnerabilities to adversarial manipulation, underscoring the need for robustness in automated review systems. Our findings highlight the importance of rigorous, error-focused evaluation to guide responsible deployment of LLM-based tools in peer review and other critical scientific workflows.