LLM在同行评审中的使用与影响:ICML 2026随机实验与调查
Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026
浏览论文内容
中文总结 AI 辅助
本研究通过ICML 2026的随机实验和调查,发现LLM使用政策对评审结果影响甚微,但存在显著不合规行为,为同行评审政策设计提供启示。
中文摘要 AI 辅助
LLM正在迅速重塑同行评审,使得理解评审者如何在实践中使用它们以及不同的LLM使用政策如何影响评审结果变得至关重要。我们通过在ICML 2026(一个涉及超过24,000篇论文和17,000名评审者的重要机器学习会议)上进行的随机实验和匿名后续调查来研究这些问题。评审者被分配到禁止所有LLM使用的保守政策或允许有限辅助的宽松政策,并在主轨道的部分论文和评审者中进行随机化。政策分配对最终论文决定、论文评分和评审者信心几乎没有影响,尽管宽松政策下的评审长度增加了5.5-7%。后续调查回复(N=1,486)揭示了对LLM的多样化态度和显著的不合规行为:22.5%的保守政策评审者报告尽管有禁令仍使用了LLM,36.5%的宽松政策评审者报告至少一次明确禁止的使用。我们讨论了对未来同行评审政策和工具设计的影响。
英文摘要
LLMs are rapidly reshaping peer review, making it important to understand how reviewers use them in practice and how different LLM-use policies affect review outcomes. We investigate these questions through a randomized experiment and an anonymous post-survey at ICML 2026, a major machine learning conference involving over 24,000 papers and 17,000 reviewers. Reviewers were assigned to either a conservative policy prohibiting all LLM use or a permissive policy allowing limited assistance, with randomization among a subset of main-track papers and reviewers. Policy assignment had near-zero effects on final paper decisions, paper scores, and reviewer confidence, although reviews under the permissive policy were 5.5-7% longer. Post-survey responses (N=1,486) revealed diverse attitudes toward LLMs and substantial noncompliance: 22.5% of conservative-policy reviewers reported using an LLM despite the prohibition, and 36.5% of permissive-policy reviewers reported at least one explicitly disallowed use. We discuss implications for future peer-review policy and tool design.
发表机构
- Microsoft Research(微软研究院)
- University of Pennsylvania(宾夕法尼亚大学)
- Google Research(谷歌研究)
- University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
- EPFL(洛桑联邦理工学院)
- Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。