arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FIRSTPASS:基于真实编辑结果的多领域多轮同行评审数据集

FIRSTPASS: A Multi-Domain, Multi-Round Peer Review Dataset Grounded in Real Editorial Outcomes

Prabhjot Singh, Somnath Luitel, Manmeet Singh, Josh Durkee

arXiv 2608.26129首次发表:更新:

AI 中文总结

该研究推出首个多领域多轮同行评审数据集FIRSTPASS,填补现有数据集学科单一的空白,为跨学科AI科学判断提供可复现的基准测试资源。

AI 中文摘要

现有的科学同行评审数据集仅针对计算机科学和机器学习 venues 训练 AI 系统,这些模型能评判消融研究,却从未见过生物审稿人要求污染控制或化学家质疑核磁共振(NMR)光谱归属。我们推出 FIRSTPASS,这是首个基于多学科高影响力期刊完整多轮编辑对话构建的大规模同行评审数据集。该数据集从《自然-通讯》2022年11月起推行的强制性透明同行评审中筛选而来,包含3668条记录,覆盖5个科学领域(生物学、化学、神经科学、物理学和地球科学),捕捉了科学验证的完整迭代结构:初始审稿人报告、作者逐点回应以及更新后的审稿人评估。每条记录带有直接源自编辑决策的结果标签(STANDARD 代表两轮评审;EXTENDED 代表三轮及以上评审),提供了以往所有语料库中缺失的真实标注。自动审核确认其内容完整性达100%,专家评审平均字数为2155,比会议 venue 的评审密集得多。所有数据、解析管道和评估脚本均已发布,以支持跨学科 AI 科学判断的可复现基准测试。

英文摘要

Scientific peer review datasets have trained AI systems exclusively on Computer Science and Machine Learning venues, producing models that critique ablation studies yet have never seen a biology reviewer demand contamination controls or a chemist question Nuclear Magnetic Resonance (NMR) spectral assignments. We introduce FIRSTPASS, the first large-scale peer review dataset built on complete multi-round editorial dialogues from a multidisciplinary high-impact journal. Curated from Nature Communications mandatory transparent peer review (instituted November 2022), FIRSTPASS comprises 3,668 records spanning five scientific domains (biology, chemistry, neuroscience, physics, and earth science), capturing the full iterative structure of scientific validation: initial referee reports, author point-by-point responses, and updated reviewer assessments. Each record carries an outcome label derived directly from editorial decisions (STANDARD for two-round review; EXTENDED for three or more rounds), providing ground truth absent in all prior corpora. An automated audit confirms 100% content integrity. Expert reviews average 2,155 words, substantially denser than conference venue reviews. All data, parsing pipelines, and evaluation scripts are released to enable reproducible benchmarking of AI scientific judgment across disciplines.

CommentsAccepted at the AI for Science Workshop at the 43rd International Conference on Machine Learning (ICML 2026), 3 pages, 1 figure, 1 table

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑