AI 中文总结
本研究利用公开评审数据构建并校准了同行评审噪声的观测估计器,应用于ICLR 2017-2025年数据,估计不一致率为18-30%,并揭示2020与2021年高翻转率的不同机制。
AI 中文摘要
如果同一篇论文由不同的评审员进行评审,会议接受/拒绝决定会有多大变化?运行第二个独立的程序委员会是回答这一问题的黄金标准,但成本过高:仅执行过两次(NeurIPS 2014和2021)。我们仅利用公开的评审数据构建了这一数量的观测估计器,对其进行两次校准,并将其应用于ICLR的九年数据(2017-2025年;36,113篇论文,134,912条评审)。该估计器使用贝叶斯有序Probit模型将评分分解为论文质量和评审噪声,通过逻辑模型将评分映射为决策,并模拟两个独立委员会(后验抽取B=1,000;委员会规模k=2,3,4)。估计的不一致率在k=2时为23-30%,在k=4时为18-24%;30-50%的已接受论文会被拒绝。外部校准:在NeurIPS 2021评审员数量级别(k=3)下,模拟的2021年不一致率为23.3% [21.7%, 25.0%],而报告值为23.0%(偏差+0.3个百分点);接受精度和委员会相关性分别在5个百分点和0.04以内一致。内部校准:在18,740篇具有4条以上评审的论文上,随机的无模型2+2评审员拆分与k=2模拟在2018年和2021-2025年的差异在1个百分点以内。纵向来看,我们发现在2017-2025年间评审噪声没有稳健的时间趋势。2020年和2021年高接受论文翻转率具有不同的机制:2020年的四点量表压缩了分数(23.7%的论文内部方差为零),反事实分析表明,粗化量表会使不一致率提高约7个百分点;2021年则结合了样本中最低的信噪比和最接近阈值的接受情况。对于LLM时代,对论文内评分方差的2023年断点检验未发现断点,但该设计几乎没有统计功效,且不存在2022年后的评审文本或置信度数据,因此未尝试进行LLM归因。
英文摘要
How much of a conference accept/reject decision would change if the same paper were reviewed by a different set of reviewers? Running a second independent program committee is the gold standard for answering this, but it is prohibitively expensive: done only twice (NeurIPS 2014 and 2021). We build an observational estimator of this quantity from public review data alone, calibrate it twice, and apply it to nine years of ICLR (2017-2025; 36,113 papers, 134,912 reviews). The estimator decomposes scores with a Bayesian ordered-probit model into paper quality and reviewer noise, maps scores to decisions with a logistic model, and simulates two independent committees (posterior draws B=1,000; committee sizes k=2,3,4). Estimated disagreement rates are 23-30% at k=2 and 18-24% at k=4; 30-50% of accepted papers would be rejected. External calibration: at the NeurIPS 2021 reviewer-count caliber (k=3), the simulated 2021 disagreement rate is 23.3% [21.7%, 25.0%] vs. reported 23.0% (bias +0.3pp); accept precision and committee correlation agree within 5pp and 0.04. Internal calibration: on 18,740 papers with 4+ reviews, random model-free 2+2 reviewer splits agree with the k=2 simulation within 1pp in 2018 and 2021-2025. Longitudinally, we find no robust time trend in reviewer noise over 2017-2025. The high accepted-paper flip rates of 2020 and 2021 have distinct mechanisms: the 2020 four-point scale compressed scores (23.7% of papers had zero within-paper variance), and a counterfactual shows coarsening the scale raises disagreement by about 7pp; 2021 instead combined the lowest signal-to-noise ratio in the sample with the most threshold-crowded acceptances. For the LLM era, a 2023 breakpoint test on within-paper score variance finds no break, but the design has almost no power, and no post-2022 review text or confidence data exist, so no LLM attribution is attempted.