arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

边界线处的偏见:同行评审中谁获得疑义利益?

Bias at the Borderline: Who Gets the Benefit of the Doubt in Peer Review?

Hazem Ibrahim, Talal Rahwan, Yasir Zaki

arXiv 2607.26280首次发表:更新:

AI 中文总结

该研究以ICLR 2019-2025年3万余篇稿件为对象,发现边界线稿件中无顶尖机构作者的接受率更低,差距集中在决策前有arXiv预印本的稿件,经预注册检验未发现对应结果差异,揭示双盲评审旨在防范的统计歧视漏洞。

AI 中文摘要

本研究针对ICLR(一个公开包含被拒稿件完整评审记录的大型机器学习会议)的同行评审展开。评审人员对每篇稿件打分;对于分数未明确结果的边界带稿件,领域主席会做出自主的接受或拒绝决定。我们探究该决定是否公平:来自知名机构、WEIRD国家(西方、受过教育、工业化、富裕、民主国家)或全男性团队的作者是否会在边际上获得疑义利益?覆盖2019-2025年ICLR的31711篇稿件,其中10416篇为边界线稿件,在评审分数相同的情况下,无前25机构作者的边界线稿件接受率低0.5至1.6个百分点。该差距出现在自主决定阶段,在预注册的2026年ICLR队列的样本外数据中再次出现,且几乎完全集中在通过决策前arXiv预印本可识别的稿件中(差距为-3.4个百分点,而无预印本的稿件差距为-0.2个百分点)。相同分数未必意味着稿件质量相当:领域主席可能对分数未涵盖的质量做出回应。我们采用稳健结果检验,仅当接受率较低的群体实现更好的下游结果时才判定存在歧视。我们测量了决策双方的5项结果(引用量、影响力、两种新颖性形式、最终发表 venues),包括对被拒稿件的首次“被错失的稿件”检验。核心结果为零:在预注册的27项检验中,经校正后无与接受率差距一致的差异;无证据表明任何群体在我们测量的结果上面临更高门槛。该零结果并非免责:决策前预印本通过政策允许的手段打破了盲审,接受差距几乎完全存在于该漏洞中,符合领域主席将已显现的机构声望作为先验使用的情况——这是结果检验可能无法检测到的统计歧视,也是双盲评审旨在防止的做法。

英文摘要

We study peer review at ICLR, a large machine-learning conference whose complete review record, including rejected submissions, is public. Reviewers score each submission; for the borderline band whose scores do not settle an outcome, an area chair makes a discretionary accept-or-reject call. We ask whether that call is even-handed: do authors from prestigious institutions, WEIRD countries, or all-male teams get the benefit of the doubt at the margin? Across ICLR 2019-2025 (31,711 submissions; 10,416 borderline), borderline papers without a top-25-institution author are accepted at a 0.5 to 1.6 percentage point lower rate at the same reviewer scores. The gap arises at the discretionary stage, reappears out-of-sample in the pre-registered ICLR 2026 cohort, and concentrates almost entirely among submissions identifiable through a pre-decision arXiv preprint (-3.4 vs. -0.2 points). Equal scores need not mean equal papers: an area chair may respond to quality the scores miss. We apply a robust outcome test, which concludes discrimination only when the group accepted at a lower rate also realizes better downstream outcomes. We measure five outcomes (citations, disruption, two forms of novelty, eventual venue) on both sides of the decision, including the first "ones that got away" test of rejected submissions. Our headline result is a null: across a pre-registered family of 27 tests, no disparity concordant with the decision-rate gap survives correction; we find no evidence that any group faced a higher bar on the outcomes we measure. That null is not an exoneration. A pre-decision preprint pierces the blind through policy-permitted means, and the acceptance gap lives almost entirely in that porosity, consistent with area chairs using revealed institutional prestige as a prior: statistical discrimination that outcome tests may not detect, and a practice double-blind review exists to prevent.

Comments27 pages, 6 figures, 11 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑