arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

机器学习同行评审中,审稿人评分在不同研究领域间不具备可比性

Reviewer Scores Are Not Comparable Across Research Areas in ML Peer Review

Binyan Xu, Xilin Dai, Fan Yang, Kehuan Zhang

arXiv 2607.27209首次发表:更新:

发表机构

The Chinese University of Hong Kong; Zhejiang University(香港中文大学; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究通过ICLR 2021-2026年50289篇论文数据发现,机器学习同行评审中同一审稿人评分在不同研究领域的论文接收概率差达8倍,因测量设计缺陷导致评分跨领域不可比,呼吁采用校准评审信号并发布主题分层接收率

AI 中文摘要

机器学习会议的同行评审日益依赖审稿人评分作为主要决策工具。随着每年投稿量从数千篇增至数万篇,尚无系统审查验证该工具是否在各研究领域间发挥统一作用,或接收结果是否受审稿人评分既未捕捉也无法控制的因素影响。本立场论文认为,接收结果受审稿人评分之外的因素影响,根本原因是测量设计缺陷,而非个体偏见。当固定数值尺度汇总结构不均的审稿人群体的质量判断时,绝对评分在不同领域间变得不可比,领域主席必须用领域先验替代基于评分的决策。利用ICLR 2021至2026年涵盖219个研究主题的50289篇论文数据,我们表明,在任意给定审稿人评分下,论文的接收概率依主题不同可相差高达8倍。我们排除了评分文化、专家审稿人标准、理性领域主席加权及质量稀释等其他解释。我们呼吁程序委员会采用固有校准的评审信号,并发布按主题分层、基于评分条件的接收率作为一级公平度量标准。

英文摘要

Peer review at ML conferences increasingly relies on reviewer scores as the primary decision instrument. As submissions have scaled from thousands to tens of thousands per year, no systematic audit has examined whether this instrument functions uniformly across research areas, or whether acceptance outcomes are in practice shaped by forces that reviewer scores neither capture nor control. This position paper argues that acceptance outcomes are shaped by forces beyond reviewer scores, and that the underlying cause is a measurement design failure, not individual bias. When a fixed numerical scale aggregates quality judgments across communities with structurally non-uniform reviewer pools, absolute scores become incomparable across areas, and area chairs must substitute community priors for score-based decisions. Using ICLR 2021--2026 data covering 50,289 papers across 219 research topics, we show that at any given reviewer score, a paper's acceptance probability varies by up to 8x depending on its topic. We rule out scoring culture, expert reviewer standards, rational area chair reweighting, and quality dilution as alternative explanations. We call on program committees to adopt inherently calibrated review signals and publish topic-stratified, score-conditional acceptance rates as a first-class fairness metric.

Comments23 pages, 8 figures, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑