arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在发表后研究评估中,评审者比论文更重要

Judges matter more than papers in post-publication research assessment

Robert Ward, Alex Jones, Lutz Bornmann

arXiv 2607.09783首次发表:更新:

AI 中文总结

研究探讨发表后研究评估中评审者与论文对评估结果的影响,通过大型数据库和多级模型发现评审者相关效应占评分变异比例大,方向性偏差解释变异小,得出评估结果受评审者影响更大及需进行噪声审计的结论。

AI 中文摘要

研究评估依赖专家评审,但人类判断有噪声,不清楚评估差异主要源于真正研究质量差异还是评审者间不必要的差异。众多研究虽强调研究评估中的分歧和偏差,但未量化评审者相关噪声与被评估作品差异的关系。本文在一个大型发表后同行评审数据库中表明,研究评估更多受评审者间差异而非被评估研究差异驱动。通过多级模型分解了评审者相关变异,发现评审者相关效应在评分中占比远超被评估论文。方向性偏差指标解释的变异不到1%。结论是评估结果更多由评审者而非论文本身塑造,结果表明在高风险科学评估中进行噪声审计的必要性。

英文摘要

Research assessment relies on expert evaluations, yet human judgement is noisy, and it is unclear whether differences in assessment arise primarily from differences in genuine research quality or from unwanted differences between evaluators. While numerous studies highlight disagreement and biases in research assessment, they have not quantified judge-related noise relative to variation in the evaluated works. Here we show, in a large post-publication peer review database, that research assessment is driven more by differences between evaluators than by difference in the evaluated research. We partition variance in 239,521 research quality ratings assigned by 12,649 judges to 193,128 papers from the H1 Connect post-publication peer review platform. Using multilevel models, we decomposed judge-related variation into differences in overall severity and differences in the weighting of scientific attributes. We found that judge-related effects accounted for substantially more variance in ratings than the evaluated papers. In our most detailed model, judge-level effects and judge-specific slopes explained 61% of the total variance, whereas combined paper and journal-level effects accounted for only 7%. By contrast, examined measures of directional bias, such as author gender and global affiliation, explained less than 1% of the variance. We conclude that assessment outcomes were shaped more by the judges than by the papers themselves. Our results demonstrate the necessity of noise audits in high-stakes scientific evaluation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑