arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24471cs.DL

基于指标的研究评价中中立性的幻觉

The illusion of neutrality in metric-based research evaluation

  • New York University Abu Dhabi(阿布扎比纽约大学)

机构由 AI 辅助整理,请以论文原文为准。

Fengyuan Liu, Hazem Ibrahim, Yasir Zaki, Talal Rahwan

AI总结:

本研究通过调查869名研究者,发现他们在评价研究成果时对场所声望和引用次数赋予几乎相等的权重,但个体差异显著,且倾向于更重视引用,呼吁对定量指标使用保持谨慎。

AI中文摘要:

场所声望和引用次数是两种被广泛使用但并非完美的研究质量信号。当这两种信号发生冲突时,评价者必须决定各自应赋予多少权重。然而,目前尚不清楚不同学科的研究者在评估研究成果时如何权衡这两种信号。为填补这一空白,我们调查了869名研究者,使用假设的部门招聘规则之间的配对选择,这些规则对场所声望和引用次数赋予不同权重,询问哪种规则能产生更好的科学成果。在795名回答者中,其选择在内部大体一致,这些选择对两种信号赋予了几乎相等的总体权重。这种表面上的平衡掩盖了显著的个体异质性:近五分之二的回答者处于最重视场所或最重视引用的区间。此外,回答者平均选择的更重视引用的规则多于他们认为其部门在招聘中实际使用的规则。通过揭示研究指标冲突时出现的主观判断,我们的发现强化了对使用定量指标评价研究时应更加谨慎的呼吁。

英文摘要:

Venue prestige and citation counts are two widely used, albeit imperfect, signals of research quality. When the two signals conflict, evaluators must decide how much weight to assign each. Yet, it remains unknown how researchers across disciplines trade off these two signals when evaluating research outcome. To fill this gap, we surveyed 869 researchers using paired choices between hypothetical departmental hiring rules that assigned different weights to venue prestige and citation counts, asking which would produce better science. Among 795 respondents whose choices were largely internally consistent, choices placed nearly equal aggregate weight on the two signals. This apparent balance concealed substantial individual heterogeneity: nearly two in five respondents occupied the most venue-heavy or citation-heavy intervals. Moreover, respondents on average chose more citation-heavy rules than they believed their departments used in hiring. By revealing the subjective judgments that arise when research indicators conflict, our findings reinforce calls for greater caution when quantitative indicators are used to evaluate research.

↑