AI 中文总结
该研究利用REF2021的近9万篇论文,发现ChatGPT对男性第一作者论文的评分略高于女性,此差异在部分学科更明显,提示AI研究评估需谨慎。
AI 中文摘要
大型语言模型(LLMs)正被考虑用于研究评估,这引发了人们对AI偏见引入的担忧。本研究利用英国2021年研究卓越框架(REF)的89744篇期刊论文,调查ChatGPT对研究质量的评分是否会因第一作者的性别而不同。为避免直接性别偏见,作者信息被隐瞒了。不过,在大多数评估单元(UoAs)中,男性第一作者的论文获得的ChatGPT评分略高,尤其是在健康、科学和工程相关学科,基于部门级代理指标,这种模式通常比REF评分的模式更明显。在大多数UoAs中,男性第一作者的论文在排名上相对于REF评分的ChatGPT增益也更有利,尽管差异通常较小。不过,在社会科学、艺术和人文领域的独立研究中,性别差异并不明显。至少从摘要复杂度来看,第一作者研究中偏向男性的模式无法用写作风格的性别差异来解释。一些ChatGPT与REF的差异也可能反映了生成REF代理评分时使用的部门平均过程。平均ChatGPT评分可能会通过领域、主题、方法、期刊背景或作者结构等其他因素间接因第一作者性别而不同。因此,这是对基于AI的研究评估需保持谨慎的另一个理由。
英文摘要
Large Language Models (LLMs) are being considered for research evaluation, raising concerns about the introduction of AI bias. This study investigates whether ChatGPT research quality scores differ by first-author gender using 89,744 journal articles from the UK Research Excellence Framework (REF) 2021. Author information was withheld from ChatGPT to avoid direct gender bias. Nevertheless, male first-authored papers had slightly higher ChatGPT scores in most Units of Assessment (UoAs), especially in health, science and engineering-related subjects, and this pattern was often stronger for ChatGPT than for REF scores, based on a departmental-level proxy. Rank-based ChatGPT gains relative to REF scores were also more favourable for male first-authored papers in most UoAs, although the differences were generally small. Gender differences were not evident for solo research in the social sciences, arts and humanities, however. The male-favouring pattern for first-authored research was not explained by gender differences in writing styles, at least as reflected in abstract complexity. Some ChatGPT-REF differences may also reflect the departmental averaging process used to generate the REF proxy scores. Average ChatGPT scores may differ by first-author gender indirectly through other factors, such as field, topic, method, journal context or authorship structure. Thus, this is an additional reason to be cautious with AI-based research evaluation.