arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ChatGPT评估已发表期刊文章质量(从PDF、标题或摘要)是否与个人评审员一样可靠?

Is ChatGPT as reliable as individual reviewers assessing the quality of published journal articles from PDFs or titles and abstracts?

Mike Thelwall, Parveen Ali

arXiv 2607.25965首次发表:更新:

AI 中文总结

研究比较ChatGPT-5.4与个人评审员对期刊文章质量评分,用专家评分作参照,涉及从标题/摘要及PDF输入的评分。结果显示与专家评分相关性大多为正,ChatGPT-5.4对PDF评估更详但评分预测未改善,其深入评论有误导性。

AI 中文摘要

虽然大语言模型(LLMs)在对已发表期刊文章的研究质量进行评分方面能力较弱到中等,但它们尚未与个体专家评审员进行比较。基于PDF的ChatGPT质量评分是否能通过更深入的评估来提高标题和摘要的评分也尚不清楚。为了解决这两个问题,本文使用了专家评分(英国评估单位[UoA]3联合健康专业的98个内部部门评级、UoA13建筑的44个评级以及UoA34的58篇图书馆与信息科学文章),并将其与ChatGPT - 5.4从标题/摘要和PDF输入中得出的评分进行比较。对于UoA3,个体评审员的评分也相互进行了比较,并与ChatGPT - 5.4进行了比较。与专家评分的等级相关性几乎都在统计上显著为正,但相关性之间的差异大多不显著,尽管微弱地表明ChatGPT - 5.4对于UoA3可能比个体评审员更可靠。此外,虽然ChatGPT - 5.4对PDF的评估比标题/摘要更详细,但其评分预测似乎并没有改善。因此,虽然结果大致证实了ChatGPT评分在对学术文档进行排名方面的价值,但其对PDF的明显更深入的评估评论在没有转化为改进的评分预测这一意义上是有误导性的。

英文摘要

Whilst Large Language Models (LLMs) have a weak to moderate ability to score published journal articles for research quality, they have not been compared with individual expert reviewers. It is also unknown whether quality scores from ChatGPT based on PDFs can improve on those from titles and abstracts through a deeper evaluation. To address both issues, this article uses expert scores (98 internal departmental ratings for UK Unit of Assessment [UoA] 3 Allied Health Professions, 44 for UoA13 Architecture, and 58 library and information science articles from UoA34), comparing them against ChatGPT-5.4 scores from both title/abstract and PDF inputs. For UoA3, individual reviewer scores were also compared against each other and ChatGPT-5.4. The rank correlations with expert scores are almost all statistically significantly positive, but differences between the correlations are mostly not, despite weakly suggesting that ChatGPT-5.4 can be more reliable than individual reviewers for UoA3. Moreover, whilst ChatGPT-5.4 provides more detailed evaluations of PDFs than of titles/abstracts, its score predictions do not seem to improve. Thus, whilst the results broadly confirm the value of ChatGPT scores for ranking academic documents, its apparently deeper evaluative comments on PDFs are misleading in the sense of not translating to improved score predictions.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑