使用大型语言模型对资助申请进行评分
Scoring Grant Applications with Large Language Models
浏览论文内容
中文总结 AI 辅助
本研究评估六个开放权重LLM对2267份英国资助申请的评分,发现其平均排名与专家评分正相关,虽不足以取代最终评审,但可辅助初筛和偏见检查。
中文摘要 AI 辅助
目的:评估资助申请既耗时又困难,这增加了学术同行评审的整体负担。虽然资助方正在探索人工智能是否能提供帮助,但目前尚无关于大型语言模型(LLMs)对当代资助申请评分准确性的已发表研究。设计/方法/途径:本研究调查了六个开放权重的大型语言模型(Gemma 3 1B/4B/12B/27B、DeepSeek R1 32B、Qwen 3 32B)能否为2267份近期英国经济与社会研究理事会(ESRC)和工程与物理科学研究理事会(EPSRC)的资助申请提供有用的评分,并将其与原始评审人和资助小组成员的评分进行比较。研究发现:尽管LLM的评分单独来看并不准确,但取平均值并转换为排名后,它们与专家平均评分呈正相关。表现最佳的LLM,Gemma 3 27B(使用不同提示词进行10次迭代),与评审人平均评分的排名相关性为中等(平均rho=0.26)。Gemma 3 27B与单个评审人的平均相关性为0.19,低于评审人之间0.24的平均相关性,这表明其评分略弱于单个评审人的评分。Gemma 3 27B与资助小组成员平均评分的排名相关性较弱(平均rho=0.17),与单个小组成员的平均相关性较低(平均rho=0.14),这远低于小组成员之间的相关性(平均rho=0.38)。虽然这些相关性似乎太弱,无法在最终小组阶段取代专家评审,但LLM评分可能有助于初始评审阶段,例如帮助识别最薄弱的提案以进行快速桌面拒稿、替换一名人工评审员,或用于三角验证以检查偏见。
英文摘要
Purpose: Assessing grant applications is time-consuming and difficult, adding to the overall burden of academic peer review. Whilst funders are exploring whether AI can help, there is no published research into the accuracy of Large Language Models (LLMs) for scoring contemporary grants. Design/methodology/approach: This study investigates whether six open-weight LLMs (Gemma 3 1B/4B/12B/27B, DeepSeek R1 32B, Qwen 3 32B) can give useful scores for 2267 recent UK Economic and Social Research Council (ESRC), and Engineering and Physical Sciences Research Council (EPSRC) grant applications, comparing them with scores from the original reviewers and funding panel members. Findings: Although the LLM scores are individually inaccurate, when averaged and converted to ranks they correlate positively with expert average scores. The best performing LLM, Gemma 3 27B (10 iterations with varied prompts), had moderate rank correlations with average reviewer scores (mean rho=0.26). Gemma 3 27B's average correlation with individual reviewers was 0.19, which is lower than the inter-reviewer mean correlation of 0.24, suggesting that it scores are slightly weaker than individual reviewer scores. Gemma 3 27B had weak rank correlations with average panel member scores (mean rho=0.17), with lower average correlations with individual panellists (mean rho=0.14), which is substantially lower than the inter-panellist correlation (mean rho=0.38). Whilst the correlations seem too weak to replace expert review at the final panel stage, LLM scores might help with the initial reviewing state, such as by helping identify the weakest proposals for fast-track desk rejections, to replace one human reviewer, or for triangulation to check for bias.
发表机构
- University of Sheffield(谢菲尔德大学)
- The Open University(开放大学)
- University of Milano-Bicocca(米兰比可卡大学)
- University of Salford(索尔福德大学)
机构由 AI 辅助整理,请以论文原文为准。