发表机构
ITMO University; Russian Presidential Academy of National Economy and Public Administration; Ivannikov Institute for System Programming of the Russian Academy of Sciences(ITMO大学; 俄罗斯总统国民经济与公共管理学院; 俄罗斯科学院伊万尼科夫系统编程研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出SAGE系统,利用Schema引导大语言模型将基金评审细则转化为结构化检查,生成与证据关联的可审计草稿,在35份申请上评估,辅助评审一致性达kappa=0.58,优于基线,提升了评审的可靠性和可审计性。
AI 中文摘要
基金评审人必须将详细的标准应用于申请表、预算和支持性文件,同时生成同事可以检查的评估。我们提出了SAGE,即Schema引导的基于方面的基金评估,这是一个将基金评审细则转化为结构化检查并将其判断与申请包中的证据联系起来的系统。我们在35份非营利基金申请上分两个阶段评估SAGE。与原始竞赛中的105份评审进行事后比较显示,序数一致性为中等(kappa = 0.29)。基金会随后在检查SAGE后进行了一项标准级别的重新评审,产生了202项评估。在此辅助轮次中,SAGE达到了kappa = 0.58,并优于每个标准一个提示的基线(在共同子集上kappa = 0.33),具有更高的秩相关性和更低的误差。一项声明级别的审计进一步识别了结构化草稿中已确认、有争议和未处理的部分。SAGE通过生成一份详细的、与证据相关联的且可审计的草稿供专家修正,从而实现了评审方法论的操作化。
英文摘要
Grant reviewers must apply detailed criteria to application forms, budgets, and supporting documents while producing assessments that colleagues can inspect. We present SAGE, Schema-Guided Aspect-Based Grant Evaluation, a system that translates a grant rubric into structured checks and links its judgements to evidence from the application package. We evaluate SAGE in two stages on 35 nonprofit grant applications. A post-factum comparison with 105 reviews from the original competition shows fair ordinal agreement (kappa = 0.29). The foundation then conducted a criterion-level re-review after inspecting SAGE, producing 202 assessments. In this assisted round, SAGE reached kappa = 0.58 and outperformed a one-prompt-per-criterion baseline (kappa = 0.33 on the common subset), with higher rank correlation and lower error. A claim-level audit further identifies confirmed, disputed, and unaddressed parts of the structured draft. SAGE operationalizes the review methodology by producing a detailed, evidence-linked, and auditable draft for expert correction.
Comments14 pages, 2 figures, 10 tables