基于人工智能的论文评审:关于人类评审优先级及其对自动评审影响的实证研究
AI-Based Thesis Assessment: An Empirical Study of Human Evaluation Priorities and Their Impact on Automated Assessment
浏览论文内容
中文总结 AI 辅助
该研究调查了84名导师确定的35项论文评审准则权重,对比人工智能系统RubiSCoT的默认权重后发现存在显著差异,整合导师权重的最优配置仅小幅降低评审偏差且无统计显著性,仅权重校准无法大幅提升AI与人类评审的一致性。
中文摘要 AI 辅助
基于评分规则的人工智能论文评审系统使用准则权重为不同评审准则分配重要性等级,这些权重通常通过专家判断确定,但几乎没有实证证据表明论文导师实际如何对评审准则进行优先级排序。因此,本研究调查了论文导师确定的论文评审准则权重,并评估其对人工智能评审的影响。我们调查了四个学科的84名论文导师,收集了35项论文评审准则的权重数据。与人工智能评审系统RubiSCoT[1]的默认准则权重对比发现,导师确定的权重与默认权重存在显著差异。为评估这些差异的实际影响,我们将导师确定的权重整合到多种校准配置中,并在包含80篇德语论文的语料库上进行评估。表现最佳的配置将人工智能生成评审与导师分配评审之间的平均相对偏差从11.18%降至10.85%,不过该改进不具备统计学意义。人类导师之间的一致性显著更强,导师间平均相对偏差为4.44%。研究结果表明,仅准则权重校准无法大幅提升人工智能评审与人类评审的一致性。
英文摘要
Rubric-based AI systems for thesis assessment use criterion weights to assign different levels of importance to evaluation criteria. These weights are typically defined through expert judgment, although little empirical evidence exists regarding how thesis supervisors actually prioritize evaluation criteria. Consequently, this study investigates supervisor-derived criterion weights in thesis assessment and evaluates their impact on AI-based assessment. We surveyed 84 thesis supervisors across four academic disciplines and collected weighting data for 35 thesis assessment criteria. Comparison with the default criterion weights of the AI assessment system RubiSCoT [1] revealed substantial divergences between supervisor-derived and default criterion weights. To evaluate the practical implications of these differences, the supervisor-derived weights were integrated into multiple calibration configurations and evaluated on a corpus of 80 German-language theses. The best-performing configuration reduced the mean relative deviation between AI-generated and supervisor-assigned evaluations from 11.18% to 10.85%, although the improvement was not statistically significant. Human supervisors showed substantially stronger agreement with each other, exhibiting a mean inter-supervisor relative deviation of 4.44%. The findings indicate that criterion-weight calibration alone does not substantially improve alignment between AI-generated and human assessments.
发表机构
- IU International University of Applied Sciences(IU应用科学大学)
机构由 AI 辅助整理,请以论文原文为准。