arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大规模AI手写物理评估评分:评分一致性与奥林匹克团队选拔结果

Large Scale AI Grading of Handwritten Physics Assessments: Score Agreement and Olympiad Team Selection Outcomes

Praveen Pathak, Siddharth Tiwary, Charudatt Kadolkar, Vijay Singh, David Rakestraw, Shirish Pathare, Anwesh Mazumdar

arXiv 2608.20521首次发表:更新:

发表机构

Homi Bhabha Centre for Science Education–TIFR; Lawrence Livermore National Laboratory; University of California, Berkeley; Indian Institute of Technology, Guwahati; Centre for Excellence in Basic Sciences(霍米·巴巴科学教育中心-塔塔基础研究院; 劳伦斯利弗莫尔国家实验室; 加州大学伯克利分校; 古瓦哈蒂印度理工学院; 基础科学卓越中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究评估基于GPT-5.5的AI对三类物理手写评估的评分表现,其与官方评分相关性高且能准确选拔奥林匹克团队,第二轮评分优化了一致性,AI评分可作为考官控制下的辅助评分工具。

AI 中文摘要

多模态AI能够读取手写物理解题内容,但高风险评分需与官方分数及结果保持一致。本研究评估了基于GPT-5.5的AI在三项评估任务中的评分表现,涉及416名考生的520份手写提交内容、10364页扫描页面,三项评估分别为全国物理奥林匹克理论考试、包含理论与实验部分的奥林匹克最终选拔营、大学量子力学考试。每份提交由AI依据官方评分标准评分两次,第二轮评分采用了经第一轮分歧分析后制定的逐页修改及证据定位指令,评分过程中AI未获取官方人工评分或AI与人工评分的对比结果。AI与官方评分的总分相关性较高,介于0.91至0.97之间;在奥林匹克最终选拔中,AI恢复的五人团队与官方选拔结果完全一致;第二轮评分提升了整体题目部分的评分一致性,尤其在第一轮分歧较大的区域,主要难点仍在于精确的部分学分评分,尤其是实验工作部分。因此,可靠的AI评分依赖于详细的评分标准,应在考官控制下作为第二评分者或审计工具使用。

英文摘要

Multimodal AI can read handwritten physics solutions, but high-stakes grading requires agreement with official scores and outcomes. This study evaluated GPT-5.5-based grading on 10364 scanned pages from 520 handwritten submissions by 416 unique candidates or students across three assessments: a national Physics Olympiad theory examination, the final Olympiad selection camp with theory and experiment components, and a university quantum-mechanics examination. Each submission was graded twice by AI using the official rubrics. The second round used revised page-by-page and evidence-location instructions developed after first-round disagreement analysis. During grading, AI did not see official human marks or AI--human comparisons. Total-score correlations with official marks were high (0.91--0.97). For the final Olympiad selection, AI recovered the same five-student team as official grading. The second round improved aggregate question-part agreement, especially where first-round disagreements were larger. The main difficulty remained exact partial-credit grading, especially in experimental work. Reliable AI grading therefore depends on detailed rubrics and should be used as a second reader or audit tool under examiner control.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑