arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07523cs.CYcs.AIcs.PL

从被评估模型到评估辅助工具:面向编程考试的基于大语言模型的难度校准多证据研究

From Evaluated Models to Evaluation Aids: A Multi-Evidence Study of LLM-Based Difficulty Calibration for Programming Examinations

Hongfei Yan, Jiangkai Xiong, Yiqing Li, Chong Chen

首次发表
浏览论文内容

中文总结 AI 辅助

该研究将大语言模型作为编程考试难度校准的辅助证据,结合多类数据验证其与考试难度的相关性,可用于题目验证等场景,但存在应用限制。

中文摘要 AI 辅助

平行班编程考试的难度差异会影响课程评估的公平性。本研究将大语言模型从基准评估目标重新定位为解释考试难度的辅助证据源,结合AI证据与汇总的学生表现、题目曝光率、在线判题流程数据及教师解读。首先,10个模型与120名学生同步完成一套含8道题的期末考试:AI通过率与学生通过率呈正相关(斯皮尔曼相关系数rho=0.866,精确p值=0.0119),基于解题的综合难度指数与学生通过率呈负相关(rho=-0.905,精确p值=0.0046)。随后,通过可审计API调用在第三方兼容OpenAI的端点上运行单一结构化审查器,其模型标签(gpt-5.6-sol)无法认证为OpenAI官方上游模型;调用元数据和原始响应已存档。在来自11个平行班期末考试的79道题目中,AI整体难度与题目级通过率的相关系数rho=-0.871,与未作答率的相关系数rho=0.800;在含26道题的纵向“数据结构与算法B”样本中,相关系数分别为-0.829和0.883。含106道题的入门课程(CS101)样本是边界情况:题目级相关性减弱至rho=-0.552,16次考试的考试级相关性接近零,学生群体构成主导了考试级结果。曝光折扣(0-0.40)和重复题目扰动测试未改变这些相关方向。因此,AI证据可作为题目验证、平行班公平性讨论及纵向质量跟踪的外部参考,但模型身份边界、单一审查器设计及审查输出不稳定性设定了明确限制:AI难度标度不得用于单个学生评估或自动成绩调整。

英文摘要

Difficulty differences across parallel-class programming examinations affect the fairness of course assessment. This study repositions large language models from benchmark evaluation targets to auxiliary evidence sources for interpreting exam difficulty, combining AI evidence with aggregated student performance, item exposure, online-judge process data, and teacher interpretation. First, ten models solved an eight-problem final exam synchronously with 120 students: AI pass rate correlated positively with student pass rate (Spearman rho = 0.866, exact p = 0.0119), and a solving-based composite difficulty index correlated negatively with it (rho = -0.905, exact p = 0.0046). A single structured reviewer was then run via auditable API calls on a third-party OpenAI-compatible endpoint whose model label (gpt-5.6-sol) cannot authenticate an official OpenAI upstream model; call metadata and raw responses are archived. Across 79 problems from 11 parallel-class final exams, AI overall difficulty correlated with problem-level pass rate at rho = -0.871 and with non-attempt rate at rho = 0.800; in a 26-problem longitudinal Data Structures and Algorithms B sample, the correlations were -0.829 and 0.883. A 106-problem introductory-course (CS101) sample marks the boundary: the problem-level correlation weakened to rho = -0.552, and the exam-level correlation across 16 exams was near zero, with cohort composition dominating exam-level outcomes. Exposure-discount (0-0.40) and duplicate-problem perturbation tests did not change these directions. AI evidence can thus serve as an external reference for problem validation, parallel-class fairness discussion, and longitudinal quality tracking, while the model-identity boundary, single-reviewer design, and review-output instability set explicit limits: AI difficulty scales must not be used for individual student evaluation or automatic grade adjustment.

发表机构

  • School of Computer Science, Peking University(北京大学计算机学院)
  • Yuanpei College, Peking University(北京大学元培学院)
  • School of Government, Beijing Normal University(北京师范大学政府管理学院)

机构由 AI 辅助整理,请以论文原文为准。

↑