AI 中文总结
本研究基于一起课堂事件的成绩数据,通过统计分析发现期中与期末成绩几乎无关联,无法证明存在未经授权的AI辅助,仅提供了评估相关争议统计证据的框架。
AI 中文摘要
2026年春季,布朗大学的一名经济学教授布置了一份开卷期中考试,在出现异常高的分数后,将期末考试改为了监考形式。在完成课程的59名学生中,平均分数从100分制的95.7分降至48.8分。讲师将分数下降归因于在期中考试中未经授权使用生成式AI,而其他人则提出了考试焦虑、期末考试难度更大、学生退课以及均值回归等原因。公开的分数数据包含每名学生的期中与期末成绩配对,呈现出两个显著模式:学生的期中与期末成绩之间的相关性仅为0.06,个体成绩变化范围从4分的提升到100分的下降。置换检验发现,学生的期中与期末成绩之间不存在统计上可检测的关联,但当关联强度符合其他解释的预测时,该检验超过90%的概率能检测到关联。一旦考虑模型复杂度,学生的期中成绩不包含关于其期末成绩任何信息的模型,对数据的拟合效果与任何允许关联的模型相当。这两个结果都不排除存在弱关联的可能,但共同表明数据并不要求存在任何关联。在模拟课堂中,每名学生的两项成绩通过其自身能力保持关联,正如其他解释所暗示的那样,上述两个模式几乎不会同时出现;只有当这种关联几乎被切断时,两个模式才会频繁同时出现。可能存在多种切断关联的机制:非学生本人完成的期中答案会切断关联,测试内容与期中显著不同的期末考试(无AI辅助)也会切断关联。这些成绩无法区分这些可能性,也无法识别出是否有学生使用了未经授权的辅助。本研究提供了一个框架,用于评估未来学术考核中关于生成式AI使用的争议中的统计证据。
英文摘要
In spring 2026, an economics professor at Brown University gave a take-home midterm and, after unusually high scores, made the final exam proctored. Among the 59 students who completed the course, average scores fell from 95.7 out of 100 to 48.8. The instructor attributed the drop to unauthorized use of generative AI on the midterm; others proposed test anxiety, a harder final, student withdrawals, and regression to the mean. The publicly released scores, one midterm-final pair per student, show two striking patterns: the correlation between a student's two scores is only 0.06, and individual changes range from a 4-point gain to a 100-point loss. Permutation tests find no statistically detectable association between students' midterm and final scores, yet would detect association of the strength the alternative explanations predict more than 90 percent of the time. Once model complexity is accounted for, a model in which a student's midterm carries no information about that student's final describes the data as well as any model that permits an association. Neither result rules out a weak association, but together they show that the data do not require any. In simulated classes where each student's two scores remain linked through that student's own proficiency, as the alternative explanations imply, the two patterns almost never appear together; they appear together regularly only when that link is nearly severed. More than one mechanism could have severed it: midterm answers that were not the students' own work would have done so, and so would a final testing substantially different material, with no assistance involved. The scores cannot distinguish these possibilities or identify which students, if any, used unauthorized assistance. Our study offers a framework for evaluating statistical evidence in future disputes over generative AI use in academic assessment.