大型大学中生成式人工智能的可用性、成绩与学生满意度
Generative AI Availability, Grades, and Student Satisfaction at a Large University
浏览论文内容
中文总结 AI 辅助
研究大型大学中GenAI对学生成绩和满意度的影响,通过教学大纲和行政数据,用差异中的差异设计及人工验证的大语言模型管道衡量课程GenAI易感性,发现GenAI对成绩和自我报告理解无显著差异影响,仅在特定假设下对兴趣有显著影响,缓解相关担忧。
中文摘要 AI 辅助
生成式人工智能(GenAI)在高等教育中的传播引发了人们的担忧,即学生将认知努力转移到人工智能上,不学习却获得高分。若“GenAI替代假说”成立,在GenAI易受影响的课程(更多依赖课后作业和论文等评估而非课堂考试的课程)中成绩应不成比例地提高。替代也可能影响学生满意度。我们使用美国一所大型大学(2015 - 2025年;156,135名学生;87,936门课程)的教学大纲和行政数据来测试替代假说。通过人工验证的大语言模型管道从教学大纲中提取评估类型来衡量课程的GenAI易感性,采用差异中的差异设计比较ChatGPT发布前后课程的结果,并将新冠疫情影响建模为持续或短暂的。我们发现GenAI的可用性对总体成绩或先前成绩较低的学生没有显著差异影响。对自我报告理解的影响同样不显著;对兴趣的影响仅在假设短暂疫情影响时显著。我们的发现缓解了对GenAI提高成绩和降低学生满意度的担忧。
英文摘要
The spread of generative AI (GenAI) in higher education has raised concerns that students offload cognitive effort to AI, earning high grades without learning. If this "GenAI substitution hypothesis" is true, grades should rise disproportionately in GenAI-susceptible courses--those relying more on assessments like take-home problem sets and essays rather than in-class exams. Substitution could also affect student satisfaction, measured here as self-reported understanding and interest in the subject, which prior research links to assessments. We test the substitution hypothesis using syllabus and administrative data from a large U.S. university (2016-2025; 138,386 students; 72,730 course offerings). We measure courses' GenAI susceptibility using a human-validated LLM pipeline to extract assessment types from syllabi, and use a differences-in-differences design comparing outcomes across courses before and after ChatGPT's release, while modeling COVID-19 pandemic effects as either persistent or transient. We find no significant differential effect of GenAI availability on grades overall or among previously lower-performing students. Effects on self-reported understanding are likewise insignificant; effects on interest are significant only assuming transient pandemic effects. Our findings temper concerns that GenAI inflates grades and reduces students' satisfaction.