arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.07097cs.AIecon.GNq-fin.EC

验证过的,而非生成的:专家验证的AI学习材料与大学课程中学习收益的分布

Verified, not generated: expert-verified AI study materials and the distribution of learning gains in a university course

Canh Thien Dang, An Nguyen

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过双重差分实验发现,专家验证的AI学习材料将判断负担从学生转移至导师,显著提升低分学生成绩并缩小差距,强调评估应关注收益分布而非仅平均效应。

中文摘要 AI 辅助

关于生成式AI在教育中的实验研究大多报告平均效应,然而实地证据表明,AI可能缩小或扩大成就差距。我们认为,其方向取决于判断负担,即学习者在学习AI输出之前必须提供的专业知识以筛选其内容,而发布前的专家验证将这一负担从学生转移到负责任的导师身上。我们在一个两队列的双重差分设计中检验了这一论点,在该设计中,一门必修的一年级大学经济学课程的一半学生获得了由基于来源的模型生成并由一名具名研究生助教检查的AI生成播客、常见问题解答和基于测验的学习指南(170名学生;340份考试成绩)。访问与50分制组件中2.34分的优势相关。低于二等上界分类边界的分数比例相对于反事实下降了24.7个百分点,在23至31分的每个阈值上效应显著,在高于31分的阈值上均不显著,且平均效应的大约四分之三来自最低五分位数组。阈值估计对从干预前队列中移除得分最低的学生具有稳健性;平均效应则不然。来自36名学生的访谈和反馈表明,验证标签给了学生一个参与AI生成材料的理由,而不会结束他们对材料的审查。仅报告平均效应的AI学习资源评估无法检测出这些资源旨在帮助的学生是否真正受益。

英文摘要

Experimental studies of generative AI in education mostly report average effects, yet field evidence shows that AI can narrow attainment gaps or widen them. We argue that the direction depends on the judgement burden, the expertise a learner must supply to screen AI output before learning from it, and that expert verification before release moves this burden from students to an accountable tutor. We test the argument in a two-cohort difference-in-differences design in which one half of a compulsory firstyear university economics course received AI-generated podcasts, FAQs and quiz-based study guides, produced with a source-grounded model and checked by a named graduate teaching assistant (170 students; 340 examination marks). Access was associated with a 2.34-mark advantage on a 50-mark component. The share of marks below the upper-second classification boundary fell by 24.7 percentage points relative to the counterfactual, effects were significant at every threshold from 23 to 31 marks and at none above, and roughly three-quarters of the average originated in the bottom quintile. The threshold estimate is robust to removing the lowest-scoring students from the pre-intervention cohort; the average effect is not. Interviews and feedback from 36 students indicate that the verification label gave students a reason to engage with AI-generated material without ending their scrutiny of it. Evaluations of AI learning resources that report only mean effects cannot detect whether the students the resources are meant to help are the ones who gain.

发表机构

  • King’s Business School(国王商学院)
  • King’s College London(伦敦国王学院)

机构由 AI 辅助整理,请以论文原文为准。

↑