arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14789cs.HC

以创新速度评估AI辅导:从业者主导的GCSE科学AI辅导平台微随机试验

Evaluating AI Tutoring at the Speed of Innovation: Practitioner-Led Micro-Randomised Trials of an AI Tutoring Platform in GCSE Science

Wayne Harrison, Rahil Khowaja, Emma Dobson, Germaine Uwimpuhwe, Steve Higgins

首次发表
浏览论文内容

中文总结 AI 辅助

针对教育AI评估滞后于技术发展的问题,本文通过教师主导的微随机试验评估AI辅导平台Medly,结果显示其提升GCSE成绩,提出快速累积评估架构。

中文摘要 AI 辅助

教育领域的人工智能(AI)系统正以与传统评估方式不相适应的速度发展。当一项大规模试验完成设计、实施、分析和发表时,所研究的技术可能已发生实质性变化。这给循证教育带来了时间上的难题:对及时证据的需求可能促使人们依赖薄弱的观察性或使用数据,而传统严格评估产生的证据可能过于缓慢,难以指导快速演变的实践。我们研究了教师主导的微随机对照试验(micro-RCTs)作为应对此问题的一种方案。实证案例是对AI辅导平台Medly在英格兰中学GCSE生物、化学和物理学科中进行的为期四周、多地点、个体随机评估。在929名完成基线评估的学生中,644名完成了后测。在主要意向性治疗(ITT)分析中,被分配到Medly的学生比进行常规自主复习的学生取得了更高的后测成绩(Hedges' g = 0.33,95%置信区间0.18至0.48)。在物理(g = 0.31)、化学(g = 0.32)和生物(g = 0.52)中均观察到正向估计,且没有证据表明因弱势状态而产生不同影响。更高的平台参与度与更高的成绩相关,但这些随机化后分析被视为探索性而非因果性。流失率相当高(30.7%),结果测量采用与课程对齐而非标准化方式,过程评估响应有限。因此,我们将这些发现视为初步结果。我们认为,微随机试验对教育AI的价值不在于用小规模研究取代确定性评估,而在于构建一种快速、累积的评估架构,在该架构中,随机化估计可以随着技术及其实施的发展而被生成、复制和更新。

英文摘要

Artificial intelligence (AI) systems in education are developing on timescales that sit uneasily with conventional evaluation. By the time a large-scale trial has been designed, delivered, analysed and published, the technology under study may have changed materially. This creates a temporal problem for evidence-informed education: the need for timely evidence can encourage reliance on weak observational or usage data, while conventional rigorous evaluation may produce evidence too slowly to guide rapidly evolving practice. We examine teacher-led micro-randomised controlled trials (micro-RCTs) as one response to this problem. The empirical case is a four-week multisite individually randomised evaluation of Medly, an AI-powered tutoring platform, in GCSE Biology, Chemistry and Physics in English secondary schools. Of 929 students completing baseline assessment, 644 completed post-testing. In the primary ITT analysis, students allocated to Medly achieved higher post-test attainment than students undertaking business-as-usual self-directed revision (Hedges' g = 0.33, 95% CI 0.18 to 0.48). Positive estimates were observed in Physics (g = 0.31), Chemistry (g = 0.32) and Biology (g = 0.52), with no evidence of differential impact by disadvantage status. Greater platform engagement was associated with higher attainment, but these post-randomisation analyses are treated as exploratory rather than causal. Attrition was substantial (30.7%), outcome measures were curriculum-aligned rather than standardised, and process evaluation response was limited. We therefore interpret the findings as preliminary. We argue that the value of micro-RCTs for educational AI lies not in replacing definitive evaluation with small studies, but in enabling a rapid, cumulative evaluation architecture in which randomised estimates can be generated, replicated and updated as technologies and their implementation evolve.

发表机构

  • What Worked Education
  • Durham University(杜伦大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑