arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SLATE:AI生成的幻灯片是否具有教育有效性?一个用于语言教学质量与学习者知识获取的基准

SLATE: Are AI-Generated Slides Educationally Effective? A Benchmark for Language Teaching Quality and Learner Knowledge Acquisition

Jingzhuo Wu, Jiajun Zhang, Liu Yi, Leqi Zheng, Yuheng Jing, Xinyuan Zhou, Quan yang

arXiv 2609.06212首次发表:更新:

发表机构

Beijing Normal University; University of Science and Technology of China; Institute of Automation, Chinese Academy of Sciences; Beijing Language and Culture University; Tsinghua University(北京师范大学; 中国科学技术大学; 中国科学院自动化研究所; 北京语言大学; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SLATE基准通过前测-后测设计评估AI生成语言教学幻灯片,发现教学设计比内容有效性更影响学习收益,且前沿模型可能产生负学习收益,呼吁转变评估范式。

AI 中文摘要

大语言模型在生成语言教学幻灯片方面已展现出卓越的能力。然而,视觉美观与实际教学效果之间仍存在关键的不匹配。为解决这一差距,我们引入了SLATE(基于幻灯片的学习评估以衡量教学有效性),这是首个通过教学有效性和学习者知识获取来评估AI生成语言教学幻灯片的基准。SLATE将来自低资源语言且网络存在感极低的语言学奥林匹克谜题转化为90个标准化的教学单元,包含1,133个可评估项目,并配以结构化的课程大纲以及匹配的近迁移和远迁移测试集。这种前测-后测设计消除了预训练知识泄漏,确保成绩提升反映的是学习而非先前回忆。使用视觉语言模型作为可扩展的学习者代理,并辅以三系统人工试点提供方向性支持,我们的结果表明,内容有效性与学习收益呈弱相关,而教学设计则呈强正相关。此外,大多数系统在近迁移和远迁移准确性之间存在显著差距,甚至前沿模型也可能产生负学习收益。SLATE揭示了工件质量与教学有效性之间的分离,呼吁在生成式教学系统的构建、评估和部署方式上实现范式转变。

英文摘要

LLMs have achieved remarkable capabilities in generating language teaching slides. However, a critical mismatch persists between visual polish and actual instructional effectiveness. To address this gap, we introduce SLATE (Slide-based Learning Assessment for Teaching Effectiveness), the first benchmark that evaluates AI-generated language teaching slides through instructional effectiveness and learner knowledge acquisition. SLATE transforms linguistics olympiad puzzles from low-resource languages with negligible web presence into 90 standardized instructional units comprising 1,133 assessable items, paired with a structured course outline and matched near- and far-transfer test sets. This pretest-posttest design eliminates pretrained knowledge leakage, ensuring gains reflect learning rather than prior recall. Using VLMs as scalable learner proxies and directionally supported by a three-system human pilot, our results show that content validity exhibits a weak association with learning gain, while pedagogical design exhibits a robust positive association. Moreover, most systems show a significant gap between near- and far-transfer accuracy, and even frontier models can produce negative learning gains. SLATE reveals a dissociation between artifact quality and instructional effectiveness, calling for a paradigm shift in how generative teaching systems are built, evaluated, and deployed.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑