发表机构
Beijing Normal University; University of Science and Technology of China; Institute of Automation, Chinese Academy of Sciences; Beijing Language and Culture University; Tsinghua University(北京师范大学; 中国科学技术大学; 中国科学院自动化研究所; 北京语言大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SLATE基准通过前测-后测设计评估AI生成语言教学幻灯片,发现教学设计比内容有效性更影响学习收益,且前沿模型可能产生负学习收益,呼吁转变评估范式。
AI 中文摘要
大语言模型在生成语言教学幻灯片方面已展现出卓越的能力。然而,视觉美观与实际教学效果之间仍存在关键的不匹配。为解决这一差距,我们引入了SLATE(基于幻灯片的学习评估以衡量教学有效性),这是首个通过教学有效性和学习者知识获取来评估AI生成语言教学幻灯片的基准。SLATE将来自低资源语言且网络存在感极低的语言学奥林匹克谜题转化为90个标准化的教学单元,包含1,133个可评估项目,并配以结构化的课程大纲以及匹配的近迁移和远迁移测试集。这种前测-后测设计消除了预训练知识泄漏,确保成绩提升反映的是学习而非先前回忆。使用视觉语言模型作为可扩展的学习者代理,并辅以三系统人工试点提供方向性支持,我们的结果表明,内容有效性与学习收益呈弱相关,而教学设计则呈强正相关。此外,大多数系统在近迁移和远迁移准确性之间存在显著差距,甚至前沿模型也可能产生负学习收益。SLATE揭示了工件质量与教学有效性之间的分离,呼吁在生成式教学系统的构建、评估和部署方式上实现范式转变。
英文摘要
LLMs have achieved remarkable capabilities in generating language teaching slides. However, a critical mismatch persists between visual polish and actual instructional effectiveness. To address this gap, we introduce SLATE (Slide-based Learning Assessment for Teaching Effectiveness), the first benchmark that evaluates AI-generated language teaching slides through instructional effectiveness and learner knowledge acquisition. SLATE transforms linguistics olympiad puzzles from low-resource languages with negligible web presence into 90 standardized instructional units comprising 1,133 assessable items, paired with a structured course outline and matched near- and far-transfer test sets. This pretest-posttest design eliminates pretrained knowledge leakage, ensuring gains reflect learning rather than prior recall. Using VLMs as scalable learner proxies and directionally supported by a three-system human pilot, our results show that content validity exhibits a weak association with learning gain, while pedagogical design exhibits a robust positive association. Moreover, most systems show a significant gap between near- and far-transfer accuracy, and even frontier models can produce negative learning gains. SLATE reveals a dissociation between artifact quality and instructional effectiveness, calling for a paradigm shift in how generative teaching systems are built, evaluated, and deployed.