arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20484cs.CLcs.AI

Edustories:来自课堂实践的真实案例研究集

Edustories: A Collection of Real-world Case Studies from Classroom Practices

Michal Štefánik, Jan Nehyba, Jirina Karasova, Martin Fico, Lucie Škarková, Markéta Košatková, David Kosatka

首次发表
浏览论文内容

中文总结 AI 辅助

针对集体课堂中AI辅助研究的空白,构建含1,492个真实课堂案例的Edustories数据集,评估LLM预测教学干预成效,发现最强模型58%准确率低于专家64%,凸显局限与潜力。

中文摘要 AI 辅助

尽管人工智能在教育领域的潜力已被广泛认可,但以往的大多数研究都侧重于对学生的个性化辅助。相比之下,全球范围内的大多数教育实践仍然发生在集体课堂环境中。为了帮助研究人员研究人工智能在集体教学中的辅助作用,我们引入了Edustories,这是一个包含1,492个由教师撰写的案例研究的数据集,描述了真实的中小学课堂情境,涉及具有挑战性的学生行为、教学干预及其结果。在许多其他应用中,Edustories能够评估大语言模型预测教师干预成功与否的能力,这对于为在职教师提供有用的反馈至关重要。我们将四个语言模型家族的最新模型与专家评估进行比较,发现当前模型在预测课堂结果方面不及人类专家;最强的模型达到了58%的准确率,而人类专家为64%。这一差距既凸显了人工智能作为在职教师助手的局限性,也展现了其新兴潜力。

英文摘要

Despite the widely recognized potential of AI in education, most prior work has focused on individualized student assistance. In contrast, the majority of educational practice worldwide still takes place in collective classroom settings. To enable researchers to study AI assistance in collective teaching, we introduce Edustories, a dataset of 1,492 teacher-written case studies describing real elementary and high-school classroom situations involving challenging student behavior, pedagogical interventions, and their outcomes. Among many other applications, Edustories enables evaluating LLMs' ability to predict the success of teacher interventions, crucial for providing practicing teachers with useful feedback. Comparing the latest models from four language-model families against expert assessments, we find that current models fall short of human expertise in predicting classroom outcomes; the strongest models reach 58% accuracy compared to 64% of human experts. This gap highlights both the limitations and the emerging potential of AI as assistants for practicing teachers.

发表机构

  • Masaryk University(马萨里克大学)
  • National Institute of Informatics(国立情报学研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑