首届“教学怪兽挑战赛”的研究发现:AI智能体的教学内容知识基准
Findings of the First Teaching Monster Challenge: A Benchmark of Pedagogical Content Knowledge in AI Agents
浏览论文内容
中文总结 AI 辅助
研究推出首个以学习者角色为评估标准的教学视频生成基准“教学怪兽挑战赛”,测试发现当前AI教学系统内容适配能力不足,自动裁判存在局限,发布相关资源供后续研究。
中文摘要 AI 辅助
当前AI智能体已能够解决问题、像学科专家一样作答、生成多模态长文本内容,但它们能否调整课程以适配特定学习者(教育领域称之为教学内容知识,即Pedagogical Content Knowledge,简称PCK)尚未有基准测试。为了衡量这一能力,我们推出了“教学怪兽挑战赛”(Teaching Monster Challenge),这是首个将学习者角色作为明确评估标准的教学视频生成基准。每个系统会被给定一个主题和一个学习者角色,必须生成完整的教学视频。所有视频由大语言模型(LLM)裁判筛选,通过群体两两投票排名,最终由专家小组确定。首届赛事显示,当前系统能较好处理教学内容,但在内容呈现和适配学习者方面弱得多。同一过程也暴露出自动裁判的局限:LLM裁判能区分表现最差的尾部系统,但对最强系统的排名不佳;最强系统从裁判处获得的分数几乎相同,导致其排名与人类偏好不符。因此,要取得进展不仅需要更好的教学系统,还需要更好的自动裁判,我们将该基准、评分标准及人类判断结果作为测试平台发布,供各方研究使用。
英文摘要
AI agents can now solve problems, answer like subject experts, and generate long-form multimodal content. However, whether they can adapt a lesson to fit a specified learner, which education calls Pedagogical Content Knowledge (PCK), has not been benchmarked. To measure it, we introduce the Teaching Monster Challenge, the first instructional video generation benchmark to treat the learner persona as an explicit evaluation criterion. Each system is given a topic and a learner persona and must generate a complete instructional video. Every video is screened by an LLM-judge, ranked by crowd pairwise voting, and finalized by an expert panel. The first edition shows that today's systems handle the content well but are far weaker at presenting it and adapting it to the learner. The same process exposes a limit of automatic judging. The LLM-judge separates a clear low-performing tail but ranks the strongest systems poorly. The strongest systems receive nearly identical scores from the judge, so its ranking of them does not match human preference. Progress therefore requires not only better teaching systems but also better automatic judges, and we release the benchmark, rubric, and human judgments as a testbed for both.