衡量基础模型的专业教育教学能力
Measuring the Professional Educational Competence of Foundation Models
浏览论文内容
中文总结 AI 辅助
本文提出EDU 1.0基准,基于美中印教师入职考试评估基础模型的教育能力,发现模型在学科教学法上普遍弱于一般教学法,强调教学内容知识是关键短板。
中文摘要 AI 辅助
基础模型在人口规模上进行辅导、评估和教学,因此需要外部定义的教育能力衡量标准。现有基准侧重于困难的学术问题或孤立的合成教育任务,而非真实的教师入职标准。我们引入了EDU 1.0(基础模型教育尽职调查),该基准使用教师入职评估作为教育能力的代理指标,而非替代人类资格的替代品。EDU 1.0包含来自美国、中国和印度的教师认证和招聘考试中的10,012道题目,包括美国Praxis系列、中国国家教师资格考试和印度Kendriya Vidyalaya Sangathan考试。它涵盖了基础素养和知识、教学原则,以及语言艺术、数学、科学、社会科学和教育实践等学科特定的教学专长。在36个基础模型变体中,最强的专有模型达到了96.5%的响应平衡分数;领先的开源权重模型落后2.1个百分点,而可在单个加速器上部署的领先系统达到了92.2%。这些汇总数据掩盖了一个共同的局限性:所有三个系统在一般教学原则上的得分均高于要求在教学法应用于特定学科内的评估。它们的学科评估分数跨度在4.1至8.8个百分点之间,且随着能力下降,它们相对于一般教学法的差距从2.6个百分点扩大到4.1个百分点。因此,突出的需求是教学内容知识,即使特定学科内容对特定学习者可教的能力,而非一般教学法或模型规模。
英文摘要
Foundation models tutor, assess, and instruct at population scale, requiring externally defined measures of educational competence. Existing benchmarks emphasize difficult academic problems or isolated synthetic educational tasks rather than authentic teacher-entry standards. We introduce EDU 1.0 (Educational Due Diligence for Foundation Models), which uses teacher-entry assessments as proxies for educational competence rather than substitutes for human qualification. EDU 1.0 comprises 10,012 questions from teacher certification and recruitment examinations in the United States, China, and India, including the U.S. Praxis series, China's National Teacher Qualification Examination, and India's Kendriya Vidyalaya Sangathan examinations. It covers foundational literacy and knowledge, pedagogical principles, and subject-specific pedagogical expertise across language arts, mathematics, science, social science, and education practice. Across 36 foundation-model variants, the strongest proprietary model reaches a response-balanced score of 96.5%; the leading open-weight model trails by 2.1 percentage points, while the leading system deployable on a single accelerator reaches 92.2%. These aggregates conceal a shared limitation: all three systems score higher on general pedagogical principles than on assessments requiring pedagogy to be applied within a discipline. Their subject-assessment scores span 4.1-8.8 points, and their shortfall relative to general pedagogy widens from 2.6 to 4.1 points as capability declines. The outstanding requirement is therefore pedagogical content knowledge, the capacity to make particular subject matter teachable to particular learners, rather than general pedagogy or model scale.
发表机构
- Shanghai Innovation Institute(上海创新研究院)
- East China Normal University(华东师范大学)
- Faculty of Education at East China Normal University (ECNU)(华东师范大学教育学院)
- Shanghai Institute of AI for Education, East China Normal University (ECNU)(华东师范大学上海教育人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。