AI 中文总结
研究针对第二语言教育中AIED系统评估不足的问题,引入L2-Bench开源基准。通过 validated分类法、基于评分标准的评估方法和评估数据集,对大型模型能力进行评估,发现Claude Opus 4.7总体最佳,难任务性能降,为教育决策提供方法,推动AI评估科学成熟。
AI 中文摘要
尽管人工智能在教育中迅速得到应用,但对人工智能驱动的教育(AIED)系统的严格评估仍严重不足,尤其是在第二语言(L2)教育中,这是最常见但评估最少的人工智能应用之一。我们引入了L2-Bench,这是一个包含1000多个任务-响应对的开源基准,以帮助对与语言学习和评估相关的大型语言模型能力进行以教学法为主导的评估。关键的是,L2-Bench衡量的是模型在学习经验设计原则应用方面的性能,而不仅仅是对这些原则的了解或广泛的学习成果。我们的贡献包括:(1)一个经过验证的分类法,包含12项能力和31项子能力,由200多名专家从业者验证(任务真实性:4.42/5.00,标准充分性:4.18/5.00);(2)一种基于评分标准的评估方法,我们认为如果进行调整,可以推广到类似的(开放式、定性的)学科;(3)一个评估数据集,能在不同的L2教育场景中产生关于模型优势、劣势和情境稳健性的可靠信号。我们发现,在大型模型中,Claude Opus 4.7总体表现最佳(85.5%),不过在一些组成任务上略逊一筹。我们还发现,在更难的任务上性能显著下降(69.9%至73.4%)。L2-Bench为教育利益相关者提供了更好的方法,以便在实际应用、使用和管理AIED方面做出更明智的决策,同时推动教育领域人工智能评估科学的成熟。
英文摘要
Despite rapid AI adoption in education, rigorous evaluation of AI-powered educational (AIED) systems remains critically underdeveloped, particularly in second language (L2) education, one of the most common yet least evaluated AI applications. We introduce L2-Bench, an open-source benchmark of 1,000+ task-response pairs to aid the pedagogy-led evaluation of LLM capabilities relating to language learning and assessment. Crucially, L2-Bench measures model performativity on the application of learning experience design principles rather than mere knowledge of those principles or broad learning outcomes. Our contributions include: (1) a validated taxonomy of 12 competencies and 31 subcompetencies validated by 200+ expert practitioners (task authenticity: 4.42/5.00, criteria adequacy: 4.18/5.00); (2) a rubric-based evaluation methodology that we believe can, if adapted, generalize to similar (open-ended, qualitative) disciplines; (3) an evaluation dataset that produces reliable signal about model strengths, weaknesses, and contextual robustness across diverse L2 education scenarios. We find that, among large models, Claude Opus 4.7 performs best overall (85.5%), though is marginally outperformed on several constituent tasks. We also find that performance drops notably on harder tasks (69.9% to 73.4%). L2-Bench provides education stakeholders better methods to make more informed decisions about real-world AIED adoption, use, and governance, while advancing the maturing science of AI evaluations for education.