arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LessonBench-V1:用于评估人工智能课程生成代理的基准数据集

LessonBench-V1: A Benchmark Dataset for Evaluating AI Lesson Generation Agents

Ravidu Suien Rammuni Silva, Ahmad Lotfi, Isibor Kennedy Ihianle, Golnaz Shahtahmassebi, Jordan J. Bird

arXiv 2607.13041首次发表:更新:

发表机构

Department of Computer Science, Nottingham Trent University, Nottingham, UK; Department of Physics and Mathematics, Nottingham Trent University, Nottingham, UK(计算机科学系,诺丁汉特伦特大学,诺丁汉,英国; 物理与数学系,诺丁汉特伦特大学,诺丁汉,英国)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对无标准化基准评估人工智能课程生成系统的问题,引入LessonBench-V1基准数据集,其含人工编写课程与逆向工程课程计划,经人工审核并综合多种教学方法生成,可系统评估相关代理,还提出三维评估管道。

AI 中文摘要

基于大语言模型的人工智能教育内容生成系统不断发展,但尚无标准化基准来系统评估它们。本研究介绍了LessonBench-V1,这是一个基准数据集,包含647个由人工编写的课程以及基于大语言模型逆向工程的课程计划,涵盖240个STEM主题。课程来自97个可靠开源资源。课程计划经人工审核,通过综合多种教学方法生成,包含3620个学习目标及教学元数据,可系统评估课程生成人工智能代理并支持进一步研究,还提出了三维评估管道。

英文摘要

Large Language Model (LLM) based AI educational content generation systems are increasingly being developed, yet no standardised benchmark exists to systematically evaluate them. This study introduces LessonBench-V1, a benchmark dataset comprising 647 human-written lessons paired with LLM-based reverse-engineered lesson plans across 240 STEM topics spanning mathematics, physics, chemistry, and computer science. The lessons are drawn from 97 trusted open sources, including LibreTexts, Brilliant.org and GeeksForGeeks. Each lesson plan is human-reviewed and produced through a pedagogically grounded methodology that synthesises Bloom's Taxonomy, Gagné's Events, Merrill's First Principles, and the 5E Instructional Model. The lesson plans capture 3,620 learning objectives with pedagogical metadata, enabling systematic, reproducible evaluation of lesson-generation AI agents and supporting further research. The study further proposes a three-dimensional evaluation pipeline for use with the dataset.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑