arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于自动评估人工智能导师的知识蒸馏

Knowledge Distillation for Automated AI Tutor Evaluation

Tahmid Al Hannan, Diego Garcia, Alex Njoroge, Suha Al Juboori, Tarek Sakakini

arXiv 2607.10647首次发表:更新:

发表机构

Folsom Lake College(福尔松湖学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

大语言模型融入教育但缺乏教学质量评估方法,研究引入FATE模型评估人工智能导师,利用前沿大语言模型的知识蒸馏生成额外监督提升评估性能,通过基准测试证明了FATE作为自动评估器的效用。

AI 中文摘要

大语言模型快速融入K-12和高等教育,却缺乏可靠的教学质量评估方法。研究社区开始探索人工智能导师自动评估领域,我们引入FATE(FLC人工智能导师评估器),这是一个专门用于评估人工智能导师的8B参数语言模型。该模型与BEA 2025共享任务的四个核心评估轨道一致,评估错误识别、错误定位、指导和可操作性方面的教学能力。由于教学评估数据有限,我们利用前沿大语言模型的知识蒸馏来生成额外监督,绝对性能提升高达22.63个百分点。最后,通过对流行商业模型生成的教学回复进行基准测试,证明了FATE作为自动评估器的效用。平均而言,Gemini 2.5 Flash表现最佳(82.被评估的人工智能导师的教学质量评估方法。核心方法是引入FATE模型,利用前沿大语言模型的知识蒸馏生成额外监督。主要贡献是提升了评估性能,并通过基准测试证明了FATE作为自动评估器的效用。

英文摘要

The rapid integration of Large Language Models (LLMs) into K-12 and higher education has outpaced the development of reliable methods for evaluating their pedagogical quality. As the research community starts to explore the space of automating evaluation of AI tutors, we introduce FATE (FLC AI Tutor Evaluator), a specialized 8B-parameter language model designed to evaluate AI tutors. Aligned with the four core evaluation tracks from the BEA 2025 Shared Task, our model assesses pedagogical ability across Mistake Identification, Mistake Location, Guidance, and Actionability. Because pedagogical evaluation is a specialized task with limited labeled data, we leverage knowledge distillation from a frontier LLM to generate additional supervision, yielding absolute performance gains up to 22.63 percentage points. Finally, we demonstrate FATE's utility as an automated evaluator by benchmarking instructional responses generated by popular commercial models, including ChatGPT, Claude, Gemini, and DeepSeek. On average, we have found that Gemini 2.5 Flash perfomed best (82.88%), then ChatGPT 5.5 Instant (80.75%), followed by DeepSeek V4 Flash (80.13%) and Claude Sonnet 4.6 (74.00%).

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑