arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EduPanel:用于教学视频的三智能体大语言模型评判器——可靠性、互补性和人类信任校准

EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration

Jia-Kai Dong, Yi-Cheng Lin, Hung-yi Lee

arXiv 2607.18529首次发表:更新:

发表机构

National Taiwan University; NTU Artificial Intelligence Center of Research Excellence(国立台湾大学; NTU人工智能卓越研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对教学视频教学质量评估需求,提出基于评分标准、以学习者为条件的EduPanel大语言模型评判器,通过专门智能体分解评估,经多项分析实现高可靠性,能辅助教育评估,提高评分准确性且让专家可检测不可靠输出。

AI 中文摘要

教学视频正成为教育的主要媒介,对其教学质量进行可扩展评估的需求日益增长。现有自动评判器无法完全满足这一需求,因为教学质量取决于多模态证据,且应针对目标学习者进行评估。我们提出了EduPanel,这是一种基于评分标准、以学习者为条件的大语言模型评判器,它通过专门智能体分解评估,为教学质量的不同方面生成可解释的评估。通过专家研究、架构消融和学习者角色分析,EduPanel实现了与人类专家中位数相当的可靠性。在专家评估中,其反馈提高了评分准确性(平均绝对误差从0.87降至0.73),同时专家仍能检测出不可靠输出(曲线下面积 = 0.77)而不是盲目接受。这些结果表明,EduPanel可以作为教育评估的有效助手,而非人类专家的替代品。

英文摘要

Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Existing automatic judges do not fully address this setting because teaching quality depends on multimodal evidence and should be evaluated with respect to the intended learner rather than as a universal property. We present EduPanel, a rubric-grounded, learner-conditioned LLM judge that decomposes evaluation across specialized agents to produce interpretable assessments for different aspects of teaching quality. Across expert studies, architecture ablations, and learner-persona analyses, EduPanel achieves reliability comparable to a median human expert. In expert evaluation, its feedback improves scoring accuracy (MAE 0.87 to 0.73), while experts remain able to detect unreliable outputs (AUC = 0.77) instead of accepting them blindly. These results suggest that EduPanel can serve as effective assistants for educational evaluation rather than replacements for human experts.

Comments19 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑