arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05411cs.AI

基于教学适宜性指数评估和改进基于大语言模型(LLM)的AI导师的教学适配性

Evaluating and Improving Pedagogical Fit in LLM-Based AI Tutors with the Pedagogical Suitability Index

Benjamin Barlog, Hudson Craig, Zedong Peng

AI总结:

该研究提出教学适宜性指数(PSI)评估LLM导师的教学适配性,评估4类LLM并发现PSI引导的反馈可显著提升弱表现案例的适配性。

AI中文摘要:

大语言模型(LLM)正越来越多地被用作AI导师,但正确的答案并不总是具有教学适宜性的答案。在课堂学习中,有效的帮助不仅取决于答案的正确性,还取决于响应是否匹配学习者当前的基础、课程的先后顺序以及概念引入的时机。现有评估主要聚焦于答案质量,导致这种教学适配性未得到充分衡量。我们提出了教学适宜性指数(Pedagogical Suitability Index, PSI),这是一个由6个基于理论的子分数构成的综合指标,用于评估LLM生成的辅导响应与学习者准备程度和课程进度的匹配程度,我们还进一步将PSI用作响应改进的结构化反馈信号。我们使用配对的标准提示和有缺陷的提示,在240个基于场景的评估中评估了4个LLM导师(ChatGPT、Gemini、Gemma4和Qwen3),随后对62个表现较弱的案例应用了PSI引导的再生协议。所测试的4个模型的基线差异总体上较小(PSI范围为0.557至0.638),开源模型和闭源模型在教学适配性方面未表现出明显的区分。在测试的提示扰动下,总体PSI基本保持稳定(差值为-0.002),但出现了子分数间的权衡。更重要的是,PSI引导的反馈显著改善了表现较弱的案例:62个案例中有51个得到了改善(占比82.3%)。对PSI选定的62个弱案例进行的针对性人工评估提供了初步证据,表明所识别的弱点具有教学意义,且许多PSI引导的再生对应于人类判断的改进。这些结果表明,对有效辅导而言,对学习者和课程的适配可能比模型类别本身更重要,且这种适配既可以衡量也可以改进。

英文摘要:

Large language models (LLMs) are increasingly used as AI tutors, but a correct answer is not always a pedagogically appropriate one. In classroom learning, effective help depends not only on correctness, but also on whether a response matches the learner's current foundation, the course sequence, and the timing of concept introduction. Existing evaluations focus mainly on answer quality, leaving this instructional fit under-measured. We present the Pedagogical Suitability Index (PSI), a composite metric of six theory-informed sub-scores that evaluates how well LLM-generated tutoring responses align with learner readiness and curricular progression, and we further use PSI as a structured feedback signal for response improvement. We evaluate four LLM tutors (ChatGPT, Gemini, Gemma4, and Qwen3) across 240 scenario-based evaluations using paired standard and defective prompts, then apply a PSI-guided regeneration protocol to 62 weak-performing cases. Baseline differences across the four tested models were modest overall (PSI range: 0.557 to 0.638), and open-weight and closed models did not exhibit a clear separation in pedagogical fit. Under the tested prompt perturbations, overall PSI remained largely stable (Delta = -0.002), though sub-score trade-offs emerged. More importantly, PSI-guided feedback substantially improved weak-performing cases: 51 of 62 cases improved (82.3%). Focused manual evaluation of the 62 PSI-selected weak cases provides initial evidence that the identified weaknesses are instructionally meaningful and that many PSI-guided regenerations correspond to human-judged improvement. These results suggest that learner- and curriculum-aware alignment may matter more for effective tutoring than model category alone, and that such alignment is both measurable and improvable.

补充信息

↑