arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12331cs.LG

模拟不参与学习的学生以评估基于LLM的导师

Simulating Disengaged Students to Evaluate LLM-based Tutors

Xianghui Meng, Jionghao Lin

首次发表
浏览论文内容

中文总结 AI 辅助

提出DAS2协议,模拟五种学习者参与状态以评估AI导师,验证了模拟有效性并揭示导师表现随状态变化。

中文摘要 AI 辅助

由计算模型生成的模拟学生为评估人类和AI导师使用的辅导策略和教学方法提供了一种实用途径。然而,此类模拟应考虑不参与行为,包括游戏系统、空转和离题行为,因为导师可能需要针对不同的学习者状态采取不同的回应。我们提出了不参与感知的学生模拟器(DAS2),这是一种可复现的部署前协议,模拟五种学习者参与状态:参与、游戏、空转、离题和混合,并评估AI导师在这些状态下的表现。使用ASSISTments09数据集,两名编码员根据匿名交互日志摘要独立标记了100个抽样辅导会话。他们达到了84%的一致性(Cohen's kappa = 0.78),在一致同意的案例中,人类共识标签与DAS2基于规则的标签在81%的案例中匹配(kappa = 0.75)。将模拟条件设定为预期的学习者状态,将模拟会话与真实会话之间的正确率差距从游戏状态的0.54降至0.20,从空转状态的0.51降至0.18。微调的Qwen2.5-7B更好地匹配了真实响应时间分布,而仅提示的GPT-4o生成了更可区分的学习者状态。对来自Claude、Llama、Gemini、Qwen和GPT家族的五个AI导师的评估显示,相对排名在学习者状态和交互长度上保持稳定,而绝对性能有所变化,揭示了导师支持中特定于状态的差异。人工验证进一步表明,自动化导师评估与人类判断并不完全一致。DAS2提供了一个部署前框架,用于在部署前评估AI导师如何响应多样化的学习者参与状态。

英文摘要

Simulated students generated by computational models provide a practical way to evaluate tutoring strategies and pedagogical approaches used by human and AI tutors. However, such simulations should account for disengaged behaviors, including gaming the system, wheel-spinning, and off-task behavior, because tutors may need different responses for different learner states. We present Disengagement-Aware Student Simulators (DAS2), a reproducible pre-deployment protocol that models five learner-engagement states: engaged, gaming, wheel-spinning, off-task, and mixed, and evaluates AI tutor performance across these states. Using ASSISTments09, two coders independently labeled 100 sampled tutoring sessions based on anonymized interaction-log summaries. They achieved 84% agreement (Cohen's kappa = 0.78), and among agreed cases, human consensus labels matched DAS2 rule-based labels in 81% of cases (kappa = 0.75). Conditioning simulations on intended learner states reduced the correctness-rate gap between simulated and authentic sessions from 0.54 to 0.20 for gaming and from 0.51 to 0.18 for wheel-spinning. Fine-tuned Qwen2.5-7B better matched authentic response-time distributions, while prompt-only GPT-4o generated more distinguishable learner states. Evaluation of five AI tutors from the Claude, Llama, Gemini, Qwen, and GPT families showed that relative rankings remained stable across learner states and interaction lengths, while absolute performance varied, revealing state-specific differences in tutor support. Human validation further showed that automated tutor evaluation does not fully align with human judgment. DAS2 provides a pre-deployment framework for evaluating how AI tutors respond to diverse learner-engagement states before deployment.

发表机构

  • The University of Hong Kong(香港大学)

机构由 AI 辅助整理,请以论文原文为准。

↑