考察引导式AI导师如何解决学生困境的差异
Examining Variation in How Guided AI Tutors Resolve Student Impasses
浏览论文内容
中文总结 AI 辅助
本研究分析引导式AI导师在化学辅导中的困境处理差异,发现困境深度增加会降低恢复几率,提问效果随困境持续衰减,而直接解决错误更有效,为实时学习分析提供轮次级信号。
中文摘要 AI 辅助
当学生遇到困难时,导师面临援助困境:过早提供帮助可能阻碍有成效的挣扎,而迟迟不提供帮助则会使学生陷入令人沮丧的持续困境(即空转)。生成式AI导师越来越多地使用限制直接给出答案的护栏,但对于此类导师在困境持续时的行为表现,目前知之甚少。我们分析了来自1,260个真实会话的20,462个学生轮次,这些会话使用一个引导式LLM化学导师,识别出6,630个困境轮次,分为三种主要类型:概念错误、表达不确定或寻求帮助。随后,我们利用这些困境模拟了三种辅导条件,以研究AI导师在困境中的指导差异:基线、不直接回答和引导式导师。对于150个困境的样本,提示特异性改变了教学法:基线导师在50.7%的回应中直接提供答案,不直接回答的导师每次都提出后续问题,而引导式导师则根据情境以多种方式回应。接着,我们分析了真实互动中的困境轨迹,发现每增加一个困境轮次,下一轮恢复的几率降低12.7%(AOR = 0.873,p < .001),早期退出者被困在递归概念引导中,未能进入执行阶段。提问的益处随着困境持续而衰减(脚本化问题×深度 AOR = 0.78;后续问题×深度 AOR = 0.83),而解决学生错误则变得更有益(AOR = 1.14);在脚本化问题失败后,重复该问题后恢复的案例占28.1%,而导师转而解决错误时恢复的案例占39.8%。对于学习分析,这些发现将困境深度和类型识别为可观察的、轮次级别的对话信号,分析系统可利用这些信号实时触发渐进式、状态敏感的援助。
英文摘要
When a student is stuck, a tutor faces the assistance dilemma: help given too early can hinder productive struggle, while help withheld too long leaves the student in a frustrating, persistent impasse (i.e., wheel spinning). Generative AI tutors increasingly use guardrails restricting answer-giving, yet little is known about how such tutors behave once an impasse persists. We analyze 20,462 student turns from 1,260 authentic sessions with a guided LLM chemistry tutor, identifying 6,630 impasse turns of three major types: conceptual errors, expressed uncertainty, or help-seeking. We then used these impasses to simulate three tutoring conditions to study variation in AI tutor guidance through impasses: baseline, no-direct-answer, and guided tutor. For a sample of 150 impasses, prompt specificity changed pedagogy: a baseline tutor provided the answer directly in 50.7% of responses, a no-direct-answer tutor asked a follow-up question every time, and the guided tutor responded in a wide variety of ways depending on the context. We then analyzed impasse trajectories in authentic interactions, finding that each additional impasse turn lowered the odds of next-turn recovery by 12.7% (AOR = 0.873, p < .001), and early dropouts were caught in recursive concept elicitation before reaching execution. The benefit of questioning decayed as impasses persisted (scripted question x depth AOR = 0.78; follow-up x depth AOR = 0.83), whereas addressing the student's error grew more beneficial (AOR = 1.14); after a failed scripted question, repeating it was followed by recovery in 28.1% of cases, compared with 39.8% when the tutor addressed the error instead. For learning analytics, these findings identify impasse depth and type as observable, turn-level dialogue signals that analytics can use to trigger graduated, state-sensitive assistance in real time.