发表机构
KAIST(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究开发五级支架水平量表,分析大学AI课程中203名学生的14637个LLM辅导响应,发现超95%为高水平帮助,支架水平与学生对话行为相关但对考试表现预测性有限,为LLM辅导提供实证基线与测量框架。
AI 中文摘要
学生越来越多地将LLM用作课程作业和问题解决的辅导工具,但目前人们对LLM在真实学习交互中作为辅导者提供的帮助程度知之甚少,这一点很重要,因为辅导响应在帮助学生完成任务的直接程度上存在很大差异。我们将这一维度操作化为支架水平,并开发了一个经过人工标注验证的五级量表,该量表根据响应提供的直接帮助程度对其进行表征。我们将该量表应用于大学AI课程中203名学生的14637个LLM响应,这些响应绝大多数集中在高水平帮助上,超过95%被归类为“解释”或“解决”。支架水平与学生后续的对话行为存在系统性关联,但除了先前的成绩和对话行为外,它对三次后续考试的表现几乎没有额外的预测信息。这些发现为辅导交互中的LLM帮助提供了实证基线,并为评估不同辅导设计如何改变这种帮助提供了测量框架。
英文摘要
Students increasingly use LLMs as tutors for coursework and problem solving. Little is known about the level of assistance LLMs provide when students use them as tutors in authentic learning interactions. This matters because tutoring responses can differ substantially in how directly they help students complete a task. We operationalize this dimension as scaffolding level and develop a five-level scale, validated against human annotations, that characterizes responses according to the degree of direct assistance they provide. We apply the scale to 14,637 LLM responses from 203 students in a university AI course. Responses are overwhelmingly concentrated at high levels of assistance, with more than 95% classified as either Explaining or Solving. Scaffolding level is systematically associated with students' subsequent conversational behavior, but provides little additional predictive information about performance on three subsequent exams beyond prior achievement and dialogue behavior. These findings provide an empirical baseline for LLM assistance in tutoring interactions and a measurement framework for evaluating how alternative tutoring designs change that assistance.