AI 中文总结
本研究探讨K-12数学辅导对话中LLM诊断学生失败模式的有效性,发现跨模型一致性高于人机一致性,警示共识不能替代有效性验证。
AI 中文摘要
在K-12数学辅导中,学生与辅导者之间的对话为学习者的解题过程和困难来源提供了丰富的证据。学习分析研究越来越依赖大型语言模型(LLMs)从对话中提取此类信息,用于各种下游任务,包括知识追踪、行为建模和学生推理错误的诊断。然而,这些模型生成的解释的有效性仍未得到充分理解。在这项探索性研究中,我们使用一个操作性诊断编码手册,考察了LLM对数学辅导对话中五种学生失败模式分类的有效性:不确定性、错误归因、运算符选择、概念缺口和程序性失误。跨模型来看,人机一致性为中等水平(kappa = .524-.597),而跨模型一致性显著更高(kappa = .755-.781;alpha = .769)。这些发现表明,跨模型一致性可能造成正确性的误导性表象,挑战了LLM之间的共识构成有效学习者解释证据的假设。对于学习分析而言,其含义是明确的:只有推断的构念有效时,可扩展的标注才有用,而模型共识不能替代有效性的独立证据。
英文摘要
In K-12 mathematics tutoring, student-tutor dialogue provides rich evidence of learners' problem-solving processes and sources of difficulty. Learning analytics research increasingly relies on large language models (LLMs) to extract such information from dialogue for a variety of downstream tasks, including knowledge tracing, behavioral modeling, and diagnosis of student reasoning errors. However, the validity of these model-generated interpretations remains insufficiently understood. In this exploratory study, we examine the validity of LLM classifications of five student failure modes in mathematics tutoring dialogue using an operational diagnostic codebook: uncertainty, misattribution, operator selection, conceptual gap, and procedural slip. Across models, human-LLM agreement was moderate (kappa = .524-.597), while cross-model agreement was substantially higher (kappa = .755-.781; alpha = .769). These findings show that cross-model agreement can create a misleading appearance of correctness, challenging the assumption that consensus among LLMs constitutes evidence of valid learner interpretation. For learning analytics, the implication is clear: scalable labeling is useful only if the inferred constructs are valid, and model consensus cannot substitute for independent evidence of that validity.
CommentsSubmitted to LAK27 as a short paper. Currently under review