发表机构
Drexel University; Nick Howley College of Engineering and Computing; School of Computer and Information Sciences(德雷塞尔大学; 尼克·豪利工程与计算学院; 计算机与信息科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有LLM学生模拟难以区分高低掌握水平的局限,提出基于SSKG的方法,通过抽样三元组掌握概率确定答题结果并生成理由,有效区分了不同掌握水平的学生。
AI 中文摘要
大型语言模型(LLM)正越来越多地用于模拟不同掌握水平的学生,这类模拟可生成合成训练数据并对辅导系统进行压力测试。然而,常见的基于提示的方法将答案决策交给LLM,即便被要求模拟低掌握水平的学生,LLM也倾向于依据其内置能力运行,导致这些方法难以区分低掌握与高掌握水平的学生。我们使用379个经大学理事会校准的SAT代数题目和5种典型掌握轮廓证明了这一局限,来自三家厂商的三款LLM(Gemini 3.1 Flash Lite、Claude Haiku 4.5和GPT-5.4-mini)在所有轮廓上的准确率达96.8%-100%。为解决该局限,我们提出一种基于随机学生知识图谱(SSKG)的方法:从一本公开代数教科书中提取课程知识图谱(CKG),将每道SAT题目解答分解为所需三元组的链条,SSKG为每个三元组分配掌握概率,通过抽样确定答题正确性,再由LLM生成与结果一致的第一人称理由。该模拟将所有轮廓上的准确率降至44.1%-85.2%,并产生清晰的单调掌握梯度。
英文摘要
Large language models (LLMs) are increasingly used to simulate students at different mastery levels. These simulations can generate synthetic training data and stress-test tutoring systems. However, common prompt-based approaches leave the answer decision to the LLM, which tends to perform according to its built-in capabilities even when instructed to simulate a student with low mastery. As a result, these approaches may have difficulty distinguishing students with low and high levels of mastery. We demonstrate this limitation using 379 College Board-calibrated SAT Algebra items and five archetypal mastery profiles. Three LLMs from three vendors (Gemini 3.1 Flash Lite, Claude Haiku 4.5, and GPT-5.4-mini) achieve 96.8-100% accuracy across all profiles. To address this limitation, we introduce a method grounded in a Stochastic Student Knowledge Graph (SSKG). A curriculum knowledge graph (CKG) is extracted from an open algebra textbook, and each SAT solution is decomposed into a chain of required triples. The SSKG assigns a mastery probability to each triple, which is sampled to determine question correctness. An LLM then generates a first-person rationale consistent with the outcome. The simulation reduces accuracy to 44.1-85.2% across profiles and produces a clear monotone mastery gradient.