发表机构
Korea University(高丽大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出JevKT方法,利用系统一级LLM在几乎无学习者记录时执行知识追踪,在七个数据集上性能优于多数深度KT模型,成本仅为系统二级方法的约1/100。
AI 中文摘要
知识追踪(KT)模型需要大量已记录的学习者,因此新课程或平台启动时会出现无可用模型的情况。在基于LLM的KT中,LLM生成答案,我们称之为系统二级(System-Two);它要么在目标数据上进行微调,要么对十个样本进行推理和投票,这一过程速度较慢且给出的概率较为粗略。我们提出的问题是:现成的系统一级(System-One)LLM(可在单次推理中直接返回输入问题的概率),能否在几乎没有或没有学习者被记录时执行KT?在七个数据集上,Jev在未使用目标平台任何数据的情况下,达到了平均AUC值0.706,高于在8个学习者上训练的28个深度KT模型中的最佳表现(0.689),也高于系统二级Thinking-KM在所有七个数据集上的表现(0.650),且API成本仅为其约1/100。添加示例和来自已记录学习者的相似学习者统计量(JevKT)后,该数值提升至0.722;JevKT在最多16个学习者时显著领先于深度KT,在最多64个学习者时平均领先,而监督式KT在64至128个学习者之间赶上。在我们测试的阅读器中,这种增益是Jev特有的,因为通过官方系统一级适配器以完全相同的输入请求查询的其他三个LLM,在所有七个数据集上的表现均低于Jev;阅读器交换和污染检查未发现输入格式或记忆数据解释该增益的证据。对于新学习者,这种优势从他们的首次交互就存在,而在所有学习者都已记录的未见过的项目上,深度KT仍处于领先。
英文摘要
Knowledge tracing (KT) models need many logged learners, so a new course or platform starts without a usable model. In LLM-based KT the LLM generates the answer, which we call System-Two; it is either fine-tuned on the target data or reasons and votes over ten samples, which is slow and gives coarse probabilities. We ask whether an off-the-shelf System-One LLM, which returns a probability for a typed question directly in a single pass, can perform KT when few or no learners are logged. On seven datasets, Jev without any data from the target platform reaches a mean AUC of .706, above the best of 28 deep KT models trained on 8 learners (.689) and above System-Two Thinking-KT on all seven datasets (.650) at about 1/100 of its API cost. Adding examples and a similar-learner statistic from the logged learners (JevKT) raises this to .722; JevKT stays significantly ahead of deep KT up to 16 learners and ahead on average up to 64, and supervised KT catches up between 64 and 128 learners. Among the readers we tested, the gain is specific to Jev, since three other LLMs queried with the byte-identical typed request through the official System-One adapter fall below it on all seven datasets, and reader swaps and contamination checks find no evidence that the input format or memorised data explain the gain. For new learners the advantage holds from their first interactions, whereas on unseen items with all learners logged, deep KT remains ahead.
Comments41 pages, 7 figures. Code and result summaries: https://anonymous.4open.science/r/jevkt-coldstart