arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一个掌握阈值并不适用于所有知识追踪模型

One Mastery Threshold Does Not Fit All Knowledge Tracing Models

Xianghui Meng, Yujing Zhang, Jionghao Lin

arXiv 2610.00095首次发表:更新:

发表机构

The University of Hong Kong(香港大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探讨不同知识追踪模型下掌握阈值设定的差异,发现同一阈值在不同模型间决策差异大,需根据模型和教学优先级重新校准,以平衡表现、练习与进阶。

AI 中文摘要

辅导系统使用掌握阈值来决定学生何时可以停止练习并继续前进,但当底层知识追踪(KT)模型发生变化时,相同的数值阈值可能导致截然不同的决策。我们检查了四个公共教育数据集中的六种KT模型,并使用进阶后表现、进阶覆盖率、练习负担以及不同先前表现组之间的差异,评估了从0.50到0.99的12个阈值。我们还在30种预定义教学设置下,确定了能够平衡表现、额外练习和进阶的阈值。贝叶斯知识追踪(BKT)对阈值变化相对不敏感,而神经模型随着阈值增加变得更加具有选择性。这在一定程度上反映了模型输出的差异:BKT估计潜在掌握概率,而神经模型估计下一次回答正确的概率,因此相同的截断值并不代表相同的掌握水平。最佳平衡阈值在不同模型和设置之间差异显著。在测试设置的一半中,神经模型和BKT选择的阈值差异超过0.10,尽管当更重视减少额外练习和允许更多学生进阶时,这一差距变小。更严格的阈值也并未可靠地减少表现差距,并可能不成比例地限制进阶,其中先前基础较强的学生进阶频率是较弱学生的3.26倍。这些结果表明,当KT模型或教学优先级发生变化时,应重新校准掌握阈值,并根据其对表现、练习、进阶和可及性的影响进行评估。

英文摘要

Tutoring systems use mastery thresholds to decide when students can stop practicing and advance, but the same numerical threshold can lead to very different decisions when the underlying knowledge tracing (KT) model changes. We examine six KT models across four public educational datasets and evaluate 12 thresholds from 0.50 to 0.99 using post-advancement performance, advancement coverage, practice burden, and disparities across prior-performance groups. We also identify thresholds that balance performance, extra practice, and advancement under 30 predefined instructional settings. Bayesian Knowledge Tracing (BKT) is relatively insensitive to threshold changes, while neural models become much more selective as thresholds increase. This partly reflects different model outputs: BKT estimates latent mastery probability, whereas neural models estimate the probability of a correct next response, so the same cutoff does not represent the same level of mastery. The best-balanced threshold varied substantially across models and settings. In half of the tested settings, neural models and BKT differed by more than 0.10 in their selected thresholds, although this gap became smaller when greater priority was placed on reducing extra practice and allowing more students to advance. Stricter thresholds also did not reliably reduce performance gaps and could disproportionately restrict advancement, with stronger-prior students advancing up to 3.26 times as often as weaker-prior students. These results show that mastery thresholds should be recalibrated when the KT model or instructional priorities change and evaluated by their effects on performance, practice, advancement, and access.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑