arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

知识追踪的稳定且忠实的解释

Stable and Faithful Explanations for Knowledge Tracing

Praveena Padi, Arun Morampudi, Ujval Sai Gopal Irrinki, Pradeep Kumar Dolabehera Kakitapelli

arXiv 2609.28502首次发表:更新:

发表机构

Georgia Institute of Technology; Amazon Web Services; University of the Cumberlands(佐治亚理工学院; 亚马逊云科技; 坎伯兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出一种验证协议,结合预测竞争力、解释稳定性和重训练忠实性,通过重建ASSISTments 2009数据并比较XGBoost与深度模型,证明解释差异源于信息而非模型,且TreeSHAP特征排名稳定有效。

AI 中文摘要

知识追踪(KT)模型不透明地预测学生表现,限制了教学行动。本研究提出了一种验证协议,同时测试预测竞争力(RQ1)、解释稳定性(RQ2)和基于重训练的忠实性(RQ3)。从ASSISTments 2009和2012中设计了跨越五个教学主题的十三个行为特征,其中历史特征由时间上先前的交互计算得出,当前响应延迟仅保留用于回顾性分析。ASSISTments 2009被重建:未修正的技能构建器版本将每个多技能交互在每个技能一行上重复,由于这些行共享一个正确性标签,它们将该标签泄漏到先前交互特征中。重建降低了模型AUC并重新排序了解释结果。使用Tree SHapley Additive exPlanations(TreeSHAP)解释的极端梯度提升(XGBoost)模型与四个深度基线(DKT、SAKT、AKT和SimpleKT)在信息匹配协议下进行比较,该协议给予深度模型相同的行为信号,并将XGBoost限制为仅使用从标识符和正确性流中可推导的信息。XGBoost在2012上达到0.777的曲线下面积(AUC),在重建的2009上达到0.786,排除当前响应延迟后,预测时AUC分别为0.771和0.775;限制在基线的信息范围内时,其表现与基线相同(0.697对0.700,0.717对0.720),将差异定位在提供的信息上,而非模型家族。排名在折叠、种子和条件方案中保持一致(Spearman rho = 0.989-1.000),移除排名最高的TreeSHAP特征比随机移除更能损害AUC,尽管分裂增益和排列排名表现相当。学生级别的示例是说明性解释,而非经过验证的建议。

英文摘要

Knowledge tracing (KT) models predict student performance opaquely, limiting pedagogical action. This study contributes a validation protocol testing predictive competitiveness (RQ1), explanation stability (RQ2) and retraining-based faithfulness (RQ3) together. Thirteen behavioral features across five pedagogical themes were engineered from ASSISTments 2009 and 2012, with history features computed from temporally preceding interactions and current response latency retained only for retrospective analysis. ASSISTments 2009 was rebuilt: the uncorrected skill-builder release duplicates each multi-skill interaction across one row per skill, and because those rows share one correctness label, they leak it into preceding-interaction features. Rebuilding lowered model AUC and reordered the explanation results. An Extreme Gradient Boosting (XGBoost) model explained with Tree SHapley Additive exPlanations (TreeSHAP) was compared against four deep baselines (DKT, SAKT, AKT and SimpleKT) under an information-matched protocol giving the deep models the same behavioral signals and restricting XGBoost to what is derivable from the identifier-and-correctness stream they consume. XGBoost reached an area under the curve (AUC) of 0.777 on 2012 and 0.786 on rebuilt 2009, with prediction-time AUCs of 0.771 and 0.775, respectively, after excluding current response latency; restricted to the baselines' information it performed as they did (0.697 against 0.700, and 0.717 against 0.720), locating the difference in information supplied, not model family. Rankings were consistent across folds, seeds and conditioning schemes (Spearman rho = 0.989-1.000), and removing top-ranked TreeSHAP features harmed AUC more than random removal, though split-gain and permutation rankings performed comparably. Student-level examples are illustrative interpretations, not validated recommendations.

Comments41 pages, 8 figures, 14 tables. Code is available at https://doi.org/10.5281/zenodo.22715077

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑