发表机构
FPT University; Van Lang University(FPT大学; 文朗大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究跨架构审计深度知识追踪模型的人口统计学偏见,发现最准确的AKT模型偏见最严重,且标准缓解措施无效。
AI 中文摘要
深度知识追踪(DKT)模型隐式地决定自适应系统认为哪些学生已掌握某项技能,然而关于其人口统计学公平性的证据几乎全部来自贝叶斯知识追踪;为现代系统提供动力的深度模型尚未接受同等规模的跨架构审计。我们填补了这一空白:在带有人口统计学元数据的两个公开数据集上,即Eedi(1590万次交互)和OULAD(预处理后16.7万条记录),以三种训练机制(标准、重加权、对抗)训练四种架构(DKT、DKVMN、SAKT、AKT),并使用ABROCA、学生级自助置信区间和置换检验进行评估,以回应近期关于公平性指标不稳定性的批评。研究得出三项发现。(i)偏见真实存在但依赖于情境:每种架构在Eedi上都显示出显著的社会经济地位ABROCA(0.018-0.023,p<0.005),经济弱势学生的每组AUC更低,而经过多重性校正后,四种架构中有三种在OULAD上性别偏见显著,但在Eedi上可忽略不计。(ii)最准确的架构偏见最严重:AKT通过项目级Rasch嵌入获得约4个AUC点的提升,并显示出最大的社会经济地位ABROCA,在配对自助法下超过所有其他架构(p≤0.002);仅消融Rasch嵌入即可同时消除准确性提升和额外偏见。(iii)标准缓解措施不可靠:在保持准确性的所有配置中,重加权和对抗性去偏基本不改变ABROCA,尽管对抗器在全反转强度下被固定在随机水平,且弱强度阳性对照排除了探针失效的可能性。
英文摘要
Deep knowledge tracing (DKT) models implicitly decide which students an adaptive system believes have mastered a skill, yet almost all evidence on their demographic fairness comes from Bayesian knowledge tracing; the deep models that power modern systems have received no comparable cross-architecture audit. We close this gap: four architectures (DKT, DKVMN, SAKT, AKT) trained under three regimes (standard, reweighting, adversarial) on two public datasets with demographic metadata, Eedi (15.9M interactions) and OULAD (167k after preprocessing), evaluated with ABROCA, student-level bootstrap confidence intervals, and permutation tests addressing recent critiques of fairness-metric instability. Three findings emerge. (i) Bias is real but context-dependent: every architecture shows a significant socioeconomic ABROCA on Eedi (0.018-0.023, $p<0.005$), with per-group AUC lower for economically disadvantaged students, while gender bias is significant on OULAD for three of four architectures after multiplicity correction yet negligible on Eedi. (ii) The most accurate architecture is the most biased: AKT gains about 4 AUC points from item-level Rasch embeddings and shows the largest socioeconomic ABROCA, exceeding every other architecture under a paired bootstrap ($p\leq0.002$); ablating only the Rasch embeddings removes the accuracy gain and the excess bias together. (iii) Standard mitigation is unreliable: reweighting and adversarial debiasing leave ABROCA essentially unchanged in every configuration that preserves accuracy, even though the adversary is pinned at chance at full reversal strength and a weak-strength positive control rules out a dead probe.
Comments15 pages, 3 figures