arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.06988cs.LG

知识追踪中基准与模型诊断的信息论评估框架

An Information-Theoretic Evaluation Framework for Benchmark and Model Diagnosis in Knowledge Tracing

Houru Jiang, Zixi Wang, Tengteng Cheng, Xueyi Li, Mingliang Hou, Jiaqi Zheng, Renqiang Luo, Teng Guo, Zitao Liu

首次发表
浏览论文内容

中文总结 AI 辅助

提出信息论评估框架,利用上下文树加权量化局部不确定性,按熵带诊断知识追踪模型性能,揭示改进非均匀并识别噪声敏感区域。

中文摘要 AI 辅助

知识追踪(KT)模型主要使用诸如曲线下面积(AUC)和准确率等聚合指标进行评估。然而,这些全局分数掩盖了剩余误差的来源,并且无法指示基准是否正在接近饱和。虽然在现实KT设置中估计全局理论性能极限具有挑战性,但量化局部可预测性是可能的。为解决此问题,我们提出了一个用于KT基准诊断的信息论评估框架。我们在项目响应历史和当前项目查询上使用上下文树加权(CTW)作为操作性的因果不确定性坐标,同时将其与完整KT信息集下未观察到的局部不可约不确定性(LIU)区分开来。通过将预测投影到这一共享的不确定性坐标上,我们评估模型在不同熵带上的性能提升,而不仅仅是在全局层面。在NIPS Task 3/4和Algebra 2005上的全面评估揭示,模型改进高度不均匀。现代KT模型在高熵区域显示出显著提升,额外的项目感知参考、对数损失和等频率分析支持了这一局部化结论。该框架还标记了表观提升需要检查噪声敏感行为的区域。通过将这些局部建模失败与真实提升一起呈现,该方法为研究残差预测结构以及当前KT基准和模型的局限性提供了诊断工具。

英文摘要

Knowledge tracing (KT) models are predominantly evaluated using aggregate metrics such as area under the curve (AUC) and accuracy. However, these global scores obscure where the remaining errors originate and fail to indicate whether a benchmark is approaching saturation. While estimating a global theoretical performance limit is challenging in realistic KT settings, it is possible to quantify local predictability. To address this, we propose an information-theoretic evaluation framework for KT benchmark diagnosis. We use Context Tree Weighting (CTW) on item-response histories and current-item queries as an operational causal uncertainty coordinate, while distinguishing it from the unobserved Local Irreducible Uncertainty (LIU) under the full KT information set. By projecting predictions onto this shared uncertainty coordinate, we evaluate model performance gains across distinct entropy bands rather than only at the global level. Comprehensive evaluations on NIPS Task 3/4 and Algebra 2005 reveal that model improvements are highly non-uniform. Modern KT models show substantial gains in high-entropy regions, and additional item-aware references, log-loss, and equal-frequency analyses support this localization. The framework also flags regions where apparent gains require checks for noise-sensitive behavior. By surfacing these local modeling failures alongside genuine gains, this approach provides a diagnostic tool for studying both residual predictive structure and the limitations of current KT benchmarks and models.

发表机构

  • Jinan University(暨南大学)
  • Guangdong Institute of Smart Education, Jinan University(暨南大学广东智慧教育研究院)
  • Jilin University(吉林大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑