arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

公共教育预测基准的有限结构可靠性:对七个数据集的四维审计

Limited Structural Reliability in Public Educational Prediction Benchmarks: A Four-Dimension Audit of Seven Datasets

Yan Ma, Lizhuo Zhang

arXiv 2609.29625首次发表:更新:

发表机构

Changsha University of Science and Technology; Hunan Agricultural University(长沙理工大学; 湖南农业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

通过四维审计七个公共教育预测数据集,发现多数存在跨组脆弱性,基准可靠性受数据结构限制而非算法,提出建模前审计作为质量门槛。

AI 中文摘要

在七个公共教育预测数据集中,有三个通过了全部四项建模前可靠性检查;其余四个要么未通过组感知泛化测试,要么缺乏运行这些测试所需的来源元数据。一个数据集最初被归类为未通过,但在从留出矩阵中排除组标识特征后得到修正,这表明该审计能够区分真正的跨组混杂与特征编码伪影。每个数据集在模型优化前都使用四项检查进行审计:基线差距、分割不稳定性、空值分离以及组感知留出下的元数据充分性。主要失败模式并非仅仅是弱的独立同分布性能,而是跨组脆弱性:在最明显的情况下,UCI Student 的独立同分布 R 平方从 0.242 降至组留出 R 平方 -0.097,而 Higher Ed 则从 0.041 崩溃至 -8.79。增加模型复杂度并未消除这一模式:集成模型改善了结构稳健的数据集,但在脆弱数据集上放大了不稳定性或在组留出下失败。一项探索性的跨数据集比较进一步表明,更强的性能概况聚集在更大、群体更丰富、性能更接近的数据集中,而随机分割性能严重高估了脆弱数据集中的可部署信号。分类指标敏感性分析得出了相同的主要结论。结果表明,教育人工智能中的基准可靠性更多地受限于数据结构、群体异质性和评估设计,而非算法选择。一个可复用的建模前审计为公共教育数据集支持强基准或部署声明之前提供了一个最低质量门槛。

英文摘要

Across seven public educational prediction datasets, three passed all four pre-modeling reliability checks; the remaining four either failed group-aware generalization tests or lacked the provenance metadata needed to run them. One dataset was initially classified as failing but corrected after excluding group-identifier features from the holdout matrix, demonstrating that the audit can distinguish genuine cross-group confounding from feature-encoding artifacts. Each dataset was audited before model optimization using four checks: baseline gap, split instability, null separation, and metadata adequacy under group-aware holdout. The dominant failure mode was not weak iid performance alone but cross-group fragility: in the clearest case, UCI Student declined from iid R-squared 0.242 to group-holdout R-squared -0.097, while Higher Ed collapsed from 0.041 to -8.79. Increasing model complexity did not remove this pattern: ensemble models improved structurally sound datasets but amplified instability or failed under group holdout on fragile ones. An exploratory cross-dataset comparison further showed that stronger profiles clustered in larger, richer-grouped, performance-proximal datasets, while random-split performance severely overstated deployable signal in fragile datasets. Classification-metric sensitivity analyses reached the same substantive conclusions. The results show that benchmark reliability in educational AI is constrained less by algorithm choice than by data structure, group heterogeneity, and evaluation design. A reusable pre-modeling audit offers a minimum quality gate before public educational datasets support strong benchmark or deployment claims.

Comments26 pages, 9 figures, 13 tables. Accepted by Scientific Reports (2026). Code and data: https://doi.org/10.5281/zenodo.21887892

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑