arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估用于早期慢性肾病预测的机器学习模型的可靠性:数据泄露和预测器稳定性的系统综述

Evaluating Reliability in Machine Learning Models for Early Chronic Kidney Disease Prediction: A Systematic Review of Data Leakage and Predictor Stability

Mashrul Hossain, Nafesa Kibria, Fahim Shahriar

arXiv 2607.11963首次发表:更新:

发表机构

East West University(东西方大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究对慢性肾病预测的机器学习模型进行系统综述,引入结构化分类法和评分框架评估可靠性,发现泄露与性能夸大有关,多数预测器缺乏稳定性,揭示很多性能提升源于方法局限而非真实预测力。

AI 中文摘要

利用机器学习进行慢性肾病的早期检测在医疗相关计算机科学领域引起了极大关注。尽管该领域进展迅速,但许多报告的研究仍不一致且可能具有误导性。一个重大缺陷是缺乏对方法学问题的有组织评估。关键问题包括数据泄露、获取患者时间记录受限以及报告的临床指标不一致。本研究对使用可解释机器学习技术的现有慢性肾病预测研究进行了系统文献综述,通过在主要学术数据库中系统搜索选出19项相关研究。为评估方法学可靠性,引入了信息泄露的结构化分类法和定量泄露评分框架来系统评估慢性肾病预测研究的可靠性。分析揭示了泄露与夸大性能之间的紧密关系。高泄露研究报告的平均准确率为95.48%,无泄露研究为80.2%,增长约15.28%。此外,跨研究特征稳定性分析表明只有一小部分预测器可一致重现,超过80%缺乏可靠性。总体而言,研究结果表明许多报告的性能提升源于方法学局限而非真正预测能力。

英文摘要

The early detection of Chronic Kidney Disease using machine learning has attracted significant interest in healthcare-related computer science. Despite rapid advancements in this field, many reported studies remain inconsistent and potentially misleading. A significant drawback is the lack of organized evaluation regarding methodological concerns. Key issues include data leakage, limited access to temporal patient records and inconsistency in reported clinical indicators. This research offers a systematic literature review of existing CKD prediction studies using interpretable machine learning techniques, where nineteen relevant studies were selected via systematic searches across major academic databases. To assess methodological reliability, this study introduces a structured taxonomy of information leakage and a quantitative leakage scoring framework to systematically evaluate reliability across CKD prediction studies. The analysis reveals a strong relationship between leakage and inflated performance. Here, High leakage-studies report an average accuracy of 95.48%, compared to 80.2% for leakage-free studies, reflecting an increase of approximately 15.28%. Furthermore, a cross-study feature stability analysis shows that only a small subset of predictors is consistently reproducible, with over 80% lacking reliability. Overall, the findings suggest that many reported performance improvements stem from methodological limitations rather than true predictive capability.

Comments17 pages, 7 Figures, Preprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑