AI 中文总结
本研究在严格泄漏控制下复现慢性肾脏病分类研究,比较多种管道,发现高准确率主要由临床近端变量和缺失模式驱动,强调需外部前瞻性验证。
AI 中文摘要
对小型临床数据集的近乎完美分类可能取决于预处理、数据泄漏、缺失数据处理和评估样本。我们在明确的泄漏控制下,复现了一项已发表的慢性肾脏病(CKD)分类研究,该研究使用UCI CKD数据集(400条记录,24个预测变量)。在任何学习型预处理之前,固定了分层的70:30开发-保留划分;插补、编码、缩放、监督特征选择和PCA仅在训练折内拟合。通过5x3重复分层交叉验证比较了九个分类器、三种特征表示和三种缺失数据策略,并仅基于训练交叉验证进行排名(并列时报告);一组冻结的决赛模型在保留集上评估一次(判别力、校准、Brier分数),视为锁定的历史内部样本。消融研究、分组SHAP值、另外三个种子以及事后敏感性分析检验了稳健性。排名最高的选定特征管道(KNN插补,Extra Trees)达到了重复交叉验证准确率0.9869和保留集准确率0.9833(精确95%区间0.9411-0.9980)。PCA(27个成分,95.96%方差)相对于可解释的原始特征表示仅提供了边际的、依赖种子的优势,且多个管道在统计上无法区分。选择过程的嵌套重采样给出了外部准确率0.9845,比相同折上的最高排名管道低0.0095;仅缺失指标产生的ROC面积为0.85-0.88,移除肾脏、泌尿和血液学字段将准确率降至0.89。泄漏控制后高内部性能仍然存在,但似乎主要由临床近端实验室和尿液分析变量以及CKD相关的缺失驱动;这不是临床验证,外部前瞻性评估仍然必要。
英文摘要
Near-perfect classification of small clinical datasets can hinge on preprocessing, data leakage, missing-data handling, and the evaluation sample. We reproduced a published chronic kidney disease (CKD) classification study on the UCI CKD dataset (400 records, 24 predictors) under explicit leakage controls. A stratified 70:30 development-holdout partition was fixed before any learned preprocessing; imputation, encoding, scaling, supervised feature selection, and PCA were fitted only within training folds. Nine classifiers, three feature representations, and three missing-data strategies were compared by 5x3 repeated stratified cross-validation and ranked on training cross-validation only (ties reported); a frozen set of finalists was evaluated once on the holdout (discrimination, calibration, Brier score), treated as a locked historical internal sample. Ablation, grouped SHAP values, three further seeds, and post hoc sensitivity analyses examined robustness. The top-ranked selected-feature pipeline (KNN imputation, Extra Trees) reached a repeated cross-validation accuracy of 0.9869 and a holdout accuracy of 0.9833 (exact 95% interval 0.9411-0.9980). PCA (27 components, 95.96% of variance) gave only a marginal, seed-dependent advantage over interpretable original-feature representations, and several pipelines were statistically indistinguishable. Nested resampling of the selection procedure gave an outer accuracy of 0.9845, 0.0095 below the top-ranked pipeline on the same folds; missingness indicators alone yielded a ROC area of 0.85-0.88, and removing renal, urinary, and hematologic fields cut accuracy to 0.89. High internal performance persisted after leakage controls but appears driven largely by clinically proximal laboratory and urinalysis variables and CKD-associated missingness; it is not clinical validation, and external, prospective evaluation remains necessary.
Comments16 pages, 9 figures