发表机构
Erzurum Technical University; Firat University(埃尔津詹技术大学; 菲拉特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文以场景化指南形式系统比较多种机器学习验证方法,指出无普适最优方案,强调独立单元须匹配部署目标且数据依赖操作须排除性能估计观测。
AI 中文摘要
模型验证评估完整学习过程在新数据上的性能。然而,无效的数据分割可能产生乐观且稳定的结果。本教程回顾了留出法验证、训练/验证/测试集设计、重复随机子采样、k折及重复分层交叉验证、留一法和留p法方案、分组感知验证以及嵌套分组交叉验证。通用机器学习原则与EEG时段、配对眼OCT图像、重复临床测量和多中心数据相关联。八个受控场景比较了有缺陷的设计与无泄漏设计:其中七个使用带可审计指标的锁定混淆矩阵,一个使用可复现的重复研究模拟。这些场景涵盖全局特征选择、归一化泄漏、依赖记录、中心混合、重复使用测试集以及估计器不稳定性。还考察了偏差、方差、指标聚合、不确定性和计算成本。提供了数据规模矩阵、决策树和报告清单。附带了可复现的MATLAB模板和对应的scikit-learn实现。结果表明,没有一种验证方法普遍最优。独立单元必须与预期部署目标相匹配。每个依赖数据的操作还必须排除用于性能估计的观测值。
英文摘要
Model validation estimates the performance of a complete learning procedure on new data. However, an invalid split can produce an optimistic and stable result. This tutorial reviews hold-out validation, train/validation/test designs, repeated random subsampling, k-fold and repeated stratified cross-validation, leave-one-out and leave-p-out schemes, group-aware validation, and nested group cross-validation. General machine-learning principles are linked to EEG epochs, paired-eye OCT images, repeated clinical measurements, and multicenter data. Eight controlled scenarios compare flawed and leakage-safe designs: seven use locked confusion matrices with auditable metrics, and one uses a reproducible repeated-study simulation. The scenarios cover global feature selection, normalization leakage, dependent records, center mixing, repeated test-set use, and estimator instability. Bias, variance, metric aggregation, uncertainty, and computational cost are also examined. A data-size matrix, a decision tree, and reporting checklists are provided. Reproducible MATLAB templates and scikit-learn counterparts are included. The results show that no validation method is universally best. The independent unit must match the intended deployment target. Every data-dependent operation must also exclude the observations used for performance estimation.
CommentsTutorial with eight controlled scenarios; includes MATLAB and Python/scikit-learn code listings