arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14705cs.CVcs.LG

深度学习图像分类器超参数优化的交叉验证方法研究

On Cross-Validation for Hyperparameter Optimization of Deep Learning Image Classifiers

Ljubomir Buturovic

首次发表
浏览论文内容

中文总结 AI 辅助

该研究对比固定留出、重排留出、5折交叉验证三种超参数优化协议,发现小样本医学图像分类中交叉验证法的绝对性能估计误差更低,推荐计算资源充足时采用该方法。

中文摘要 AI 辅助

超参数优化(HPO)会显著影响深度学习(DL)图像分类器的性能,但目前对于如何生成驱动超参数优化的验证信号,尤其是在医学成像等领域常见的小样本场景下,几乎没有实证指导。我们在绝对性能估计误差(AEE;即最优配置的验证集AUROC与测试集AUROC之间的绝对差值)这一指标上,对比了三种超参数优化协议:固定留出法(F)、重排留出法(R)和5折交叉验证法(C)。所有协议的搜索空间、采样器、训练流程、架构和测试集均保持一致。我们在三个公开数据集上评估了这些协议,涵盖两类场景:二分类医学成像数据集(RSNA肺炎胸片、二值化HAM10000皮肤病变数据集)和200分类自然成像数据集(Tiny ImageNet),涉及不同规模的开发集样本量n,以及两种骨干网络(所有数据集均采用ResNet-18,RSNA数据集采用Vision Transformer(ViT-S/16))。在医学数据集上,所有点估计结果均显示交叉验证法优于两种留出法,且AEE的降低幅度在小样本量时最大,随n增大而减小;在保守的家族式调整后,该模式仍保持稳健。在Tiny ImageNet数据集上,三种协议的AEE均可忽略不计,测试集AUROC整体相近。在12种医学场景中,固定留出法的平均AEE低于重排留出法,不过这一次要发现的支持度不够统一。对于小样本医学图像分类,当计算资源允许时,我们推荐采用基于交叉验证的超参数优化,因为它以额外计算量为代价,能获得更可靠的开发阶段后续测试性能估计。

英文摘要

Hyperparameter optimization (HPO) can materially affect the performance of deep learning (DL) image classifiers, but there is little empirical guidance on how to derive the validation signal that drives it, especially for the small sample sizes common in fields such as medical imaging. We compared three HPO protocols in terms of {\em absolute performance-estimation error} (AEE; the absolute difference between the winning configuration's validation AUROC and its test AUROC): fixed holdout (F), reshuffled holdout (R), and 5-fold cross-validation (C). The search space, sampler, training procedure, architecture, and test set were held identical across protocols. We evaluated the protocols on three public datasets spanning two regimes: binary medical imaging (RSNA pneumonia radiographs and binarized HAM10000 skin lesions) and 200-class natural imaging (Tiny ImageNet), across a range of development set sizes $n$ and two backbones (ResNet-18 on all datasets, Vision Transformer (ViT-S/16) on RSNA). On the medical datasets, every point estimate favored cross-validation over both holdout protocols, with reductions in AEE largest at small sample sizes and diminishing as $n$ increased. This pattern remained robust under conservative family-wise adjustment. On Tiny ImageNet, AEE was negligible under all three protocols. Test AUROC was generally similar among protocols. Fixed holdout had lower mean AEE than reshuffled holdout in 11 of 12 medical conditions, although this secondary finding was less uniformly supported. For small-sample medical image classification, we recommend cross-validation-based HPO when computational resources permit because it trades additional computation for a more reliable development-time estimate of subsequent test performance.

↑