发表机构
National University of Sciences and Technology (NUST)(国立科技大学(NUST))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过实证比较,发现使用DCGAN生成的合成图像增强脑肿瘤MRI分类数据集,在保持分类器和测试集不变时,并未提升Swin Transformer模型的准确率(均为96%),且ROC-AUC略有下降,提示合成增强需谨慎评估其分布保真度与下游效用。
AI 中文摘要
生成对抗网络(GAN)越来越多地被用于扩充医学影像数据集,但合成图像不一定能为下游分类带来益处。本研究探讨了在分类器和评估集保持不变的条件下,类别特定的深度卷积生成对抗网络(DCGAN)增强是否能够改善脑肿瘤分类。实验在7,200张脑部磁共振成像(MRI)扫描图像上进行,涵盖四个类别:胶质瘤、脑膜瘤、垂体瘤和无肿瘤。每个类别使用1,400张真实图像进行训练,并预留400张用于测试。基线Swin Transformer分类器仅使用真实训练图像进行训练,并与使用相同真实图像加上每类500张DCGAN生成图像增强的第二模型进行比较。两种条件均在相同的保留测试集上进行评估。两个模型达到了相同的96%总体准确率,而宏F1分数基本保持不变,ROC-AUC在增强后从0.987略微下降至0.982。类别层面的分析显示错误分布发生了小幅调整,而非一致性的性能提升。在采用的评估设置下,真实图像与合成图像之间的FID值介于209.15至314.27之间,表明存在显著的分布差异。这些结果表明,不应假定合成增强能够改善医学图像分类,而应同时评估其分布保真度和下游任务效用。
英文摘要
Generative adversarial networks (GANs) are increasingly used to augment medical imaging datasets, but synthetic images do not necessarily provide downstream classification benefits. This study investigates whether class-specific Deep Convolutional Generative Adversarial Network (DCGAN) augmentation improves brain tumor classification when the classifier and evaluation set are held constant. Experiments were conducted on 7,200 brain magnetic resonance imaging (MRI) scans across four classes: glioma, meningioma, pituitary tumor, and no tumor. For each class, 1,400 real images were used for training and 400 were reserved for testing. A baseline Swin Transformer classifier was trained using only the real training images and compared with a second model trained using the same real images augmented with 500 DCGAN-generated images per class. Both conditions were evaluated on the identical held-out test set. The two models achieved the same overall accuracy of 96%, while macro F1 remained effectively unchanged and ROC-AUC decreased slightly from 0.987 to 0.982 after augmentation. Class-level analysis showed small redistributions in errors rather than a consistent performance gain. FID values between real and synthetic images ranged from 209.15 to 314.27, indicating substantial distributional differences under the adopted evaluation setup. These results suggest that synthetic augmentation should not be assumed to improve medical image classification and should instead be evaluated for both distributional fidelity and downstream task utility.