发表机构
University of Missouri; Government Degree College (Autonomous), Siddipet; Amar Biotech Private Limited(密苏里大学; 西迪佩特政府自治学院; 阿马尔生物科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出三层污染框架,审计三个公开脑肿瘤MRI分类语料库,发现严重的数据集泄漏,但去重后性能稳定,表明数据集完整性是独立于准确率的基准质量维度。
AI 中文摘要
从MRI自动分类脑肿瘤是医学影像中深度学习的一个大量发表的应用程序,在公开基准上报告的准确率通常超过98%。然而,准确率并不能捕捉基准质量的一个关键维度:数据集完整性,定义为测试数据在图像、患者和采集源层面独立于训练数据。我们引入了一个包含重复、患者和源标签泄漏的三层污染框架,以评估该文献所依赖的公开语料库。我们针对胸部X光片阴性对照审计了三个最广泛使用的语料库,并量化了每一层对九种架构和三种评估条件下测量性能的影响。污染在每一层都很严重:主要语料库的官方测试分割中有28.8%在其自身训练分割中有近孪生,第二个语料库有22.3%的测试图像字节相同地泄漏,95.5%的可追踪测试图像与训练共享患者,且包含无解剖结构的文件头特征以0.959的平衡准确率区分肿瘤与非肿瘤,与微调的ResNet骨干网络持平。出乎意料的结果是,移除所有已识别的泄漏测试图像后,平衡准确率基本保持不变:去重后的稳定性能并不能证明基准的完整性。我们的发现确立了数据集完整性作为基准质量的一个独立、可测量的轴,而稳定的排行榜无法证明这一点。对于生物医学研究,仅凭这些语料库上报告的准确率并不能证明模型学会了识别肿瘤而不是利用数据集特定的线索。我们发布了受污染文件列表、恢复的患者标识符和去重后的分割。
英文摘要
Automated classification of brain tumors from MRI is a heavily published application of deep learning in medical imaging, with reported accuracies on public benchmarks routinely exceeding 98%. However, accuracy does not capture a critical dimension of benchmark quality: dataset integrity, defined as the independence of test from training data at the image, patient, and acquisition-source levels. We introduce a three-layer contamination framework comprising duplicate, patient, and source-label leakage to assess the public corpora on which this literature rests. We audit the three most widely used corpora against a chest-radiograph negative control and quantify each layer's effect on measured performance across nine architectures and three evaluation conditions. Contamination is severe at every layer: 28.8% of the dominant corpus's official test split has a near-twin in its own training split, a second corpus leaks 22.3% of its test images byte-identically, 95.5% of traceable test images share a patient with training, and file-header features containing no anatomy separate tumor from no-tumor at 0.959 balanced accuracy, at parity with fine-tuned ResNet backbones. The unexpected result is that removing every identified leaked test image leaves balanced accuracy essentially unchanged: stable performance after deduplication does not establish benchmark integrity. Our findings establish dataset integrity as a distinct, measurable axis of benchmark quality that a stable leaderboard cannot certify. For biomedical research, reported accuracy on these corpora alone does not establish that a model has learned to recognize tumors rather than exploit dataset-specific cues. We release the contaminated-file lists, recovered patient identifiers, and deduplicated splits.