arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37848cs.CVcs.LG

评估选择影响生物医学机器学习结论:一项儿科肺炎基准案例研究

Evaluation Choices Shape Biomedical ML Claims: A Pediatric Pneumonia Benchmark Case Study

Bhanu Prakash Vangala, Sowmya Guda, Latha Peddi, Navya Vangala

首次发表
浏览论文内容

中文总结 AI 辅助

本研究以儿科肺炎基准为例,揭示评估选择(如划分、训练策略、阈值等)显著影响生物医学ML结论,并提出七项报告建议以提升结果可解释性。

中文摘要 AI 辅助

生物医学机器学习论文常将模型性能压缩为一个首要数字。该数字可能看似模型固有属性,即便其高度依赖于基准的评估方式。我们在广泛使用的Kermany儿科胸部X光数据集上,使用九种图像分类器和受控评估协议研究此问题。在相同协议下,八个预训练骨干网络之间的AUROC差异仅为0.026。相比之下,改变骨干网络是否冻结或微调平均使AUROC变化0.044,而改变决策阈值平均使平衡准确率变化0.090。官方测试划分与训练池也存在可测量的差异:一个划分分类器区分两者的AUC为0.697,对正常X光片则升至0.898。最引人注目的是,仅使用文件属性(不含图像解剖信息)的分类器在训练池内达到0.992的平衡准确率,但在官方测试划分上降至0.496。基于验证集拟合的阈值和校准也无法完美迁移。这些结果表明,当划分、训练策略、阈值、指标、校准和不确定性未与高分基准一同报告时,同一高分可支持不同结论。我们最后提出七项报告建议,每项均与研究中测得的效应相关联。

英文摘要

Biomedical machine learning papers often compress model performance into one headline number. That number can look like a property of the model even when it depends strongly on how the benchmark was evaluated. We study this problem on the widely used Kermany pediatric chest radiograph dataset using nine image classifiers and a controlled evaluation protocol. Under the same protocol, the eight pretrained backbones differ by only 0.026 AUROC. In contrast, changing whether the backbone is frozen or fine-tuned changes AUROC by 0.044 on average, and changing the decision threshold changes balanced accuracy by 0.090 on average. The official test split is also measurably different from the training pool: a partition classifier distinguishes them at AUC 0.697, rising to 0.898 for normal radiographs. Most strikingly, a classifier using only file properties, with no image anatomy, reaches 0.992 balanced accuracy within the training pool but falls to 0.496 on the official test split. Validation-fitted thresholds and calibration also transfer imperfectly. These results show that a high benchmark score can support different conclusions when the split, training policy, threshold, metric, calibration, and uncertainty are not communicated with it. We end with a seven-item reporting recommendation in which each item is tied to an effect measured in the study

发表机构

  • University of Missouri(密苏里大学)
  • Government Degree College(政府学位学院)
  • Amar Bio Tech Pvt Ltd(阿马尔生物技术私人有限公司)

机构由 AI 辅助整理,请以论文原文为准。

↑