发表机构
University of Würzburg(维尔茨堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出MEDSEGBENCHMARKER框架,通过整合多类评估相关功能实现2D医学图像分割的受控基准,案例研究显示评估设置的微小选择会影响基准结论,该框架可提升基准的明确性与可复现性。
AI 中文摘要
尽管医学图像分割(MIS)领域进展迅速,但由于数据集异质性、评估协议不一致及架构快速演进,分割模型的公平可复现比较仍具挑战性。尤其常见的比较隐含假设模型排名不受数据划分、预处理、指标聚合、不确定性估计及计算约束的影响,而缺乏可扩展的统一评估框架进一步限制了对新模型、数据集和训练范式的系统研究。本文提出MEDSEGBENCHMARKER(MSB),一种用于2D医学图像分割受控基准的配置驱动框架,其整合了重复与近重复图像检测、组感知数据划分、YAML研究规范、可恢复训练、超参数优化、交叉验证及基于检查点的评估。MSB不仅保留聚合性能指标,还导出样本级和类别级的像素计数、预测结果及评估上下文,这些基础产物支持事后分析而无需重复推理。本文通过包含三个异质性2D数据集的案例研究验证MSB,对多个医学图像分割模型及通用视觉模型在256像素和512像素输入分辨率下进行评估。尽管不同聚合策略间的排名相关性较高,但对相同预测结果的重新聚合在6个数据集-分辨率设置中的3个改变了排名第一的架构;提高输入分辨率会产生模型和数据集依赖的性能增益与损失,需结合经验测量的推理复杂度考虑。这些结果表明,评估和实验设置中看似微小的选择会影响基准结论,可在GitHub获取的MSB为使基准条件和评估选择明确且可复现提供了实用且可扩展的基础。
英文摘要
Despite rapid advances in MIS, fair and reproducible comparisons of segmentation models remain challenging due to heterogeneous datasets, inconsistent evaluation protocols, and rapidly evolving architectures. In particular, comparisons often implicitly assume that model rankings are invariant to data partitioning, preprocessing, metric aggregation, uncertainty estimation, and computational constraints. The lack of extensible and unified evaluation frameworks further limits systematic investigation of new models, datasets, and training paradigms. We present MEDSEGBENCHMARKER (MSB), a configuration-driven framework for controlled benchmarking of 2D MIS. It integrates duplicate and near-duplicate image detection, group-aware data splitting, YAML study specifications, resumable training, hyperparameter optimization, cross-validation, and checkpoint-based evaluation. Rather than retaining only aggregate performance measures, MSB exports sample- and class-level pixel counts and predictions together with the evaluation context. These elementary artifacts enable post-hoc analyses without repeated inference. We demonstrate MSB in a case study involving three heterogeneous 2D datasets and multiple MIS and general-purpose vision models evaluated at 256- and 512-pixel input resolutions. Reaggregation of identical predictions changes the top-ranked architecture in three of six dataset-resolution settings, despite high rank correlations between aggregation strategies. Increasing input resolution produces model- and dataset-dependent performance gains and losses that must be considered alongside empirically measured inference complexity. These results show that seemingly minor choices in evaluation and experimental setup can affect benchmark conclusions. MSB, available at GitHub, provides a practical and extensible basis for making benchmark conditions and evaluation choices explicit and reproducible.
Comments10 pages, 4 figures;