arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越分布内指标:先天性心脏病分割的系统性分布外评估

Beyond In-Distribution Metrics: A Systematic Out-of-Distribution Evaluation of Congenital Heart Disease Segmentation

Aniketh Vijesh, Shrisharanyan Vasu, Abhijit Ramesh, Clare Pomeroy-Ward, Harikrishnan Anil Maya, Sarin Xavier, Mahesh Kappanayil, Gilad Gressel

arXiv 2609.17068首次发表:更新:

发表机构

Amrita Vishwa Vidyapeetham; Amrita Institute of Medical Sciences and Research Centre, Amrita Vishwa Vidyapeetham(阿姆里塔大学; 阿姆里塔医学科学与研究中心,阿姆里塔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究首次系统评估先天性心脏病分割的分布外泛化,发现nnU-Net分布内性能高但跨队列差,SwinUNETR更鲁棒,并建议采用跨数据集测试。

AI 中文摘要

先天性心脏病(CHD)的诊断和手术规划通常需要患者特定的三维解剖模型,但手动分割劳动强度大,尤其是在复杂解剖结构中。尽管深度学习方法可以自动化这一过程,但它们通常在分布内进行评估,尽管在扫描仪、协议、机构、人群和成像模态方面存在临床相关的偏移。我们提出了,据我们所知,首次对CHD分割中分布外(OOD)泛化进行系统性评估,使用ImageCHD作为保留的目标队列。我们在CT和CMR联合训练、仅CT训练、自监督预训练和有限目标域适应下比较了代表性分割架构。分布内性能被证明是跨队列鲁棒性的不良指标:nnU-Net在验证集上达到最高Dice(0.77),但在ImageCHD上降至0.51,而SwinUNETR的泛化能力显著更好,达到0.67 Dice。MAE和JEPA预训练仅提供适度的额外收益,表明在此设置中,架构对鲁棒性的贡献大于所测试的预训练策略。当引入有限的目标域监督时,所有SwinUNETR变体仅用11个标记的ImageCHD病例就超过0.76 Dice。这些发现表明,传统的分布内评估可能掩盖临床重要的泛化失败,并支持将显式跨数据集测试作为CHD分割评估的关键组成部分。

英文摘要

Congenital heart disease (CHD) diagnosis and surgical planning often require patient-specific 3D anatomical models, but manual segmentation is labor-intensive, particularly in complex anatomies. Although deep-learning methods can automate this process, they are typically evaluated in-distribution, despite clinically relevant shifts in scanner, protocol, institution, population, and imaging modality. We present, to our knowledge, the first systematic evaluation of out-of-distribution (OOD) generalization in CHD segmentation, using ImageCHD as a held-out target cohort. We compare representative segmentation architectures under combined CT and CMR training, CT-only training, self-supervised pretraining, and limited target-domain adaptation. In-distribution performance proves to be a poor indicator of cross-cohort robustness: nnU-Net achieves the highest validation Dice (0.77) but falls to 0.51 on ImageCHD, while SwinUNETR generalizes substantially better, reaching 0.67 Dice. MAE and JEPA pretraining provide only modest additional benefit, suggesting that architecture contributes more to robustness than the tested pretraining strategies in this setting. When limited target-domain supervision is introduced, all SwinUNETR variants exceed 0.76 Dice with only 11 labeled ImageCHD cases. These findings demonstrate that conventional in-distribution evaluation can obscure clinically important generalization failures and support explicit cross-dataset testing as a key component of CHD segmentation evaluation.

Comments12 pages, 6 figures, 2 tables. Accepted at STACOM 2026, held in conjunction with MICCAI 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑