发表机构
Carnegie Mellon University; Harvard Medical School(卡内基梅隆大学; 哈佛医学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
通过消融实验分离架构与正则化组件,发现正则化损失是深度学习医学图像配准精度的主要驱动因素,能大幅减少不真实变形且几乎不增加推理成本。
AI 中文摘要
深度学习配准方法通常会在基础网络上叠加两类增强:架构性增强(如仿射预对齐阶段)和训练目标增强(如正则化损失)。论文往往同时采用两者,因此不清楚哪类增强在起作用。我进行了一项受控消融实验以区分它们。使用OASIS脑部MRI数据集(394个训练对象,20个测试对象),我训练了同一配准流程的四个变体:带有基本相似性损失的基线3D U-Net、带有完整正则化套件的相同U-Net、带有基本损失的仿射加可变形架构,以及带有完整套件的仿射架构。我评估了配准精度(MSE、NCC、SSIM)、变形质量(雅可比行列式保持、位移统计、解剖合理性评分)和计算成本。仅正则化就贡献了大部分增益:在MSE改进指标上相对增益21.3%(从1.78%到2.16%,P<.001),在NCC改进上相对增益21.8%,同时将最大变形从53.1单位降至0.51单位,减少了99.0%,且几乎不增加计算成本(推理时间-0.06%)。组合模型产生了最大的精度增益,为25.8%(从1.78%到2.24%),并将解剖合理性从0.596提高到0.930,推理时间成本适中增加+9.8%。梯度相关性从基线的0.742上升到完全增强模型的0.980。所有增强变体在合理的变形约束下均达到亚体素精度。在此设置中,正则化损失是主要驱动因素,在推理时免费提供精度增益和几乎所有的变形控制,而仿射架构以可接受的成本增加了较小的互补性益处。不真实变形减少99%解决了临床部署的一个已知障碍。
英文摘要
Deep learning registration methods routinely stack two kinds of enhancement on a base network: architectural additions such as affine pre-alignment stages, and training-objective additions such as regularization losses. Papers tend to adopt both at once, so it is unclear which is doing the work. I ran a controlled ablation to separate them. Using the OASIS brain MRI dataset (394 training subjects, 20 test subjects), I trained four variants of the same registration pipeline: a baseline 3D U-Net with basic similarity losses, the same U-Net with a full regularization suite, an affine-plus-deformable architecture with basic losses, and the affine architecture with the full suite. I evaluated registration accuracy (MSE, NCC, SSIM), deformation quality (Jacobian determinant preservation, displacement statistics, an anatomical plausibility score), and computational cost. Regularization alone accounted for most of the gain: a 21.3% relative gain on the MSE-improvement metric (1.78% to 2.16%, P<.001) and a 21.8% relative gain in NCC improvement, while cutting maximum deformation from 53.1 to 0.51 units, a 99.0% reduction, at essentially no computational cost (-0.06% inference time). The combined model produced the largest accuracy gain, 25.8% (1.78% to 2.24%), and raised anatomical plausibility from 0.596 to 0.930, at a moderate +9.8% inference-time cost. Gradient correlation rose from 0.742 at baseline to 0.980 for the fully enhanced model. All enhanced variants reached sub-voxel accuracy under plausible deformation constraints. Regularization losses are the primary driver in this setting, delivering the accuracy gains and almost all of the deformation control for free at inference time, while the affine architecture adds a smaller complementary benefit at acceptable cost. The 99% reduction in unrealistic deformations addresses a known barrier to clinical deployment.
Comments15 pages, 1 figure