当在线适应有害时:参数冻结的测试时集成用于持续医学图像分割
When Online Adaptation Hurts: Parameter-Frozen Test-Time Ensembling for Continual Medical Image Segmentation
浏览论文内容
中文总结 AI 辅助
针对医学图像分割在域变化下性能下降的问题,提出参数冻结的测试时集成(PIE)方法,通过多视图预测平均增强推理,在心脏MRI上优于在线适应基线,揭示了适应可能有害的失败模式。
中文摘要 AI 辅助
医学图像分割器在站点、扫描仪供应商或协议发生变化时,性能往往会下降。持续测试时适应(CTTA)在没有目标标签的情况下解决了这个问题,但在非平稳流上更新模型可能不可行,并可能导致大量错误。我们研究了一种更合理且有意义的替代方案:参数冻结的推理增强(PIE)。我们使用一个在源域训练的、学习了解剖结构保持的尺度和翻转视图的分割器,将它们的预测映射回原始位置,并平均概率。我们不修改模型的权重或归一化统计量。在来自M&Ms的心脏MRI流上,该流在供应商A上训练,并依次在供应商B、C和D上评估,PIE的平均Dice为0.7786,而源域仅推理为0.7680,其他五种在线适应基线为0.7388--0.7416。受控消融实验表明,性能在28个视图时饱和,置信度加权、类别先验校正、连通分量滤波、形态学细化和切片间平滑没有效果或导致负迁移。心脏MRI和眼底图像的定性结果也一致表明,冻结集成保留了更薄和嵌套的解剖结构。这些结果为医学CTTA提供了一个强大且稳定的基线,并揭示了一个重要的失败模式:适应和手工精炼可能不如精心设计的推理可靠。
英文摘要
Medical image segmenters often get worse when sites, scanner vendors, or protocols change. Continual test-time adaptation (CTTA) addresses this problem without target labels, but it can be impossible to update a model on a non-stationary stream and can lead to a lot of errors. We examine a more reasonable and meaningful alternative: parameter-frozen inference enhancement(PIE). We use a source-trained segmenter that learns about anatomy-preserving scale and flip views, maps their predictions back to the native location, and averages the probabilities. We do not modify the weights of the model or the normalization statistics. On a cardiac MRI stream from M\&Ms, which is trained on vendor A and evaluated sequentially on vendors B, C, and D, PIE has 0.7786 mean Dice, compared to 0.7680 for source-only inference and 0.7388--0.7416 for five other online-adaptation baselines. The controlled ablations show that performance saturates at 28 views, and confidence weighting, class-prior correction, connected-component filtering, morphological refinement, and inter-slice smoothing have no effect or cause negative transfer. Qualitative results on cardiac MRI and fundus images are also consistent with the frozen ensemble keeping thinner and nested anatomical structures. These results provide a strong, stable baseline for medical CTTA and expose an important failure mode: adaptation and handcrafted refinement can be less reliable than carefully designed inference.
发表机构
- Hong Kong Baptist University(香港浸会大学)
机构由 AI 辅助整理,请以论文原文为准。