发表机构
College of Biomedical Engineering, Fudan University; Faculty of Information Science and Computing, University of Macau(复旦大学生物医学工程学院; 澳门大学信息科学与计算学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对域外MRI分割的域偏移导致概率校准失效的问题,提出基于反向扩散一致性的CARD方法,利用生成式形状先验的分歧实现像素级校准,显著降低了多器官MRI域偏移下的校准误差。
AI 中文摘要
概率校准使模型置信度与预测准确率对齐,便于临床医生识别不可靠的分割区域。但在域偏移场景下,该对齐会失效——伪影与未见过的扫描方案会导致模型产生高置信度的错误。现有事后方法在测试时基于预测熵、logit模式或增强响应调整校正,然而这些代理指标均来自终端预测,而终端预测正是受域偏移破坏的对象。这促使我们探索终端预测之外的可靠性证据,类别扩散可通过两种方式提供此类证据:其一,生成式形状先验在外观受破坏时保持容量受限的参考结构完整,其与主分割器的分歧可凸显主模型的错误;其二,每一步反向扩散都会生成类别分布,从而将持续分歧与瞬时差异区分开。对整个轨迹进行聚合后,该分歧与Dice系数的相关度达0.788,而匹配的判别式对照方法仅为0.521。因此,我们提出CARD(Calibration via Agreement in Reverse Diffusion,基于反向扩散一致性的校准方法),它将该分歧的时间聚合结果映射到一个温度场,该温度场会应用于所有类别的每个像素,从而在不改变分割结果的情况下调整置信度。在心脏、前列腺和脑部MRI的域偏移场景中,CARD在49项对比中的45项里,均比各场景下最强基线方法降低了校准误差。
英文摘要
Probability calibration aligns model confidence with predictive accuracy, enabling clinicians to identify unreliable segmentation regions. This alignment breaks down under domain shift, where artifacts and unseen protocols produce confident errors. Existing post-hoc methods adapt the correction at test time, conditioning on predictive entropy, the logit pattern, or augmentation response, but each proxy is read from the terminal prediction, the very quantity that shift corrupts. This motivates reliability evidence beyond the terminal prediction, which categorical diffusion provides in two ways. First, a generative shape prior keeps a capacity-limited reference intact when appearance is corrupted, so its disagreement with the primary segmentor highlights primary-model errors. Second, every reverse step yields a class distribution, separating persistent disagreement from transient discrepancy. Aggregated over the trajectory, this disagreement correlates with Dice at 0.788, against 0.521 for a matched discriminative control. We therefore propose CARD (Calibration via Agreement in Reverse Diffusion), which maps the temporal aggregate of this disagreement to a temperature field applied per pixel across all classes, so that confidence changes while the segmentation does not. Across cardiac, prostate and brain MRI shifts, CARD lowers calibration error in 45 of 49 comparisons against the strongest baseline in each setting.
Comments10 pages, 5 figures