arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新思考语义分割中的事后校准

Rethinking Post-Hoc Calibration in Semantic Segmentation

Tristan Kirscher, Kim-Celine Kahl, Balint Kovacs, Maximilian R. Rokuss, Klaus Maier-Hein, Xavier Coubez, Philippe Meyer, Sylvain Faisan

arXiv 2607.01902首次发表:更新:

发表机构

ICube Laboratory, CNRS UMR 7357, University of Strasbourg; CLCC Institut Strauss; German Cancer Research Center (DKFZ); Faculty of Mathematics and Computer Science, University of Heidelberg; Medical Faculty Heidelberg, Heidelberg University; Pattern Analysis and Learning Group, Dept. of Radiation Oncology, Heidelberg University Hospital(斯特拉斯堡大学ICube实验室,CNRS UMR 7357; 斯特劳斯研究所; 德国癌症研究中心; 海德堡大学数学与计算机科学学院; 海德堡大学医学院; 海德堡大学医院放射肿瘤科模式分析与学习组)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对语义分割中事后校准存在的平移不变性和决策保持问题,提出平移不变校准器和决策保持校准器,在保持校准性能的同时避免分割退化。

AI 中文摘要

可靠的置信度估计在语义分割中至关重要,尤其是在安全关键场景中,过度自信的错误可能误导下游决策。然而,现代分割模型往往校准不佳。事后校准提供了一种无需重新训练分割模型即可修正置信度估计的实用方法,但其在密集预测中的应用存在常被忽视的结构性问题。我们研究了两个此类问题。首先,对所有logits添加常数不会改变softmax概率,但几种标准校准器仍可能依赖于这个任意偏移。因此,编码相同预测分布的两个logit表示可能产生不同的校准概率。我们将平移不变(TI)校准器定义为输出在此类偏移下不变的校准器,刻画了哪些常见校准器满足该性质,并构建了平移敏感校准器的TI对应版本,以隔离消除表示依赖性的影响。其次,事后校准通常通过最小化基于似然的目标来拟合,而分割模型则使用任务特定指标(如Dice)进行训练。这种不匹配可能导致校准改变类别排序并降低部署的分割图。我们研究了在argmax保持和顺序保持约束下的决策保持校准。由于施加这些约束会将仿射softmax校准器退化为温度缩放,我们引入了类别条件仿射校准器,这些校准器可以在保持更强表达性的同时实现argmax保持或顺序保持,从而量化决策保持引起的校准-分割权衡。在自然图像和医学分割基准测试以及基于腐败的协变量偏移下,匹配比较表明,TI变体通常改善校准指标,而决策保持变体防止分割退化并保持强校准性能。这些结果为语义分割中定义良好的事后校准流程提供了实用设计原则。

英文摘要

Reliable confidence estimates are essential in semantic segmentation, yet modern models often remain miscalibrated. We investigate two overlooked issues in post-hoc calibration. First, adding a constant to all logits leaves softmax probabilities unchanged, but several standard calibrators depend on this arbitrary offset. In segmentation, this offset can vary across pixels or voxels, introducing spatially varying representation dependence. We characterize translation-invariant (TI) calibrators and construct TI counterparts of shift-sensitive methods. Second, calibrating with cross-entropy can degrade segmentation quality due to mismatched training and calibration objectives and limited calibration data. We investigate decision-preserving calibration under argmax- and order-preservation constraints. Since these constraints restrict affine softmax calibrators to temperature scaling, we introduce more expressive class-conditional affine calibrators that preserve decisions. Across natural-image and medical segmentation benchmarks, including corruption-based covariate shift, TI variants generally improve calibration, while decision-preserving variants prevent segmentation degradation by construction and retain strong calibration performance. Our findings provide practical design principles for post-hoc calibration in semantic segmentation.

CommentsAccepted at Transactions on Machine Learning Research (TMLR)

Journal refTransactions on Machine Learning Research (2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑