SCALE:嵌入空间中通过一致性标注进行合成校准
SCALE: Synthetic Calibration via Agreement Labeling in Embedding Space
- Dartmouth College(达特茅斯学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对病理学基础模型校准不足的问题,提出利用嵌入空间中锚点插值生成合成一致性信号进行校准,在多个数据集上显著提升低一致性病例的校准性能,同时保持判别能力。
AI中文摘要:
计算病理学的基础模型通常使用AUC和准确率进行评估,而校准往往未经测试。这一点很重要,因为一个模型可能在平均意义上准确,但对即使对病理学家来说也困难的病例赋予过度自信的概率。我们研究了八个病理学基础模型的校准情况。使用病理学家一致性作为诊断难度的衡量标准,我们发现校准误差在低一致性病例中始终高于高一致性病例。这种模式仅从总体期望校准误差(ECE)中并不明显。随后,我们提出了合成一致性校准,一种无需收集多标注者标签即可改善校准的方法。给定一个训练好的线性探针,我们选择高置信度嵌入作为类锚点,并在来自相反类别的锚点之间进行插值。插值权重编码了诊断模糊性的连续概念,我们将其用作合成一致性信号,以一致性感知的标签平滑重新训练探针。在包含七位病理学家标注的MHIST上,合成一致性校准恢复了基于真实病理学家一致性的标签平滑所获得的大部分校准改进,同时相对于未校准基线显著降低了低一致性ECE。判别指标得以保留。在PatchCamelyon和BreakHis(无多标注者标签的公开组织病理学数据集)上,该方法改善了所评估基础模型的校准,而依赖标注者的方法在无额外专家标注的情况下无法使用。
英文摘要:
Foundation models for computational pathology are usually evaluated using AUC and accuracy, while calibration is often left untested. This matters because a model can be accurate on average but still assign overly confident probabilities to cases that are difficult even for pathologists. We study calibration across eight pathology foundation models. Using pathologist agreement as a measure of diagnostic difficulty, we find that calibration error is consistently higher on low-agreement cases than on high-agreement cases. This pattern is not apparent from aggregate expected calibration error (ECE) alone. We then propose synthetic agreement calibration, a method for improving calibration without collecting multi-annotator labels. Given a trained linear probe, we select high-confidence embeddings as class anchors and interpolate between anchors from opposite classes. The interpolation weights encode a continuous notion of diagnostic ambiguity, which we use as a synthetic agreement signal to retrain the probe with agreement-aware label smoothing. On MHIST, which includes annotations from seven pathologists, synthetic agreement calibration recovers most of the calibration improvement obtained by label smoothing based on real pathologist agreement, while substantially reducing low-agreement ECE relative to the uncalibrated baseline. Discrimination metrics are preserved. On PatchCamelyon and BreakHis, public histopathology datasets without multi-annotator labels, the method improves calibration across the evaluated foundation models, whereas annotator-dependent approaches cannot be used without additional expert annotation.