发表机构
Radboud University Medical Center(拉德堡德大学医学中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究病理学基础模型受非生物学变异影响的问题,提出交叉混杂稳健性余量(CRoMa),通过直接比较距离评估模型稳健性,将其重塑为队列范围的余量分布,为模型选择和稳健性评估提供原则基础。
AI 中文摘要
病理学基础模型正迈向临床应用,但易受不同中心系统性非生物学变异影响。组织制备、染色和扫描差异在其表征中显著编码,导致捷径学习并削弱跨队列和机构的泛化能力。稳健性指数(RI)量化局部表征几何由生物学还是非生物学变异主导,但其基于计数的公式丢弃了距离信息。研究表明增加距离权重改变不大,因更深层次限制在于RI的汇总、固定邻域设计。为此引入交叉混杂稳健性余量(CRoMa),它直接比较到交叉混杂生物学匹配和同混杂生物学干扰物的距离,将稳健性重塑为队列范围的余量分布而非单个汇总分数。通过在三个基准上评估20个切片级编码器和在第四个基准上评估4个载玻片级编码器的冻结表征,发现中位数CRoMa排名在数据集间大致一致,但模型内存在显著异质性。每个切片编码器都有一个由混杂因素主导的下尾,其普遍性和严重性在模型间差异明显。较高的CRoMa还与监督适应后较小的捷径诱导性能下降相关。CRoMa将表征几何转化为预期下游捷径易感性的分布稳健性读数,为稳健性评估和模型选择提供了原则基础。
英文摘要
Pathology foundation models encode non-biological variation introduced by tissue preparation, staining and scanning, enabling shortcut learning that undermines generalisation across institutions. The Robustness Index (RI) was proposed to assess whether local representation geometry is dominated by biological or non-biological variation. However, its construction suffers from structural limitations that make cross-model comparison unreliable, calling for a more principled metric. We introduce the Cross-confounder Robustness Margin (CRoMa), a signed, per-sample margin that measures whether samples sharing the same biology but different confounder lie closer than samples sharing the same confounder but different biology. It is defined for every sample, allowing models to be compared on the same cohort and robustness to be analysed as a distribution rather than reduced to a single pooled score. We evaluated CRoMa across 20 tile-level encoders on three benchmarks. Rankings by median CRoMa were highly consistent across benchmarks (Spearman rho ~ 0.90), yet every encoder retained confounder-dominated samples, whose prevalence and severity varied markedly. Similar patterns emerged for four slide-level encoders evaluated on a separate benchmark, extending the analysis beyond tile-level representations. Higher median CRoMa was associated with smaller shortcut-induced performance losses in downstream linear probes, supporting its use as a representation-level indicator of shortcut susceptibility.
CommentsPreprint