计算病理学中测试时自适应的解释稳定性:大规模基准测试
Explanation Stability of Test-Time Adaptation in Computational Pathology: A Large-Scale Benchmark
浏览论文内容
中文总结 AI 辅助
该研究构建了计算病理学TTA的大规模基准,发现不同TTA方法对模型解释的影响差异显著,解释稳定性与自适应质量弱相关,强调了解释稳定性是TTA的重要可靠性维度。
中文摘要 AI 辅助
测试时自适应(TTA)已成为将部署的模型适配到未标记目标数据的实用方法,该设置在计算病理学中尤为相关,因为染色、扫描仪和队列偏移是常规情况。尽管大多数TTA方法通过对准确率的影响进行评估,但临床使用还取决于模型的解释在自适应后是否保持可靠。在本文中,我们仔细研究了这一在很大程度上未被测量的影响。我们在两个组织病理学基准Camelyon17和NCT CRC-HE上,针对从卷积网络到视觉Transformer的五种架构、一个病理学基础模型、17种TTA方法以及四种归因族,研究了TTA下的解释稳定性。在2958次自适应运行中,我们观察到一个清晰且系统的模式:不同TTA方法在改变模型解释的程度上差异显著,冻结骨干方法几乎不改变归因,而像CoTTA和RoTTA这样的持续方法则导致最大的漂移。这种效应并非均匀分布:卷积网络比Transformer和基础模型骨干敏感得多,且解释漂移随自适应强度增加而增大,但对批量大小基本不敏感。令人惊讶的是,解释稳定性与自适应质量仅弱相关。一些方法几乎完美保留解释,却降低了校准或准确率,产生了仅通过准确率或仅解释评估会遗漏的隐性失效。这些发现表明,解释稳定性是计算病理学中TTA的一个独特可靠性维度。我们发布了指标、协议和完整基准,以支持未来对不仅准确、而且稳定且可临床审计的自适应方法的研究。代码:this https URL
英文摘要
Test-time adaptation (TTA) has become a practical way to adapt deployed models to unlabeled target data, a setting that is especially relevant in computational pathology where staining, scanner, and cohort shifts are routine. While most TTA methods are evaluated by their effect on accuracy, clinical use also depends on whether the model's explanations remain reliable after adaptation. In this paper, we take a closer look at this largely unmeasured effect. We study explanation stability under TTA across two histopathology benchmarks, Camelyon17 and NCT CRC-HE, using five architectures ranging from convolutional networks to vision transformers and a pathology foundation model, seventeen TTA methods, and four attribution families. Across 2,958 adaptation runs, we observe a clear and systematic pattern: TTA methods differ sharply in how much they move model explanations, with frozen-backbone methods leaving attributions almost unchanged and continual methods such as CoTTA and RoTTA causing the largest drift. This effect is not uniform. Convolutional networks are substantially more sensitive than transformer and foundation-model backbones, and explanation drift increases with adaptation strength while remaining largely insensitive to batch size. Surprisingly, explanation stability is only weakly coupled to adaptation quality. Some methods preserve explanations almost perfectly while degrading calibration or accuracy, producing silent failures that would be missed by accuracy-only or explanation-only evaluation. These findings show that explanation stability is a distinct reliability axis for TTA in computational pathology. We release the metric, protocol, and full benchmark to support future work on adaptation methods that are not only accurate, but also stable and clinically auditable. Code: https://github.com/bahumanyarg11/tta-explanation-stability-pipeline
发表机构
- R.V. College of Engineering(R.V.工程学院)
- University of Nottingham(诺丁汉大学)
机构由 AI 辅助整理,请以论文原文为准。