HistoFID-校准跨病理学基础模型的弗雷歇距离评估
HistoFID- Calibrating Frechet-distance evaluation across pathology foundation models
浏览论文内容
中文总结 AI 辅助
研究数字病理学中用组织学基础模型取代 Inception 网络后 FID 结果受影响的问题,通过特定比率恢复可比性,划分编码器组,揭示其对生成模型评估结论的影响,还评估了 TuroCompress 编解码器并发布相关资源。
中文摘要 AI 辅助
弗雷歇 inception 距离(FID)通过将高斯分布拟合到固定网络的特征并测量两个高斯分布之间的距离来比较两个图像集。在数字病理学中,通常用组织学基础模型取代 inception 网络,假设域编码器能给出更有意义的分数。研究表明这种选择会改变结果。对于一组固定的切片集,原始弗雷歇距离在六个常见编码器之间变化约三十倍,且排序不遵循嵌入维度。利用内部队列(约 50 万张苏木精和伊红及免疫组化切片)和公共 TCGA BRCA 队列(100 张切片),对 Inception-v3、Phikon-v2、CONCH、UNI2-h、Virchow2 和 Prov-GigaPath 进行基准测试。将每个距离表示为与编码器自身队列内下限的比率可恢复可比性,在队列内将编码器间变异系数降低约 89%,跨队列降低 58%。编码器分为敏感组和不变组,这一划分决定了哪个生成模型被判定更真实。在切片级别,注意力池化编码器能记录池化补丁距离无法看到的每张切片组成,使距离提高约 320 倍。还评估了 TuroCompress 编解码器,其在测试编解码器中以最小文件大小达到最高重建保真度。最后发布了归一化协议、每个编码器的扰动面板和特征提取。
英文摘要
The Frechet Inception Distance (FID) compares two image sets by fitting a Gaussian to the features of a fixed network and measuring the distance between the two Gaussians. In digital pathology the Inception network is routinely replaced by a histology foundation model, on the assumption that a domain encoder gives a more meaningful score. We show that this choice changes the result. For one fixed pair of tile sets, the raw Frechet distance varies about thirty-fold across six common encoders, and the ordering does not follow embedding dimension, so a raw score cannot be read without naming the encoder. Using a held-out in-house cohort (about 500,000 H&E and immunohistochemistry tiles from 2,119 slides) and a public TCGA BRCA cohort (100 slides), we benchmark Inception-v3, Phikon-v2, CONCH, UNI2-h, Virchow2 and Prov-GigaPath across within-cohort baselines, cross-cohort drift, controlled perturbations, compression, stain normalization, and two generative models. Expressing each distance as a ratio to the encoder's own within-cohort floor restores comparability, cutting the across-encoder coefficient of variation by about 89% within cohort and 58% across cohorts. The encoders separate into a sensitive group (CONCH, Phikon-v2, Inception-v3) and an invariant group (UNI2-h, Virchow2, Prov-GigaPath), and this split decides which generative model is judged more realistic, so the encoder can change the conclusion of a generative evaluation. At the slide level, an attention-pooling encoder registers per-slide composition that a pooled patch distance cannot see, raising the distance about 320-fold on matched cohorts. Using the same protocol we evaluate TuroCompress, a proprietary pathology codec, which reaches the highest reconstruction fidelity at the smallest file size among codecs tested. We release the normalization protocol, the per-encoder perturbation panel, and the feature extracts.