AI 中文总结
本研究揭示多模态大语言模型在跨域组织学相似性判断中优于病理学基础模型,通过MOSAIC基准评估17个模型,发现病理编码器存在域内偏差,而LLMs通过语义视觉比较更鲁棒,为跨机构检索提供新方案。
AI 中文摘要
最先进的病理学基础模型,在数百万张组织学切片上训练,当比较跨越切片或机构边界时,可能无法保持组织相似性。我们表明,通用多模态大语言模型,未经病理学基础模型训练,在跨域组织学相似性判断中始终优于这些专门模型。使用我们发布的相对相似性框架MOSAIC(跨机构和队列的模型相似性评估)基准,我们评估了6个数据集上的17个模型,发现病理学编码器常常将同机构、不同疾病的切片评为比同疾病、不同机构的切片更相似,这是一种在标准域内评估中不可见的临床危险失败模式。大语言模型似乎不太容易受到这种失败的影响,可能是因为它们对形态和组织结构进行语义视觉比较,而不是依赖与采集背景相关的捷径特征。扩大训练数据并不能解决病理学编码器的问题,这表明问题在于学习目标而非数据覆盖。我们的结果揭示了当前病理学基础模型的基本鲁棒性差距,并将多模态大语言模型确立为跨机构检索、数据集协调和多站点质量控制的可行替代方案。代码和数据将在接受后发布。
英文摘要
State-of-the-art pathology foundation models, trained on millions of histology tiles, can fail to preserve tissue similarity when comparisons cross slide or institution boundaries. We show that general-purpose multimodal LLMs, without being trained as pathology foundation models, consistently outperform these specialized models in cross-domain histological similarity judgments. Using a relative similarity framework that we release as the MOSAIC (Model Similarity Assessment across Institutions and Cohorts) benchmark, we evaluate 17 models across 6 datasets and find that pathology encoders often rank same-institution, different-disease tiles as more similar than same-disease, different-institution tiles, a clinically dangerous failure mode invisible to standard within-domain evaluations. LLMs appear less susceptible to this failure, likely because they perform semantic visual comparison of morphology and tissue architecture rather than relying on shortcut features tied to acquisition context. Scaling training data does not resolve the problem for pathology encoders, implicating the learning objective rather than data coverage. Our results expose a fundamental robustness gap in current pathology foundation models and establish multimodal LLMs as a viable alternative for cross-institutional retrieval, dataset harmonization, and multi-site quality control. Code and data will be released upon acceptance.
CommentsTo appear in NeurIPS 2026 (https://neurips.cc/virtual/2026/poster/152201)