arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27710cs.CVcs.LG

FFM-CP:用于少样本计算病理学的视觉-语言基础模型跨骨干融合

FFM-CP: Cross-Backbone Fusion of Vision-Language Foundation Models for Few-Shot Computational Pathology

  • Giessen University(吉森大学)
  • University Medicine Gottingen(哥廷根大学医学中心)
  • Technical University of Denmark(丹麦技术大学)
  • Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
  • Max Planck Institute for Biology of Ageing(马克斯·普朗克衰老生物学研究所)
  • National Taiwan University(国立台湾大学)
  • VNU University of Medicine and Pharmacy(越南国立大学医学与药学大学)
  • University of Medicine and Pharmacy, Hue University(顺化大学医学与药学大学)
  • German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)
  • University of Stuttgart(斯图加特大学)

机构由 AI 辅助整理,请以论文原文为准。

Anh-Tien Nguyen, Trung DQ. Dang, Nghiem Tuong Diep, Bui Ngoc Han Nguyen, Tan-Ha Mai, Miriam Cindy Maurer, Phuong Hoa Nguyen, Thi Thuy Uyen Nguyen, Youngjun Park… 展开作者

Anh-Tien Nguyen, Trung DQ. Dang, Nghiem Tuong Diep, Bui Ngoc Han Nguyen, Tan-Ha Mai, Miriam Cindy Maurer, Phuong Hoa Nguyen, Thi Thuy Uyen Nguyen, Youngjun Park, Daniel Sonntag, Duy Minh Ho Nguyen, Anne-Christin Hauschild

AI总结:

针对病理学基础模型单一表现不佳且标注稀缺的问题,提出FFM-CP框架,通过Procrustes对齐多模型表示并利用统一图跨骨干融合,在6个数据集上多数比较中优于最强单模型。

AI中文摘要:

病理学视觉-语言基础模型在不同疾病和任务上的表现各异,没有单一模型能始终表现最佳。专家病理标注的高成本也可能限制可用于任务特定适配的标注数据。结合互补的预训练表示是应对这些局限的一种潜在方法,然而从少量标注样本中学习有效融合仍具挑战性。我们提出了少样本计算病理学融合基础模型(FFM-CP),这是一个在少样本学习设置下结合多个病理学视觉-语言模型的框架。该框架首先利用从对应支持图像估计的闭式正交Procrustes变换来对齐异构表示。这种对齐在不训练额外对齐网络的情况下保留了模型内部的特征几何结构。在对齐空间内,一个统一图通过联合细化支持图像特征以及视觉和文本类原型,实现跨骨干的信息交换。这些细化后的表示支持互补的文本原型分支和案例检索分支,分别捕获语义类知识和类内视觉变异。每个分支学习组合来自所有有序骨干对的预测,允许由一个模型编码的查询利用另一个模型所表示的证据。我们在六个组织病理学数据集上,以每类4、8和16个样本的设置评估了三种骨干组合。在54次比较中,FFM-CP在50次中取得了比每个融合集中最强的单独适配成员更高的平均宏F1分数。这些发现表明,当标注有限时,结合互补的预训练表示可以改善组织病理学分类。

英文摘要:

Pathology vision-language foundation models vary in performance across diseases and tasks, with no single model consistently performing best. The high cost of expert pathology annotation can also limit the labeled data available for task-specific adaptation. Combining complementary pretrained representations is a potential approach to these limitations, yet learning an effective fusion from few labeled examples remains challenging. We introduce Few-shot Fusion Foundation Models of Computational Pathology (FFM-CP), which is a framework that combines multiple pathology vision-language models in the few-shot learning setting. The framework first aligns heterogeneous representations using a closed-form Orthogonal Procrustes transformation estimated from corresponding support images. This alignment preserves within-model feature geometry without training an additional alignment network. Within the aligned space, a unified graph enables information exchange across backbones by jointly refining support-image features and visual and textual class prototypes. These refined representations support complementary text-prototype and case-retrieval branches that capture semantic class knowledge and within-class visual variation, respectively. Each branch learns to combine predictions from all ordered backbone pairs, allowing queries encoded by one model to draw on evidence represented by another. We evaluate three backbone combinations on six histopathology datasets at 4, 8, and 16 shots per class. FFM-CP achieves higher mean macro-F1 than the strongest individually adapted member of each fused set in 50 of 54 comparisons. These findings suggest that combining complementary pretrained representations can improve histopathological classification when annotations are limited.

↑