发表机构
College of Informatics, Korea University(韩国大学信息学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对病理学基础模型计算成本高问题,提出拉瓜迪亚框架,通过多阶段流程,在临床语言指导下整合多模型知识,进行自适应蒸馏。实验表明其87M参数学生模型性能出色,凸显临床语言对构建高效可靠数字病理系统的作用。
AI 中文摘要
病理学基础模型(PFMs)能提供强大的全切片图像(WSI)表示,但计算成本巨大。知识蒸馏(KD)可创建高效学生模型,现有的多教师方法常采用次优的均匀加权,忽略组织异质性。我们提出拉瓜迪亚(LaGuadia)框架,通过在临床语言指导下动态整合多个PFMs的专业知识,开发紧凑的病理学图像编码器。其采用多阶段流程,先从病理报告中提取视觉可观察的临床关键词,再通过视觉语言元教师(MedSigLIP)将视觉特征与关键词对齐以提供密集语义指导,最后进行自适应KD,根据教师与临床叙述的语义对齐加权。在WSI字幕、视觉问答和幻灯片级分类任务上的实验表明,一个87M参数的拉瓜迪亚学生模型匹配或超越了如GigaPath和UNI等基础规模模型,实现了强大的事实一致性和稳健泛化。这些结果凸显了临床语言作为构建高效可靠数字病理系统的有效语义锚点。代码可在该https网址获取。
英文摘要
Pathology Foundation Models (PFMs) offer powerful Whole Slide Image (WSI) representations but suffer from massive computational costs. While Knowledge Distillation (KD) can create efficient student models, existing multi-teacher methods often use suboptimal uniform weighting that ignores tissue heterogeneity. We propose LaGuadia (Language-Guided Adaptive DistillAtion), a framework that develops a compact pathology image encoder by dynamically integrating expertise from multiple PFMs under clinical linguistic guidance. Our approach utilizes a multi-stage pipeline: first, extracting visually observable clinical keywords from pathology reports; second, aligning visual features with these keywords via a Vision-Language meta-teacher (MedSigLIP) to provide dense semantic guidance; and finally, performing adaptive KD where teacher contributions are weighted based on their semantic alignment with the clinical narrative. Experiments on WSI captioning, visual question answering, and slide-level classification tasks demonstrate that an 87M parameter LaGuadia student model matches or exceeds foundation-scale models such as GigaPath and UNI, achieving strong factual consistency and robust generalization. These results highlight clinical language as an effective semantic anchor for building efficient and reliable digital pathology systems. Code is available at https://github.com/hvcl/LaGuadia.