Lumen:面向零样本计算病理学的预训练视觉与语言编码器参数高效对齐
Lumen: Parameter-Efficient Alignment of Pretrained Vision and Language Encoders for Zero-Shot Computational Pathology
浏览论文内容
中文总结 AI 辅助
Lumen通过低秩适配器对齐冻结的Virchow2和BioMedBERT,仅训练0.40%参数,在零样本病理学基准上取得领先性能,验证了参数高效对齐的有效性。
中文摘要 AI 辅助
病理学视觉语言模型通常通过在成对的图像-标题数据上预训练或微调大型编码器来构建。我们探究了是否可以通过对冻结的单模态基础模型进行参数高效对齐来组装病理学视觉语言模型,从而保持其预训练表示不变。为此,我们提出了Lumen,它使用秩为4的适配器和投影头对齐冻结的Virchow2和BioMedBERT骨干网络,仅在公开的QUILT-1M语料库上训练总参数的0.40%。在九个公开的零样本补丁基准上,Lumen取得了最高的平均机会校正平衡准确率,为0.546,而最强基线为0.461(配对差异0.086,95%置信区间0.042-0.136)。在淋巴结转移检测中,Lumen在4,214张内部保留切片上达到了0.964的AUROC(95%置信区间0.956-0.971),在跨越九个外部队列和六个器官的2,368张切片上达到了0.955的AUROC(95%置信区间0.942-0.966)。在内部校准阈值下,它超越了所有视觉语言基线,内部平衡准确率为0.909(95%置信区间0.896-0.923),外部为0.915(95%置信区间0.902-0.929)。Lumen在各项评估中表现优异,但跨模态检索除外,在该任务中排名第三,仅次于CONCH和PathGen-L/14。对两个编码器进行完全微调并未给Lumen带来相对于低秩适应的一致优势,尽管它改善了检索。因此,对齐冻结的单模态基础模型在补丁和切片级别上产生了强大且可迁移的性能,同时仅训练了总参数的一小部分。
英文摘要
Pathology vision-language models are commonly built by pretraining or fine-tuning large encoders on paired image-caption data. We asked whether a pathology vision-language model can instead be assembled by parameter-efficient alignment of frozen unimodal foundation models, leaving their pretrained representations untouched. Here we present Lumen, which aligns frozen Virchow2 and BioMedBERT backbones using rank-4 adapters and projection heads, training only 0.40% of the total parameters on the public QUILT-1M corpus. Across nine public zero-shot patch benchmarks, Lumen achieved the highest mean chance-corrected balanced accuracy, 0.546 versus 0.461 for the strongest baseline (paired difference 0.086, 95% CI 0.042-0.136). On lymph-node metastasis detection, Lumen reached an AUROC of 0.964 (95% CI 0.956-0.971) on 4,214 held-out internal slides and 0.955 (95% CI 0.942-0.966) on 2,368 slides across nine external cohorts and six organs. At the internally calibrated threshold, it outperformed all vision-language baselines, with a balanced accuracy of 0.909 (95% CI 0.896-0.923) internally and 0.915 (95% CI 0.902-0.929) externally. Lumen performed competitively across the evaluations, with the exception of cross-modal retrieval, where it ranked third behind CONCH and PathGen-L/14. Fully fine-tuning both encoders gave Lumen no consistent benefit over low-rank adaptation, although it improved retrieval. Aligning frozen unimodal foundation models therefore yields strong and transferable performance at patch and slide level while training only a small fraction of the parameters.
发表机构
- University of Bern(伯尔尼大学)
- University Hospital Cologne(科隆大学医院)
- University Hospital of Bern(伯尔尼大学医院)
机构由 AI 辅助整理,请以论文原文为准。