arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23860cs.AI

面向多模态大语言模型的医学视觉编码器预训练

Pretraining of Medical Visual Encoders Toward Multi-modal Large Language Models

Tianyou Jiang

首次发表
浏览论文内容

中文总结 AI 辅助

针对MLLM中视觉编码器与自回归LLM的语义接口差距,提出MedMLIP框架,通过报告生成预训练视觉编码器并采用局部关系蒸馏避免视觉崩溃,实验证明跨LLM迁移的有效性。

中文摘要 AI 辅助

多模态大语言模型(MLLMs)通常复用通过CLIP预训练的视觉编码器,尽管这些ViT的特征最终由自回归大语言模型(LLMs)消费。我们将这种不匹配称为语义接口差距,并引入MedMLIP框架,该框架通过冻结的LLM进行报告生成来预训练视觉编码器,同时采用局部关系蒸馏(LRD)以保留视觉补丁之间的关系,避免视觉崩溃。我们在IU-Xray和Open-PMC-300K上预训练MedMLIP,并在VQA-RAD和SLAKE上评估所得编码器。仅迁移ViT,而指导LLM和投影器被替换,从而评估跨LLM的可迁移性。我们的跨LLM迁移实验证明了针对自回归LLM接口预训练视觉编码器同时努力保留更细粒度视觉信息的价值。代码和预训练模型可在该https URL获取。

英文摘要

Multimodal Large Language Models (MLLMs) commonly reuse visual encoders pretrained with CLIP, although the features of these ViTs are ultimately consumed by autoregressive LLMs. We refer to this mismatch as the semantic-interface gap and introduce MedMLIP, a framework that pretrains the visual encoder through report generation with a frozen LLM, while employing Local Relational Distillation (LRD) to preserve relationships among visual patches to avoid visual collapse. We pretrain MedMLIP on IU-Xray and Open-PMC-300K and evaluate the resulting encoders on VQA-RAD and SLAKE. Only the ViT is transferred, while the guiding LLM and projector are replaced, allowing us to assess cross-LLM transferability. Our cross-LLM transfer experiments demonstrate the value of pretraining visual encoders for their autoregressive LLM interface while trying to preserve more fine-grained visual information. Code and the pretrained model are available at https://github.com/SkyCol/MedMLIP

发表机构

  • University of Bern(伯尔尼大学)

机构由 AI 辅助整理,请以论文原文为准。

↑