发表机构
School of Computer Science, Hangzhou Dianzi University(杭州电子科技大学计算机学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有医学视觉-语言预训练的三大瓶颈,本文提出SCALPEL框架,通过临床报告对比微调、非对称对齐策略及解剖-否定感知目标,在三大医学基准上实现了多项任务的最优性能。
AI 中文摘要
视觉-语言预训练(VLP)是医学多模态表示学习的基石。然而,现有医学VLP框架在处理冗长、术语密集的临床报告时,常受限于轻量文本编码器有限的上下文窗口和浅层表示能力。尽管集成医学大语言模型(LLM)能提供前所未有的临床推理能力,但会引入三大瓶颈:(i)生成式LLM在标准对比目标下出现各向同性表示坍塌;(ii)联合端到端训练需大批次,内存开销过高;(iii) vanilla对比损失忽略细粒度解剖学偏侧性和否定修饰词,引发医学幻觉。为应对这些挑战,我们提出SCALPEL,即基于大语言模型驱动的编码器学习的医学视觉-语言语义跨模态对齐框架。首先,临床报告对比微调通过领域特定临床文本适配,将生成式LLM转换为各向同性编码器;其次,采用非对称对齐策略,利用离线特征缓存实现高效训练;关键是,我们提出解剖-否定感知目标,明确惩罚涉及偏侧性混淆或错误否定的不匹配图像-文本对。在MIMIC-CXR、CheXpert和IU X-Ray基准上的大量实验表明,SCALPEL在跨模态检索、零样本疾病分类和医学视觉问答任务中达到了当前最优性能。
英文摘要
Vision-language pre-training (VLP) serves as a cornerstone for medical multimodal representation learning. However, existing medical VLP frameworks are often constrained by the limited context windows and shallow representational capacities of lightweight text encoders when processing lengthy, terminology-dense clinical reports. While integrating medical large language models (LLMs) offers unprecedented clinical reasoning capabilities, it introduces three major bottlenecks: (i) the anisotropic representational collapse of generative LLMs under standard contrastive objectives, (ii) the prohibitive memory overhead of joint end-to-end training with large batch sizes, and (iii) the medical hallucinations induced by vanilla contrastive losses that ignore fine-grained anatomical laterality and negation modifiers. To address these challenges, we propose \textbf{SCALPEL}, a \textbf{S}emantic \textbf{C}ross-modal \textbf{A}lignment framework via \textbf{L}LM-\textbf{P}owered \textbf{E}ncoder \textbf{L}earning. First, Clinical Report Contrastive fine-tuning converts a generative LLM into an isotropic encoder via domain-specific clinical text adaptation. Second, an asymmetric alignment strategy leverages offline feature caching to enable efficient training. Critically, we formulate an Anatomy-Negation Aware Objective that explicitly penalizes mismatched image-text pairs involving laterality confusion or false negations. Extensive experiments across MIMIC-CXR, CheXpert, and IU X-Ray benchmarks demonstrate that SCALPEL achieves state-of-the-art performance in cross-modal retrieval, zero-shot disease classification and medical visual question answering.
Comments14 pages, 3 figures, accepted by PRCV2026