发表机构
Lahore University of Management Sciences; The University of Western Australia; Information Technology University of the Punjab(拉合尔管理科学大学; 西澳大学; 旁遮普信息技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对医学视觉语言模型在解剖结构和空间定位上的不足,提出定位透镜增强方法,通过数据、架构和对齐三层面改进,提升Med-VQA准确率最高达6.2%,且模型无关、可无缝集成。
AI 中文摘要
医学视觉语言模型(Med-VLMs)在临床任务中展现了强大的能力。然而,它们常常难以理解解剖结构和空间定位,而这些对于医学推理至关重要。为了解决这一问题,我们提出了一种面向定位感知的Med-VLM流水线增强方法,在三个层面引入改进:数据、架构和对齐。首先,我们引入了定位透镜(localization lens),这是一组经过专家验证的表示,能够提供更丰富的解剖和位置上下文。然而,由于这些表示增加了输入复杂度,我们在模型架构中集成了像素重排(pixel shuffle),以过滤和细化表示,增强空间信息处理,同时保持解剖连续性。最后,为了有效地将定位透镜表示与文本特征对齐,我们在标准损失函数之外引入了解耦对比损失(DCL)。这确保了更好的特征区分性和鲁棒性,尤其是在数据有限的医疗环境中。通过在医学视觉问答(Med-VQA)数据集上的广泛评估,我们展示了我们的方法在不同Med-VLM架构上提升了定位驱动的性能。我们对基于定位的问题的分析进一步揭示,解剖和空间推理的改进直接提升了Med-VQA的整体准确率,最高达6.2%。所提出的方法是模型无关的,可以无缝集成到现有的Med-VLM流水线中。数据集、代码和训练模型将在该https URL上公开提供。
英文摘要
Medical Vision-Language Models (Med-VLMs) have demonstrated strong capabilities in clinical tasks. However, they often struggle to understand anatomical structures and spatial positioning, which are crucial for medical reasoning. To address this, we propose a localization-aware enhancement to the Med-VLM pipeline, introducing improvements at three levels: data,architecture, and alignment. First, we introduce localization lens, a set of expert-validated representations that provide richer anatomical and positional context. However, as these representations increase input complexity, we integrate pixel shuffle within the model architecture to filter and refine representations, enhancing spatial information processing while preserving anatomical continuity. Lastly, to effectively align the localization lens representations with textual features, we incorporate decoupled contrastive loss (DCL) alongside the standard loss function. This ensures better feature discrimination and robustness, particularly in data limited medical settings. Through extensive evaluations on medical visual question answering (Med-VQA) datasets, we show that our methodology improves localization-driven performance across different Med-VLM architectures. Our analysis of localization-based questions further reveals that improvements in anatomy and spatial reasoning directly enhance the overall accuracy of Med-VQA upto 6.2%. The proposed approach is model-agnostic and can be seamlessly integrated into existing Med-VLM pipelines. The dataset, code, and trained models will be made publicly available at https://github.com/CVLABLUMS/localizationlens.
Comments10 pages, 1 figure, Medical Image Computing and Computer Assisted Intervention (MICCAI)