医学视觉语言模型如何未能充分利用其视觉编码器:皮肤病学视角
How Medical VLMs Underutilize Their Vision Encoders: A Dermatology Perspective
- Tulane University(杜兰大学)
- Oak Ridge National Laboratory(橡树岭国家实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究针对医学视觉语言模型在皮肤病学中未充分利用视觉编码器的问题,提出结合无标签提示与低标签编码器辅助重排序的干预方法,并通过机制分析验证其有效性。
AI中文摘要:
医学视觉语言模型(VLMs)在临床图像理解方面展现出显著前景,能够提供具有可解释推理的准确诊断。然而,其强大的视觉编码器与完整多模态模型之间存在关键的性能差距:在皮肤病学中,即使两者均使用零目标任务标签,MedSigLIP编码器的性能平均比MedGemma高出10.26个百分点;少样本线性探测进一步提供了强视觉表征的证据。这一差距促使我们研究视觉信息在端到端诊断中如何被使用,以及为何看似合理的预测可能缺乏图像证据的支撑。以皮肤病学作为主要测试平台,我们系统性地研究了这一现象的三种假设。我们进一步对模型内部注意力模式进行了机制性分析,表明一种简单的“先描述后决策”提示策略在生成过程中将视觉注意力提高了30-40%。任务特定的微调改善了皮肤病学分类,但在我们的评估中降低了跨领域医学问答性能。为应对这些挑战,我们将无标签提示与低标签编码器辅助重排序相结合,同时保持VLM冻结。我们在皮肤病学中跨五个VLM骨干验证了这些干预措施,并提供了跨额外医学模态的支持性表征和注意力分析。
英文摘要:
Medical Vision-Language Models (VLMs) show significant promise for clinical image understanding, offering accurate diagnosis with interpretable reasoning. However, a critical performance gap exists between their strong vision encoders and the full multimodal model: in dermatology, the MedSigLIP encoder outperforms MedGemma by an average of 10.26 percentage points even when both use zero target-task labels; few-shot linear probing provides further evidence of strong visual representations. This gap motivates an investigation of how visual information is used in end-to-end diagnosis and why plausible-sounding predictions can lack grounding in image evidence. Using dermatology as our primary testbed, we systematically investigate three hypotheses for this phenomenon. We further provide a mechanistic analysis of the model's internal attention patterns, showing that a simple describe-then-decide prompting strategy increases vision attention by 30-40% during generation. Task-specific fine-tuning improves dermatology classification but reduces cross-domain medical question-answering performance in our evaluation. To address these challenges, we combine label-free prompting with low-label encoder-assisted reranking while keeping the VLM frozen. We validate the interventions across five VLM backbones in dermatology and provide supporting representation and attention analyses across additional medical modalities.