发表机构
University of Nottingham Ningbo China; University of Nottingham(宁波诺丁汉大学; 诺丁汉大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对医学图像分割中视觉-语言跨模态对齐的不确定性问题,提出DistMedVL概率框架,含PCM-Adapter等模块,在8个基准上以6.3M参数实现优于SOTA的性能,数据效率等更优。
AI 中文摘要
视觉与文本表征的跨模态对齐是多模态医学图像理解的基础,但在真实临床场景下,两种模态均存在不确定性,阻碍了该领域的发展。现有视觉-语言分割方法依赖确定性跨模态匹配,忽略了模糊边界带来的偶然不确定性以及训练数据有限导致的认知不确定性,在域偏移场景下性能脆弱。为解决这一问题,我们提出DistMedVL,这是一种概率视觉-语言框架,在冻结编码器的基础上引入轻量级概率跨模态适配器(PCM-Adapter),以显式建模表征不确定性。具体而言,PCM-Adapter包含两个顺序模块,用于渐进式概率对齐:我们首先设计了马氏距离对齐模块(MAM),将文本标记建模为高斯分布,并通过马氏距离计算图像块-文本兼容性,得到方差条件匹配结果,从而降低不可靠特征维度的权重;此外,我们设计了分布流模块(DFM),用于估计模态级置信参数,并执行视觉引导的文本分布优化,以适配不同成像模态间的分布差异。在8个医学分割基准上开展的大量实验表明,DistMedVL仅需630万可训练参数,即可优于现有最先进方法,展现出更优的数据效率、扰动鲁棒性及跨数据集泛化能力。
英文摘要
Cross-modal alignment of visual and textual representations is fundamental to multimodal medical image understanding, yet remains hindered by uncertainty in both modalities under real-world clinical conditions. Existing vision-language segmentation methods rely on deterministic cross-modal matching, which overlooks aleatoric uncertainty from ambiguous boundaries and epistemic uncertainty from limited training data, leading to fragile performance under domain shift. To address this issue, we propose DistMedVL, a probabilistic vision-language framework that introduces a lightweight Probabilistic Cross-Modal Adapter (PCM-Adapter) upon frozen encoders to explicitly model representational uncertainty. Specifically, the PCM-Adapter comprises two sequential modules for progressive probabilistic alignment. We first devise a Mahalanobis Alignment Module (MAM) that models textual tokens as Gaussian distributions and computes patch-text compatibility via Mahalanobis distance, yielding variance-conditioned matching that downweights unreliable feature dimensions. Moreover, we devise a Distribution Flow Module (DFM) that estimates modality-wise confidence parameters and performs vision-guided refinement of textual distributions, accommodating distributional variation across imaging modalities. Extensive experiments across eight medical segmentation benchmarks demonstrate that DistMedVL outperforms state-of-the-art methods with only 6.3M trainable parameters, exhibiting superior data efficiency, perturbation robustness and cross-dataset generalization.
Comments10 pages, 5 figures