arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.02663cs.CVcs.AI

通过证据解耦表征医学视觉-语言分割中的文本分支敏感性

Characterizing Text Branch Sensitivity in Medical Vision-Language Segmentation via Evidence Decoupling

发表机构西南科技大学信息与控制工程学院
查看机构详情
  • School of Information and Control Engineering, Southwest University of Science and Technology(西南科技大学信息与控制工程学院)

机构由 AI 辅助整理,请以论文原文为准。

Ziquan Liu, Zhewei Zhu, Xuyang Shi

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过提出证据解耦解码器(EDD),探究医学视觉-语言分割中对文本分支的敏感性,发现不同数据集对文本扰动的依赖程度存在显著差异,为多模态医学图像分割的模型设计提供了实用参考。

中文摘要 AI 辅助

预训练视觉-语言模型(VLMs)通过结合临床文本,在医学图像分割任务中展现出良好性能,但目前仍不清楚文本信息对像素级预测的实际贡献程度。本研究系统探究了文本在多模态医学图像分割中的作用,首先分析了几种常用的融合策略,发现分割性能对融合模块的选择在很大程度上不敏感。为进一步理解模态间的交互作用,我们提出了一种基于证据深度学习和深度监督的证据解耦解码器(EDD),该解码器可作为内部表征分析工具,在解码过程中分解图像证据和文本调制证据,同时保持具有竞争力的分割性能。实验结果表明,不同数据集对文本扰动的敏感性存在显著差异:在BUSI和BTMRI数据集上,移除文本会导致性能灾难性下降,表明模型强烈依赖文本输入;在ISIC和Kvasir-SEG数据集上,文本的影响相对较小。我们还发现,文本主要通过全局语义调制而非独立空间定位影响预测,且驱动文本敏感性的特定语义组件在不同数据集上存在差异。这些发现为深入理解多模态医学图像分割中的模态交互提供了更深刻的认识,并为未来的模型设计提供了实用见解。

英文摘要

Pretrained vision-language models (VLMs) have shown promising performance in medical image segmentation by incorporating clinical text. However, it remains unclear how much textual information actually contributes to pixel-level predictions. In this work, we systematically investigate the role of text in multimodal medical image segmentation. We first analyze several commonly used fusion strategies and find that segmentation performance is largely insensitive to the choice of fusion module. To further understand modality interactions, we propose an Evidence Decoupling Decoder (EDD) based on evidential deep learning and deep supervision. EDD serves as an internal representation analysis tool that decomposes image evidence and text-modulated evidence throughout the decoding process while maintaining competitive segmentation performance. Experimental results show that the sensitivity to text perturbation varies substantially across datasets. On BUSI and BTMRI, removing text causes catastrophic performance drops, indicating strong model reliance on textual input. On ISIC and Kvasir-SEG, text exerts relatively marginal influence. We further find that text affects predictions mainly through global semantic modulation rather than independent spatial localization, and that the specific semantic components driving text sensitivity differ across datasets. These findings provide a deeper understanding of modality interaction in multimodal medical image segmentation and offer practical insights for future model design.

补充信息

↑