MedPlex:用于临床基础医学分割的深度视觉-语言协同适配
MedPlex: Deep Vision-Language Co-Adaptation for Clinically Grounded Medical Segmentation
浏览论文内容
中文总结 AI 辅助
MedPlex是一种端到端VLM框架,通过双向融合和两级概念对齐,实现医学图像分割的视觉-语言协同适配,在CT、MR的多类医学分割任务中达到最优性能。
中文摘要 AI 辅助
医学图像分割在很大程度上仍被视为仅依赖视觉的问题,尽管临床解读通常依赖于解剖结构、位置、外观及周围语境的文本知识。现有视觉-语言模型(VLM)范式下的文本引导分割方法,常将语言仅作为后期的条件信号,限制了其对视觉表征学习的影响。本文提出MedPlex(Medical Plexus of Vision and Language,医学视觉与语言 plexus 框架),这是一种端到端的VLM框架,使文本引导成为分割学习中连续的、基于临床的组成部分。通过双向融合(Bi-Fusion),视觉和文本表征在编码层级中协同演化。MedPlex进一步引入类别级和区域级概念对齐,以互补粒度组织共享表征:类别级对齐将每个解剖目标锚定到聚合的临床概念轮廓,区域级对齐则通过类别特定的视觉证据保留形状、位置、外观和纹理等个体概念。如此,语言在整个编码器中提供结构化监督,而非仅作为后期线索。MedPlex在CT和MR基准的多器官、心脏亚结构及肿瘤分割任务中达到了当前最优性能,包括使用真实自由文本临床监督的设置。代码:this https URL。
英文摘要
Medical image segmentation is still largely treated as a vision-only problem, although clinical interpretation often relies on textual knowledge of anatomy, location, appearance, and surrounding context. Existing text-guided segmentation methods within the Vision-Language Model (VLM) paradigm often use language only as a late conditioning signal, limiting its influence on visual representation learning. We introduce MedPlex (Medical Plexus of Vision and Language), an end-to-end VLM framework that makes text guidance a continuous, clinically grounded component of segmentation learning. Through Bi-Fusion (Bidirectional Fusion), visual and textual representations evolve jointly across the encoding hierarchy. MedPlex further introduces class-level and region-level concept alignment to organize the shared representation at complementary granularities. Class-level alignment anchors each anatomical target to an aggregated clinical concept profile, while region-level alignment preserves individual concepts, such as shape, location, appearance, and texture, through class-specific visual evidence. In this way, language provides structured supervision throughout the encoder rather than serving only as a late-stage cue. MedPlex achieves state-of-the-art performance across CT and MR benchmarks for multi-organ, cardiac substructure, and tumor segmentation, including settings with real free-text clinical supervision. Code: https://github.com/rafiibnsultan/MedPlex.
发表机构
- Wayne State University(韦恩州立大学)
- Henry Ford Health(亨利福特医疗集团)
- Institute for AI and Data Science Wayne State University(韦恩州立大学人工智能与数据科学研究院)
机构由 AI 辅助整理,请以论文原文为准。