UniMedSeg:用于多范式2D/3D医学图像分割的统一上下文学习
UniMedSeg: Unified In-Context Learning for Multi-Paradigm 2D/3D Medical Image Segmentation
- Harbin Institute of Technology at Shenzhen(哈尔滨工业大学(深圳))
- Peng Cheng Laboratory(鹏城实验室)
- University Hospital Tübingen(图宾根大学医院)
- German Center for Mental Health(德国心理健康中心)
- Chinese University of Hong Kong(香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究针对医学图像分割基础模型存在的问题,提出以Transformer为中心的UniMedSeg框架,通过映射多种信息到共享序列空间联合学习异构医学监督,引入解耦分割注意力克服内存瓶颈,在多类分割任务中实现最优性能且无需特定任务微调。
AI中文摘要:
医学图像分割基础模型期望能在不同临床场景中通用,但现有通用方法因提示范式和空间维度而碎片化。视觉上下文学习、交互式分割和语言引导分割通常由特定范式模型处理,2D和3D图像也单独建模。这种隔离阻碍了异构注释和数据被单个可扩展模型联合吸收,限制了跨范式知识转移。为解决此瓶颈,我们提出UniMedSeg,一个以Transformer为中心的通用分割框架,将视觉示例、几何交互、语言指令和2D/3D图像映射到共享序列空间,通过统一上下文接口联合学习异构医学监督,无需特定提示或维度分支。为克服视觉上下文导致的长序列内存瓶颈,我们引入解耦分割注意力,将注意力复杂度降至线性,同时保持硬件友好计算和聚焦的上下文-目标交互。在从27个公共数据集策划的大型语料库上进行广泛训练和评估,UniMedSeg在视觉上下文、交互式和语言引导分割中实现了无特定任务微调的最优性能,证明了在不同保留任务上的强大泛化能力。代码和模型权重可在指定网址公开获取。
英文摘要:
Medical image segmentation foundation models are expected to generalize across diverse clinical scenarios, yet existing universal methods remain fragmented by prompt paradigms and spatial dimensions. Visual in-context learning, interactive segmentation, and language-guided segmentation are typically handled by paradigm-specific models, while 2D and 3D images are also modeled separately. Such isolation prevents heterogeneous annotations and data from being jointly absorbed by a single scalable model and limits cross-paradigm knowledge transfer. To address this bottleneck, we propose UniMedSeg, a Transformer-centric universal segmentation framework that maps visual examples, geometric interactions, language instructions, and 2D/3D images into a shared sequence space, enabling heterogeneous medical supervision to be jointly learned through a unified in-context interface without prompt- or dimension-specific branches. To overcome the long-sequence memory bottleneck caused by visual contexts, we introduce Decoupled Split Attention, which reduces attention complexity to linear while preserving hardware-friendly computation and focused context-target interaction. Extensively trained and evaluated on a large corpus curated from 27 public datasets, UniMedSeg achieves state-of-the-art performance across visual in-context, interactive, and language-guided segmentation without task-specific fine-tuning, demonstrating strong generalization on diverse held-out tasks. The code and model weights are publicly available at https://github.com/Lii1228/UniMedSeg