发表机构
Advanced Technology Group; GE HealthCare(先进技术集团; GE医疗)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对分割基础模型在临床任务中表现不足的问题,提出FS-CPL方法,通过视觉定位学习概念提示,在四个医学影像基准上取得显著性能提升,且与骨干网络无关。
AI 中文摘要
可提示的分割基础模型(FMs)如SAM3和Medical SAM3,承诺通过自然语言接口实现医学影像的少样本、交互式指定分割,但它们在临床任务上的表现远未达到这一承诺。我们认为,这一不足并非源于医学预训练不足或提示措辞不完善,而是一种结构性局限,在配对图像-文本监督稀缺的领域(如大多数临床模态)中,这种局限将持续存在。我们进一步假设,该局限具体针对自然语言作为控制信号:直接从目标分布学习的视觉定位提示,应能在无需额外图像-文本数据或骨干网络再训练的情况下恢复损失的性能。我们提出少样本概念提示学习(FS-CPL),该方法通过掩码监督从包含K个图像-掩码对的小支持集中学习连续概念提示嵌入p*∈R^(T×d),且编码器-解码器骨干网络保持冻结。在涵盖超声和内窥镜检查的四个公共基准(BUSI、HC18、TN3K、CVC-Clinic)上,FS-CPL相较于常规文本提示实现了最高+0.62的绝对Dice系数提升,且它与骨干网络无关:它同时提升了普通SAM3和经特定领域预训练的Medical SAM3的性能,表明视觉概念提示与领域内预训练具有互补性。
英文摘要
Promptable segmentation foundation models (FMs) such as SAM3 and Medical SAM3 promise few-shot, interactively-specified segmentation for medical imaging through a natural language interface, yet their performance on clinical tasks falls well short of this promise. We posit that this shortfall is not an artefact of insufficient medical pretraining or imperfect prompt phrasing, but a structural limitation that will persist in any domain where paired image-text supervision is scarce, as it is across most clinical modalities. We further hypothesize that the limitation is specific to natural language as a control signal: a visually grounded prompt, learned directly from the target distribution, should recover the lost performance without additional image-text data or backbone retraining. We propose Few-Shot Concept Prompt Learning (FS-CPL), which learns a continuous concept prompt embedding $\mathbf{p}^* \in \mathbb{R}^{T \times d}$ from a small support set of $K$ image--mask pairs via mask supervision, with the encoder-decoder backbone frozen. Across four public benchmarks spanning ultrasound and endoscopy (BUSI, HC18, TN3K, CVC-Clinic), FS-CPL delivers absolute Dice improvements of up to $+0.62$ over canonical text prompts and is \emph{backbone-agnostic}: it lifts both vanilla SAM3 and the domain-specifically pretrained Medical SAM3, showing that visual concept prompting is complementary to in-domain pretraining.