arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.29200cs.CV

UltraSAM3:一种用于通用超声图像分割的概念驱动基础模型

UltraSAM3: A Concept-Driven Foundation Model for Universal Ultrasound Image Segmentation

Bo Xu, Quanhao Zhu, Rui Lin, Boling Zhu, Chenyuan Wang, Hongfei Lin, Feng Xia, Chenhua Ji

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有超声分割方法的任务局限或需专家视觉提示问题,提出概念驱动的UltraSAM3模型及指令引导智能体,在多基准测试中性能优于同类模型,提升了临床交互的实用性。

中文摘要 AI 辅助

超声成像因便携性、低成本和实时性在临床实践中愈发普及,这使得超声图像分割变得重要。然而,超声图像与CT、MRI等其他医学成像模态存在显著差异,常受斑点噪声、低对比度、声影和模糊边界影响。现有超声分割方法仍主要局限于任务特定模型或基于视觉提示的基础模型,要么针对特定任务定制,要么需要专家提供视觉提示,不利于灵活临床使用。为应对这些挑战,我们提出UltraSAM3,一种用于通用超声图像分割的概念驱动基础模型。与传统模型不同,UltraSAM3通过将SAM3适配为超声特定的图像-掩码-概念三元组,实现基于文本的目标指定。该模型在涵盖37个公开数据集和13个解剖类别的大规模超声分割语料库上训练,使其能够在不同器官和病变中对齐超声视觉模式与临床有意义的概念。为进一步提升真实临床交互下的可用性,我们提出一种指令引导智能体,将复杂自然语言查询解析为UltraSAM3的简洁超声概念提示。大量实验表明,UltraSAM3在多器官超声基准、外部数据集及视觉提示增强设置中,始终优于代表性的概念驱动和文本驱动生物医学分割模型;此外,该智能体提升了复杂用户指令下的分割鲁棒性。这些结果表明,超声特定的概念适配对构建可泛化、可交互的超声分割基础模型是有效的。

英文摘要

Ultrasound imaging has become increasingly widespread in clinical practice due to its portability, low cost and real-time capability, making ultrasound image segmentation important. However, ultrasound images differ substantially from CT, MRI, and other medical imaging modalities, as they are often affected by speckle noise, low contrast, acoustic shadows and ambiguous boundaries. Existing ultrasound segmentation methods are still mainly limited to task-specific models or visual-prompt-based foundation models, which are either tailored to particular tasks or require expert-provided visual prompts, making them inconvenient for flexible clinical use. To address these challenges, we propose UltraSAM3, a concept-driven foundation model for universal ultrasound image segmentation. Unlike conventional models, UltraSAM3 enables text-based target specification by adapting SAM3 to ultrasound-specific image--mask--concept triplets. The model is trained on a large-scale ultrasound segmentation corpus covering 37 public datasets and 13 anatomical categories, allowing it to align ultrasound visual patterns with clinically meaningful concepts across diverse organs and lesions. To further improve usability under realistic clinical interaction, we propose an instruction-guided agent that parses complex natural language queries into concise ultrasound concept prompts for UltraSAM3. Extensive experiments demonstrate that UltraSAM3 consistently outperforms representative concept- and text-driven biomedical segmentation models on multi-organ ultrasound benchmarks, external datasets, and visual-prompt-enhanced settings. Moreover, the agent improves segmentation robustness for complex user instructions. These results indicate that ultrasound-specific concept adaptation is effective for building generalizable and interactive ultrasound segmentation foundation models.

发表机构

  • Dalian University of Technology Affiliated Center Hospital(大连大学附属中心医院)

机构由 AI 辅助整理,请以论文原文为准。

↑