发表机构
Rochester Institute of Technology(罗切斯特理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出基于SAM3与微调Qwen的两阶段框架,无需标注即可实现内镜图像的手术器械实例级分割,在EndoVis数据集上性能优于直接使用SAM3,为无标注分割提供了新方向。
AI 中文摘要
手术器械分割是计算机辅助干预的基础任务,但现有多数方法依赖像素级标注或手动空间提示,限制了可扩展性与自动化程度。最新推出的Segment Anything Model 3(SAM3)提供了通过文本提示实现无标注自动分割的途径,但因存在较大领域差距,器械名称作为文本提示无法直接使用。为克服这些局限,本文提出一种两阶段框架,无需真实掩码或手动交互即可实现实例级分割。第一阶段,利用与自然语言对齐的通用提示词“tool”,通过SAM3的零样本能力生成二值掩码;第二阶段,将这些掩码与视觉语言模型Qwen结合,Qwen经SAM3生成的掩码区域微调后用于器械分类,从而将掩码扩展至实例级。在EndoVis 2017和2018数据集上的评估结果显示,尽管本文的两阶段方法未达到当前全监督方法的性能,但显著优于直接使用SAM3进行器械实例级文本提示分割的效果。总体而言,本文的发现既凸显了SAM3的局限也展现了其潜力,为实现无标注手术器械分割指明了有前景的方向。
英文摘要
Surgical instrument segmentation is a fundamental task for computer-assisted interventions, yet most existing methods rely on pixel-level annotations or manual spatial prompts, which limit scalability and automation. The recently introduced Segment Anything Model 3 (SAM3) offers a pathway to annotation-free, automatic segmentation via text-based prompting; however, the instrument name as a text prompt could not be directly used due to a large domain gap. To overcome these limitations, we propose a two-stage framework that achieves instance-level segmentation without requiring ground truth masks or manual interaction. In the first stage, we leverage a natural-language-aligned generic prompt - "tool" - to produce binary masks using SAM3's zero-shot capability. In the second stage, these masks are extended to instance-level by integrating a vision-language model (Qwen) that is fine-tuned on SAM3-generated masked regions for instrument classification. We evaluate our approach on the EndoVis 2017 and 2018 datasets. Results show that, while our two-stage approach does not reach the performance of current fully supervised methods, it significantly outperforms the direct use of SAM3 for instance-level instrument segmentation with text prompts. Overall, our findings highlight both the limitations and potential of SAM3, suggesting a promising direction toward annotation-free surgical instrument segmentation.
CommentsAccepted at the Medical Image Understanding and Analysis (MIUA)