arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于宫腔镜手术场景分割的自举式视觉-语言模型

Bootstrapping Vision-Language Model for Hysteroscopic Surgical Scene Segmentation

Jun Huang, Meiyi Chen, Zijie Yue, Yuhang Xiao, Fang Li, Hanli Wang, Xiaowen Tong, Yi Guo, Miaojing Shi

arXiv 2608.09302首次发表:更新:

AI 中文总结

本研究提出首个基于VLM的宫腔镜手术场景分割方法VLM-hyster,通过类别特定文本提示与掩码蒸馏分支提升性能,在自行构建的4020张图像数据集上表现优于现有模型,获多中心验证,具临床应用潜力。

AI 中文摘要

宫腔镜手术场景分割对于理解宫腔镜术中环境以及计算机辅助干预具有关键作用。然而,该任务面临独特挑战,不同病灶间形态相似度高,且手术视频中存在镜面反射、运动模糊、液体遮挡等伪影。本研究提出首个基于视觉-语言模型(VLM)的宫腔镜手术场景分割方法,可对宫腔镜手术场景中的15个代表性类别进行像素级定位。我们的VLM-hyster具有分割骨干网络,利用预训练图像编码器提取鲁棒视觉特征,结合基于Transformer的解码器实现密集预测。此外,我们设计了类别特定文本提示,并引入掩码蒸馏分支以过滤与文本提示低相关的视觉特征,使模型能更有效聚焦于类别特定图像区域,从而提升分割性能。我们收集了包含4020张带详细掩码标注的高分辨率图像的多中心宫腔镜手术场景数据集,用于模型训练与评估。实验结果表明,VLM-hyster显著优于现有最先进的AI模型;此外,经妇科医生的广泛评估以及多中心前瞻性验证,证明了VLM-hyster的鲁棒性与泛化性。结果表明,VLM-hyster在实现AI辅助宫腔镜手术中手术器械与病灶定位方面具有巨大潜力。代码可在此httpsURL获取。

英文摘要

Hysteroscopic surgical scene segmentation plays a pivotal role in understanding the hysteroscopic intraoperative environment as well as computer-assisted intervention. However, this task presents unique challenges due to the high morphological similarity among different lesions and the presence of artifacts such as specular reflections, motion blur, and fluid occlusions in surgical videos. In this work, we propose the first vision-language model (VLM)-based hysteroscopic surgical scene segmentation method, which performs pixel-wise localization for fifteen representative categories in hysteroscopic surgical scenes. Our VLM-hyster has a segmentation backbone that utilizes the pretrained image encoder for robust visual feature extraction, coupled with a transformer-based decoder for dense prediction. Moreover, we design category-specific text prompts and incorporate a masked distillation branch to filter out visual features with low correlation to the text prompts, enabling the model to focus more effectively on category-specific image regions and thereby enhancing segmentation performance. We collect a large multicentric hysteroscopic surgical scene dataset, containing 4,020 high-resolution images with detailed mask annotations, for model training and evaluation. Experimental results demonstrate that VLM-hyster substantially outperforms state-of-the-art AI models. Furthermore, extensive assessments by gynecologists, as well as multicentre and prospective validations, demonstrate VLM-hyster's robustness and generalizability. The results suggest that VLM-hyster earns considerable potential in enabling AI-assisted localization of surgical instruments and lesions in hysteroscopic surgeries. Code is available at https://github.com/viscom-tongji/VLM-hyster.

CommentsAccept by Biomedical Signal Processing and Control

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑