arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于手术器械分割的SAM分层原型-记忆适应方法

Hierarchical Prototype-Memory Adaptation of SAM for Surgical Instrument Segmentation

Xinning Yao, Jingjing Wang, Jinghua Yue, Xiaoyan Luo, Fugen Zhou, Bo Liu

arXiv 2608.24541首次发表:更新:

发表机构

Beihang University; Image Processing Center(北京航空航天大学; 图像处理中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对SAM适配手术器械分割时的鲁棒性与尺度耦合问题,本文提出HPMA框架,通过分层原型记忆与尺度匹配耦合机制,在EndoVis2017/2018数据集上取得最优性能。

AI 中文摘要

手术器械分割(SIS)是计算机辅助手术的基础,可靠的器械掩码可实现精准的场景理解与临床辅助。近期通过提示学习将基础模型如分割一切模型(SAM)适配到外科领域已取得令人鼓舞的结果,但这些适配模型在复杂外科场景下的性能受限于次优的适配机制:仅通过下游分割损失优化提示或原型易使其退化为任务特定参数,而非作为持久稳定的类别记忆,从而降低其对术中复杂变化的鲁棒性;此外,通过单一提示通路传递多尺度视觉线索会形成瓶颈,阻碍有效的尺度匹配耦合。为解决这些局限,本文提出HPMA,一种用于SAM的分层原型-记忆适应框架。具体而言,HPMA从标注的外科场景构建冻结的多尺度视觉原型记忆库,并通过轻量适配器将其集成到SAM的特征空间以保留稳定的类别证据;为最大化多尺度线索的效用,引入尺度匹配耦合机制:全局原型校准类别级提示特征,结构原型引导解码器对象查询,局部原型通过局部对齐目标对齐高分辨率特征图。在公开的EndoVis2017与EndoVis2018数据集上的大量实验表明,本文方法实现了最优性能,优于现有的基础模型适配方法。

英文摘要

Surgical instrument segmentation (SIS) is fundamental for computer-assisted surgery, where reliable instrument masks enable precise scene understanding and clinical assistance. Recently, adapting foundation models like the Segment Anything Model (SAM) to the surgical domain via prompt-learning has shown encouraging results. However, the performance of these adapted models under challenging surgical conditions is constrained by suboptimal adaptation mechanisms. Specifically, optimizing prompts or prototypes purely via downstream segmentation loss tends to cause them to degenerate into task-specific parameters rather than serving as persistent, stable category memory, thereby degrading their robustness against complex intraoperative variations. Moreover, routing multi-scale visual cues through a single prompt pathway creates a bottleneck that hinders effective scale-matched coupling. To address these limitations, we propose HPMA, a Hierarchical Prototype-Memory Adaptation framework for SAM. Specifically, HPMA constructs a frozen, multi-scale visual prototype memory bank from annotated surgical scenes and integrates it into SAM's feature space using lightweight adapters to preserve stable category evidence. To maximize the utility of multi-scale cues, we introduce a scale-matched coupling mechanism where global prototypes calibrate class-level prompt features, structural prototypes guide decoder object queries, and local prototypes align high-resolution feature maps through a local alignment objective. Extensive experiments on the public EndoVis2017 and EndoVis2018 datasets demonstrate that our approach achieves state-of-the-art performance, outperforming existing foundation model adaptation methods.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑