arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24671cs.CV

ReGround-Surg:用于指代手术视频分割的可靠性引导锚点 grounding

ReGround-Surg: Reliability-Guided Anchor Grounding for Referring Surgical Video Segmentation

Jiaxin Wen, Ming Yin, Lu Liu, Zeyu Fu

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对手术视频分割中初始锚点错误导致跟踪误差传播的问题,提出 ReGround-Surg 框架,通过可靠性引导锚点 grounding 改进 SAM2,在 Ref-EndoVis 数据集上实现优于 SOTA 的性能且速度损失极小。

中文摘要 AI 辅助

指代手术视频分割要求根据自然语言表达式在视频帧中分割目标器械或组织区域。近期基于 Segment Anything Model 2(SAM2)的两阶段方法(如 ReSurgSAM2)先在初始或选定帧中 grounding 指代目标,再通过跟踪传播选定掩码。尽管有效,其性能高度依赖初始 grounding 掩码的质量:一旦选择错误锚点,后续跟踪易传播错误。由于手术视频中存在视觉相似器械、遮挡及复杂组织-器械交互,该问题尤为棘手。为解决此问题,本文提出 ReGround-Surg,一种轻量型可靠性引导锚点 grounding 框架,用于改进基于 SAM2 的指代手术视频分割。该框架首先从指代表达式和当前帧视觉特征预测文本条件空间可靠性图,该图随后被用于两个互补分支:门控侧适配器在文本-视觉融合前增强与表达式相关的视觉区域,而可靠性加权视觉-文本注意力模块在提示标记聚合时抑制偏离目标的视觉证据。在 Ref-EndoVis17 和 Ref-EndoVis18 上的实验表明,在三个评估拆分中,该方法均优于现有最优方法,且速度降低可忽略不计。代码公开于此 https URL。

英文摘要

Referring surgical video segmentation requires segmenting a target instrument or tissue region across video frames according to a natural language expression. Recent Segment Anything Model 2 (SAM2) based two-stage methods (e.g., ReSurgSAM2) first ground the referred target in an initial or selected frame, then propagate the selected mask via tracking. Although effective, their performance is highly sensitive to the quality of the initial grounded mask: once an incorrect anchor is selected, subsequent tracking tends to propagate the error. This issue is especially challenging in surgical videos due to visually similar instruments, occlusion, and complex tissue-tool interactions. To address this issue, we propose ReGround-Surg, a lightweight reliability-guided anchor grounding framework to improve SAM2-based referring surgical video segmentation. It first predicts a text-conditioned spatial reliability map from the referring expression and current-frame visual features. The map is then reused in two complementary branches: a Gated Side Adapter enhances expression-relevant visual regions before text-to-vision fusion, while a Reliability-Weighted Vision-to-Text Attention module suppresses off-target visual evidence during prompt-token aggregation. Experiments on Ref-EndoVis17 and Ref-EndoVis18 show consistent improvements over state-of-the-art methods across three evaluation splits with negligible speed reduction. Code is publicly available at https://github.com/JiaxinWen1/ReGround-Surg.

发表机构

  • University of Exeter(埃克塞特大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑