arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.04676cs.CV

SurgNarrator:面向手术视频理解的生成式检索框架

SurgNarrator: A Generative Retrieval Framework for Surgical Video Understanding

Yuqing Feng, Jiawei Ma, Kevin Qinghong Lin, Kun Yuan, Nicolas Padoy, Daniel S. Elson, Anh Nguyen, Stamatia Giannarou, Baoru Huang

首次发表
浏览论文内容

中文总结 AI 辅助

SurgNarrator是专为手术视频理解设计的生成式检索框架,通过构建手术词汇表、适配Qwen3-VL-Embedding-8B并采用分层检索策略,在零样本设置下于12个基准上实现性能提升且延迟大幅降低。

中文摘要 AI 辅助

手术流程以结构化且重复的临床事件形式展开,通过术中手术视频对其进行实时理解,对术中决策及支持至关重要。然而,现有视频理解方法存在权衡问题:自回归视频-语言模型支持全面推理,但对时间敏感的临床应用而言不实用;对比模型延迟低,但难以处理复杂场景理解。最近,生成式检索已被用于通用领域的视频理解,但将其迁移至手术领域并非易事,因为近乎相同的视觉外观可能对应语义不同的事件,且涉及的术语具有高度手术特异性。为此,我们提出SurgNarrator,一种专为手术视频理解定制的新型生成式检索框架。我们从手术标注文本中构建精心整理的手术中心词汇表,以定义具有临床意义的检索空间;随后对预训练的Qwen3-VL-Embedding-8B进行适配,通过时间感知的对比目标学习具有区分性的临床表征。推理阶段,采用分层、流程感知的检索策略将搜索空间缩小至相关流程类型,实现快速且有效的响应。我们在12个基准上以零样本设置对方法进行全面评估,相较于现有最优基线取得了持续的性能提升,同时与生成式基线相比,输出阶段延迟降低了两个数量级以上。

英文摘要

Surgical procedures unfold as structured and recurring clinical events, whose real-time understanding via intraoperative surgical videos is critical for intraoperative decision-making and support. However, existing video understanding methods force a trade-off: autoregressive video-language models support comprehensive reasoning but are not practical for time-sensitive clinical applications, whereas contrastive models offer low latency but struggle with complex scene understanding. Recently, generative retrieval has been explored for general-domain video understanding, but transferring it to surgery is not trivial because near-identical visual appearances may indicate semantically distinct events, and the terminology involved is highly surgery-specific. To this end, we propose SurgNarrator, a new generative retrieval framework tailored for surgical video understanding. We construct a well-curated surgery-centric vocabulary from surgical captions to define a clinically meaningful retrieval space. We then adapt the pre-trained Qwen3-VL-Embedding-8B to learn discriminative clinical representations with a temporally-aware contrastive objective. During inference, a hierarchical, procedure-aware retrieval strategy narrows the search space to the relevant procedure type, delivering fast and effective responses. Our method is comprehensively evaluated on twelve benchmarks in a zero-shot setting and achieves consistent performance gains over state-of-the-art baselines, while reducing output-stage latency by more than two orders of magnitude compared with the generative baseline.

发表机构

  • University of Liverpool(利物浦大学)
  • City University of Hong Kong(香港城市大学)
  • University of Oxford(牛津大学)
  • University of Strasbourg(斯特拉斯堡大学)
  • IHU Strasbourg(斯特拉斯堡大学医院研究所)
  • Imperial College London(帝国理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑