SurgNarrator:面向手术视频理解的生成式检索框架
SurgNarrator: A Generative Retrieval Framework for Surgical Video Understanding
浏览论文内容
中文总结 AI 辅助
SurgNarrator是专为手术视频理解设计的生成式检索框架,通过构建手术词汇表、适配Qwen3-VL-Embedding-8B并采用分层检索策略,在零样本设置下于12个基准上实现性能提升且延迟大幅降低。
中文摘要 AI 辅助
手术流程以结构化且重复的临床事件形式展开,通过术中手术视频对其进行实时理解,对术中决策及支持至关重要。然而,现有视频理解方法存在权衡问题:自回归视频-语言模型支持全面推理,但对时间敏感的临床应用而言不实用;对比模型延迟低,但难以处理复杂场景理解。最近,生成式检索已被用于通用领域的视频理解,但将其迁移至手术领域并非易事,因为近乎相同的视觉外观可能对应语义不同的事件,且涉及的术语具有高度手术特异性。为此,我们提出SurgNarrator,一种专为手术视频理解定制的新型生成式检索框架。我们从手术标注文本中构建精心整理的手术中心词汇表,以定义具有临床意义的检索空间;随后对预训练的Qwen3-VL-Embedding-8B进行适配,通过时间感知的对比目标学习具有区分性的临床表征。推理阶段,采用分层、流程感知的检索策略将搜索空间缩小至相关流程类型,实现快速且有效的响应。我们在12个基准上以零样本设置对方法进行全面评估,相较于现有最优基线取得了持续的性能提升,同时与生成式基线相比,输出阶段延迟降低了两个数量级以上。
英文摘要
Surgical procedures unfold as structured and recurring clinical events, whose real-time understanding via intraoperative surgical videos is critical for intraoperative decision-making and support. However, existing video understanding methods force a trade-off: autoregressive video-language models support comprehensive reasoning but are not practical for time-sensitive clinical applications, whereas contrastive models offer low latency but struggle with complex scene understanding. Recently, generative retrieval has been explored for general-domain video understanding, but transferring it to surgery is not trivial because near-identical visual appearances may indicate semantically distinct events, and the terminology involved is highly surgery-specific. To this end, we propose SurgNarrator, a new generative retrieval framework tailored for surgical video understanding. We construct a well-curated surgery-centric vocabulary from surgical captions to define a clinically meaningful retrieval space. We then adapt the pre-trained Qwen3-VL-Embedding-8B to learn discriminative clinical representations with a temporally-aware contrastive objective. During inference, a hierarchical, procedure-aware retrieval strategy narrows the search space to the relevant procedure type, delivering fast and effective responses. Our method is comprehensively evaluated on twelve benchmarks in a zero-shot setting and achieves consistent performance gains over state-of-the-art baselines, while reducing output-stage latency by more than two orders of magnitude compared with the generative baseline.
发表机构
- University of Liverpool(利物浦大学)
- City University of Hong Kong(香港城市大学)
- University of Oxford(牛津大学)
- University of Strasbourg(斯特拉斯堡大学)
- IHU Strasbourg(斯特拉斯堡大学医院研究所)
- Imperial College London(帝国理工学院)
机构由 AI 辅助整理,请以论文原文为准。