arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25651cs.HC

SurgGraph:基于几何场景图的定量腹腔镜视频理解

SurgGraph: Quantitative Laparoscopic Video Understanding via Geometry-Grounded Scene Graphs

Jingying Wang, Rosiana Natalie, Marquise D Singleterry, Filippos Bellos, Brian George, Gurjit Sandhu, Jason J Corso, Anhong Guo, Vitaliy Popov, Xu Wang

首次发表
浏览论文内容

中文总结 AI 辅助

SurgGraph通过无需训练的流程,利用分割掩膜和深度图将手术视频中的临床关系编码为定量场景图,并构建SurgGraphQA应用,经医学教育研究验证了其显著的学习提升效果。

中文摘要 AI 辅助

手术视频是向受训者传授解剖学、工具使用和操作技能的主要资源。然而,大规模地从中学习需要能够理解手术场景的系统。现有方法存在不足:视觉语言模型缺乏细粒度的领域推理能力,任务特定模型无法泛化,而先前的场景图遗漏了临床上有意义的细节。我们提出了SurgGraph,一种无需训练、可从手术视频生成定量场景图的流程。基于分割掩膜和深度图,SurgGraph将每个临床有意义的关联(附着、遮挡、分离、工具动作)编码为<主语,动词,宾语,数值>元组,其数值量化了关联随时间变化的程度。技术评估显示,其场景理解比最先进的手术视觉语言模型基线更为精确。随后,我们构建了SurgGraphQA,一个概念验证的学习应用,可检索有意义和边界案例的示例,并生成视觉解释和反馈。一项涉及17名医学生和2名住院外科医生的研究显示了显著的学习收益,证明了其教育价值。

英文摘要

Surgical videos are a primary resource for teaching trainees anatomy, tool usage, and procedural skills. Yet learning from them at scale requires systems that understand surgical scenes. Existing approaches fall short: vision-language models lack fine-grained domain reasoning, task-specific models do not generalize, and prior scene graphs omit clinically meaningful detail. We present SurgGraph, a training-free pipeline that generates quantitative scene graphs from surgical videos. Operating on segmentation masks and depth maps, SurgGraph encodes each clinically meaningful relation (attachment, occlusion, separation, tool actions) as a <subject, verb, object, value> tuple whose numeric value quantifies the relation's extent over time. Technical evaluations show more precise scene understanding than state-of-the-art surgical VLM baselines. We then build SurgGraphQA, a proof-of-concept learning application that retrieves meaningful and boundary-case exemplars and generates visual explanations and feedback. A study with 17 medical students and 2 resident surgeons shows significant learning gains, demonstrating its educational value.

发表机构

  • University of Michigan(密歇根大学)

机构由 AI 辅助整理,请以论文原文为准。

↑