arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16469cs.CV

适用于手术室的可消毒场景图生成

Sterilizable Scene Graph Generation for Operating Rooms

Nick Lemke, Ssharvien Kumar Sivakumar, Antoine P. Sanner, John Kalkhof, Henry John Krumb, Ghazal Ghazaei, Anirban Mukhopadhyay

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出首个基于NCA的轻量型场景图生成框架SG-NCA,适配手术室无风扇设备,参数较基准少55倍,性能相当,可用于术中场景理解及下游任务。

中文摘要 AI 辅助

从手术视频中生成场景图,可通过建模对象及其语义关系实现对手术场景的整体结构化理解。尽管近期取得了进展,但现有最优方法依赖参数量大的深度学习模型,受硬件体积、卫生约束、延迟及数据隐私问题限制,难以在手术室(OR)部署。据所知,本文提出的是首个基于神经细胞自动机(NCA)的场景图生成方法,也是首个能学习结构化表示的NCA框架。本文引入SG-NCA,这是一种基于神经细胞自动机(NCA)的轻量型场景图生成框架,专为符合手术室卫生规范的无风扇设备推理设计。SG-NCA是首个将基于NCA的多类别分割用于高效对象检测与特征提取,并结合轻量型关系预测器的场景图生成模型。本文在白内障手术和胆囊切除术视频上评估SG-NCA,结果显示其性能与既定基准相当,参数数量减少55倍。本文展示了其在更适合手术室的无风扇边缘设备上的部署,并验证了手术视频字幕生成等下游应用,凸显SG-NCA在实现经济、隐私保护且适配手术室的术中场景理解方面的潜力。

英文摘要

Scene graph generation from surgical video enables a holistic and structured understanding of surgical scenes by modeling objects and their semantic relationships. Despite recent advances, state-of-the-art approaches rely on large, parameter-heavy deep learning models that are impractical for deployment in the operating room (OR) due to hardware footprint, hygiene constraints, latency, and data privacy concerns. To the best of our knowledge, this is the first scene graph generation method built on NCAs and the first NCA framework capable of learning structured representations. We introduce SG-NCA, a lightweight scene graph generation framework based on Neural Cellular Automata (NCA), designed for inference in fanless devices critical for OR hygiene protocols. SG-NCA is the first scene graph generation combining NCA-based multiclass segmentation for efficient object detection and feature extraction with a lightweight relation predictor. We evaluate SG-NCA on videos of cataract surgery and cholecystectomy, demonstrating performance comparable to established baselines while requiring 55x fewer parameters. We showcase deployment on fanless edge devices better suited for the OR and demonstrate downstream applications such as surgical video captioning, highlighting SG-NCA's potential for affordable, privacy-preserving, and OR-ready intraoperative scene understanding.

发表机构

  • Technical University of Darmstadt(达姆施塔特工业大学)
  • ImFusion GmbH(ImFusion有限公司)
  • Carl Zeiss AG(卡尔蔡司股份公司)
  • University Medical Center Mainz(美因茨大学医学中心)
  • Inria Center at University Côte d’Azur(蔚蓝海岸大学法国国家信息与自动化研究所中心)

机构由 AI 辅助整理,请以论文原文为准。

↑