arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15336cs.CV

SAGE-OR:面向手术室的半监督自适应场景图生成

SAGE-OR: Semi-supervised Adaptive Scene Graph Generation for Operating Rooms

Brandon Leblanc, Charalambos Poullis

首次发表
浏览论文内容

中文总结 AI 辅助

SAGE-OR是面向手术室的半监督自适应场景图生成框架,采用解耦范式,轻量高效,在4D-OR基准上F1达76%,经无监督手部增强后达86%,可低成本适配新手术场景。

中文摘要 AI 辅助

当前手术场景图生成方法依赖密集多模态监督与专用硬件(同步RGB-D传感器、校准装置),导致数据集构建成本高昂,且所有现有基准均局限于模拟环境。本文提出SAGE-OR,这是一种以特征为中心的框架,将传统的“先检测后推理”范式替换为解耦的“表示-推理”范式:定位来自冻结的基础模型,隐式编码于预计算特征中,无需任何定位监督即可使用;轻量级图变换器则对缓存特征执行关系推理。我们采用半监督公式,结合通用分割提示,以消除定位监督,同时通过额外的提示驱动实体(如标注中缺失的手部)实现无监督上下文增强。通用提示用于诱导近乎完美的召回率,而精度则交由下游基于注意力的推理处理,可通过提示级修改轻松适配新实体。该设计实现了一个15M参数的轻量级图变换器,训练耗时1.4小时,每帧关系推理约1ms,峰值内存低于2GB,适用于手术室所用的边缘硬件;特征提取作为单独的缓存阶段离线运行,每帧耗时4.27秒。在4D-OR基准上,核心模型达到76%的F1值,与完全监督的4D-OR基线表现相当,同时消除了所有定位标注;无监督手部增强将该值提升至86%,与需要密集多模态监督的最先进(SOTA)方法仅差4个百分点,为适配新手术场景提供了实用途径,且仅需关系和类别标签之外的标注即可完成。

英文摘要

Current surgical scene graph generation methods depend on dense multi-modal supervision and specialized hardware (synchronized RGB-D sensors, calibration rigs), making dataset construction expensive and restricting all existing benchmarks to simulated environments. We propose SAGE-OR, a feature-centric framework that replaces the traditional detect-then-reason paradigm with a decoupled representation-reasoning paradigm in which localization is derived from frozen foundation models, encoded implicitly in pre-computed features, and used without any localization supervision, while a lightweight graph transformer performs relational reasoning over cached features. We employ a semi-supervised formulation with general-purpose segmentation prompts to eliminate localization supervision while enabling unsupervised context augmentation through additional prompt-driven entities, such as hands, which are absent from annotations. General-purpose prompts are used to induce near-perfect recall, while precision is delegated to downstream attention-based reasoning, enabling simple adaptation to new entities via prompt-level modification. This design enables a lightweight 15M-parameter graph transformer that trains in 1.4 hours and runs relational inference at $\sim$1ms per frame with peak memory under 2GB, suitable for edge hardware used in the operating room; feature extraction runs offline as a separate caching stage (4.27s per frame). On the 4D-OR benchmark, the core model achieves 76% F1, matching the fully supervised 4D-OR baseline while eliminating all localization annotations, and unsupervised hand augmentation raises this to 86%, within 4 points of state-of-the-art (SOTA) methods requiring dense multi-modal supervision, providing a practical pathway for adaptation to new surgical settings without annotation other than relationship and class labels.

发表机构

  • Concordia University(康考迪亚大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑