arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SPIRIT:面向手术动作三元组识别的器械-组织交互时空成对关系建模

SPIRIT: Spatio-temporal Pairwise Relational Modeling of Instrument-Tissue Interactions for Surgical Action Triplet Recognition

Saurav Sharma, Lorenzo Arboit, Nabani Banik, Sarah Meuli, Julia Alekseenko, Jan Liechti, Franziska Heitzinger, Michela Orsi, Didier Mutter, Daniel Gero, Philipp C. Nett, Beat P. Muller, Joel L. Lavanchy, Nicolas Padoy

arXiv 2608.02188首次发表:更新:

发表机构

University of Strasbourg; CNRS; INSERM; ICube; IHU Strasbourg; University of Pavia; University Digestive Health Care Center – Clarunis; Università di Roma Tor Vergata; University Hospital of Strasbourg; University Hospital Zurich; University Hospital Bern; University of Basel(斯特拉斯堡大学; 法国国家科学研究中心; 法国国家健康与医学研究院; ICube实验室; 斯特拉斯堡大学医院研究所; 帕维亚大学; 克拉鲁尼斯大学消化保健中心; 罗马第二大学; 斯特拉斯堡大学医院; 苏黎世大学医院; 伯尔尼大学医院; 巴塞尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出SPIRIT框架,构建多中心手术数据集MultiBypass-4C-T40,通过建模三元组成分的成对关系提升多中心手术动作三元组识别性能。

AI 中文摘要

对手术活动的细粒度理解对于手术室的上下文感知辅助至关重要,包括安全监测、不良事件识别和技能评估。手术动作三元组被定义为<器械、动作、目标>形式的元组,可对器械-组织交互提供结构化描述。然而,一个关键的未解决问题是如何学习在不同医疗机构间仍保持可靠的三元组表示,因为不同医疗机构的手术视频在采集条件、外科医生风格、工具使用和组织处理上存在差异,而现有的三元组数据集不支持明确的中心间迁移评估。为解决该问题,我们提出SPIRIT,这是一个用于手术动作三元组识别的结构化框架,旨在学习在不同中心间迁移更可靠的交互表示。SPIRIT没有将每个三元组视为单一类别标签,而是先学习器械、动作和目标的时空表示,再对它们的成对关系进行建模,最后将其组合为一致的三元组预测,同时使用多头蒸馏来稳定学习过程。为评估该设置,我们构建了MultiBypass-4C-T40,这是一个针对四个地理上不同中心的Roux-en-Y胃旁路手术中密集手术动作三元组识别的多中心数据集,附带辅助阶段和步骤标注。在多种评估协议下,SPIRIT始终优于近期的强基线方法,凸显了显式关系推理对多中心三元组识别的价值。代码将在该httpsURL提供。

英文摘要

Fine-grained understanding of surgical activity is essential for context-aware assistance in the operating room, including safety monitoring, adverse event identification, and skill assessment. Surgical action triplets, defined as tuples of the form <instrument, verb, target>, provide a structured description of instrument-tissue interactions. A key open problem, however, is how to learn triplet representations that remain reliable across institutions, where surgical video varies in acquisition conditions, surgeon style, tool usage, and tissue handling, while existing triplet datasets do not support explicit evaluation of center-wise transfer. To address this problem, we propose \textbf{SPIRIT}, a structured framework for surgical action triplet recognition designed to learn interaction representations that transfer more reliably across centers. Instead of treating each triplet as a flat class label, SPIRIT first learns spatio-temporal representations for instruments, verbs, and targets, then models their pairwise relations, and finally composes them into coherent triplet predictions, with multi-head distillation used to stabilize learning. To evaluate this setting, we establish \textbf{MultiBypass-4C-T40}, a multi-centric dataset for dense surgical action triplet recognition in Roux-en-Y gastric bypass across four geographically distinct centers, with auxiliary phase and step annotations. Across multiple evaluation protocols, SPIRIT consistently outperforms strong recent baselines, highlighting the value of explicit relational reasoning for multi-centric triplet recognition. Code will be available at https://github.com/CAMMA-public/multibypass-4c-t40.

Comments31 pages, 8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑