发表机构
Beihang University; Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences; Shenzhen University of Advanced Technology(北京航空航天大学; 中国科学院深圳先进技术研究院; 深圳理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对手术视频三元组识别中难以统一建模帧内标签依赖与帧间时序语义的问题,提出TRCoRSurg框架,引入MS-CAMRE和BTRFA模块并定义TCER指标,在两个数据集上实现最优性能且一致性错误率显著降低。
AI 中文摘要
理解复杂手术场景需要识别器械、动作、目标等多个相互依赖的实体,同时保持它们随时间的关系一致性。现有手术三元组识别方法难以在统一框架内同时建模帧内标签依赖和帧间时序语义。为解决这些局限,本文提出一种整合空间、关系和时序线索的统一框架以实现鲁棒的手术三元组识别。具体而言,首先通过多尺度编码器提取类别特定的空间先验,再经带有多尺度类别激活图引导的关系提取(MS-CAMRE)的标签关系建模模块优化这些先验,使模型能捕获三元组组件间的静态共现模式和动态上下文依赖。此外,双向时序关系融合注意力(BTRFA)模块协调时序与关系表示以实现连贯的时序推理。本文还引入新评估指标三元组一致性错误率(TCER),定量衡量模型保持三元组间因果和语义一致性的能力。在CholecT45和ProstaTD数据集上的大量实验表明,本文方法达到了最优性能,AP_IVT分别提升5.1%和7.8%;且根据TCER,本文方法在两个数据集上的相对降幅分别超过36%和25%,证明了框架在时序关系协同推理中的有效性。
英文摘要
Understanding complex surgical scenes requires recognizing multiple interdependent entities, such as instruments, actions, and targets, while maintaining their relational consistency across time. Existing surgical triplet recognition methods struggle to jointly model intra-frame label dependencies and inter-frame temporal semantics in a unified manner. To address these limitations, we propose a unified framework that integrates spatial, relational, and temporal cues for robust surgical triplet recognition. Specifically, class-specific spatial priors are first extracted through a multi-scale encoder. These priors are then refined by a Label Correlation Modeling module with multi-scale class activation map-guided relational extraction (MS-CAMRE), enabling the model to capture both static co-occurrence patterns and dynamic contextual dependencies among triplet components. Furthermore, a Bidirectional Temporal-Relational Fusion Attention (BTRFA) module harmonizes temporal and relational representations to achieve coherent temporal reasoning. We also introduce a new evaluation metric, the Triplet Consistency Error Rate (TCER), which quantitatively measures the model's ability to preserve causal and semantic consistency across triplets. Extensive experiments on the CholecT45 and ProstaTD datasets show that our method achieves state-of-the-art performance, improving AP_IVT by 5.1 percent and 7.8 percent, respectively. Moreover, according to TCER, our approach achieves relative reductions of more than 36 percent and 25 percent on the two datasets, respectively, demonstrating the effectiveness of our framework in temporal-relational co-reasoning.
Commentscode: https://github.com/Neesky/TRCoRSurg