arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于语义RGB-深度信息的显微立体视频手术技能评估

Semantic RGB--Depth Based Surgical Skill Assessment in Microscopic Stereo Videos

Jecia Z. Y. Mao, Sue M. Cho, Francis X. Creighton, Deepa Galaiya, Russell H. Taylor, Manish Sahu

arXiv 2610.01205首次发表:更新:

发表机构

Johns Hopkins University; Johns Hopkins University School of Medicine(约翰霍普金斯大学; 约翰霍普金斯大学医学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有视频手术技能评估忽视三维空间信息的问题,提出语义RGB-深度融合框架,结合深度融合与分层注意力,在33例喉显微手术中技能分类F1达0.938,优于单一模态。

AI 中文摘要

显微外科技术技能的客观评估对于基于能力的培训和质量管理至关重要,然而现有的基于视频的方法主要依赖RGB图像,因此忽视了表征器械-解剖结构相互作用的三维空间关系。尽管立体手术显微镜提供了互补的深度信息,但传统的立体匹配算法在高倍成像条件下可能产生稀疏且不可靠的深度估计,限制了其在自动化技能评估中的应用。本研究提出了一种基于语义RGB-深度信息的框架,用于从显微立体视频中进行手术技能评估。一种基于回归的深度融合方法将稀疏的度量立体深度与稠密的单目深度估计相结合,生成手术场景的稠密几何表示。该表示与对应于各个手术器械和周围解剖结构的语义分解RGB流相集成。一种分层注意力架构联合编码这些流,以捕捉不同训练水平的外科医生在器械使用和器械-解剖结构相互作用方面的判别性模式。该框架在由六名外科医生(包括主治外科医生和外科住院医师)执行的33例离体经口显微喉手术中进行了评估,采用留一外科医生交叉验证方法。所提出的语义RGB-深度模型在技能水平分类中达到了0.938的F1分数,而语义RGB和语义深度分别为0.696和0.929。这些结果表明,几何信息可以改善基于显微立体视频的自动化手术技能评估。学习到的空间、时间和语义注意力模式也支持对模型所关注的场景区域、视频片段和语义流进行定性检查。

英文摘要

Objective assessment of microsurgical technical skill is essential for competency-based training and quality assurance, yet existing video-based approaches predominantly rely on RGB images and therefore overlook the 3D spatial relationships that characterize instrument-anatomy interactions. Although stereo operating microscopes provide complementary depth information, conventional stereo matching algorithms can produce sparse and unreliable depth estimates under high-magnification imaging conditions, limiting their use for automated skill assessment. This work presents a semantic RGB-Depth framework for surgical skill assessment from microscopic stereo videos. A regression-based depth fusion method combines sparse metric stereo depth with dense monocular depth estimates to generate a dense geometric representation of the surgical scene. This representation is integrated with semantically decomposed RGB streams corresponding to individual surgical instruments and surrounding anatomy. A hierarchical attention architecture jointly encodes these streams to capture discriminative patterns of instrument use and instrument-anatomy interaction across surgeons at different training levels. The framework was evaluated on 33 ex vivo transoral microlaryngeal procedures performed by six surgeons, comprising attending surgeons and surgical residents, using leave-one-surgeon-out cross-validation. The proposed semantic RGB-Depth model achieved an F1 score of 0.938 for skill-level classification, compared with 0.696 for semantic RGB and 0.929 for semantic depth. These results suggest that geometric information can improve automated surgical skill assessment from microscopic stereo videos. The learned spatial, temporal, and semantic attention patterns also support qualitative examination of the scene regions, video segments, and semantic streams emphasized by the model.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑