用于玻璃体视网膜手术阶段识别的叙事、显微镜及iOCT图像的多模态共享潜在表示
Multimodal Shared Latent Representation of Narration, Microscope and iOCT Images for Phase Recognition in Vitreoretinal Surgery
浏览论文内容
中文总结 AI 辅助
该研究提出以显微镜视图为共享锚点的多模态框架,结合手术叙事与iOCT,用双路MS-TCN++实现玻璃体视网膜手术的宏观与微观阶段识别,提升了宏观阶段识别性能,为细粒度测量提供了探索途径。
中文摘要 AI 辅助
手术阶段识别是玻璃体视网膜手术中上下文感知计算机辅助反馈的关键,但同步多模态术中数据(尤其是显微镜视图和术中OCT)的稀缺,限制了模拟外科医生自然进行的多模态整合的方法。相比之下,手术叙事在网上大量可用,可提供丰富的语义监督。现有工作主要探索成对对比学习(如术中OCT-显微镜或显微镜-叙事),而三种模态的联合建模在很大程度上未被探索。我们引入一个框架,以显微镜视图作为共享锚点,在无需完全同步的三模态数据集的情况下,桥接手术叙事和术中OCT(iOCT),利用真实的显微镜-叙事视频以及同步显微镜视频与工具对齐的iOCT对的合成数据集。对比对齐将结构先验从合成域转移到缺乏iOCT的真实视频,双路MS-TCN++整合所得嵌入以进行宏观和微观阶段的联合预测。在真实玻璃体视网膜手术上评估,我们的框架将宏观阶段识别的平均F1值从0.38的零样本基线提升至0.53,还为估计仅在真实显微镜视频中无法直接观察到的细粒度器械-组织测量提供了探索性途径;这些微观阶段估计在合成数据上进行了定量验证,在真实手术中仅进行了定性展示。据我们所知,这是首个将显微镜视图、iOCT B扫描和手术叙事统一到共享潜在空间以用于手术阶段识别的工作。
英文摘要
Surgical phase recognition is key to context-aware computer-assisted feedback in vitreoretinal procedures, yet the scarcity of synchronized multimodal intraoperative data, particularly microscope views and intraoperative OCT, limits approaches that aim to replicate the multimodal integration surgeons perform naturally. Surgical narration, by contrast, is abundantly available online and offers rich semantic supervision. Prior work has mainly explored pairwise contrastive learning (e.g., intraoperative OCT-microscope or microscope-narration), leaving the joint modeling of all three modalities largely unexplored. We introduce a framework that uses microscope views as a shared anchor to bridge surgical narrations and intraoperative OCT (iOCT) without requiring a fully synchronized tri-modal dataset, leveraging real microscope-narration videos and a synthetic dataset of synchronized microscope video and tool-aligned iOCT pairs. Contrastive alignment transfers structural priors from the synthetic domain to real videos lacking iOCT, and a dual-head MS-TCN++ integrates the resulting embeddings for joint macro- and micro-phase prediction. Evaluated on real vitreoretinal surgeries, our framework improves macro-phase recognition over a zero-shot baseline (mean F1 0.38 to 0.53) and provides an exploratory route to estimating fine-grained instrument-tissue measurements that are not directly observable in real microscope video alone; these micro-phase estimates are validated quantitatively on synthetic data and shown only qualitatively on real surgery. To our knowledge, this is the first work to unify microscope view, iOCT B-scans, and surgical narrations in a shared latent space for surgical phase recognition.
发表机构
- Technical University of Munich(慕尼黑工业大学)
- SynthesEyes GmbH(SynthesEyes公司)
- LMU University Hospital(慕尼黑大学附属医院)
机构由 AI 辅助整理,请以论文原文为准。