arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向支气管镜检查的几何感知相机定位

Geometry-Aware Camera Localization for Bronchoscopy

Lumin Chen, Qingyao Tian, Jinpeng Li, Haoyu Jiang, Huai Liao, Xinyan Huang, Hongbin Liu, Dong Yi

arXiv 2608.07116首次发表:更新:

发表机构

Centre for Artificial Intelligence and Robotics, Hong Kong Institute of Science & Innovation, Chinese Academy of Sciences; Institute of Automation, Chinese Academy of Sciences; The First Affiliated Hospital, Sun Yat-sen University(中国科学院香港创新研究院人工智能与机器人中心; 中国科学院自动化研究所; 中山大学附属第一医院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对支气管镜相机定位的精度、实时性及几何先验利用问题,提出GABL框架,采用图引导粗精定位与Transformer跟踪模型,误差显著降低且推理速度提升4倍。

AI 中文摘要

支气管镜检查中的相机定位仍是一项具有挑战性的问题,原因在于其严苛的精度要求、实时约束以及有限的训练数据。与自然场景相比,受限的解剖结构要求达到毫米级精度,而术中导航则需要低延迟推理。然而,现有方法往往无法有效利用术前几何先验,限制了其鲁棒性和精度。为解决这些局限,我们提出了一种统一的几何感知支气管镜定位框架(GABL),该框架可有效融合术前结构先验与配对的术中视频,以估计6自由度(6-DoF)相机位姿。具体而言,为解决复杂气道中的视觉歧义问题,我们提出了一种图引导的由粗到精定位方案,可有效利用结构先验实现精确的位姿估计。此外,为缓解位姿抖动并弥合视觉-结构差距,我们将基于Transformer的跟踪模型与一种新颖的RGB-深度匹配目标相结合,共同强制时空一致性和几何一致性。大量实验表明,与现有最先进方法相比,我们的方法在平移误差和旋转误差上分别显著降低了8.37%和31.76%,同时实现了4倍的推理速度提升(33.6 FPS),可用于鲁棒的实时支气管镜定位。项目网站:this https URL。

英文摘要

Camera localization in bronchoscopy remains a challenging problem due to stringent accuracy requirements, real-time constraints, and limited training data. Compared to natural scenes, the confined anatomical structures demand millimeter-level precision, while intraoperative guidance necessitates low-latency inference. However, existing methods often fail to effectively exploit preoperative geometric priors, limiting their robustness and accuracy. To address these limitations, we propose a unified geometry-aware bronchoscope localization framework (GABL) that effectively fuses preoperative structural priors with paired intraoperative video to estimate 6-DoF camera poses. Specifically, to address visual ambiguity in complex airways, we propose a graph-guided coarse-to-fine localization scheme that effectively leverages structural priors for precise pose estimation. Furthermore, to mitigate pose jitter and bridge the visual-structural gap, we integrate a Transformer-based tracking model with a novel RGB-depth matching objective, jointly enforcing spatio-temporal and geometric consistency. Extensive experiments demonstrate that our method yields remarkable reductions of 8.37% and 31.76% in translation and rotation errors over the prior state-of-the-art, alongside 4 times inference speedup (33.6 FPS) for robust real-time bronchoscope localization. Project website: https://paulili08.github.io/GABL/.

CommentsAccepted by ACM MM2026

DOI:10.1145/3767308.3835259

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑