Colon3R:基于单目结肠镜视频的跨域三维重建
Colon3R: Cross-Domain 3D Reconstruction from Monocular Colonoscopic Video
浏览论文内容
中文总结 AI 辅助
针对单目结肠镜三维重建的跨域难题,提出基于VGGT的半监督框架Colon3R,利用分层准刚性可靠性选择监督并保持源域几何,在深度、点图和位姿估计上超越现有方法。
中文摘要 AI 辅助
单目结肠镜三维重建对于手术机器人结肠镜检查至关重要,但由于弱纹理、镜面反射、视野重叠有限以及非刚性组织运动,这一任务仍然具有挑战性。传统的多视图三维重建方法依赖于稳定的对应关系和近似刚性假设,而这些条件在结肠镜检查中往往不成立。现有的内窥镜方法通常依赖特定领域的监督,然而目前缺乏足够的体内标注数据来将几何基础模型适配到临床结肠镜检查中。我们提出了Colon3R,一个基于预训练VGGT的跨域半监督框架,该框架将来自标注体模和模拟数据的耦合相机、深度和点图几何迁移到未标注的体内结肠镜检查中,而无需目标域的几何标注。与仅从体模和模拟数据学习的源域微调不同,Colon3R通过教师派生的跨视图监督直接利用未标注的体内视频。我们提出的分层准刚性可靠性机制在序列、有向对和像素三个层级选择可靠的监督,同时源域保持适配在目标域适配过程中保留已学习的耦合几何。大量实验表明,我们的方法在深度、点图和相机位姿估计方面优于现有最先进方法,取得了卓越的整体性能。在真实体内结肠镜上的定性比较进一步显示,在临床域偏移下,我们的方法比竞争方法重建出更完整且几何一致性更高的结果。代码将在论文被接收后公开。
英文摘要
Monocular colonoscopic 3D reconstruction is important for surgical robotic colonoscopy, but remains challenging due to weak texture, specular reflections, limited view overlap, and non-rigid tissue motion. Conventional multi-view 3D reconstruction methods rely on stable correspondences and approximate rigidity, which are often violated in colonoscopy. Existing endoscopic methods often rely on domain-specific supervision, whereas there are not enough in-vivo labeled data available to adapt geometry foundation models to clinical colonoscopy. We present Colon3R, a cross-domain semi-supervised framework built on pretrained VGGT that transfers coupled camera, depth, and pointmap geometry from labeled phantom and simulated data to unlabeled in-vivo colonoscopy without requiring target-domain geometric annotations. Unlike source-only fine-tuning, which learns only from phantom and simulated data, Colon3R directly exploits unlabeled in-vivo video through teacher-derived cross-view supervision. Our proposed hierarchical quasi-rigid reliability selects reliable supervision at the sequence, directed-pair, and pixel levels, while source-preserving adaptation retains the learned coupled geometry during target-domain adaptation. Extensive experiments demonstrate that our method achieves superior overall performance over state-of-the-art approaches in depth, pointmap, and camera pose estimation. Qualitative comparisons on real in-vivo colonoscopy further show substantially more complete and geometrically consistent reconstructions than competing methods under clinical domain shift. The code will be public available after the paper is accepted.