arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.02717cs.CVcs.RO

MV-dVRK:面向空间手术感知的多视点基准

MV-dVRK: A Multi-Viewpoint Benchmark for Spatial Surgical Perception

Guido Caccianiga, Sergey Prokudin, Yutong Chen, Bernard Javot, Rachael L'Orsa, Omer Burak Aladağ, Yarden Sharon, Jens Rolinger, Ivan Capobianco, Anton Deguet, S… 展开作者

Guido Caccianiga, Sergey Prokudin, Yutong Chen, Bernard Javot, Rachael L'Orsa, Omer Burak Aladağ, Yarden Sharon, Jens Rolinger, Ivan Capobianco, Anton Deguet, Siyu Tang, Katherine J. Kuchenbecker

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出首个离体手术多视点基准MV-dVRK,含静态与动态手术数据,用于对比不同三维重建方法,发现多视点优化方法在1mm容差下覆盖67%真实表面点,优于前馈基础模型的43%。

中文摘要 AI 辅助

大规模训练与精细优化技术已大幅提升稀疏多视点三维重建性能,但这类方法虽与手术高度相关,却从未在真实内镜图像上接受过严格评估。当前临床遥控机器人在患者体内部署单目立体相机,导致多视点数据极为稀缺。本文提出MV-dVRK,首个离体手术数据集,结合多曝光同步立体视点、精确表面几何与相机位姿。该基准的静态子集提供经工业3D扫描仪验证的密集SfM参考几何,以及真实相机位姿与稀疏视点测试集。我们利用MV-dVRK系统对比零样本单目、立体、多立体及多视点三维重建方法随视点数量增加的表现:使用两台内镜时,多立体重建覆盖率最高;加入第三视点后,基于优化的多视点方法表现最佳,在1毫米容差内覆盖67%的真实表面点,且恢复高精度相对相机位姿;相比之下,前馈基础模型在相同设置下仅覆盖43%的真实表面。MV-dVRK还包含10个涵盖多种手术任务的动态序列,具备不断提升的运动复杂度与组织形变,为未来多视点手术感知研究提供基础。该项目可访问:this https URL。

英文摘要

Large-scale training and refined optimization techniques have greatly improved sparse multi-view 3D reconstruction. Despite their relevance to surgery, such methods have never before been rigorously evaluated on real endoscopic images. Current clinical telerobots deploy a single stereo camera inside the patient, making multi-viewpoint data extremely rare. This paper presents MV-dVRK, the first ex-vivo surgical dataset to combine multiple exposure-synchronized stereo viewpoints with accurate surface geometry and camera poses. The static subset of the benchmark provides dense SfM reference geometry, validated against an industrial 3D scanner, together with ground-truth camera poses and sparse-view test sets. We use MV-dVRK to systematically compare zero-shot monocular, stereo, multi-stereo, and multi-view 3D reconstruction methods as the number of viewpoints increases. With two endoscopes, multi-stereo reconstruction achieves the highest coverage. With a third viewpoint, optimization-based multi-view methods perform best, covering 67% of ground-truth surface points within a 1 mm tolerance and recovering highly accurate relative camera poses. By contrast, feed-forward foundation models cover only 43% of the ground-truth surface in the same setting. MV-dVRK also includes ten dynamic sequences spanning multiple surgical tasks, with increasing kinematic complexity and tissue deformation, providing a basis for future research in multi-viewpoint surgical perception. The project is available at: https://mv-dvrk.is.mpg.de.

发表机构

  • Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所)
  • ETH Zürich(苏黎世联邦理工学院)
  • Erbe Group(爱尔博集团)
  • Tübingen University Hospital(蒂宾根大学医院)
  • Johns Hopkins University(约翰霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

↑