面向视觉三维力估计的以对象为中心的重建
Object-Centered Reconstruction for Vision-Based 3D Force Estimation
- Vanderbilt University(范德堡大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
针对手术机器人缺乏力感知的问题,提出基于立体内窥镜视频、以对象为中心重建组织点云并跟踪变形以估计三维力的方法,在模型和离体猪肠上误差显著低于相机坐标系方法,并初步验证了体内可行性。
中文摘要 AI 辅助
在机器人结直肠手术中,过大的力可能损伤组织并增加吻合口漏的风险。尽管达芬奇5(da Vinci 5)系统提供了力感知功能,但该功能在早期的达芬奇系统以及许多其他手术机器人平台上并不可用。在本工作中,我们提出了一种基于视觉的流程,用于从立体内窥镜视频中的软组织变形估计三维交互力。我们在以对象为中心的坐标系中动态重建组织点云,利用几何约束跟踪组织点,并通过神经网络预测三维力向量。我们依次在橡胶手套模型、离体猪结肠以及体内结直肠手术视频序列上评估该流程。在内窥镜视野中组织方向和位置变化以及不同相机视角下,所提方法在模型和猪结肠上分别实现了平均均方根误差(RMSE)为0.77 N和1.30 N。与相机坐标系表示相比,以对象为中心的表示将平均RMSE分别降低了51.3%和56.7%,而几何约束跟踪相比CoTracker将RMSE分别降低了19.8%和25.3%。我们进一步在体内结直肠手术序列上定性展示了基于视觉的力估计的可行性,这标志着向基于视觉、无传感器的力估计的临床转化迈出了一步。
英文摘要
Excessive force may damage tissue and increase the risk of anastomotic leakage in robotic colorectal surgery. Although the da Vinci 5 provides force sensing, this capability is unavailable on earlier da Vinci systems and many other surgical robotic platforms. In this work, we present a vision-based pipeline for estimating 3D interaction forces from soft-tissue deformation in stereo endoscopic video. We dynamically reconstruct the tissue point cloud in an object-centered coordinate frame, track tissue points with geometric constraints, and predict the 3D force vector with a neural network. We progressively evaluate the pipeline on rubber-glove phantoms, ex vivo porcine colons, and in vivo colorectal surgical video sequences. Under varying tissue orientations and positions within the endoscopic view, as well as different camera viewpoints, the proposed method achieves average root mean square error (RMSEs) of 0.77 N and 1.30 N on the phantom and porcine colon, respectively. Compared with the camera-frame representation, the object-centered representation reduces average RMSE by 51.3% and 56.7%, while geometry-constrained tracking reduces RMSE by 19.8% and 25.3% compared with CoTracker. We further qualitatively demonstrate the feasibility of vision-based force estimation on an in vivo colorectal surgical sequence, as a step toward clinical translation of vision-based, sensorless force estimation.