arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于AI的单次结构光深度重建用于实时腹腔镜手术导航

AI-based single-shot structured-light depth reconstruction for real-time laparoscopic surgical guidance

Wayne Wonseok Rodgers, Xiangyi Le, Seonghoon Jang, Shuwen Wei, Justin Opfermann, Michael Kam, Axel Krieger, Jin U. Kang

arXiv 2608.05109首次发表:更新:

发表机构

Johns Hopkins University(约翰斯·霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究开发了一种基于AI的无同步单次结构光深度重建方法,用于腹腔镜手术导航,其在体模实验中达到了高精度和26.0 Hz的帧率,性能优于基线模型和现成单目深度模型。

AI 中文摘要

意义:术中准确的深度感知对于自主和半自主机器人腹腔镜手术至关重要。传统条纹投影轮廓术可实现毫米级精度,但通常需要多次采集、数字微镜器件投影以及投影仪-相机同步,这使得将其集成到紧凑型腹腔镜系统中变得复杂。目标:开发一种无同步、单次拍摄的深度传感平台,该平台采用被动LED照明的二元掩模和带有定制U-Net深度头的VQ-VAE先验。方法:将紧凑型投影模块耦合到双通道腹腔镜的一个通道,第二个通道对条纹照明的目标成像。使用Zivid 3D相机获取722对体模图像的参考深度,将Zivid深度图重投影到SSLE图像帧中以进行监督训练和评估。VQ-VAE将每个输入编码为离散潜在表示,潜在空间U-Net预测深度,无需单独的掩模预测分支。结果:使用固定的训练/验证/测试划分,所提出的模型实现了3.70 mm的MAE、0.0326的AbsRel、delta=1.1的准确率0.962以及delta=1.1²的准确率0.970。它的MAE低于双U-Net MaskNet+DepthNet基线,并且在MAE、AbsRel和阈值准确率方面优于现成的单目深度模型。该管道在NVIDIA A100 GPU上,对301个连续帧以26.0 Hz的速率运行。结论:LED照明的二元图案平台结合潜在空间深度重建,可实现无同步、视频速率的内窥镜深度估计。结果证明了无需显式分割阶段即可进行Zivid参考的体模重建,同时强调了数据集大小和SSLE-Zivid校准精度的重要性。

英文摘要

Significance. Accurate intraoperative depth perception is important for autonomous and semi-autonomous robotic laparoscopic surgery. Conventional fringe projection profilometry can achieve millimeter-scale accuracy but often requires multi-shot acquisition, digital-micromirror-device projection, and projector-camera synchronization, complicating integration into compact laparoscopic systems. Aim. To develop a synchronization-free, single-shot depth-sensing platform using a passive LED-illuminated binary mask and a VQ-VAE prior with a custom U-Net depth head. Approach. A compact projection module was coupled to one channel of a dual-channel laparoscope, while the second channel imaged the fringe-illuminated target. A Zivid 3D camera acquired reference depth for 722 paired phantom images. Zivid depth maps were reprojected into the SSLE image frame for supervised training and evaluation. The VQ-VAE encoded each input into a discrete latent representation, and a latent-space U-Net predicted depth without a separate mask-prediction branch. Results. Using a fixed train/validation/test split, the proposed model achieved an MAE of 3.70 mm, AbsRel of 0.0326, delta=1.1 accuracy of 0.962, and delta=1.1^2 accuracy of 0.970. It achieved lower MAE than the dual U-Net MaskNet + DepthNet baseline and outperformed off-the-shelf monocular depth models in MAE, AbsRel, and threshold accuracy. The pipeline operated at 26.0 Hz over 301 consecutive frames on an NVIDIA A100 GPU. Conclusions. The LED-illuminated binary-pattern platform with latent-space depth reconstruction enables synchronization-free, video-rate endoscopic depth estimation. Results demonstrate Zivid-referenced phantom reconstruction without an explicit segmentation stage, while emphasizing the importance of dataset size and SSLE-Zivid calibration accuracy.

Comments17 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑