发表机构
The University of Hong Kong; The Chinese University of Hong Kong; Technical University of Munich; Hong Kong Polytechnic University(香港大学; 香港中文大学; 慕尼黑工业大学; 香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出Recon2Servo框架,利用DINOv3编码器和双向关系模块从B模式图像学习图像到运动推理,实现六自由度超声探头视觉伺服,并通过重建引导闭环适应提升泛化与精度。
AI 中文摘要
超声视觉伺服对于自主机器人超声至关重要,然而,由于离面运动线索有限且模糊,从二维B模式图像进行六自由度探头控制仍然具有挑战性。现有方法通常依赖解剖先验或手工制作的视觉特征,限制了其在不同成像目标上的泛化能力。受无追踪器三维超声重建的启发,我们提出了Recon2Servo,一种直接从B模式图像学习图像到运动推理的视觉伺服框架,用于六自由度探头控制。采用带低秩适应的DINOv3编码器和双向关系模块,估计当前图像与目标图像之间的相对探头位姿,以引导迭代闭环目标视图对齐。该框架结合了监督相对位姿学习、重建引导的闭环适应和有界残差位姿校正,以改善伺服过程中的运动推理。在公共数据集和使用不同超声系统从12名健康志愿者收集的内部数据集上的评估,证明了其在重建体积伺服中的有效性。补充视频中提供了在人体前臂上进行目标视图对齐和动态跟踪的额外真实机器人演示:此https URL。
英文摘要
Ultrasound visual servoing is essential for autonomous robotic ultrasound, yet 6-DoF probe control from 2D B-mode images remains challenging due to limited and ambiguous out-of-plane motion cues. Existing methods typically rely on anatomical priors or handcrafted visual features, limiting their generalizability across imaging targets. Inspired by trackerless 3D ultrasound reconstruction, we propose Recon2Servo, a visual servoing framework that learns image-to-motion inference directly from B-mode images for 6-DoF probe control. A DINOv3 encoder with low-rank adaptation and a bidirectional relation module estimate the relative probe pose between current and target images to guide iterative closed-loop target-view alignment. The framework combines supervised relative-pose learning, reconstruction-guided closed-loop adaptation, and bounded residual pose correction to improve motion inference during servoing. Evaluations on a public dataset and an in-house dataset collected from 12 healthy volunteers using different ultrasound systems demonstrate its effectiveness in reconstructed-volume servoing. Additional real-robot demonstrations of target-view alignment and dynamic tracking on a human forearm are provided in the supplementary video: https://youtu.be/qwsOdI-GMYk.