发表机构
Department of Mechanical and Automation Engineering, Chinese University of Hong Kong; Department of Biomedical Engineering, City University of Hong Kong; Centre for Artificial Intelligence and Robotics Hong Kong Institute of Science & Innovation, Chinese Academy of Sciences; Institute of Automation, Chinese Academy of Sciences; Department of Surgical and Interventional Engineering, King’s College London; Department of Vascular Ultrasound, Xuanwu Hospital, the Capital Medical University(香港中文大学机械与自动化工程学系; 香港城市大学生物医学工程学系; 中国科学院香港创新研究院人工智能与机器人中心; 中国科学院自动化研究所; 伦敦国王大学外科与介入工程学系; 首都医科大学宣武医院血管超声科)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对机器人超声目标平面探头引导问题,提出基于模型的两阶段学习管道,包括潜在条件扩散世界模型和目标条件时间变换器,在自收集数据集及实际闭环实验中取得较好成果,证明学习到的超声动力学对训练目标导向机器人探头导航有潜力。
AI 中文摘要
我们提出了一种用于机器人超声中目标平面探头引导的动作条件世界模型框架,重点是颈部超声扫描。自主超声任务通常需要大量探头运动轨迹进行训练,但收集高质量演示劳动强度大,且因超声外观依赖接触、组织变形和视图相关声学伪像而难以构建显式模拟器。我们用基于模型的两阶段学习管道解决此问题。首先,潜在条件扩散世界模型根据最近的上下文帧、探头运动和时间偏移预测未来超声观察。其次,目标条件时间变换器预测有序探头运动并使用来自冻结世界模型的奖励进行微调。在自收集数据集上的实验表明,世界模型在目标导向扫描中保留了与动作相关的解剖结构。在实际闭环实验中,该框架对颈动脉引导的成功率为70.0%,对甲状腺引导的成功率为65.0%。这些结果证明了学习到的超声动力学在训练目标导向机器人探头导航方面的潜力。
英文摘要
We present an action-conditioned world model framework for goal plane probe guidance in robotic ultrasound, with a focus on neck ultrasound scanning. Autonomous ultrasound tasks often require large numbers of probe-motion trajectories for training, but collecting high-quality demonstrations is labor-intensive and explicit simulators are difficult to build because ultrasound appearance depends on contact, tissue deformation, and view-dependent acoustic artifacts. We address this problem with a two-stage model-based learning pipeline. First, a latent conditional diffusion world model predicts future ultrasound observations from recent context frames, probe motions and temporal offset. Second, a goal-conditioned temporal transformer predicts ordered probe motions and is fine-tuned using rewards from the frozen world model. Experiments on the self-collected dataset show that the world model preserves action-dependent anatomical structure on target-directed scans. In real-world closed loop experiments, the framework achieves success rates of 70.0\% for carotid guidance and 65.0\% for thyroid guidance. These results demonstrate the potential of learned ultrasound dynamics for training goal-directed robotic probe navigation.