发表机构
Eindhoven University of Technology(埃因霍温理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出检索诱导潜在转移模型,结合V-JEPA 2.1世界模型与参考数据库检索,实现超声操作者多步规划指导,在颈动脉超声中达到86%回顾性目标视图可达率,优于基线,并支持无需探头跟踪的实时部署。
AI 中文摘要
超声检查应用广泛,但图像采集质量高度依赖于操作者的知识与专业技能。由于检查需求超过受过培训的超声技师供应,操作者指导系统旨在通过指导经验较少的用户如何移动探头朝向目标视图来弥合这一差距。在本文中,我们提出了一种用于超声采集动力学的检索诱导潜在转移模型,将超声操作者指导表述为在世界模型中的多步规划与检索。使用V-JEPA 2.1骨干网络,观测首先被编码到潜在空间中,其中解剖学相关的视图彼此接近。随后,我们从包含编码潜在状态及相应探头位置和方向的参考数据库中检索相似视图。我们不学习参数化转移函数,而是直接使用数据库中物理执行的转移来建立我们的非参数化、检索诱导转移模型,该模型支持后退时域规划。在部署时,仅从实时超声图像流生成指导,无需任何探头跟踪硬件。应用于颈动脉超声时,所提出的规划器在86%的回顾性闭环情节中达到目标视图,而代表性基线分别为52%和43%,在每个目标视图上均优于两者,包括具有挑战性的纵向颈内动脉和颈外动脉视图。一项在未见志愿者上进行的预期可行性研究,通过蒸馏在CPU上实时运行,达到83%的目标视图可达性。由于规划由与任何可编码目标潜在状态的接近度驱动,同一世界模型可以导航回任何先前获取的患者特定帧,支持可重复的纵向成像,例如用于围手术期或随访监测。
英文摘要
Ultrasound is widely used, but acquisition quality is heavily dependent on the operator's knowledge and expertise. With demand for examinations outpacing the supply of trained sonographers, operator-guidance systems aim to close this gap by instructing a less trained user how to move the probe toward a target view. In this paper, we propose a retrieval-induced latent transition model for ultrasound acquisition dynamics, formulating ultrasound operator guidance as multi-step planning and retrieval in a world model. Using a V-JEPA 2.1 backbone, observations are first encoded into a latent space where anatomically related views lie close together. We then retrieve similar views from a reference database containing encoded latent states and corresponding probe positions and orientations. Rather than learning a parametric transition function, we directly use physically executed transitions from the database to establish our nonparametric, retrieval-induced transition model that supports receding-horizon planning. At deployment, guidance is generated from the live ultrasound image feed alone, without any probe tracking hardware. Applied to carotid ultrasound, the proposed planner reaches the target view in 86% of retrospective closed-loop episodes, versus 52% and 43% for representative baselines, outperforming both on every target view, including the challenging longitudinal internal and external carotid artery views. A prospective feasibility study on unseen volunteers, run in real time on a CPU using distillation, reaches 83% target-view reachability. Because planning is driven by proximity to any encodable goal latent, the same world model can navigate back to any previously acquired, patient-specific frame, supporting reproducible longitudinal imaging for e.g. perioperative or follow-up monitoring.
Comments11 pages, 8 figures, 5 tables