UltraWorld:利用声学采样图从无跟踪临床视频学习交互式超声世界模型
UltraWorld: Learning Interactive Ultrasound World Models from Untracked Clinical Videos with Acoustic Sampling Map
浏览论文内容
中文总结 AI 辅助
UltraWorld通过自蒸馏临床超声视频构建交互式世界模型,利用声学采样图提升动作跟随,在模拟规划中显著降低误差。
中文摘要 AI 辅助
世界模型能够通过从局部观测预测探头运动的结果,实现自主超声扫描。学习这种动作-观测关系通常依赖于同步的视频-位姿对,这些数据在大规模采集时成本高昂,且在常规临床记录中基本不可用。可靠的动作跟随还需要对超声的横截面采样几何进行建模。我们提出了UltraWorld,一种自蒸馏方法,将临床超声视频中的先验知识转移为交互式世界模型,无需真实动作标注。从临床视频出发,我们将视频基础模型调整为以参考图像和解剖掩码为条件的超声生成器。沿可编程轨迹通过3D解剖结构采样的解剖掩码,为合成动作-视频对提供空间指导。然后,我们使用这些合成对将生成器自蒸馏为世界模型,该模型从局部观测和动作预测未来观测,在推理时无需解剖掩码或其他3D资源。为了进一步提高动作跟随能力,我们引入了声学采样图(AsMap),它将探头位姿和成像设置表示为逐像素的3D采样位置、波束方向和深度。实验证明了预测保真度和动作跟随能力的提升。在九个模拟闭环局部规划场景中,与视觉伺服相比,UltraWorld将最终到目标的平均距离和方向误差分别降低了29%和38%。项目页面:此https URL。
英文摘要
World models can enable autonomous ultrasound scanning by predicting the outcomes of probe motions from local observations. Learning this action--observation relationship typically relies on synchronized video--pose pairs, which are costly to collect at scale and largely unavailable in routine clinical recordings. Reliable action following further requires modeling ultrasound's cross-sectional sampling geometry. We present UltraWorld, a self-distillation recipe that transfers priors from clinical ultrasound videos into interactive world models without real action annotations. Starting from clinical videos, we adapt a video foundation model into an ultrasound generator conditioned on reference images and anatomical masks. Anatomical masks sampled along programmable trajectories through 3D anatomy provide spatial guidance for synthesizing action--video pairs. We then use these synthetic pairs to self-distill the generator into a world model that predicts future observations from local observations and actions, without requiring anatomical masks or other 3D assets at inference time. To further improve action following, we introduce the Acoustic Sampling Map (AsMap), which represents probe poses and imaging settings as pixel-wise 3D sampling positions, beam directions, and depths. Experiments demonstrate improved prediction fidelity and action following. Across nine simulated closed-loop local planning episodes, UltraWorld reduces the mean final distance to the goal and orientation error by 29\% and 38\%, respectively, compared with visual servoing. Project Page: https://ultraworld-project.github.io/.
发表机构
- The Eighth Affiliated Hospital, Sun Yat-sen University(中山大学附属第八医院)
- The Chinese University of Hong Kong(香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。