arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

边扫描边想象:用于机器人超声导航的场景图世界模型

Scanning While Imagining: A Scene-Graph World Model for Robotic Ultrasound Navigation

Xuesong Li, Shuai Chen, Feng Li, Zhongliang Jiang, Nassir Navab, Yuan Bi

arXiv 2609.32837首次发表:更新:

发表机构

Munich Center for Machine Learning(慕尼黑机器学习中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出SonoGraph-WM场景图世界模型,通过预测解剖变化和规划最短路径,实现机器人超声探头导航,在CT和体模实验中取得高成功率。

AI 中文摘要

超声(US)采集依赖于操作者解读解剖结构并预测探头移动时视图如何变化的能力。许多机器人超声导航方法在未显式预测这些解剖变化的情况下选择动作。我们提出SonoGraph-WM,一种用于预期性探头导航的、以动作和目标为条件的世界模型。该模型将解剖结构表示为场景图(SGs),捕捉可见结构、其几何形状和空间关系,而无需合成超声图像。给定SG和探头位姿的历史,一个统一的Transformer联合预测未来的SG和位姿。一个滚动时域规划器递归地想象候选轨迹,选择到达目标图的最短预测路径,并在短执行时域内跟随该路径,然后根据新的观测重新规划。为减少对带跟踪和带解剖标注的超声序列的依赖,我们沿表面约束的探头轨迹,从计算机断层扫描(CT)标签图生成对齐的SG-位姿训练数据。在四个保留的CT病例中,空间关系F1在20个预测步骤内保持在93%以上,闭环导航在胆囊和胰腺上分别达到77.50%和75.00%的成功率(使用来自标注的SG)。在机器人-体模导航实验中,使用来自标签图的SG,规划器在73.7%的试验中到达目标视图。这些发现支持CT监督的解剖世界建模用于探头规划,并强调频繁观测更新对可靠导航的重要性。项目页面:此https URL

英文摘要

Ultrasound (US) acquisition depends on the operator's ability to interpret anatomy and anticipate how the view will change with probe motion. Many robotic US navigation methods select actions without explicitly predicting these anatomical changes. We propose SonoGraph-WM, an action- and goal-conditioned world model for anticipatory probe navigation. The model represents anatomy as scene graphs (SGs), capturing visible structures, their geometry, and spatial relationships without synthesizing US images. Given a history of SGs and probe poses, a unified Transformer jointly predicts future SGs and poses. A receding-horizon planner recursively imagines candidate trajectories, selects the shortest predicted path reaching a goal graph, and follows it over a short execution horizon before replanning from new observations. To reduce reliance on tracked and anatomically annotated US sequences, we generate aligned SG--pose training data from computed tomography (CT) label maps along surface-constrained probe trajectories. On four held-out CT cases, spatial relation F1 remains above 93% over 20 prediction steps, and closed-loop navigation achieves 77.50% and 75.00% success for the gallbladder and pancreas, respectively, using annotation-derived SGs. In robot--phantom navigation experiments with label-map-derived SGs, the planner reached the target view in 73.7% of trials. These findings support CT-supervised anatomical world modeling for probe planning and highlight the importance of frequent observation updates for reliable navigation. Project Page: https://noseefood.github.io/us-sonograph-wm/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑