arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22093cs.RO

EndoNav:面向语言引导的机器人内窥镜检查的语义到几何的 grounding(语义几何关联)

EndoNav: Semantic-to-Geometric Grounding for Language-Guided Robotic Endoscopic Examination

Jecia Z. Y. Mao, Hisashi Ishida, Kathryn Jung, Masaru Ishii, Russell H. Taylor, Manish Sahu

首次发表
浏览论文内容

中文总结 AI 辅助

研究人员提出 EndoNav 框架,可将外科医生高级命令转换为患者特异性鼻窦解剖内的自主内窥镜可视化行为,其可视化效果接近住院外科医生水平,验证了该方法的可行性。

中文摘要 AI 辅助

在狭小解剖空间内进行的微创手术依赖于持续的内窥镜可视化。当前的机器人内窥镜系统可稳定或重新定位内窥镜,但不具备提供有效可视化辅助的相关上下文。我们提出 EndoNav,这是一个基于解剖学的自然语言框架,可将外科医生的高级命令转换为患者特异性鼻窦解剖结构内的自主内窥镜可视化行为。外科医生的口头命令经转录后,由以患者特异性解剖场景表示为条件的内窥镜视点代理进行解释。该视点代理并非直接生成机器人运动,而是生成结构化可视化目标,这些目标被转换为目标视点和检查轨迹,随后通过受几何约束的内窥镜运动规划和关节空间控制执行。我们使用基于 CT 的三个解剖模型进行的结构化三次鼻窦检查来评估 EndoNav。对于一具尸体标本,将自主可视化与两名住院外科医生进行的鼻窦检查进行比较。EndoNav 相对于两名外科医生的检查分别实现了 87.04% 和 84.37% 的平均可视化 IoU,而外科医生间的 IoU 为 87.44%,同时分别恢复了外科医生观察到的 92.91% 和 93.20% 的解剖表面。这些结果证明了将高级解剖学命令转换为患者特异性几何目标并将其转换为受解剖学约束的机器人可视化行为的可行性。

英文摘要

Minimally invasive procedures performed within confined anatomical spaces depend on continuous endoscopic visualization. Current robotic endoscope systems can stabilize or reposition an endoscope, but they do not possess relevant context to provide effective visualization assistance. We present EndoNav, an anatomy-grounded natural-language framework that translates high-level surgeon commands into autonomous endoscopic visualization behaviors within patient-specific sinonasal anatomy. Spoken surgeon commands are transcribed and interpreted by an endoscopic viewpoint agent conditioned on a patient-specific anatomical scene representation. Rather than generating robot motion directly, the viewpoint agent generates structured visualization objectives that are converted into target viewpoints and inspection trajectories, which are then executed through geometry-constrained endoscope motion planning and joint-space control. We evaluate EndoNav using a structured three-pass sinus examination across three CT-derived anatomical models. For one cadaveric specimen, autonomous visualization is compared with sinus examinations performed by two resident surgeons. EndoNav achieved mean visualization IoUs of 87.04% and 84.37% relative to the two surgeon examinations, compared with an inter-surgeon IoU of 87.44%, while recovering 92.91% and 93.20% of surgeon-observed anatomical surfaces, respectively. These results demonstrate the feasibility of grounding high-level anatomical commands into patient-specific geometric objectives and translating them into anatomically constrained robotic visualization behaviors.

↑