AI 中文总结
EndoWAM是首个用于通用机器人内窥镜导航的世界-动作模型,通过未来接地机制结合轻量型扩散Transformer与离散动作专家,在EndoMotion数据集上优于基线,具强零样本泛化能力。
AI 中文摘要
自主内窥镜导航可减轻临床医生的操作负担,但由于组织变形、瞬时遮挡和快速变化的视点,鲁棒控制仍然具有挑战性。现有基于学习的策略通常从当前观测值预测动作,未明确建模未来动态,限制了其在安全关键场景中的鲁棒性和可靠性。世界-动作模型(World Action Models, WAMs)是一种有前景的替代方案,它将预测性视觉动态与动作生成相结合,但由于训练数据有限、视点多样性受限、解剖结构可变形以及推理延迟高,将其扩展到机器人内窥镜领域仍然具有挑战性。我们提出EndoWAM,据我们所知,这是首个用于通用机器人内窥镜导航的WAM。EndoWAM引入了未来接地机制,该机制从视频世界模型的中间去噪特征中预测未来观测值中与任务相关的目标区域。具体而言,EndoWAM通过共享预测表示将用于未来目标区域预测的轻量型扩散Transformer与离散动作专家相结合。该设计将目标感知监督注入预测性动态建模中,提高了对视觉退化和视点变化的鲁棒性,同时通过单次去噪传递实现了实时控制。我们还引入了EndoMotion,这是一个机器人内窥镜运动数据集,涵盖三种解剖结构不同的手术:输尿管镜检查、食管镜检查和内镜逆行胰胆管造影(ERCP)。EndoWAM始终优于所有基线和替代接地策略,同时展示了对未见过的视点、环境和目标的强零样本泛化能力。这些结果确立了EndoWAM作为一种预测性、目标接地框架,适用于视觉受限的内窥镜环境中准确、通用且长时程的导航。
英文摘要
Autonomous endoscopic navigation can reduce clinicians' operational burden, yet robust control remains challenging due to tissue deformation, transient occlusions, and rapidly changing viewpoints. Existing learning-based policies typically predict actions from current observations without explicitly modeling future dynamics, limiting their robustness and reliability in safety-critical settings. World Action Models (WAMs) offer a promising alternative by coupling predictive visual dynamics with action generation, but extending them to robotic endoscopy remains challenging due to limited training data, restricted viewpoint diversity, deformable anatomy, and high inference latency. We present EndoWAM, which is, to our knowledge, the first WAM for generalizable robotic endoscopic navigation. EndoWAM introduces future grounding, which predicts task-relevant target regions in future observations from intermediate denoising features of a video world model. Specifically, EndoWAM couples a lightweight diffusion transformer for future target-region prediction with a discrete action expert through a shared predictive representation. This design injects target-aware supervision into predictive dynamics modeling, improving robustness to visual degradation and viewpoint changes while enabling real-time control in a single denoising pass. We further introduce EndoMotion, a robotic endoscopic motion dataset spanning three anatomically distinct procedures: ureteroscopy, esophagoscopy, and endoscopic retrograde cholangiopancreatography (ERCP). EndoWAM consistently outperforms all baselines and alternative grounding strategies, while demonstrating strong zero-shot generalization to unseen viewpoints, environments, and targets. These results establish EndoWAM as a predictive, target-grounded framework for accurate, generalizable, and long-horizon navigation in visually constrained endoscopic environments.