SurgVIL:利用开源外科视频扩展外科机器人模仿学习
SurgVIL: Scaling Surgical Robot Imitation Learning with Open-source Surgical Videos
查看机构详情
- Johns Hopkins University(约翰斯·霍普金斯大学)
- NVIDIA(英伟达公司)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
SurgVIL框架结合带运动学标签的体模机器人演示与开源外科视频,以弱监督补充运动学标签,在达芬奇机器人任务上提升了策略对真实组织等场景的泛化能力,为外科机器人模仿学习提供了可扩展方案。
中文摘要 AI 辅助
基于学习的外科机器人自主操作需要大规模的同步视频与机器人动作演示数据,但在临床或真实组织场景中这类数据极为稀缺,因为在受控研究系统之外通常无法获取机器人运动学数据。相比之下,在研究平台上采集的体模数据能提供准确的动作标签,但缺乏真实组织的视觉多样性。我们提出SurgVIL,这是一种利用开源外科视频扩展外科机器人模仿学习的框架。SurgVIL将带有运动学标签的体模机器人演示与来自开源数据集及在线资源的外科视频相结合,用于策略学习。由于这些视频缺少机器人运动标签,我们将近似运动学估计作为弱监督。我们在两项达芬奇(da Vinci)机器人任务上评估SurgVIL:夹针与胆囊切除术切割。在ACT、π₀和GR00T-H主干网络上,添加外科视频显著提升了对真实组织和分布外场景的泛化能力,表明存在一条从体模训练迈向可泛化外科机器人策略的可扩展路径。
英文摘要
Learning-based surgical robot autonomy requires large-scale demonstrations with synchronized videos and robot actions, but such data are exceedingly rare in clinical or realistic tissue settings because robot kinematics are typically inaccessible outside controlled research systems. In contrast, phantom data collected on research platforms provide accurate action labels but lack the visual diversity of real tissue. We propose SurgVIL, a framework for scaling surgical robot imitation learning using open-source surgical videos. SurgVIL combines kinematically labeled phantom robot demonstrations with surgical videos from open-source datasets and online sources for policy learning. Since these videos lack robot motion labels, we estimate approximate kinematics as weak supervision. We evaluate SurgVIL on two da Vinci robot tasks: needle pick-up and cholecystectomy cutting. Across ACT, $π_0$, and GR00T-H backbones, adding surgical videos substantially improves generalization to real-tissue and out-of-distribution settings, suggesting a scalable path from phantom training toward generalizable surgical robot policies.