arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Surgical WAM:面向数据高效型手术机器人学习的世界-动作模型

Surgical WAM: A World-Action Model for Data-Efficient Surgical Robot Learning

Wenrui Bao, Tianyun Jiang, Zhiben Chen, Ser-Nam Lim, Peter D. Peng, Yuzhang Shang

arXiv 2608.11204首次发表:更新:

AI 中文总结

本文提出Surgical WAM模型,通过无动作视频预训练结合带动作标注数据微调,在四项模拟手术操纵任务中将平均成功率从63.5%提至77.8%,为数据高效的手术机器人学习提供了可行方案。

AI 中文摘要

学习可靠的手术操纵策略受限于带动作标注的演示数据稀缺:远程操作手术机器人(如dVRK)带同步运动学数据的轨迹采集成本高昂,而手术任务要求精准的接触处理、长程推理及双手协调。与同步视频-运动学轨迹相比,内窥镜视频成本更低且数量更充足,利用它的自然方式是学习手术场景的世界模型。然而,现有手术世界模型主要将视频用于模拟或策略评估,极少将学习到的动力学转化为闭环控制。这一差距引出核心问题:在带动作标注演示数据预算固定的情况下,无动作视频预训练能否改善闭环手术操纵?为回答该问题,我们提出Surgical World-Action Model(Surgical WAM),这是基于Cosmos Policy构建的统一生成模型,可联合预测未来内窥镜观测结果与可执行手术机器人动作块。Surgical WAM首先从无动作视频中学习手术视觉动力学,随后在固定的带动作标注预算上进行微调;部署时,它作为闭环、后退时域控制器,执行每个预测动作块的短前缀,并根据生成的观测结果重新规划。在四个模拟手术操纵任务的套件上,视频预训练将平均成功率从63.5%提升至77.8%,其中PegTransfer任务的绝对提升达20个百分点,在接触密集型和双手任务上的提升最为显著。这些结果表明,无动作视频可为有限动作监督下的手术机器人控制学习提供可迁移的视觉动力学先验,确立了数据高效型视频预训练作为扩大手术机器人学习规模的实用路径。

英文摘要

Learning reliable surgical manipulation policies is bottlenecked by the scarcity of action-labeled demonstrations: teleoperated surgical robot (e.g., dVRK) trajectories with synchronized kinematics are costly to collect, while surgical tasks demand precise contact handling, long-horizon reasoning, and bimanual coordination. Endoscopic video is comparatively inexpensive and abundant relative to synchronized video--kinematics trajectories, and a natural way to exploit it is to learn world models of surgical scenes. However, existing surgical world models use video primarily for simulation or policy evaluation, and rarely translate the learned dynamics into closed-loop control. This gap raises our central question: under a fixed budget of action-labeled demonstrations, does action-free video pretraining improve closed-loop surgical manipulation? To answer it, we introduce the Surgical World-Action Model (Surgical WAM), a unified generative model built on Cosmos Policy that jointly predicts future endoscopic observations and executable surgical robot action chunks. Surgical WAM first learns surgical visual dynamics from action-free video and is then fine-tuned on the fixed action-labeled budget; at deployment, it acts as a closed-loop, receding-horizon controller that executes a short prefix of each predicted action chunk and replans from the resulting observation. On a suite of four simulated surgical manipulation tasks, video pretraining improves the average success rate from 63.5% to 77.8%, including an absolute gain of 20 percentage points on PegTransfer, with the largest improvements on contact-rich and bimanual tasks. These results demonstrate that action-free video provides transferable visual dynamics priors for learning surgical robot control with limited action supervision, positioning data-efficient video pretraining as a practical path toward scaling up surgical robot learning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑