可靠双臂手术子任务操作的轨迹发散 horizon 决策
Trajectory Divergence Horizon Decision for Reliable Dual-Arm Surgical Subtask Manipulation
查看机构详情
- The Chinese University of Hong Kong (CUHK)(香港中文大学)
- Huawei Technologies Co. Ltd.(华为技术有限公司)
- Shenzhen Loop Area Institute(深圳河套学院)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对现有VLA策略在手术操作中易累积误差的问题,提出TDHD机制,构建双臂手术基准并收集600个演示,实验显示其可提升器械和组织操作的成功率。
中文摘要 AI 辅助
随着临床工作量增加,手术机器人系统正被越来越多地采用,这推动了针对重复性操作子任务的自主解决方案。基于学习的控制器相比基于规则和分析的方法提升了泛化能力,但大多数控制器针对单个任务训练,难以跨流程复用。视觉-语言-动作(VLA)模型提供了整合视觉感知、语言 grounding 和动作生成的统一框架,为实现更具组合性的手术自主操作提供了有前景的路径。然而,现有 VLA 策略依赖固定长度的开环动作序列,场景条件变化会导致误差累积,给手术操作带来潜在风险。为缓解该问题,我们将手术 VLA 的部署形式化为自适应执行 horizon 决策问题,并提出轨迹发散 horizon 决策(TDHD)——一种测试时机制,通过测量小噪声扰动下两个流匹配生成轨迹的发散度来估计逐步动作的可靠性,并采用双阈值规则截断执行以触发及时重规划。我们进一步建立了类似达芬奇(da Vinci)的真实双臂基准,配备同步多视角感知和语言指令,并在器械(到达、抓取、重抓取)和组织(到达、抬起、切除)操作套件中收集了600个遥操作演示。在真实硬件上每个任务设置进行20次试验,TDHD相比最新的VLA基线持续提升性能:器械操作的成功率从55%提升至60%,组织操作的成功率从55%提升至80%,在最终操作阶段的提升最为显著。这些结果凸显了自适应执行控制对手术机器人操作中VLA模型可靠部署的重要性。
英文摘要
Surgical robotic systems are increasingly being adopted as clinical workload rises, motivating autonomous solutions for repetitive manipulation subtasks. Learning-based controllers improve generalization compared with rule-based and analytic approaches, but most are trained for individual tasks and remain difficult to reuse across procedures. Vision-Language-Action (VLA) models provide a unified framework that integrates visual perception, language grounding, and action generation, offering a promising path toward more composable surgical autonomy. However, existing VLA policies rely on fixed-length open-loop action sequences, where changing scene conditions can lead to accumulated errors and potential risks in surgical manipulation. To mitigate this issue, we formulate surgical VLA deployment as an adaptive execution-horizon decision problem and propose Trajectory Divergence Horizon Decision (TDHD), a test-time mechanism that estimates step-wise action reliability by measuring the divergence between two flow-matching-generated trajectories under small noise perturbations and truncates execution using a dual-threshold rule to trigger timely replanning. We further establish a real-world da Vinci-like dual-arm benchmark with synchronized multi-view perception and language instructions, and collect 600 teleoperated demonstrations across needle (reach, pick, regrasp) and tissue (reach, lift, resection) manipulation suites. On real hardware with 20 trials per task setting, TDHD consistently improves performance over the latest VLA baselines: success increases from 55\% to 60\% for needle manipulation and from 55\% to 80\% for tissue manipulation, with the largest gains observed in the final manipulation stages. These results highlight the importance of adaptive execution control for reliable deployment of VLA models in surgical robotic manipulation.