OmniRAS:标准化机器人辅助手术领域的基础模型训练与评估
OmniRAS: Standardizing Foundation Model Training and Evaluation in Robot-Assisted Surgery
浏览论文内容
中文总结 AI 辅助
本文针对机器人辅助手术领域基础模型稀缺的问题,提出OmniRAS系列V-JEPA-2.1编码器,发布专属数据集并完成大规模预训练与多任务评估,取得最优适配结果。
中文摘要 AI 辅助
机器人辅助手术领域的基础模型数量较少,部分原因是难以收集大型机器人手术视频语料库,且现有模型大多仅在腹腔镜基准上进行评估,多数现有模型的评估基于少量公开基准,且主要聚焦于腹腔镜手术。本文提出OmniRAS,这是一组用于机器人辅助手术的10亿参数和20亿参数规模的V-JEPA-2.1编码器,并详细说明其训练过程:首先,发布两个密集标注的机器人胆囊切除术数据集:OmniRAS-PR以及多标签YT-Chole工具-动作-目标任务,这是首个针对机器人胆囊切除术的三元组式标注,同时提供数据划分、探测协议,以及验证共享阶段本体的标注者间一致性研究;其次,记录在多达256个计算节点上进行的持续预训练,全局批次大小为6144,涵盖19个来源、总计约2650小时的手术视频(其中51%为机器人手术视频),并分析计算资源与数据组成;最后,在三元组识别、阶段与步骤识别、动作分割、检测等6项任务上,针对原始V-JEPA-2.1及专用手术模型进行评估,评估采用冻结编码器和最终四个块微调两种设置,在3个随机种子下共完成254次下游运行,其中109次为部分骨干微调。结果显示,最优OmniRAS模型在所有任务类别中均取得最强的适配结果,而冻结设置下的性能差异较小。
英文摘要
Few foundation models exist for robot-assisted surgery, partly because large robotic-surgery video corpora are difficult to assemble and existing models are evaluated mostly on laparoscopic benchmarks. Further, most existing models are evaluated on a small set of public benchmarks, mostly focused on laparoscopic surgery. We present OmniRAS, a family of 1B- and 2B-parameter V-JEPA-2.1 encoders for robot-assisted surgery, and detail their training. First, we release two densely annotated robotic-cholecystectomy datasets: OmniRAS-PR and a multi-label YT-Chole tool-verb-target task, the first triplet-style annotation for robotic cholecystectomy, together with splits, probe protocols, and an inter-rater study validating the shared phase ontology. Second, we document continued pretraining at up to 256 compute nodes with global batch 6,144 over 19 sources totaling approximately 2,650 hours of surgical video, 51% robotic, and analyze compute and data composition. Third, we evaluate against raw V-JEPA-2.1 and specialized surgical models on six tasks spanning triplet, phase, and step recognition, action segmentation, and detection, under frozen-encoder and final-four-block fine-tuning regimes. Across three seeds, this yields 254 downstream runs, including 109 with partial backbone fine-tuning. The best OmniRAS models achieve the strongest adapted results across all task families, while frozen differences are smaller.
发表机构
- University of Illinois Chicago(伊利诺伊大学芝加哥分校)
- Argonne National Laboratory(阿贡国家实验室)
- Department of Surgery, University of Illinois Chicago(伊利诺伊大学芝加哥大学外科学院)
- Affiliated Hospital of Qingdao University(青岛大学附属医院)
机构由 AI 辅助整理,请以论文原文为准。