基于冻结预训练视频基础模型的外科技能视频评估
Video Based Assessment of Surgical Skills Using Frozen Pretrained Video Foundation Models
- FAMU-FSU College of Engineering(佛罗里达农工大学-佛罗里达州立大学工程学院)
- Department of Biomedical Engineering, Rensselaer Polytechnic Institute(伦斯勒理工学院生物医学工程系)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出VBA-Net+框架,利用冻结预训练视频基础模型进行FLS技能评分回归与通过/失败分类,在缝合和图案切割数据集上实现高R²和AUC,优于帧级基线,为视频评估提供基准。
AI中文摘要:
基于视频的自动化外科技能评估已快速发展,但在参与者级别泛化到未见受训者的连续标准化评分预测的严格评估仍然有限。我们提出了VBA-Net+,一个仅使用视频的框架,用于腹腔镜手术基础(FLS)评分回归和通过/失败分类,使用预训练视频基础模型作为冻结特征提取器。我们使用VideoPrism、V-JEPA2和VideoMAE v2评估了两个FLS数据集,即缝合和图案切割,并以帧级SimCLR作为基线。一个轻量级的全卷积头在嵌入上离线训练,并在标准化评估协议内使用参与者级别的留一用户交叉验证(LOUO)进行评估。对于连续评分预测,最佳表示在缝合上达到R²=0.6367,在图案切割上达到0.9261。对于在官方FLS阈值下的通过/失败分类,接收者操作特征曲线下面积(AUC)分别达到0.9073和0.9906。冻结视频编码器流程通常优于帧级SimCLR流程,尤其是在缝合方面,为仅视频的FLS评估提供了基准。
英文摘要:
Automated video-based surgical skill assessment has advanced rapidly, yet rigorous evaluation of continuous standardized score prediction under participant-level generalization to unseen trainees remains limited. We introduce VBA-Net+, a video-only framework for Fundamentals of Laparoscopic Surgery (FLS) score regression and pass-fail classification using pretrained video foundation models as frozen feature extractors. We evaluate two FLS datasets, suturing and pattern cutting, using VideoPrism, V-JEPA2, and VideoMAE v2, with a frame-level SimCLR baseline. A lightweight fully convolutional head is trained on embeddings offline and evaluated using participant-level leave-one-user-out (LOUO) cross-validation within the standardized assessment protocol. For continuous score prediction, the best representation achieves $R^2$=0.6367 for suturing and 0.9261 for pattern cutting. For pass-fail classification at official FLS thresholds, area under the receiver operating characteristic curve (AUC) reaches 0.9073 and 0.9906, respectively. Frozen video-encoder pipelines generally outperformed the frame-level SimCLR pipeline, particularly for suturing, providing a benchmark for video-only FLS assessment.