超越规模:手术技能模型是否能学习到跨评估准则的可迁移表征?
Looking Beyond the Scale: Do Surgical Skill Models Learn Transferable Representations Across Assessment Rubrics?
浏览论文内容
中文总结 AI 辅助
本文研究手术技能模型的跨评估准则可迁移性,发现JIGSAWS预训练骨干在LASANA上迁移可行但反向迁移失败,任务特定头部主导技能预测,视觉成分非唯一预测因素。
中文摘要 AI 辅助
基于视觉的手术技能评估已在领域内取得良好效果,但一个核心问题尚未被探讨:这些模型是学习到了可迁移的手术熟练程度表征,还是仅编码了数据集特有的视觉模式?本文使用LASANA和JIGSAWS数据集,系统分析了限制GOALS与OSATS评估准则间跨域技能迁移的因素。每种评估方法都有针对性诊断目的:端到端训练用于测试有监督技能学习是否可直接迁移;自适应锐度感知最小化(ASAM)用于探究更平坦的损失景观是否能提升泛化能力;基于增强的自监督学习与对比学习用于评估域不变预训练是否能将技能与视觉上下文解耦。使用JIGSAWS的不相交参与者留出测试集,双向评估迁移效果。结果显示存在不对称性:在JIGSAWS上预训练的骨干网络在LASANA上的CCC值达0.77至0.80,与端到端基线接近,说明当目标域提供一致监督时,跨准则迁移是可行的。所有方法向JIGSAWS的迁移均失败,可能源于标注不一致。使用Kinetics预训练骨干网络的对照实验表明,任务特定头部承担了大部分技能预测负担,而骨干网络仅需提供足够的时空特征。这些发现为基于视觉的技能评估提供了新视角:技能表征是否跨评分系统迁移这一核心问题此前未被研究,结果表明视觉成分在技能预测中占主导但非唯一因素,未来需进一步将可迁移技能特征与特定视觉域绑定的特征明确区分。
英文摘要
Vision-based surgical skill assessment has shown strong in-domain results, yet a fundamental question remains unasked: do these models learn transferable representations of surgical proficiency, or do they merely encode dataset-specific visual patterns? This paper systematically analyzes what limits cross-domain skill transfer between the GOALS and OSATS assessment scales using the LASANA and JIGSAWS datasets. Each evaluated method serves a targeted diagnostic purpose: end-to-end training to test whether supervised skill learning transfers directly, Adaptive Sharpness-Aware Minimization (ASAM) to probe whether flatter loss landscapes improve generalization, and augmentation-based self-supervised and contrastive learning to assess whether domain-invariant pretraining decouples skill from visual context. Transfer is evaluated in both directions using a disjoint-participant held-out test set for JIGSAWS. Results reveal an asymmetry: backbones pretrained on JIGSAWS achieve CCC values of 0.77 to 0.80 on LASANA, closely matching the end-to-end baseline, showing cross-rubric transfer is feasible when the target domain provides consistent supervision. Transfer to JIGSAWS fails across all methods, likely due to annotation inconsistencies. Control experiments with a Kinetics-pretrained backbone suggest task-specific heads carry the majority of the skill prediction burden, while the backbone need only provide adequate spatiotemporal features. These findings offer a new perspective on vision-based skill assessment: the central question of whether skill representations transfer across scoring systems has not been previously investigated. Results indicate the visual component is dominant but not solely responsible for skill prediction; further work is needed to conclusively disentangle transferable skill features from those bound to a specific visual domain.
发表机构
- National Center for Tumor Diseases (NCT)(国家肿瘤疾病中心)
- DKFZ(德国癌症研究中心)
- Faculty of Medicine and University Hospital Carl Gustav Carus(卡尔·古斯塔夫·卡鲁斯医学院及大学医院)
- TUD Dresden University of Technology(德累斯顿工业大学)
- Helmholtz-Zentrum Dresden-Rossendorf (HZDR)(德累斯顿-罗森多夫亥姆霍兹中心)
- BMFTR Research Hub 6G-Life(BMFTR 6G生活研究中心)
- Deutsches Krebsforschungszentrum (DKFZ)(德国癌症研究中心)
- The Centre for Tactile Internet with Human-in-the-Loop (CeTI)(人在回路触觉互联网中心)
机构由 AI 辅助整理,请以论文原文为准。