arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SurgSkill-Bench:多模态外科手术技能评估基准

SurgSkill-Bench: A Benchmark for Multimodal Surgical Skill Assessment

Chaohui Dang, Zheheng Jiang, James Glasbey, David Luke, Theodoros Arvanitis, Le Zhang

arXiv 2608.30872首次发表:更新:

发表机构

University of Birmingham; University of Leicester; University Hospitals North Midlands(伯明翰大学; 莱斯特大学; 中北部大学医院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究构建了多模态外科手术技能评估基准SurgSkill-Bench,设计两种评估设置,通过实验验证内容自适应采样和专家评论辅助可提升评分预测性能,最佳平均AUROC达0.88。

AI 中文摘要

对外科手术技术技能的客观评估对外科培训和结构化反馈至关重要,但当前工作流程仍依赖劳动密集型的专家评审。现有自动化方法主要聚焦于视觉输入,在联合研究手术操作表现、结构化技能评分与专家反馈方面支持有限。我们推出SurgSkill-Bench,这是一个初始的视频-评分-文本基准式数据集,包含214个外科手术训练模拟视频、六维度OSATS评分以及专家自由文本评论。我们定义了两种评估设置:仅基于视频的OSATS预测(用于自动化评估)和事后专家评论辅助预测(其中评估者评论作为辅助信息)。我们使用代表性的冻结视觉骨干网络、内容自适应关键帧采样以及简单的视频-文本协同注意力融合模块开展受控基准实验。在内部视频级验证中,内容自适应采样提升了该数据集的仅视频预测性能,而评估者评论在辅助设置中提供了额外的与评分相关的信号。在数据集特定的中位数二分法下,最佳平均AUROC达到0.88。我们进一步讨论了与数据集规模、元数据完整性以及评论辅助预测解释相关的评估约束,代码将在后续公开发布。

英文摘要

Objective assessment of surgical technical skill is important for surgical training and structured feedback, but current workflows remain dependent on labor-intensive expert review. Existing automated approaches primarily focus on visual inputs and provide limited support for jointly studying operative performance, structured skill scores, and evaluator feedback. We introduce SurgSkill-Bench, an initial video-score-text benchmark-style dataset containing 214 surgical training simulation videos, six-dimensional OSATS scores, and expert free-text comments. We define two evaluation settings: video-only OSATS prediction for automated assessment and post hoc expert-comment-assisted prediction, where evaluator comments are available as auxiliary information. We provide controlled baseline experiments using representative frozen visual backbones, content-adaptive key-frame sampling, and a simple video-text co-attention fusion module. Under internal video-level validation, content-adaptive sampling improves video-only performance in this dataset, while evaluator comments provide additional score-related signal in the assisted setting. The best mean AUROC reaches 0.88 under dataset-specific median dichotomization. We further discuss evaluation constraints related to dataset scale, metadata completeness, and the interpretation of comment-assisted prediction. Code will be released publicly at a later date.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑