SurgVidLM: Towards Multi-grained Surgical Video Understanding with Large Language Model
SurgVidLM:迈向多粒度外科视频理解的大型语言模型
机构 * The Chinese University of Hong Kong(香港中文大学) ; Sun Yat-sen University(中山大学) ; University of Strasbourg(斯特拉斯堡大学) ; Technical University of Munich(慕尼黑技术大学) ; Monash University(墨尔本大学) ; Centre for Artificial Intelligence and Robotics, HKISI-CAS(人工智能与机器人中心,HKISI-CAS) ; Shanghai AI Laboratory(上海人工智能实验室)
专题命中 视频理解 :video understanding(title,abstract);video language model(abstract);video reasoning(abstract);分类 cs.CV
AI总结 SurgVidLM通过多粒度分析提升外科视频理解能力,结合全局与局部机制实现更精确的手术流程解析。