手术视频语言模型的分层感知评估与双曲基线
A Hierarchy-Aware Video-Language Model Evaluation and Hyperbolic Baseline for Surgery
- University of Amsterdam(阿姆斯特丹大学)
- Amsterdam UMC(阿姆斯特丹大学医学中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对手术视频理解中忽略层级结构的评估问题,本文提出分层感知评估套件SurgHiBench和双曲模型HyperSurg,实验表明双曲几何能改善预测正确性并随层级树状度提升增益。
AI中文摘要:
外科手术遵循从阶段到步骤的层级结构,然而用于识别这些结构的视频语言模型却采用扁平化的逐级指标进行评估,忽略了跨层级的一致性和错误结构。在本文中,我们针对这一问题做出了两项贡献:(i)我们引入了SurgHiBench,这是首个面向手术视频理解的分层感知评估套件,包含三项任务,分别衡量不同粒度级别上的识别、一致性和严重性。我们评估了一个通用型CLIP模型、一个欧几里得手术模型,以及作为第二项贡献的(ii)HyperSurg,这是一种新的双曲模型,通过蕴含锥体强制阶段-步骤的包含关系,并在四个(现有的)跨越三种手术类型的数据集上进行评估。该套件揭示,两个具有相同准确率的模型可能产生严重程度差异很大的预测错误,范围从正确阶段内的兄弟类别混淆到无关的跨阶段预测。双曲几何将预测推向正确的手术邻域,并且这些增益随每个数据集注释层级结构的树状程度而扩展,为分层感知几何何时有帮助提供了原则性指标。
英文摘要:
Surgical procedures follow a phase-to-step hierarchy, yet the video-language models used to recognize them are evaluated with flat per-level metrics that ignore cross-level coherence and error structure. In this paper we make two contributions to address this problem, (i) we introduce SurgHiBench, the first hierarchy-aware evaluation suite for surgical video understanding, with three tasks measuring recognition, consistency, and severity across granularity levels. We evaluate a general-purpose CLIP model, a Euclidean surgical model, and, as second contribution: (ii) HyperSurg, a new hyperbolic model that enforces phase-step containment via entailment cones, across four (existing) datasets spanning three procedure types. The suite reveals that two models with the same accuracy can produce predictions of very different error severity, ranging from sibling confusions within the correct phase to unrelated cross-phase predictions. Hyperbolic geometry shifts predictions toward the correct procedural neighborhood, and these gains scale with the tree-likeness of each dataset's annotation hierarchy, providing a principled indicator when hierarchy-aware geometry helps.