发表机构
German Cancer Research Center (DKFZ) Heidelberg; Heidelberg University; National Center for Tumor Diseases (NCT); Helmholtz Imaging; Stanford University; University of Pennsylvania(德国癌症研究中心(DKFZ)海德堡; 海德堡大学; 国家肿瘤疾病中心(NCT); 亥姆霍兹影像; 斯坦福大学; 宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有视频理解评估忽视长时间跨度时间一致性的问题,提出临床基础数据集HeiCo-FOCUS,含3万VQA对覆盖五项能力,多轨评估显示前沿模型在时间定位上平均准确率仅19.7%,凸显挑战并促进长视频可靠推理模型发展。
AI 中文摘要
近期视觉语言模型(VLM)的进展已在广泛的基准任务中推动了视频理解的快速进步。然而,现有评估主要聚焦于短期推理,未能评估一项关键能力:在长时间跨度上维持累积时间一致性。为弥补这一评估缺口,我们引入了HeiCo-FOCUS,一个通过外科手术中异物情境理解任务来评估长上下文视频理解的临床基础数据集。该数据集基于海德堡结直肠手术数据集构建,要求模型在持续数小时的手术过程中,连续跟踪多个物体,包括其插入、操作、遮挡和移除。HeiCo-FOCUS包含30,000对视觉问答(VQA)数据,覆盖五项核心能力:物体识别、时间定位、聚合、事件与程序理解以及复杂推理。该数据集通过严格的多阶段标注流程构建,涉及大规模众包标注和39名外科领域专家,以确保高质量和临床相关性。为系统性地探测模型行为,我们引入了一个多轨评估框架,从单帧到完整手术过程逐步增加时间和上下文需求。对十个前沿VLM的实验表明,HeiCo-FOCUS任务远未解决:仅约一半的模型明显优于纯文本基线。在视频轨道上,模型在事件与程序理解方面表现最佳(所有模型的平均准确率为56.5%),而时间定位对所有评估模型而言仍特别具有挑战性(平均准确率为19.7%)。因此,我们期望HeiCo-FOCUS能成为开发能够在数小时视频上进行可靠、时间一致推理的模型的催化剂。
英文摘要
Recent advances in Vision-Language Models (VLMs) have led to rapid progress in video understanding across a wide range of benchmark tasks. However, existing evaluations largely focus on short-term reasoning, failing to assess a critical capability: maintaining cumulative temporal consistency over extended time horizons. To close this evaluation gap, we introduce HeiCo-FOCUS, a clinically grounded dataset for evaluating long-context video understanding through the task of Foreign Object Contextual Understanding in Surgery. Built on a dataset of Heidelberg Colorectal surgeries, this task requires models to continuously track multiple objects as they are inserted, manipulated, occluded, and removed over procedures lasting up to hours. HeiCo-FOCUS comprises 30,000 visual question answering (VQA) pairs covering five core capabilities: object recognition, temporal grounding, aggregation, event and procedural understanding, and complex reasoning. The dataset was constructed through a rigorous multi-stage annotation pipeline involving large-scale crowd annotation and 39 surgical domain experts to ensure high quality and clinical relevance. To systematically probe model behavior, we introduce a multi-track evaluation framework that progressively increases temporal and contextual demands from single frames to full procedures. Experiments with ten frontier VLMs show that HeiCo-FOCUS tasks are far from solved: only around half of the models clearly outperform a text-only baseline. Across the video tracks, models perform best on event and procedural understanding (mean Accuracy: 56.5% across all models), while temporal grounding remains particularly challenging for all evaluated models (mean Accuracy: 19.7%). We therefore expect HeiCo-FOCUS to serve as a catalyst for the development of models capable of reliable, temporally consistent reasoning over hours-long videos.
Comments28 pages, 9 figures, 5 tables. Code: https://github.com/IMSY-DKFZ/orena-focus