发表机构
NVIDIA; Clemson University(英伟达; 克莱姆森大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VANTAGE-Bench评估视觉语言模型在基础设施AI(固定摄像头场景)中的性能差距,涵盖多任务和视频跟踪,发现时间能力最弱,开放权重模型在2D定位领先。
AI 中文摘要
随着视觉语言模型(VLMs)向物理部署迈进,研究重点一直放在以主体为中心的消费视频上评估的面向行动的具身AI。这忽视了一类普遍的物理AI:基础设施AI,它依赖固定摄像头进行开放循环洞察,如安全监控和操作日志。我们引入了VANTAGE-Bench,一个衡量这种“基础设施AI差距”的基准。它涵盖三个运营领域(物流、交通和智能空间),统一了跨语义、空间、时间和时空能力的图像和视频评估,并超越了多项选择,扩展到八种任务形式,包括密集字幕和时空定位。它增加了单目标跟踪的单遍轨迹协议,并且据我们所知,这是首次在固定摄像头基础设施视频上进行此类评估,与专业跟踪器进行评分。标注涵盖3,346个媒体资产的三种模式:3,342个视频任务标注,4,281个图像定位标注,以及27,404个检测框。对17个模型进行零样本评估,我们发现相对于以消费者为中心的基准的不足是集中的,而非普遍的。事件验证、指代表达和时间定位在每个模型规模下都下降约9到24个百分点,而视频问答保持在VideoMME的5.3个百分点以内,2D空间指向与BLINK相比没有差距。时间支柱在绝对项上最弱:没有系统在时间定位上超过55.7 mIoU,或在密集视频字幕上超过37.3 SODA_c。在跟踪方面,前沿模型在短时间范围内接近专业跟踪器约5个百分点,但随着时间范围延长而分离。开放权重模型在2D目标定位上完全领先,因此规模和专有访问都不能解释这种模式。数据、评估工具和排行榜:此https URL。
英文摘要
As Vision-Language Models (VLMs) advance toward physical deployment, the focus has remained on action-oriented Embodied AI evaluated on subject-centric consumer video. This overlooks a pervasive class of Physical AI: Infrastructure AI, which relies on fixed cameras for open-loop insights like safety monitoring and operational logging. We introduce VANTAGE-Bench, a benchmark measuring this "Infrastructure AI Gap." It spans three operational domains (Logistics, Transportation, and Smart Spaces), unifies image and video evaluation across semantic, spatial, temporal, and spatio-temporal capabilities, and moves beyond multiple-choice to eight task formulations including dense captioning and spatio-temporal grounding. It adds a single-pass trajectory protocol for Single Object Tracking and, to our knowledge, the first such evaluation on fixed-camera infrastructure video, scored against specialist trackers. Annotation spans three regimes over 3,346 media assets: 3,342 video-task annotations, 4,281 image-grounding annotations, and 27,404 detection boxes. Evaluating 17 models zero-shot, we find the shortfall relative to consumer-centric benchmarks is concentrated, not general. Event verification, referring expressions, and temporal localization fall roughly 9 to 24 points at every model scale, while video question answering stays within 5.3 points of VideoMME and 2D spatial pointing shows no shortfall against BLINK. The temporal pillar is weakest in absolute terms: no system exceeds 55.7 mIoU on temporal localization or 37.3 SODA_c on dense video captioning. On tracking, frontier models come within roughly 5 points of specialist trackers over short horizons but separate as the horizon extends. Open-weight models lead 2D object localization outright, so neither scale nor proprietary access explains the pattern. Data, evaluation harness, and leaderboard: https://vantage-bench.org/
Comments23 pages, 2 figures, 14 tables. Project page: https://vantage-bench.org/; dataset: https://huggingface.co/datasets/nvidia/PhysicalAI-VANTAGE-Bench;