发表机构
The University of Texas at Austin(德克萨斯大学奥斯汀分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出CrossView多相机视频问答基准,评估发现GPT-5.2等模型跨相机推理准确率低,开源模型表现更差,该基准可用于测试模型联合处理多视角的能力。
AI 中文摘要
视频理解基准长期以来聚焦于单相机场景,现代多模态语言模型在图像和视频任务上表现出色。但现实世界依赖多相机网络:自动驾驶汽车、安保系统和机器人均从多个同步视角采集数据。本文认为这并非单相机问题的简单扩展,而是本质不同的问题。跨相机推理需要处理随视角数量扩展的上下文、解决仅部分相机可见的遮挡、判断哪些视角重要,并整合重叠或分歧视角的证据。当前模型难以应对这些挑战,且缺乏针对该问题的系统基准。本文提出CrossView,一个涵盖自动驾驶、安保监控、第一人称/第三人称视频及机器人领域的多相机视频问答基准。对GPT-5.2等专有模型和Qwen3-VL等开源模型的评估显示,它们准确率均较低,且开源模型差距显著。性能与模型联合处理多视角的能力高度相关,这使CrossView成为严格的多相机视频基准。本文在该https URL开源代码和数据集。
英文摘要
Video understanding benchmarks have long centered on single-camera settings, where modern multi-modal language models achieve strong performance across image and video tasks. Yet, the real world runs on multi-camera networks: autonomous vehicles, security systems, and robots all gather data across many simultaneous views. We argue that this is not simply "more" of the single-camera problem; it is fundamentally different. Multi-camera reasoning requires handling context that scales with the number of views, resolving occlusions visible from only a subset of cameras, judging which views matter, and integrating evidence across perspectives that may overlap or diverge. Current models struggle with exactly these challenges, yet no benchmark systematically targets them. We introduce CrossView, a multi-camera video question-answering benchmark spanning autonomous driving, security surveillance, egocentric/exocentric video, and robotics. Evaluation of proprietary models, such as GPT-5.2, and open-source models, like Qwen3-VL, reveals consistently low accuracy, with open-source models trailing by a wide margin. Performance scales strongly with a model's ability to jointly process multiple viewpoints, positioning CrossView as a rigorous benchmark for multi-camera video. We open-source our code and dataset at https://utaustin-swarmlab.github.io/CrossView.
CommentsECCV 2026