arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06475cs.CV

Thinking with Cameras: 通过动态视角控制进行主动视觉推理的监控视频理解

Thinking with Cameras: Active Visual Reasoning via Dynamic Viewpoint Control for Surveillance Video Understanding

  • School of Computer Science and Engineering, Central South University(中南大学计算机科学与工程学院)
  • SenseTime Research(商汤科技研究院)

机构由 AI 辅助整理,请以论文原文为准。

Xiao Zhang, Wang Zeng, Sheng Jin, Wentao Liu, Chen Qian, Shichao Kan

AI总结:

针对监控视频理解中固定视角被动观察的局限,提出CamVLM框架,通过动态视角控制实现主动视觉推理,并构建大规模数据集和强化学习策略优化,显著提升性能。

AI中文摘要:

大型视觉语言模型(LVLMs)近年来在通用视频理解方面取得了显著进展。然而,由于缺乏大规模领域特定数据集以及固定视角被动观察的限制,它们在监控视频中的应用仍然具有挑战性。在监控场景中,当目标距离较远、体积较小、被遮挡或移出当前摄像头视野时,关键视觉证据很容易被遗漏。在这项工作中,我们引入了CamVLM,一个用于“Thinking with Cameras”的新框架,该框架使LVLMs能够通过动态视角控制主动获取视觉证据,而不是被动分析固定的视频流。我们首先构建了CCTV-Anomaly,一个包含10个异常类别、共14,459个视频的大规模监控视频理解数据集,并配有详细的描述和事件标注。我们进一步将视角控制表述为一个主动视觉感知问题,并构建了CamTrack-53K,一个用于学习摄像头动作的以对象为中心的视角轨迹数据集。此外,我们提出了一种基于强化学习的视角策略优化框架,该框架将摄像头控制建模为顺序决策过程,并学习超越监督轨迹模仿的长时程观察策略。大量实验表明,CamVLM在被动观察和动态视角设置下均达到了最先进的性能,验证了基于主动摄像头的推理在监控视频理解中的有效性。我们的数据集、模型和代码将在提供的网址上公开。

英文摘要:

Large vision-language models (LVLMs) have recently achieved remarkable progress in general-purpose video understanding. However, their application to real-world surveillance remains challenging due to the lack of large-scale domain-specific datasets and the limitation of passive observation from fixed viewpoints. In surveillance scenarios, critical visual evidence can be easily missed when targets are distant, small, occluded, or move beyond the current camera view. In this work, we introduce CamVLM, a new framework for Thinking with Cameras, which enables LVLMs to actively acquire visual evidence in real-world surveillance by continuously controlling camera viewpoints. We first construct CCTV-Anomaly, a large-scale surveillance video understanding dataset containing 14,133 videos across 10 anomaly categories, with detailed captions and event annotations. We further formulate viewpoint control as an active visual perception problem and build CamTrack-53K, an object-centric viewpoint trajectory dataset for learning camera actions. Moreover, we propose a reinforcement learning based viewpoint policy optimization framework, which models camera control as a sequential decision-making process and learns long-horizon observation strategies beyond supervised trajectory imitation. Extensive experiments demonstrate that CamVLM achieves state-of-the-art performance under both passive observation and dynamic viewpoint settings, validating the effectiveness of active camera-based reasoning for surveillance video understanding. Our datasets, model, and code will be available at https://github.com/xiaozhang79/CamVLM.

↑