arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11674cs.CV

VESSI——面向监控与调查的VLM增强支持框架

VESSI - VLM-Enhanced Support for Surveillance and Investigations

Saverio Cavasin, Pietro Tedeschi, Mattia Tamiazzo, Alessandro Brighente, Simone Milani, Mauro Conti

首次发表
浏览论文内容

中文总结 AI 辅助

针对传统视频监控分析系统的不足,提出基于VLM的VESSI框架,结合CMUS评估方法,实验显示其能标记超66%相关视频、减少超85%审查时间,提升监控分析能力。

中文摘要 AI 辅助

自动化视频监控分析已成为情报基础设施和执法机构的关键组成部分。传统系统缺乏用于全面态势感知和取证任务的语义模块,限制了其有意义地解释事件或支持事后调查的能力,这减缓了运营洞察力并增加了人类分析师的负担。视觉语言模型(VLM)的最新进展为弥合这一差距提供了有前景的途径。为解决该问题,我们提出VLM增强的监控与调查支持框架(VESSI),这是一种基于VLM的框架,旨在通过对视频序列进行提示驱动的查询来增强自动化视频监控分析,其中显著视觉特征被转换为文本描述。我们使用四个最先进的模型测试了我们的框架。由于该任务的大多数数据集是未标记的,我们还提出了复合模型效用分数(CMUS)来评估VLM性能。实验结果表明,我们的解决方案显著提高了人类操作员的分析能力,并增强了自动化监控系统的灵活性。在我们的评估中,最可靠的模型在超过66%的视频中标记了潜在相关活动,同时将审查时间减少了超过85%,在选择性和效率之间提供了实用的平衡。由无参考CMUS评估产生的模型排序被正常视频CMUS评估重现,并且与从5909个人工参考帧获得的误报率排序匹配。这种一致性支持该分数在评估环境中的运营使用。

英文摘要

Automated video surveillance analysis has become a critical component of intelligence infrastructures and Law Enforcement agencies. Traditional systems lack the semantic module for comprehensive situational awareness and forensic tasks, limiting their ability to interpret events meaningfully or support post-incident investigations. This slows operational insight and increases the burden on human analysts. Recent advances in Vision-Language Models (VLMs) offer promising pathways to bridge this gap. To address this, we propose VLM-Enhanced Support for Surveillance and Investigations (VESSI), a VLM-based framework designed to enhance automated video surveillance analysis through prompt-driven interrogation of video sequences where salient visual features are converted into textual descriptions. We test our framework with four state-of-the-art models. Since most datasets for this task are unlabeled, we also propose the Composite Model Utility Score (CMUS) to assess VLM performance. Experimental results show that our solution substantially improves the analysis capabilities of human operators and enhances the flexibility of automated surveillance systems. In our evaluation, the most reliable model flagged potentially relevant activity in more than 66% of the videos while reducing review time by more than 85%, offering a practical balance between selectivity and efficiency. The model ordering produced by the reference-free CMUS evaluation was reproduced by the normal-video CMUS evaluation and matched the false-positive-rate ordering obtained from 5,909 manually referenced frames. This agreement supports the operational use of the score within the evaluated setting.

发表机构

  • University of Padua(帕多瓦大学)
  • CY4GATE S.p.A.(CY4GATE股份公司)
  • Örebro University(厄勒布鲁大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑