arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2406.12235cs.CV

Holmes-VAD:通过多模态大语言模型实现无偏且可解释的视频异常检测

Holmes-VAD: Towards Unbiased and Explainable Video Anomaly Detection via Multi-modal LLM

  • Huazhong University of Science and Technology(华中科技大学)
  • Baidu Inc.(百度公司)
  • University of Michigan(密歇根大学)

机构由 AI 辅助整理,请以论文原文为准。

Huaxin Zhang, Xiaohao Xu, Xiang Wang, Jialong Zuo, Chuchu Han, Xiaonan Huang, Changxin Gao, Yuehuan Wang, Nong Sang

更新

AI总结:

Holmes-VAD提出首个大规模多模态VAD指令微调基准VAD-Instruct50k,结合轻量级时间采样器与多模态大语言模型,实现无偏、可解释的视频异常定位与解释生成。

AI中文摘要:

面向开放式的视频异常检测(VAD),现有方法在面对具有挑战性或未见事件时往往表现出有偏的检测结果,并且缺乏可解释性。为解决这些不足,我们提出Holmes-VAD,这是一个新颖的框架,利用精确的时间监督和丰富的多模态指令来实现准确的异常定位和全面的解释。首先,为了构建无偏且可解释的VAD系统,我们构建了首个大规模多模态VAD指令微调基准,即VAD-Instruct50k。该数据集采用精心设计的半自动标注范式创建。对收集到的未修剪视频应用高效的单帧标注,然后利用鲁棒的现成视频字幕生成器和大型语言模型(LLM)将其合成为对异常和正常视频片段的高质量分析。在VAD-Instruct50k数据集的基础上,我们开发了一个面向可解释视频异常检测的定制化解决方案。我们训练一个轻量级时间采样器来选择具有高异常响应的帧,并微调一个多模态大语言模型(LLM)以生成解释性内容。大量实验结果验证了所提出的Holmes-VAD的通用性和可解释性,使其成为面向真实世界视频异常分析的一种新颖的可解释技术。为了支持社区,我们的基准和模型将在https://holmesvad.github.io公开提供。

英文摘要:

Towards open-ended Video Anomaly Detection (VAD), existing methods often exhibit biased detection when faced with challenging or unseen events and lack interpretability. To address these drawbacks, we propose Holmes-VAD, a novel framework that leverages precise temporal supervision and rich multimodal instructions to enable accurate anomaly localization and comprehensive explanations. Firstly, towards unbiased and explainable VAD system, we construct the first large-scale multimodal VAD instruction-tuning benchmark, i.e., VAD-Instruct50k. This dataset is created using a carefully designed semi-automatic labeling paradigm. Efficient single-frame annotations are applied to the collected untrimmed videos, which are then synthesized into high-quality analyses of both abnormal and normal video clips using a robust off-the-shelf video captioner and a large language model (LLM). Building upon the VAD-Instruct50k dataset, we develop a customized solution for interpretable video anomaly detection. We train a lightweight temporal sampler to select frames with high anomaly response and fine-tune a multimodal large language model (LLM) to generate explanatory content. Extensive experimental results validate the generality and interpretability of the proposed Holmes-VAD, establishing it as a novel interpretable technique for real-world video anomaly analysis. To support the community, our benchmark and model will be publicly available at https://holmesvad.github.io.

补充信息

↑