基于大型视觉语言模型的上下文结构视频异常检测
Context-structured Video Anomaly Detection with Large Vision-Language Models
浏览论文内容
中文总结 AI 辅助
研究视频异常检测难题,提出CSI-VAD方法,将视频分解为环境、对象、时间三个上下文并分别推理,免训练且不依赖文本提示和数据集调整,实验证明该方法优于基线并具竞争力。
中文摘要 AI 辅助
训练视频异常检测器具有挑战性,因为标注多样且罕见的异常事件存在困难和成本。尽管近期大型视觉语言模型实现了免训练推理,但现有方法大多依赖对采样视频的整体推理,可能会错过特定上下文的异常线索。本文提出了CSI-VAD,一种免训练的视频异常检测器,可跨不同上下文识别异常事件。其关键思想是将每个视频分解为三个不同的上下文(环境、对象、时间),并在单独分支中进行特定上下文推理。由于仅基于特定上下文视觉线索进行异常判断,无需预定义描述异常事件的文本提示或特定数据集调整。在UCF-Crime和UBnormal上的实验表明,CSI-VAD持续优于直接整体基线,并与现有方法具有竞争力,显示了结构化上下文分解在免训练视频异常检测中的优势。
英文摘要
Training video anomaly detectors is challenging due to the difficulty and cost of annotating diverse and rare abnormal events. Although recent large vision-language models enable training-free inference, existing approaches mostly rely on holistic inference over sampled video and may miss context-specific anomaly cues. In this paper, we present CSI-VAD, a training-free video anomaly detector that identifies abnormal events across diverse contexts. The key idea is to decompose each video into three distinct contexts (environment, objects, time) and perform context-specific inference in separate branches. Because we ground anomaly judgments solely in context-specific visual cues, we do not require predefined text prompts describing abnormal events or dataset-specific tuning. Experiments on UCF-Crime and UBnormal show that CSI-VAD consistently improves over the direct holistic baseline and achieves competitive performance against existing methods, showing the advantage of structured context decomposition for training-free video anomaly detection.