TD-VAD:通过文本驱动学习打破视频异常检测中的视觉依赖
TD-VAD: Breaking Visual Dependence in Video Anomaly Detection with Text-Driven Learning
- School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院)
- School of Artificial Intelligence Engineering, Hefei Institute of Technology(合肥工业大学人工智能工程学院)
- China Mobile Zijin Innovation Institute(中国移动紫金创新研究院)
- School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对现有视频异常检测依赖视觉数据的问题,提出TD-VAD方法,利用LLM生成的文本描述训练模型,结合事件演化因果注意力模块与CLIP编码器,在XD-Violence和UCF-Crime上性能大幅优于现有方法。
AI中文摘要:
视觉数据是现有视频异常检测(VAD)方法训练的典型前提。然而,由于异常数据稀少且异常事件种类繁多,获取足够的带标注异常数据用于训练颇具挑战性且无法扩展。在本研究中,我们主张将文本视为视频序列应用于VAD模型的有效性,并提出一种新颖的文本驱动视频异常检测(TD-VAD)方法以打破视觉依赖。与异常视频数据相比,异常事件的文本描述易于收集,且其类别标签可直接得出。具体而言,我们的方法利用由大语言模型(LLM)生成的具有时间特征的类视频文本描述来训练VAD模型,无需依赖目标域的异常数据。为捕捉事件的长程与短程时间逻辑,我们设计了事件演化因果注意力模块,以建模跨时间的上下文依赖关系。在推理阶段,考虑到文本与视频序列间的域差距,我们使用冻结的CLIP编码器提取视频帧的嵌入,以对齐文本模态同时保留关键视觉信息。在两个大规模VAD数据集XD-Violence和UCF-Crime上开展的综合实验表明,我们的方法大幅优于现有的单类和无监督VAD方法。
英文摘要:
Visual data is typically a prerequisite for training existing video anomaly detection (VAD) methods. However, obtaining sufficient annotated anomaly data for training is challenging and not scalable due to the rarity of anomaly data and the wide variety of abnormal events. In this work, we advocate that the effectiveness of treating texts as video sequences for the VAD model and propose a novel Text-Driven Video Anomaly Detection (TD-VAD) approach to break visual dependence. In contrast to the anomaly video data, text descriptions of abnormal events are easy to collect, and their class labels can be directly derived. Specifically, our method utilizes video-like text descriptions with temporal characteristics generated by LLM to train a VAD model, without any reliance on target-domain anomaly data. To capture the long- and short-range temporal logic of events, we design the event evolution causal attention module to model contextual dependencies across time. During inference, considering the domain gap between the texts and video sequences, we use the frozen CLIP encoder to extract embeddings of video frames to align the text modality while retaining crucial visual information. Comprehensive experiments on two large-scale VAD datasets, XD-Violence and UCF-Crime, demonstrate that our method outperforms prior one-class and unsupervised VAD methods by a large margin.