发表机构
National Yang Ming Chiao Tung University(国立阳明交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
NOVA提出免训练零样本视频异常检测框架,通过常态感知提示构建和视觉常态锚点强化常态侧,在UCF-Crime和XD-Violence上取得最优性能。
AI 中文摘要
免训练零样本视频异常检测(ZS-VAD)利用视觉-语言模型(VLM)从预定义的异常词汇中定位异常实例,而无需提供任何视频。现有的基于CLIP的方法往往强调异常侧语义,而与之竞争的常态侧则未被仔细设计。我们指出现有解决方案中的两个关键局限:(i)决策边界模糊:正常提示可能包含与异常语义接近的模糊动词,如“running”,从而降低了VLM嵌入空间中正常与异常的区分度;(ii)模态差距:文本正常锚点特征与视觉帧之间的对齐不良。我们提出NOVA,一个免训练的ZS-VAD框架,在语言和视觉两个层面强化常态侧。NOVA引入了常态感知提示构建(NA),该模块排除与异常相邻的动词,并将正常描述偏向静态、低运动场景。为克服文本-视觉模态差距,NOVA构建了视觉常态锚点(VNA),该锚点从每个测试视频的初始帧中创建加权视觉正常锚点,提供视频特定的正常参考,而无需任务特定训练或标注。NOVA在UCF-Crime上达到89.86%的AUC,在XD-Violence上达到95.07%的AUC和84.82%的AP,在可比较的免训练零样本方法中达到了最先进的性能。
英文摘要
Training-free zero-shot video anomaly detection (ZS-VAD) leverages vision-language models (VLMs) to localize anomaly instances from a predefined anomaly vocabulary, without providing any video. Existing CLIP-based methods often emphasize anomaly-side semantics, while the competing normality side remains less carefully formulated. We identify two key limitations in existing solutions: (i) blurred decision boundary: normal prompts may contain ambiguous verbs, such as running, that are semantically close to anomalies, reducing normal and abnormal separation in the VLM embedding space; and (ii) modality gap: poor alignment between features of textual normal anchors and visual frames. We propose NOVA, a training-free ZS-VAD framework that strengthens the normal side at both linguistic and visual levels. NOVA introduces Normality-Aware Prompt Construction (NA), which excludes anomaly-adjacent verbs and biases normal descriptions toward static, low-motion scenes. To overcome the text-vision modality gap, NOVA constructs a Visual Normality Anchor (VNA), which creates a weighted visual normal anchor from the initial frames of each test video, providing a video-specific normal reference without task-specific training or annotations. NOVA achieves 89.86 percent AUC on UCF-Crime and 95.07 percent AUC and 84.82 percent AP on XD-Violence, reaching state-of-the-art performance among comparable training-free zero-shot methods.