发表机构
Nanjing University; Harbin Institute of Technology; Nanjing University of Science and Technology(南京大学; 哈尔滨工业大学; 南京理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Probe-VAD提出序数二元探测框架,从冻结视觉语言模型中直接探测严重性偏好,生成连续异常分数,实现免训练、低成本的视频异常检测,性能优越。
AI 中文摘要
视频异常检测(VAD)旨在定位未修剪视频中的异常事件。视觉语言模型(VLMs)为免训练的VAD提供了丰富的视觉理解能力,但现有方法在视觉理解与异常评分之间施加了限制性的接口。基于字幕的流程将视觉证据压缩为文本,可能丢弃细微线索,而直接生成数字则迫使模型通过一组预定义分数来表达其判断。此类接口可能模糊异常严重程度的细微差异,导致视觉上不同的片段获得相似的表示或分数,从而限制了异常排序的分辨率。我们提出了Probe-VAD,一个序数二元探测框架,直接从冻结的VLM中探测严重性偏好。给定原始视频片段,Probe-VAD查询十个有序的严重性阈值,并提取受约束的YES/NO延续似然。其归一化偏好形成累积严重性剖面,从中将尾部证据聚合成连续的异常分数,并通过等渗投影强制序数一致性。在公开VAD基准上的实验表明,该方法以低计算成本实现了优越性能。Probe-VAD提供了一种简单接口,可将冻结VLM的视觉理解转化为连续、对排序敏感的异常分数,无需特定任务的训练或基于字幕的压缩。代码可在以下网址获取:此https URL。
英文摘要
Video anomaly detection (VAD) aims to localize anomalous events in untrimmed videos. Vision-language models (VLMs) provide rich visual understanding for training-free VAD, but existing approaches impose restrictive interfaces between visual understanding and anomaly scoring. Caption-based pipelines compress visual evidence into text, potentially discarding subtle cues, while direct numerical generation forces the model to express its judgment through a small set of predefined scores. Such interfaces can obscure subtle differences in anomaly severity, causing visually distinct clips to receive similar representations or scores and thereby limiting the resolution of anomaly ranking. We propose \textbf{Probe-VAD}, an ordinal binary-probing framework that directly probes severity preferences from a frozen VLM. Given raw video clips, Probe-VAD queries ten ordered severity thresholds and extracts constrained \textit{YES}/\textit{NO} continuation likelihoods. Their normalized preferences form a cumulative severity profile, from which tail evidence is aggregated into a continuous anomaly score, with isotonic projection enforcing ordinal consistency. Experiments on public VAD benchmarks demonstrate superior performance with low computational cost. Probe-VAD provides a simple interface for translating frozen VLM visual understanding into continuous, rank-sensitive anomaly scores without task-specific training or caption-based compression. Code is available at: https://github.com/yvestine/COVAS-VAD.
CommentsUnder Review