发表机构
Centre for Vision, Speech and Signal Processing (CVSSP); University of Surrey; Surrey Institute for People-Centred AI (PAI)(视觉、语音与信号处理中心; 萨里大学; 萨里以人为本人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Cog-VADU提出无需训练的认知推理框架,通过异常检测思维链提示和跨模态重排序,实现视频异常检测的零样本性能提升与可解释性。
AI 中文摘要
视频异常检测(VAD)旨在时间上定位视频中的异常事件。现有的大多数方法依赖于特定数据集的训练和精心策划的标注,限制了其在开放场景中的泛化能力。近期基于大型视觉-语言模型(LVLMs)的零样本方法缓解了这种依赖性,但往往缺乏时间连续性和结构化推理。我们提出了Cog-VADU,一个完全无需训练的框架,将VAD重新定义为顺序认知推理任务。Cog-VADU引入了异常检测思维链提示(CoADTP),将LVLM展开为跨视频片段的循环推理链。通过随时间传播结构化推理,模型保持隐式时间记忆,从而能够稳健地区分复杂异常与高运动正常活动。为提高可靠性,我们进一步设计了一个跨模态重排序阶段,将文本推理与视觉嵌入对齐,强制语义一致性和时间连贯性,以获得精细且稳定的预测。在多个公开VAD基准上的大量实验表明,Cog-VADU实现了具有竞争力的零样本性能。此外,跨模型评估显示,CoADTP以模型无关的方式持续增强基于推理的异常检测,为现实应用提供可解释且可泛化的异常理解。
英文摘要
Video Anomaly Detection (VAD) aims to temporally localize abnormal events in videos. Most existing approaches rely on dataset-specific training and curated annotations, limiting generalization in open-set scenarios. Recent zero-shot methods based on Large Vision- Language Models (LVLMs) alleviate this dependency but often lack temporal continuity and structured reasoning. We propose Cog-VADU, a fully training-free framework that reformulates VAD as a sequential cognitive reasoning task. Cog-VADU introduces Chain-of- Anomaly Detection Thought Prompting (CoADTP), which unrolls an LVLM into a recurrent reasoning chain across video segments. By propagating structured rationales over time, the model maintains implicit temporal memory, enabling robust discrimination between com- plex anomalies and high-motion normal activities. To improve reliability, we further design a cross-modal re-ranking stage that aligns textual rationales with visual embeddings, enforcing semantic consistency and temporal coherence for refined and stable predictions. Extensive experiments on multiple public VAD benchmarks demonstrate that Cog-VADU achieves competitive zero-shot performance. Moreover, cross-model evaluations show that CoADTP consistently enhances reasoning-based anomaly detection in a model-agnostic manner, pro- viding interpretable and generalizable anomaly understanding for real-world applications.
CommentsPublished in Transactions on Machine Learning Research (TMLR), 2026. 39 pages
Journal refTransactions on Machine Learning Research, August 2026