arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向视频异常检测与异常预测的统一分数匹配范式

A Unified Score Matching Paradigm for Video Anomaly Detection and Anticipation

Congqi Cao, Zhenhe Liang, Hanwen Zhang, Yifan Zhao, Qinyi Lv, Lingtong Min, Yanning Zhang

arXiv 2610.11149首次发表:更新:

发表机构

National Engineering Laboratory for Integrated Aero-Space-Ground-Ocean Big Data Application Technology; School of Computer Science, Northwestern Polytechnical University; School of Electronics and Information, Northwestern Polytechnical University(空天地海一体化大数据应用技术国家工程实验室; 西北工业大学计算机学院; 西北工业大学电子信息学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出统一分数驱动框架Uni-DSM,通过同一分数公式的不同推理与监督范式,将视频异常检测与预测任务统一,在多基准数据集上实现最优性能且效率高。

AI 中文摘要

视频异常检测(Video Anomaly Detection, VAD)是计算机视觉领域一项基础且关乎安全的任务。近期的生成式方法从分布视角检测异常,但仍受限于局部异常模式。与此同时,视频异常预测(Video Anomaly Anticipation, VAA)作为事后检测的主动延伸,引入了额外挑战。具体而言,VAD中依赖真实帧的对比推理范式不适用于VAA,阻碍了其发展。为应对这些挑战,我们提出了一种基于去噪分数匹配(denoising score matching, DSM)的统一分数驱动框架,命名为Uni-DSM,该框架通过对学习到的数据分布进行似然估计和分数函数来建模异常模式。在该统一框架内,我们采用共享的噪声条件分数Transformer骨干网络,结合场景相关嵌入和运动感知加权进行分布级建模。Uni-DSM未引入独立架构,而是通过基于同一分数公式构建的不同推理与监督范式,将VAD和VAA统一起来。对于VAD,我们实例化了自回归去噪分数匹配(autoregressive denoising score matching, ADSM)机制,该机制通过自回归去噪逐步累积异常证据,实现了对视觉线索之外局部模式的增强感知。对于VAA,我们通过引入轻量级辅助解码器和新颖的自蒸馏去噪分数匹配(self-distilled denoising score matching, SDSM)机制来扩展同一架构。我们的方法通过从输出差异构建监督,而非依赖不可用的未来真实值,实现了适用于早期异常预测的高效训练。在多个基准数据集上的大量实验表明,我们的方法在VAD和VAA中均达到了最先进的性能,同时保持了高计算效率,建立了从异常检测到预测的统一且可扩展的流程。

英文摘要

Video anomaly detection (VAD) is a fundamental and safety-critical task in computer vision. Recent generative approaches detect anomalies from a distributional perspective, but remain limited by local anomaly modes. Meanwhile, video anomaly anticipation (VAA), as a proactive extension beyond post-hoc detection, introduces additional challenges. In particular, the contrastive inference paradigm in VAD, which relies on ground-truth frames, is not applicable to VAA, hindering its development. To address these challenges, we propose a unified score-driven framework, termed Uni-DSM, based on denoising score matching (DSM), which models anomaly patterns through likelihood estimation and score functions over the learned data distribution. Within this unified framework, we adopt a shared noise-conditioned score transformer backbone with scene-dependent embeddings and motion-aware weighting for distribution-level modeling. Instead of introducing separate architectures, Uni-DSM unifies VAD and VAA through different inference and supervision paradigms built upon the same score-based formulation. For VAD, we instantiate an autoregressive denoising score matching (ADSM) mechanism, which progressively accumulates anomalous evidence via autoregressive denoising, enabling enhanced perception of local modes beyond visual cues. For VAA, we extend the same architecture by incorporating a lightweight auxiliary decoder and a novel self-distilled denoising score matching (SDSM) mechanism. By constructing supervision from output discrepancies instead of relying on unavailable future ground truth, our method achieves efficient training suitable or early anomaly anticipation. Extensive experiments on multiple benchmark datasets demonstrate state-of-the-art performance in both VAD and VAA while maintaining high efficiency, establishing a unified and scalable pipeline from anomaly detection to anticipation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑