arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VQ-VAD:面向以人为中心的视频异常检测的向量量化运动表示学习

VQ-VAD: Vector-quantized Motion Representation Learning for Human-centric Video Anomaly Detection

Narges Rashvand, Ghazal Alinezhad Noghre, Shanle Yao, Gabriel Maldonado, Hamed Tabkhi

arXiv 2608.05069首次发表:更新:

AI 中文总结

针对现有基于姿态的视频异常检测方法的局限,提出VQ-VAD框架,通过向量量化学习离散运动表示,在多评估设置的基准测试中取得优异性能,代码已公开。

AI 中文摘要

视频异常检测(VAD)本质上极具挑战性,原因在于异常样本稀缺,且监控视频中存在巨大的视觉变化,包括光照、视角和人类外观的变化。为减轻视觉噪声并解决隐私问题,近期研究转向基于姿态的VAD,该方法聚焦于运动动态而非原始视频数据。然而,现有的基于姿态的方法在连续潜在空间中对人类行为进行建模,限制了其学习紧凑运动模式的能力,而这种模式是鲁棒行为分析所必需的。我们通过提出Vector-Quantized Video Anomaly Detection(VQ-VAD)解决了这一问题,这是一种新型以人为中心的异常检测框架,用于学习离散运动表示。VQ-VAD对最初为图像生成开发的Vector-Quantized GAN(VQ-GAN)进行调整,使其能够处理关键点序列并构建正常行为的运动码本。仅在正常运动序列上进行训练,VQ-VAD通过识别无法映射到学习码本的观测运动序列的高重建误差来检测异常。我们在四个异常检测基准上的三种互补评估设置(包括域内、跨域和跨数据集泛化)中进行了广泛实验。VQ-VAD实现了出色的域内准确率(在HR-SHT[15]上为81.83%),从CMU Panoptic[14]进行的有效跨域迁移(在HR-SHT[15]上无需重新训练即可达到76.69%),以及具有竞争力的跨数据集鲁棒性。本研究的代码库可在此https URL获取。

英文摘要

Video Anomaly Detection (VAD) is inherently challenging due to the scarcity of anomalies and the large visual variability in surveillance footage, including changes in lighting, viewpoint, and human appearance. To mitigate visual noise and address privacy concerns, recent work has shifted to pose-based VAD, which focuses on motion dynamics rather than raw video data. However, existing pose-based approaches model human behavior in continuous latent spaces, limiting their ability to learn compact motion patterns necessary for robust behavior analysis. We address this by proposing Vector-Quantized Video Anomaly Detection (VQ-VAD), a novel human-centric anomaly detection framework that learns discrete motion representations. VQ-VAD adapts Vector-Quantized GAN (VQ-GAN), originally developed for image generation, to operate on keypoint sequences and construct a motion codebook of normal behavior. Trained exclusively on normal motion sequences, VQ-VAD detects anomalies by identifying high reconstruction errors when an observed motion sequence cannot be mapped to the learned codebook. We conduct extensive experiments across three complementary evaluation settings, including in-domain, cross-domain, and cross-dataset generalization, on four anomaly detection benchmarks. VQ-VAD achieves strong in-domain accuracy (81.83% on HR-SHT [15]), effective cross-domain transfer from CMU Panoptic [14] (76.69% on HR-SHT [15] without retraining), and competitive cross-dataset robustness. The code base for this work is available at https://github.com/TeCSAR-UNCC/VQ-VAD.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑