发表机构
Chulalongkorn University(朱拉隆功大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出基于YOLO v11n-pose与CLIP的轻量两阶段实时视频异常检测框架,在多个数据集上实现高吞吐量与良好检测性能,减少了对额外模块的依赖。
AI 中文摘要
本文提出一种用于实时视频异常检测的轻量两阶段框架。第一阶段采用YOLO v11n-pose在单次前向传播中检测人体并提取17个骨骼关键点;第二阶段通过CLIP ViT-B/32对每个裁剪后的人体区域进行编码,并计算其与预定义异常行为文本描述的余弦相似度。该架构无需光流、独立姿态估计器及基于密度的评分模块。在CUHK Avenue、ShanghaiTech Campus及朱拉隆功大学采集的自定义室内数据集上进行实验,结果显示在NVIDIA Titan XP GPU上的端到端吞吐量约为51 FPS,较多特征基线提升3.36倍,同时保持帧级AUROC值分别为89.26%、70.26%和84.13%。
英文摘要
We propose a lightweight two-stage framework for real-time video anomaly detection. The first stage employs YOLO v11n-pose to detect persons and extract seventeen skeletal keypoints in a single forward pass. The second stage encodes each cropped person region through CLIP ViT-B/32 and computes cosine similarity against predefined textual descriptions of anomalous behaviors. This architecture eliminates the need for optical flow, standalone pose estimators, and density-based scoring modules. Experiments on CUHK Avenue, ShanghaiTech Campus, and a custom indoor dataset collected at Chulalongkorn University demonstrate an end-to-end throughput of approximately 51 FPS on an NVIDIA Titan XP GPU, a 3.36x speedup over the multi-feature baseline, while maintaining frame-level AUROC values of 89.26%, 70.26%, and 84.13%, respectively.
Comments5 pages, 2 figures, 3 tables. Presented at the 9th IEEE International Conference on Multimedia Information Processing and Retrieval (MIPR 2026); accepted for publication in IEEE Xplore