发表机构
College of Pharmaceutical Engineering of Traditional Chinese Medicine, Tianjin University of Traditional Chinese Medicine; School of Artificial Intelligence, Beijing Normal University; Guangzhou Institute of Technology, Xidian University; State Key Laboratory of Component-based Chinese Medicine, Tianjin University of Traditional Chinese Medicine(天津中医药大学中药制药工程学院; 北京师范大学人工智能学院; 西安电子科技大学广州研究院; 天津中医药大学组分中药国家重点实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对视频异常检测在复杂条件下的局限性,提出EVAD框架,利用事件相机与传统视频。构建大规模基准数据集,设计对比多模态预训练框架及自适应融合模块,实验验证了该方法在现实场景中检测VAD的有效性。
AI 中文摘要
视频异常检测(VAD)对自动监控至关重要,但仅依靠可见光视频在光照变化、快速运动和复杂背景等具有挑战性的条件下仍很脆弱。为解决这些限制,我们提出了EVAD,一个联合利用传统视频和生物启发式事件相机捕获的事件流的事件增强VAD框架。事件传感器以高时间分辨率异步捕获亮度变化,对运动模糊和极端光照具有鲁棒性,并提供与基于视频的视觉信息互补的运动显著线索。为支持多模态VAD研究,我们构建了一个大规模的可见事件基准,包含在不同光照水平、运动模式和背景复杂性下收集的63亿个事件和376368个视频帧,填补了基于事件的异常检测中现实且可扩展数据集的空白。基于此数据集,我们设计了一个对比多模态预训练框架,通过对齐事件流、可见视频及文本描述的语义嵌入来学习判别性事件表示。一个自适应融合模块动态整合基于事件的时间线索和基于视频的空间语义,提高对环境干扰的鲁棒性。在基准和提出的TJUTCM Pha数据集上的实验表明,EVAD始终优于其他方法,验证了基于事件的传感在现实世界场景中对VAD的有效性。
英文摘要
Video anomaly detection (VAD) is critical for automated surveillance but remains fragile under challenging conditions such as illumination variations, fast motion, and complex backgrounds when relying solely on visible light videos. To address these limitations, we propose EVAD, an event enhanced VAD framework that jointly exploits conventional video and event streams captured by bio inspired event cameras. Event sensors asynchronously capture brightness changes with high temporal resolution, offering robustness to motion blur and extreme lighting, and providing motion salient cues complementary to video based visual information. To support multi modal VAD research, we construct a large scale visible event benchmark comprising 6.3 billion events and 376,368 video frames collected under diverse illumination levels, motion patterns, and background complexities, filling the gap of realistic and scalable datasets for event based anomaly detection. Building upon this dataset, we design a contrastive multi modal pretraining framework to learn discriminative event representations by aligning semantic embeddings across event streams, visible videos, and textual descriptions. An adaptive fusion module then dynamically integrates event based temporal cues with video based spatial semantics, improving robustness to environmental disturbances. Experiments on benchmarks and the proposed TJUTCM Pha dataset demonstrate that E VAD consistently outperforms methods, validating the effectiveness of event-based sensing for VAD in real world scenarios.