arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向弱监督视频异常检测的自适应多粒度时序建模

Adaptive Multi-Granularity Temporal Modeling for Weakly Supervised Video Anomaly Detection

Changyi Li, Yu Xiao

arXiv 2609.05066首次发表:更新:

发表机构

Aalto University(阿尔托大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对弱监督视频异常检测中现有方法对异常时序变化适应性不足的问题,提出含TRM、ESM及自适应相似度融合策略的框架,在两个基准上性能优于现有最先进方法。

AI 中文摘要

随着视频监控数据的规模远超人工标注能力,弱监督视频异常检测(WSVAD)已成为关键研究前沿。现有多数方法在多实例学习(MIL)框架内构建WSVAD,该框架依赖固定的、手工设计的时序先验来监督异常评分。然而,此类框架对真实视频中异常持续时间和时序动态的广泛变化适应性有限,常导致不稳定或不可靠的片段级预测。为解决这一局限,我们提出一种用于WSVAD的自适应时序建模框架,明确考虑跨多个时序粒度的视频动态变化。首先,我们引入时序精修模块(TRM),该模块利用动态位置编码和可学习类令牌来建模长程时序依赖,同时提炼稳定的全局视频级表示。其次,为捕捉频率和持续时间各异的异常事件,我们开发自适应事件分割模块(ESM),该模块通过时序不连续性分析识别事件边界,并将片段特征聚合成具有区分性的事件级表示。最后,针对片段级和事件级预测,我们提出一种基于自适应相似度的融合策略,该策略将异常评分动态整合到视频级预测中,用全局语义相关性取代固定的top-k聚合启发式方法。在两个基准上的大量实验表明,所提框架始终优于现有最先进的方法。

英文摘要

As the scale of video surveillance data outpaces manual annotation capacities, weakly supervised video anomaly detection (WSVAD) has emerged as a critical research frontier. Most existing approaches formulate WSVAD within a Multiple Instance Learning (MIL) framework that relies on rigid, hand-crafted temporal priors to supervise anomaly scoring. However, such formulations exhibit limited adaptability to the wide variation in anomaly durations and temporal dynamics observed in real-world videos, often leading to unstable or unreliable snippet-level predictions. To address this limitation, we propose an adaptive temporal modeling framework for WSVAD that explicitly accounts for variations in video dynamics across multiple temporal granularities. First, we introduce a Temporal Refinement Module (TRM) that leverages dynamic positional encoding and a learnable class token to model long-range temporal dependencies while distilling a stable global video-level representation. Second, to capture anomalous events with varying frequency and duration, we develop an adaptive Event Segmentation Module (ESM) that identifies event boundaries through temporal discontinuity analysis and aggregates snippet features into discriminative event-level representations. Finally, for snippet-level and event-level predictions, we propose an adaptive similarity-based fusion strategy that dynamically integrates anomaly scores into video-level predictions, replacing fixed top-k aggregation heuristics with global semantic relevance. Extensive experiments on two benchmarks demonstrate that the proposed framework consistently outperforms state-of-the-art methods.

CommentsAccepted by PRCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑