arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于理论支撑状态空间核心的严格因果流视频异常检测

Strictly Causal Streaming Video Anomaly Detection with a Theoretically-Grounded State-Space Core

Yogesh Kumar

arXiv 2608.24810首次发表:更新:

发表机构

Indian Institute of Technology Jodhpur(焦特布尔印度理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出带衰减门的对角线性状态空间循环的严格因果流异常检测器,在Apple M3 Pro硬件上实测高效,在两个数据集上达一定帧级AUC,门控效果依赖数据集大小,后续将缩小准确性差距并扩展评估。

AI 中文摘要

近期研究将Mamba风格状态空间模型(SSMs)应用于视频异常检测,但现有方法仍依赖内部缓冲片段或窗口,缺乏对时间记忆与检测延迟关系的理论解释,且仅通过GPU吞吐量而非这些方法目标部署的边缘硬件来评估效率。我们提出一种严格因果流异常检测器,其固定大小状态在每帧输入时以O(1)时间和内存更新,无前瞻且无片段缓冲。该检测器的时间核心是带输入和状态依赖衰减门的对角线性状态空间循环,通过冻结视觉骨干的因果下一个嵌入预测进行自监督训练。我们推导了循环衰减频谱与检测延迟及可可靠捕获的最短异常之间的闭式关系,并在UCSD Ped2和CUHK Avenue数据集上进行了实证验证。从学习到的基础衰减(57至59帧)预测的收敛延迟上限远高于实测检测延迟(1.6和18.4帧),表明是事件边界门而非基础衰减决定响应性。我们进一步在Apple M3 Pro硬件上直接测量端到端延迟和吞吐量,分别为每帧0.74毫秒和0.77毫秒(超过1300 FPS),而非模拟GPU数值。在未调优初始配置下,该方法在Ped2和Avenue数据集上达到67.9%和70.2%的帧级AUC,在准确性上落后于先前非因果SSM基线。对衰减率、状态大小和门控的消融实验显示,门的贡献依赖于数据集大小,在较小的Ped2训练集上会损害准确性,但在较大的Avenue数据集上有帮助。缩小这一准确性差距并将评估扩展至第三个更大基准是近期下一步工作。

英文摘要

Recent work has applied Mamba style state space models (SSMs) to video anomaly detection, yet existing approaches still rely on buffering clips or windows internally, lack a theoretical account of how temporal memory relates to detection latency, and benchmark efficiency only through GPU throughput rather than the edge hardware these methods are intended to target. We introduce a strictly causal streaming anomaly detector whose fixed size state is updated in O(1) time and memory per incoming frame, with no lookahead and no clip buffering. Its temporal core is a diagonal linear state space recurrence with an input and state dependent decay gate, trained self supervised through causal next embedding prediction on a frozen visual backbone. We derive a closed form relationship between the recurrence decay spectrum and both detection delay and the shortest anomaly it can reliably capture, then validate empirically on UCSD Ped2 and CUHK Avenue. The settling delay bound predicted from the learned base decay (57 to 59 frames) sits far above the measured detection delay (1.6 and 18.4 frames), showing that the event boundary gate, not the base decay, governs responsiveness. We further report end to end latency and throughput measured directly on Apple M3 Pro hardware, 0.74 ms and 0.77 ms per frame (over 1300 FPS), rather than simulated GPU numbers. With an untuned initial configuration the method reaches 67.9 percent and 70.2 percent frame level AUC on Ped2 and Avenue, trailing prior non causal SSM baselines in accuracy. Ablations over decay rate, state size, and gating reveal that the gate contribution is dataset size dependent, hurting accuracy on the smaller Ped2 training set but helping on the larger Avenue one. Closing this accuracy gap and extending evaluation to a third, larger benchmark are immediate next steps.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑