发表机构
National Tsing Hua University; NVIDIA(国立清华大学; 英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ReactVAU提出慢-快解耦框架,以轻量快速检测与重量慢速推理协同,实现流式视频异常理解,兼顾性能与效率。
AI 中文摘要
在本文中,我们提出了ReactVAU,一个用于实时流式视频异常理解(VAU)的慢-快解耦框架。现有的VAU方法依赖于全局时间采样的离线推理,这违反了因果性,并阻碍了其在实时监控流中的部署。相反,通用的流式视频模型满足因果访问,但在记忆压缩过程中会稀释罕见的瞬态异常,并且通常在长时间的正常间隔内统一调用重量级多模态大语言模型(MLLM)。ReactVAU通过三个协同组件解决了这一差距:一个基于空间网格折叠(SGF)的轻量级快速检测模块,用于连续异常过滤;一个异常感知持久记忆(AAPM),保护关键视觉线索免受时间衰减;以及一个重量级慢速推理模块,在正常流期间保持休眠,仅在可疑事件发生时被唤醒,以进行语义验证和因果描述。在多个基准上的大量实验表明,ReactVAU在严格的流式约束下运行,同时在异常检测和因果推理方面均达到有竞争力的性能,并通过最小化重量级MLLM的调用显著提高了计算效率。项目页面可在以下https URL获取。
英文摘要
In this paper, we propose ReactVAU, a Slow-Fast Decoupled Framework for real-time streaming Video Anomaly Understanding (VAU). Existing VAU methods rely on offline inference with global temporal sampling, which violates causality and prevents deployment in live surveillance streams. Conversely, general streaming video models satisfy causal access but dilute rare transient anomalies during memory compression and often invoke heavyweight MLLMs uniformly over long normal intervals. React VAU addresses this gap with three synergistic components: a lightweight Fast Detection Module based on Spatial Grid Folding (SGF) for continuous anomaly filtering; an Anomaly-Aware Persistent Memory (AAPM) that protects critical visual cues from temporal decay; and a heavyweight Slow Reasoning Module that remains dormant during normal streams and is awakened only by suspicious events for semantic verification and causal description. Extensive experiments on multiple benchmarks demonstrate that ReactVAU operates under strict streaming constraints while simultaneously achieving competitive performance in both anomaly detection and causal reasoning, alongside significantly enhanced computational efficiency by minimizing heavyweight MLLM invocations. Project page is available at https://huiyuiui.github.io/ReactVAU/
CommentsAccepted to ECCV 2026. Project page: https://huiyuiui.github.io/ReactVAU/