发表机构
Beijing Jiaotong University; Institute of Automation, Chinese Academy of Sciences(北京交通大学; 中国科学院自动化研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对视频异常检测的“何时-何事”脱节问题,提出无训练框架GtS及工具增强型智能体方法,扩展基准并引入联合评估指标,实现了更高的异常检测准确性与推理速度。
AI 中文摘要
视频异常检测(VAD)旨在识别异常事件并定位其时间区间。现有方法存在“何时-何事”脱节问题:传统基于DNN的方法可定位异常发生时间,但缺乏语义理解;而基于LLM的方法能解释发生何事,却忽视精确的时间定位。我们将此归因于缺乏统一的推理范式。受人类检查监控视频的方式启发——全局扫视形成时间假设、审视可疑片段、迭代思考以修正错误——我们从两个角度研究这种全局到局部的范式。我们首先提出GtS(扫视后审视),这是一种无训练框架,利用静态和动态文本引导实现从粗到细的异常定位与理解,平衡了准确性与速度。为突破冻结外部模块带来的瓶颈,我们进一步提出一种工具增强型智能体VAD方法:多模态大语言模型学习调用视频裁剪工具、检查密集重采样的帧,并通过冷启动监督微调及联合答案-定位奖励的强化学习来自我修正定位错误的假设。为训练与评估,我们将先前的VAGU基准扩展为VAGU-T(视频异常定位、理解与思考),包含7567个真实世界视频,覆盖21个异常类别,具备经人类验证的定位、解释、问答对及思维链工具调用轨迹。我们还引入JeAUG指标,联合评估语义可解释性与时间精度。实验表明,GtS大幅超越无训练基线,而智能体模型兼具更高准确性与更快推理速度。
英文摘要
Video Anomaly Detection (VAD) aims to identify anomalous events and localize their temporal intervals. Existing approaches exhibit a "when-what" dissociation: traditional DNN-based methods localize when anomalies occur but lack semantic understanding, whereas LLM-based methods explain what happens but neglect precise temporal grounding. We attribute this to the absence of a unified reasoning paradigm. Inspired by how humans inspect surveillance videos - glancing globally to form temporal hypotheses, scrutinizing suspicious segments, and thinking iteratively to correct errors - we study this global-to-local paradigm from two perspectives. We first propose Glance then Scrutinize (GtS), a training-free framework using static and dynamic textual guidance for coarse-to-fine anomaly grounding and understanding, balancing accuracy and speed. To break the ceiling imposed by frozen external modules, we further propose a tool-augmented agentic VAD method, where a multimodal large language model learns to invoke a video cropping tool, inspect densely resampled frames, and self-correct mislocalized hypotheses, via cold-start supervised fine-tuning followed by reinforcement learning with a joint answer-grounding reward. For training and evaluation, we extend our prior VAGU benchmark into VAGU-T (Video Anomaly Grounding, Understanding, and Thinking), comprising 7,567 real-world videos over 21 anomaly categories with human-validated grounding, explanations, QA pairs, and chain-of-thought tool-calling traces. We further introduce JeAUG, a metric jointly evaluating semantic interpretability and temporal precision. Experiments show that GtS substantially surpasses training-free baselines, while the agentic model delivers both higher accuracy and faster inference.
Comments34 pages, 8 figures, 8 tables. Journal extension of our AAAI 2026 paper (arXiv:2507.21507)