AI 中文总结
本文提出无需训练的多智能体框架AgenticVAU,通过四个专门智能体协作完成视频异常理解的探索-验证过程,在VAU-Bench数据集上优于零样本推理及强化学习基线方法。
AI 中文摘要
视频异常理解(VAU)聚焦于全面解读视频中的异常事件,要求模型识别异常发生情况、挖掘支撑证据并解释潜在原因,而非仅完成简单的异常检测。现有VAU方法常依赖专门训练或有限观测,限制了泛化能力或证据覆盖范围;单智能体替代方案虽支持自适应视频观测,但仍将探索、观测与决策整合在统一推理过程中,角色专业化程度有限且结构化证据协调不足。为解决这些局限,本文提出AgenticVAU,这是一个无需训练的多智能体框架,将VAU建模为探索-验证过程:系统先发现潜在异常,再通过针对性观测验证。为此,引入四个专门智能体分别处理视觉规则构建、搜索规划、视频观测与最终决策,这些智能体通过锚点注册表(绑定各观测结果的共享证据内存)通信。在该智能体框架引导下,AgenticVAU交替进行广泛的时间探索、密集的局部验证与跨区间对比,直至收集到充足证据。本文在VAU-Bench的ECVA、UCF-Crime和MSAD子集上开展大量实验,结果显示AgenticVAU优于零样本推理及基于强化学习的基线方法,证明了多智能体协作对视频异常理解的价值。
英文摘要
Video anomaly understanding (VAU) focuses on comprehensively interpreting abnormal events in videos, requiring models to identify anomalous occurrences, discover their supporting evidence, and explain the underlying causes beyond simple anomaly detection. Existing VAU methods often rely on specialized training or limited observations, restricting generalization or evidence coverage. Although single-agent alternatives support adaptive video observation, they still integrate exploration, observation, and decision-making within a unified reasoning process, offering limited role specialization and structured evidence coordination. To address these limitations, we present AgenticVAU, a training-free multi-agent framework that casts VAU as an explore--verify process, where the system first discovers potential anomalies and then verifies them through targeted observations. To achieve this, four specialized agents are introduced to handle visual-rule construction, search planning, video observation, and final decision, respectively. These agents communicate through an anchor registry, a shared evidence memory that binds each observation. Guided by this agent framework, AgenticVAU interleaves broad temporal exploration, dense local verification, and cross-interval comparison until sufficient evidence is collected. We conduct extensive experiments on the ECVA, UCF-Crime, and MSAD subsets of VAU-Bench, the results show that AgenticVAU outperforms zero-shot inference and reinforcement learning-based baselines, demonstrating the value of multi-agent collaboration for video anomaly understanding.