发表机构
State Key Laboratory of Networking and Switching Technology; Beijing University of Posts and Telecommunications; School of Digital Media & Design Art(网络与交换技术国家重点实验室; 北京邮电大学; 数字媒体与设计艺术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视频异常检测中传统方法泛化差、多模态智能体工具编排难的问题,提出过程监督强化学习框架VTO,结合VAD-Tool工具集,在工具调度上较基线提升10.2%。
AI 中文摘要
视频异常检测(VAD)是一项关键却极具挑战性的任务,原因在于真实场景的复杂多样性。传统深度学习方法因在不同场景间泛化能力差而存在根本局限;尽管多模态智能体为VAD提供了颇具前景的工具学习范式,但当前依赖监督微调的系统难以应对复杂的工具编排,而标准强化学习常因粗粒度的结果奖励导致提前终止。为解决这些挑战,我们提出VTO,这是一个过程监督强化学习框架。VTO超越了静态工具使用,使智能体能够动态探索并与环境交互。具体而言,我们引入了一个基于基础模型的认知评估器,以提供上下文感知的语义反馈,该评估器被无缝整合到过程监督认知对齐中,从而提供细粒度的、分步骤的监督。通过明确惩罚逻辑截断并奖励完整的因果链,智能体优化其多步骤推理策略以实现相互关联的工具编排。为支持我们提出的框架,我们精心构建了VAD-Tool,这是一个分层可视化工具集,包含12种专门的视觉工具,范围从实体跟踪到高风险危害检测,并建立了相应的基准用于严格的多步骤推理评估。在VAD-Tool上进行的大量实验表明,VTO显著优于基线方法,在工具调度方面实现了高达10.2%的绝对准确率提升。代码和数据可在此https URL获取。
英文摘要
Video anomaly detection (VAD) is a critical yet challenging task due to the complex and diverse nature of real-world scenarios. Traditional deep learning approaches are fundamentally limited by poor generalization across diverse scenarios. While multimodal agents offer a promising tool-learning paradigm for VAD, current systems relying on supervised fine-tuning struggle with complex orchestration, and standard reinforcement learning often causes premature termination due to coarse-grained outcome rewards. To address these challenges, we propose VTO, a process-supervised reinforcement learning framework. Moving beyond static tool usage, VTO enables the agent to dynamically explore and interact with the environment. Specifically, we introduce a foundation model-driven cognitive evaluator to provide context-aware semantic feedback, which is seamlessly integrated into a Process-Supervised Cognitive Alignment that delivers fine-grained, step-wise supervision. By explicitly penalizing logical truncation and rewarding complete causal chains, the agent optimizes its multi-step reasoning policy for interrelated tool orchestration. To support our proposed framework, we meticulously crafted VAD-Tool, a hierarchical visual tool set comprising 12 specialized vision tools spanning from entity tracking to high-stakes hazard detection, and established the corresponding benchmark for rigorous multi-step reasoning evaluation. Extensive experiments on VAD-Tool demonstrate that VTO significantly outperforms baselines, achieving up to a 10.2\% absolute accuracy improvement in tool scheduling. Code and data are available at https://github.com/MICLAB-BUPT/VTO.
CommentsAccepted by ACM MM 2026