发表机构
Carnegie Mellon University; University of Alabama at Birmingham; Griffith University(卡内基梅隆大学; 阿拉巴马大学伯明翰分校; 格里菲斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对工业视频异常检测,现有基于VLM的方法在工业场景中性能不佳的问题,提出无训练的智能框架O-VAD,通过跟踪对象时空动态和推理状态轨迹来检测异常,在多数据集实验中优于多种方法且能提供可解释报告。
AI 中文摘要
工业视频异常检测(IVAD)旨在识别工业过程中的异常对象和事件,对现代制造和质量控制系统至关重要。现有的基于视觉语言模型(VLM)的异常推理方法在一般领域能检测开放式异常,但在工业场景中性能下降。为此引入无训练的智能框架O-VAD,无需特定领域知识,强调对象状态演变。它跟踪检测对象的时空动态和潜在变换,通过对象时间状态轨迹推理识别异常对象。该方法克服了依赖正常片段再训练或注入领域知识的现有方法的局限性。在三个IVAD数据集上的大量实验表明,O-VAD优于前沿VLM、智能框架和在各自数据集上微调的传统VAD方法,还能提供异常过程和类型的可解释报告。
英文摘要
Industrial Video Anomaly Detection (IVAD) aims to identify anomalous objects and events in an industrial process, which is crucial for modern manufacturing and quality control systems. Existing VLM-based anomaly reasoning methods are capable of detecting open-ended anomalies in general domains. However, their performance declines in industrial settings characterized by intricate object transformations, strict physics, and procedural constraints. To tackle the complexity of such interaction-intensive detection, we introduce a training-free agentic framework for anomaly detection free of domain-specific knowledge, emphasizing object state evolution like humans inspectors. It is designed to track spatial-temporal dynamics and underlying transformations of detected objects over time, and then reason over the object-wise temporal state trajectories to identify abnormal objects in grounded frames. Our method overcomes limitations of prior approaches that rely on retraining on normal clips or injecting domain knowledge as context for test-time inference. Extensive experiments on three IVAD datasets demonstrate that our method outperforms frontier VLMs, agentic frameworks, and traditional VAD methods fine-tuned on the respective datasets, while providing interpretable reports over anomaly processes and types.
CommentsAccepted to ECCV 2026