AI 中文总结
该研究引入TAU-Bench这一以跟踪为核心的基准,用于联合评估异常实例跟踪与细粒度视频异常理解,发现现有视觉-语言模型存在语义推理与视觉接地的差距,为构建更可靠的视频异常理解系统提供了关键评估方向。
AI 中文摘要
人类通过连贯的感知过程理解异常事件,在此过程中,他们识别焦点实例,跟踪其在事件展开时的行为,并解释其为何违背周围场景的预期。视频异常理解(VAU)旨在赋予模型类似的能力,超越判断视频是否异常,转向解释事件如何发展以及为何重要。尽管近期的视觉-语言模型(VLMs)能够生成详细且合理的异常描述,但其语义流畅性并不能确保这些解释在时间上仍基于正确的异常实例。现有基准通常通过独立协议评估跟踪和语义理解,使得此类实例-语义不一致性在很大程度上未被测量。因此,我们引入TAU-Bench,一个以跟踪为核心的基准,用于联合评估异常实例跟踪和细粒度异常理解。TAU-Bench包含1118个视频、1454个跟踪序列和202438个像素级掩码,涵盖49个事件类别和45个场景类别,以及将实例级识别、事件级理解和场景级推理关联起来的以跟踪为核心的注释。为大规模构建TAU-Bench,我们开发了一个自动化数据引擎,整合了异常适用性过滤、异常实例跟踪构建、分层字幕注释和人工质量控制。对代表性VLM家族的评估显示,生成合理异常解释的模型可能仍然无法可靠地定位和跟踪正确的实例,揭示了语义推理与视觉接地之间存在的持续差距。这些发现因此强调,以实例为基础的评估是迈向更忠实、更可靠的VAU系统的重要一步。
英文摘要
Humans understand anomalous events through a coherent perceptual process in which they identify the focal instance, follow its behavior as the event unfolds, and interpret why it violates the expectations of the surrounding scene. Video anomaly understanding (VAU) seeks to endow models with a similar capability, moving beyond deciding whether a video is anomalous toward explaining how the event develops and why it matters. Although recent vision--language models (VLMs) can generate detailed and plausible anomaly descriptions, their semantic fluency does not ensure that these interpretations remain grounded in the correct anomaly instance over time. Existing benchmarks typically evaluate tracking and semantic understanding through separate protocols, leaving such instance--semantic inconsistency largely unmeasured. We therefore introduce TAU-Bench, a track-centric benchmark for jointly evaluating anomaly instance tracking and fine-grained anomaly understanding. TAU-Bench contains 1,118 videos, 1,454 tracks, and 202,438 pixel-level masks spanning 49 event and 45 scene categories, together with track-centric annotations that connect instance-level identification, event-level understanding, and scene-level reasoning. To build TAU-Bench at scale, we developed an automated data engine integrating anomaly suitability filtering, anomaly instance track construction, hierarchical caption annotation, and human quality control. Evaluations across representative VLM families show that models producing plausible anomaly interpretations may still fail to localize and track the correct instance reliably, revealing a persistent gap between semantic reasoning and visual grounding. These findings therefore highlight instance-grounded evaluation as an important step toward more faithful and reliable VAU systems.