arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EvoSherlock:面向视频中未见长尾安全关键事件的智能体终身演化

EvoSherlock: Towards Agentic Lifelong Evolution for Unseen Long-Tailed Security-Critical Events in Videos

Zixin Fan, Jiahong Lu, Changsheng Zheng, Yu Hong, Jingjing Wang

arXiv 2609.19201首次发表:更新:

发表机构

Soochow University; Qufu Normal University(苏州大学; 曲阜师范大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对视频中长尾安全关键事件的终身演化任务,提出EvoSherlock,通过智能体控制器协调因果视频生成与因果解耦对齐模块,从稀缺样本中持续分类和定位新事件,并缓解灾难性遗忘。

AI 中文摘要

现有的面向安全视频理解(SVU)系统假设了一个“封闭世界”,即静态类别集、充足标签以及所有事件类型事先已知的前提。现实世界中的安全关键事件打破了这些假设:它们遵循长尾分布,新类型不断涌现,且关键安全事件可能仅提供少量样本。我们将这一差距形式化为“视频中长尾安全关键事件的终身演化任务(L²-SCE)”,这是一个新任务,要求视觉语言模型(VLMs)在遗忘先前学习事件的前提下,从稀缺样本中持续分类和时序定位新出现的安全关键事件。此外,该任务揭示了两个关键挑战:(1)事件内稀缺性,即极端数据稀缺可能削弱新事件的分类和时序定位;(2)事件间干扰,即跨事件特征纠缠和表示漂移可能加剧灾难性遗忘。在此基础上,我们提出了EvoSherlock,一种由智能体控制器(Agentic Controller)以自反思闭环控制端到端编排的因果增强方法,包括两个核心模块:事件内因果视频生成模块(CVG)和事件间因果解耦与对齐模块(CDA),分别应对上述两个挑战。特别地,本文构建了一个L²-SCE数据集以模拟真实世界的增量条件。在我们基准上的大量实验证明了EvoSherlock相对于多个先进基线的优势。这些结果验证了所提出的L²-SCE任务的重要性以及EvoSherlock在从稀缺样本中分类和时序定位新兴安全关键事件方面的有效性。

英文摘要

Existing Security-oriented Video Understanding (SVU) systems assume a \emph{closed world}, \ie static category sets, abundant labels, and the premise that all event types are known upfront. Real-world security-critical events break these assumptions: they follow long-tailed distributions, new types emerge continuously, and critical security events may offer only a few samples. We formalize this gap as \textbf{Lifelong Evolving Task for Long-Tailed Security-Critical Events in Videos ({\boldmath$L^2$}-SCE)}, a new task that requires VLMs to continually classify and temporally localize newly emerging security-critical events from scarce samples without forgetting previously learned events. Furthermore, \task reveals two critical challenges: (1)~\textbf{Intra-Event Scarcity}, where extreme data scarcity may weaken both classification and temporal localization for new events, and (2)~\textbf{Inter-Event Interference}, where cross-event feature entanglement and representation drift may strengthen catastrophic forgetting. On this basis, we propose \textbf{\method}, a causal-enhanced approach orchestrated end-to-end by an \textbf{Agentic Controller} with self-reflective closed-loop control, which includes two core modules: the Intra-Event \textbf{C}ausal \textbf{V}ideo \textbf{G}eneration module (\textbf{CVG}) and the Inter-Event \textbf{C}ausal \textbf{D}ecoupling and \textbf{A}lignment module (\textbf{CDA}), to address the above two challenges, respectively. Especially, this paper constructs a \task dataset to simulate real-world incremental conditions. Extensive experiments on our benchmark demonstrate the advantages of \method over several advanced baselines. These justify the importance of the proposed \task and the effectiveness of \method in classifying and temporally localizing emerging security-critical events from scarce samples.

CommentsAccepted to ACM Multimedia 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑