FedVAR:用于视频异常识别的原型对齐联邦框架
FedVAR: Prototype-Aligned Federated Framework for Video Anomaly Recognition
浏览论文内容
中文总结 AI 辅助
针对联邦视频异常识别中的语义错位问题,本文提出FedVAR框架,通过视觉语言模型的原型对齐机制缓解错位,实验显示其性能优于现有联邦基线。
中文摘要 AI 辅助
在工业物联网(IIoT)与信息物理系统(CPS)时代,联邦学习(FL)为视频异常识别(VAR)提供了一种有前景的去中心化智能范式,该任务对于维护高保真数字孪生、保障关键任务环境的安全至关重要。然而,分布式边缘客户端之间固有的数据异质性会引发语义错位这一根本挑战,即客户端对“正常”与“异常”事件学习到的特征表示存在差异。在VAR场景中,该问题尤为突出——多样且细粒度的异常类别使每个客户端对异常形成了不同的语义解读。现有联邦方法主要聚焦于二分类异常检测,未解决该语义错位问题,无法实现有效的细粒度识别。本文提出FedVAR,一种专为VAR设计的弱监督FL框架,利用视觉语言模型(VLMs)的丰富表示,采用基于原型的对齐机制,为所有客户端创建共享语义锚点,以重新中心化并对齐其视觉与文本特征空间。该过程强制去中心化网络中“正常”表示的一致性,直接缓解语义错位,实现了通信开销极小的异常方向向量的鲁棒提示学习。我们在具有挑战性的基准上,针对各类非独立同分布(non-IID)数据划分方案、 unseen域及 novel异常类别开展了大量实验,结果表明FedVAR始终优于现有最先进的联邦基线,为基于视频的CPS构建了一种鲁棒的分布式智能框架。
英文摘要
In the era of Industrial Internet of Things (IIoT) and Cyber-Physical Systems (CPS), Federated Learning (FL) offers a promising decentralized intelligence paradigm for Video Anomaly Recognition (VAR). This task is vital for maintaining high-fidelity Digital Twins and ensuring safety in mission-critical environments. However, the inherent data heterogeneity across distributed edge clients leads to a fundamental challenge known as semantic misalignment, where clients learn divergent feature representations of "normal" and "abnormal" events. The problem becomes particularly pronounced in VAR, where the presence of diverse and fine-grained anomaly categories leads each client to develop distinct semantic interpretations of abnormality. Existing federated methods primarily focus on binary anomaly detection and fail to address this misalignment, preventing effective fine-grained recognition. In this paper, we introduce FedVAR, a weakly-supervised FL framework explicitly designed for VAR. Leveraging the rich representations of Vision-Language Models (VLMs), FedVAR employs a prototype-based alignment mechanism that creates a shared semantic anchor for all clients to re-center and align their visual and textual feature spaces. This process enforces a consistent representation of "normality" across the decentralized network, directly mitigating semantic misalignment and enabling robust prompt-learning of anomaly direction vectors with minimal communication overhead. We conduct extensive experiments on challenging benchmarks under various non-IID data partitioning schemes, unseen domains, and novel anomaly classes. The results demonstrate that FedVAR consistently outperforms state-of-the-art federated baselines, establishing a robust framework for distributed intelligence in video-based CPS.
发表机构
- Chungbuk National University(忠北国立大学)
- Electronics and Telecommunications Research Institute (ETRI)(电子通信研究院)
- University of Central Florida(中佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。