通过在异常样本污染下具有选择性聚合的联邦学习实现稳健的无监督网络入侵检测
Robust Unsupervised Network Intrusion Detection via Federated Learning with Selective Aggregation under Anomalous Sample Contamination
浏览论文内容
中文总结 AI 辅助
针对实际中训练数据含异常致网络入侵检测性能下降问题,提出一种基于联邦学习的稳健训练方法,利用联邦学习局限性减弱异常数据影响,引入选择性聚合机制排除异常更新,实验证明该方法在含异常数据环境中性能更优且随异常比例增加仍稳定。
中文摘要 AI 辅助
网络入侵检测系统(NIDS)对物联网(IoT)环境至关重要,因为针对物联网设备的恶意软件日益复杂。无监督学习方法因摆脱对标记数据集的依赖而颇具前景。但实际中训练数据完全干净的假设常被违背,尤其是直接从部署的网络设备收集数据样本时,训练数据集中可能存在异常。这种污染会降低检测性能,因此需要能在受污染的未标记训练数据下有效运行的稳健无监督NIDS方法。我们提出一种即使在存在未标记异常时也有效的异常检测稳健训练方法。该方法有两个主要组件:一是利用联邦学习(FL)代表性不足的已知局限性,减弱少数受损客户端异常数据的影响;二是在模型聚合期间引入选择性聚合机制,通过期望最大化(EM)算法量化本地客户端模型与全局参考之间的“距离”,检测并排除模型更新与多数显著不同的客户端组,确保异常更新不影响全局模型。在多个NIDS数据集上的实验表明,我们的方法在异常数据污染环境中优于现有方法,且随着异常比例增加仍能保持检测性能。
英文摘要
Network intrusion detection systems (NIDS) have become essential for Internet of Things (IoT) environments, as malware targeting IoT devices continues to evolve in sophistication. Unsupervised learning approaches offer a promising direction by removing the dependency on labeled datasets. However, the common assumption that training data are entirely clean is often violated in practice, particularly when data samples are collected directly from deployed network devices, where anomalies are likely to be present in the training datasets. Such contamination degrades detection performance and highlights the need for robust unsupervised NIDS methods capable of operating effectively under contaminated unlabeled training data. To address this issue, we propose a robust training methodology for anomaly detection (AD) that remains effective even in the presence of unlabeled anomalies. Our method consists of two primary components. First, we exploit a known limitation of federated learning (FL), namely its tendency to underrepresent minority data. By leveraging this characteristic, we attenuate the influence of anomalous data originating from a small number of compromised clients. Second, we introduce a selective aggregation mechanism during model aggregation, which quantifies the "distance" between local client models and a global reference. Specifically, we employ the Expectation-Maximization (EM) algorithm to detect and exclude client groups whose model updates significantly diverge from the majority. This selective aggregation ensures that anomalous updates do not compromise the global model. Experiments conducted on multiple NIDS datasets demonstrate that our method outperforms existing approaches in environments contaminated with anomalous data. Furthermore, the proposed method maintains its detection performance even as the proportion of anomalies increases.
发表机构
- School of Engineering, Institute of Science Tokyo(东京科学大学工程学院)
机构由 AI 辅助整理,请以论文原文为准。