发表机构
State Key Laboratory of Complex & Critical Software Environment, Beihang University(复杂与关键软件环境国家重点实验室,北京航空航天大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对恶意流量检测中数据多样性不足和分布外泛化差的问题,提出改进的频域混合数据增强方法,理论分析混合机制,增强序列特征多样性,实验证明优于现有方法。
AI 中文摘要
网络流量的强动态性常常迫使恶意流量检测模型处理分布外数据。通常,基于深度学习的恶意流量检测模型需要大量高质量的训练数据。然而,由于标注难度高和资源消耗大等挑战,现有数据集往往存在多样性不足的问题,且无法捕捉不断演变的流量模式,导致训练出的模型分布外泛化能力较差。数据增强已被广泛采用以提高数据多样性和模型泛化能力。近年来,频域混合增强因其在有效扰动数据的同时保留关键结构信息而展现出良好的性能。该方法显示出增强恶意流量检测模型的潜力。然而,现有研究缺乏对混合机制的理论解释,且未适应网络流量的特点。在本文中,我们首先对当前频域混合方法进行理论分析,揭示其基本原理和局限性。我们进一步提出一种改进的基于频域混合的网络流量数据增强方法,该方法增强了网络流量中序列特征的多样性,并提高了恶意流量检测模型的分布外泛化能力。在多个人工和真实世界数据集上的大量实验表明,我们的方法在不同网络环境下显著提高了检测性能,并优于其他数据增强方法。
英文摘要
The strong dynamics of network traffic often force malicious traffic detection models to handle out-of-distribution data. Typically, deep learning-based malicious traffic detection models require a large amount of high-quality training data. However, owing to challenges such as high labeling difficulty and resource consumption, existing datasets often suffer from insufficient diversity and fail to capture evolving traffic patterns, leading to poor out-of-distribution generalization ability of the trained models. Data augmentation has been widely adopted to improve data diversity and model generalization. Recently, frequency-domain mixing augmentation has shown promising performance because it effectively perturbs data while preserving key structural information. This approach shows potential for enhancing malicious traffic detection models. However, existing studies lack theoretical interpretation of the mixing mechanism, and do not adapt to the characteristics of network traffic. In this paper, we first conduct a theoretical analysis of the current frequency-domain mixing method, revealing its underlying principles and limitations. We further propose an improved frequency-domain mixing-based data augmentation method for network traffic data, which enhances the diversity of sequence features in network traffic and improves the out-of-distribution generalization of malicious traffic detection models. Extensive experiments on multiple artificial and real-world datasets demonstrate that our method substantially improves detection performance across diverse network environments and outperforms other data augmentation approaches.
CommentsIn Proceedings of the 2026 ACM SIGSAC Conference on Computer and Communications Security