arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PERO:面向加密流量分类的高效鲁棒预训练后 foundation 模型

PERO: Efficient Robust Post-Training Foundation Models for Encrypted Traffic Classification

Wumei Du, Jiarong Wen, Kaiyu Zhang, Zi Yang, Yiqin Lv, Longfei Zhang, Dong Liang, Zheng Xie

arXiv 2608.15504首次发表:更新:

发表机构

National University of Defense Technology(国防科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对加密流量分类中鲁棒优化成本高的问题,提出PERO框架,通过轻量代理估计风险并选择高风险样本更新模型,在保持性能的同时降低了计算与内存成本。

AI 中文摘要

加密流量分类对网络安全至关重要,但实际部署中对罕见但高损失的错误(如恶意流量分类错误)极为敏感。加密流量 foundation 模型作为一种有前景的通用技术,可实现出色的整体性能。然而,采用经验风险最小化等标准目标函数常忽略高风险尾部事件,常用性能指标也难以反映风险敏感场景下的鲁棒性局限。将条件风险价值等鲁棒优化目标直接应用于预训练后阶段,对大模型而言计算成本过高,因为识别高损失样本会消耗大量计算资源。为此,我们提出 Pre-Evaluation Robust Optimization(PERO,预评估鲁棒优化),这是一种面向加密流量 foundation 模型的高效鲁棒预训练后框架。PERO 采用轻量代理估计样本级风险,并选择高风险样本子集更新 foundation 模型,将风险估计与昂贵的大模型优化解耦。在典型加密流量数据集上的大量实验表明,PERO 相比现有优秀的鲁棒预训练后方法,具有相当或更优的鲁棒性和平均性能,同时大幅降低了计算和内存成本。

英文摘要

Encrypted traffic classification is vital for network security, yet real-world deployments are inherently sensitive to rare but high-loss errors such as misclassification of malicious traffic. The encrypted traffic foundation model, as a promising general-purpose technique, can achieve impressive overall performance. However, employing standard objectives such as empirical risk minimization often overlooks high-risk tail events, and commonly used performance metrics hardly reflect robustness limitations in risk-sensitive scenarios. Directly applying robust optimization objectives, such as conditional value-at-risk, to post-training is computationally prohibitive for large models, as identifying high-loss samples exhausts substantial computation. To this end, we propose Pre-Evaluation Robust Optimization (PERO), an efficient robust post-training framework for encrypted traffic foundation models. PERO employs a lightweight proxy to estimate sample-wise risk and selects a subset of high-risk samples to update the foundation model, decoupling risk estimation from expensive large-model optimization. Extensive experiments on typical encrypted traffic datasets show that PERO achieves competitive or superior robustness and average performance compared to outstanding robust post-training methods, while significantly reducing computational and memory costs.

Comments16 pages, 6 figures, 6 tables, conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑