发表机构
Institute for Network Sciences and Cyberspace, Tsinghua University; State Key Laboratory of Internet Architecture; Renmin University of China; Department of Computer Science and Technology, Tsinghua University; Huawei; Beijing National Research Center for Information Science and Technology (BNRist), Tsinghua University(清华大学网络科学与网络空间研究院; 互联网架构国家重点实验室; 中国人民大学; 清华大学计算机科学与技术系; 华为; 清华大学北京信息科学与技术国家研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
NetLexicon提出离散预训练框架,通过向量量化构建流量词典,结合状态转移预测与统计特征对齐目标,在加密流量分析中提升性能与效率,实现可解释表示。
AI 中文摘要
加密网络流量分析需要有效表示可观察的通信行为。现有的预训练方法通常改编NLP/CV目标和序列架构,这激发了对能够捕捉流量特定交互模式的学习目标的需求。我们提出了NetLexicon,一个离散预训练框架,从无标记流量中学习可复用的行为状态。它通过向量量化将上下文流量窗口转换为离散状态,构建了一个紧凑的流量词典。我们设计了两个互补的预训练目标。状态转移预测(STP)根据观察到的历史预测后续序列结构和数据包特征,而统计特征对齐(SFA)将学习到的状态锚定在窗口级流量统计中。两者共同引导词典捕捉重复出现的通信行为及其演变。我们在四个基准上评估了NetLexicon,涵盖Web应用识别、服务类型识别和恶意软件检测。NetLexicon在每个基准上相比最强基线将Macro-F1提升了最多25.5个百分点,并将每个epoch的微调时间相比评估的基线减少了最多23.6倍。进一步分析表明,学习到的离散状态捕捉了数据包大小、时序和数据传输中的可识别模式。这些结果表明,将可观察的行为结构纳入预训练支持加密流量分析的有效、高效且可解释的表示。
英文摘要
Encrypted Web traffic analysis requires effective representations of observable communication behavior. Existing pretraining methods often adapt NLP/CV objectives and sequence architectures, motivating learning objectives that capture traffic-specific interaction patterns. We present NetLexicon, a discrete pretraining framework that learns reusable behavioral states from unlabeled traffic. It converts contextual traffic windows into discrete states through vector quantization, constructing a compact traffic lexicon. We design two complementary pretraining objectives. State Transition Prediction (STP) forecasts subsequent sequence structure and packet features from observed history, while Statistical Feature Alignment (SFA) grounds learned states in window-level traffic statistics. Together, they guide the lexicon to capture recurring communication behaviors and their evolution. We evaluate NetLexicon on four benchmarks covering Web application identification, service type identification, and malware detection. NetLexicon improves Macro-F1 by up to 25.5 percentage points over the strongest baseline on each benchmark and reduces fine-tuning time per epoch by up to 23.6 times relative to the evaluated baselines. Further analysis shows that the learned discrete states capture recognizable patterns in packet size, timing, and data transfer. These results demonstrate that incorporating observable behavioral structure into pretraining supports effective, efficient, and interpretable representations for encrypted traffic analysis.