基于LLM的多模态推理用于加密流量解释:一个基准
Multimodal Reasoning with LLM for Encrypted Traffic Interpretation: A Benchmark
- School of Microelectronics and Communication Engineering, Chongqing University(重庆大学微电子与通信工程学院)
- School of Data Science, Lingnan University(岭南大学数据科学学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出BGTD基准和mmTraffic框架,通过结合原始字节与结构化注释,实现可解释的加密流量解释,生成高保真的人可读报告,同时保持高分类准确率。
AI中文摘要:
网络流量作为关键媒体格式,对现代互联网基础设施的安全和通信至关重要。尽管现有方法性能优异,但面临两个关键瓶颈:(1) 无法捕捉多维语义,超越单模态序列模式;(2) 黑箱性质,仅提供类别标签,缺乏可审计的推理过程。本文提出Byte-Grounded Traffic Description (BGTD)基准,结合原始字节与结构化专家注释,提供必要的行为特征和可验证的证据链,用于多模态推理。基于BGTD,本文提出端到端的交通-语言表示框架mmTraffic,连接物理流量编码与语义解释。为缓解模态干扰和生成幻觉,mmTraffic采用联合优化的感知-认知架构。通过整合以感知为中心的流量编码器和以认知为中心的LLM生成器,mmTraffic实现精细化的流量解释,保证类别预测。大量实验表明,mmTraffic能够自主生成高保真、人可读且证据支持的流量解释报告,同时保持与专用单模态模型(如NetMamba)相当的高分类准确率。源代码可在https://github.com/lgzhangzlg/Multimodal-Reasoning-with-LLM-for-Encrypted-Traffic-Interpretation-A-Benchmark获取。
英文摘要:
Network traffic, as a key media format, is crucial for ensuring security and communications in modern internet infrastructure. While existing methods offer excellent performance, they face two key bottlenecks: (1) They fail to capture multidimensional semantics beyond unimodal sequence patterns. (2) Their black box property, i.e., providing only category labels, lacks an auditable reasoning process. We identify a key factor that existing network traffic datasets are primarily designed for classification and inherently lack rich semantic annotations, failing to generate human-readable evidence report. To address data scarcity, this paper proposes a Byte-Grounded Traffic Description (BGTD) benchmark for the first time, combining raw bytes with structured expert annotations. BGTD provides necessary behavioral features and verifiable chains of evidence for multimodal reasoning towards explainable encrypted traffic interpretation. Built upon BGTD, this paper proposes an end-to-end traffic-language representation framework (mmTraffic), a multimodal reasoning architecture bridging physical traffic encoding and semantic interpretation. In order to alleviate modality interference and generative hallucinations, mmTraffic adopts a jointly-optimized perception-cognition architecture. By incorporating a perception-centered traffic encoder and a cognition-centered LLM generator, mmTraffic achieves refined traffic interpretation with guaranteed category prediction. Extensive experiments demonstrate that mmTraffic autonomously generates high-fidelity, human-readable, and evidence-grounded traffic interpretation reports, while maintaining highly competitive classification accuracy comparing to specialized unimodal model (e.g., NetMamba). The source code is available at https://github.com/lgzhangzlg/Multimodal-Reasoning-with-LLM-for-Encrypted-Traffic-Interpretation-A-Benchmark