发表机构
Orange(Orange公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对LLM分析TLS流量时的上下文窗口溢出问题,提出PCAP-LM表示,实现812倍压缩,在取证问答任务中准确率达99.3%,但存在TCP重传24%漏检等局限。
AI 中文摘要
大型语言模型(LLM)为网络流量分析提供了强大的推理能力,但标准捕获格式及其文本等效形式过于冗长,会使LLM的上下文窗口溢出两个数量级。本文提出PCAP-LM,这是一种以流为中心的LLM原生文本表示,其作用是有损知识提取步骤而非标准压缩工具:原始捕获数据使用PacketGlyphs转码为语义摘要,PacketGlyphs是本文提出的新型ASCII字母表,可编码数据包方向、TCP/TLS状态、对数刻度大小和数据包间延迟。结合受限PMI-BPE分词器和 motif 游程编码,重复的行为模式被大幅压缩。@REFS侧索引保留了对原始数据包的无损下钻功能。在5G/4G TLS 1.3批量下载流量的同质语料库上评估时,BPE词汇表在159个token时完全饱和,相比tshark -V实现了812倍的尺寸缩减,可将完整捕获数据放入单个LLM上下文窗口。在对30个预留文件的取证问答评估中,前沿LLM从PCAP-LM文档获得99.3%的准确率,而在token预算匹配的tshark -V前缀上仅为51.0%。这种有损设计存在已知盲点,最显著的是TCP重传的24%漏检率,扩展到异构混合协议环境需要重新训练词汇表。
英文摘要
Large language models (LLMs) offer powerful reasoning capabilities for network traffic analysis, but standard capture formats and their textual equivalents are prohibitively verbose, overflowing LLM context windows by two orders of magnitude. We present PCAP-LM, a flow-centric, LLM-native text representation that acts as a lossy knowledge extraction step rather than a standard compression tool: raw captures are transcoded into semantic summaries using PacketGlyphs - a novel ASCII alphabet coined in this paper that encodes packet direction, TCP/TLS state, log-scale size, and inter-packet delay. Combined with a constrained PMI-BPE tokenizer and motif run-length encoding, repetitive behavioural patterns are aggressively collapsed. A @REFS side-index preserves lossless drill-down into the original packets. Evaluated on a homogeneous corpus of 5G/4G TLS 1.3 bulk-download traffic, the BPE vocabulary fully saturates at 159 tokens, achieving an 812x size reduction over tshark -V and fitting entire captures within a single LLM context window. In a forensic question-answering evaluation over 30 held-out files, a frontier LLM achieves 99.3% accuracy from PCAP-LM documents versus 51.0% from a token-budget-matched tshark -V prefix. The lossy design introduces known blind spots - most notably a 24% false-negative rate for TCP retransmissions - and extending to heterogeneous mixed-protocol environments will require vocabulary retraining.
Comments6 pages