发表机构
Lucerne University of Applied Sciences and Arts (HSLU)(卢塞恩应用科学与艺术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究将JEPA风格预测学习用于JA4衍生网络指纹,构建JA4-JEPA模型,在特定子域和数据上训练,通过冻结kNN探测器评估,结果显示该方法能产生有用嵌入,即便跨源视图重叠不完整。
AI 中文摘要
I-JEPA和V-JEPA通过将潜在预测与目标编码器输出进行匹配来学习,而非再生原始输入,这在图像和视频领域效果良好。我们探究此目标对紧凑网络指纹是否有效。构建了JA4-JEPA,这是一个基于Transformer的模型,在从JA4DB和CIC-IDS-2017提取的JA4、JA4H、JA4S和JA4X子域上训练。训练数据结合了来自两个来源的约39.7万个样本。通过冻结kNN探测器在TLS、DNS和SSH的协议族分类上评估学习到的表示。在39416个保留样本上,模型余弦相似度达0.9899,kNN准确率达0.9220。结果表明JEPA风格的预测学习能从JA4衍生指纹中产生有用嵌入。
英文摘要
I-JEPA and V-JEPA learn by matching latent predictions to target encoder outputs rather than regenerating the original input, and this has worked well for images and video. We explore whether the same objective works for compact network fingerprints. We built JA4-JEPA, a Transformer-based model trained on JA4, JA4H, JA4S, and JA4X subfields drawn from JA4DB and CIC-IDS- 2017. The training data combines roughly 397K samples from both sources, though no single sample contains all four view families. We evaluated the learned representations with a frozen kNN probe on protocol-family classification across TLS, DNS, and SSH. On 39,416 heldout samples the model achieved a cosine similarity of 0.9899 and a kNN accuracy of 0.9220. These results indicate that JEPA-style predictive learning can produce useful embeddings from JA4-derived fingerprints, even with incomplete view overlap across sources. Keywords: JA4, network fingerprinting, JEPA, predictive representation learning, self-supervised learning
Commentsthe manuscript requires substantial revision following further validation