发表机构
School of Electronic Information and Electrical Engineering, Shanghai Jiao Tong University; School of Cyber Science and Engineering, Southeast University(上海交通大学电子信息与电气工程学院; 东南大学网络空间安全学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出流量原生基础模型SCSM,通过分割、组合、缩放和掩码构建对比预训练视图,在时间漂移和跨领域任务上显著提升网站指纹识别可迁移性。
AI 中文摘要
网站指纹识别从加密流量元数据中推断用户访问的网站。然而,在固定采集条件下训练的模型往往会随着网站集合、采集时间、网络路径、浏览器或防御措施的变化而性能下降。现有的可迁移攻击要么依赖于对单个轨迹的手工扰动,要么将面向语言的架构适配到流量上,限制了其捕获流量原生语义的能力。为解决这些局限性,我们提出了SCSM,一种用于可迁移网站指纹识别的流量原生基础模型。具体而言,SCSM通过分割、组合、缩放和掩码操作,从同一组无标签轨迹中构建预训练视图对。这些操作产生多样化的可观测模式,同时保留真实轨迹片段中的底层数据包事件和局部流量动态。相应的窗口化流量计数矩阵作为对比预训练的输入,用于训练一个状态空间编码器,无需网站标注。预训练编码器随后在一个小型标注支持集上进行微调,并在推理时跨时间尺度聚合预测。实验结果表明,在GTT23上的六个时间漂移任务中,SCSM的平均top-3准确率超过最强基线16.4%,在四个跨领域数据集上的平均宏F1分数超过最强基线13.8%。代码和数据集将在该https URL上提供。
英文摘要
Website fingerprinting infers the websites visited by users from encrypted traffic metadata. However, models trained under fixed collection conditions often degrade as website sets, collection times, network paths, browsers, or defenses change. Existing transferable attacks either rely on handcrafted perturbations of individual traces or adapt language-oriented architectures to traffic, limiting their ability to capture traffic-native semantics. To address these limitations, we propose SCSM, a traffic-native foundation model for transferable website fingerprinting. Specifically, SCSM constructs pairs of pretraining views from the same group of unlabeled traces through Segmentation, Combination, Scaling, and Masking. These operations produce diverse observable patterns while preserving the underlying packet events and local traffic dynamics of real trace fragments. The corresponding windowed traffic counting matrices serve as inputs for contrastive pretraining of a state-space encoder without website annotations. The pretrained encoder is then fine-tuned on a small labeled support set, and predictions are aggregated across temporal scales at inference. Experimental results demonstrate that SCSM surpasses the strongest baselines by 16.4\% in average top-3 accuracy over six temporal-drift tasks on GTT23 and by 13.8\% in average macro-F1 over four cross-domain datasets. The code and datasets will be made available at https://github.com/SJTU-dxw/WF-SCSM.