通过模型无关的流量定制缓解网站指纹识别中代理诱导的流量漂移
Mitigating Proxy-Induced Traffic Drift in Website Fingerprinting via Model-Agnostic Traffic Tailoring
浏览论文内容
中文总结 AI 辅助
针对网站指纹识别中代理协议多样性导致的流量漂移问题,提出模型无关预处理框架PA3,通过协议漂移指纹定制流量,使WF模型在未见过协议上的F1分数平均提升0.12,最高达0.96以上。
中文摘要 AI 辅助
基于深度学习的网站指纹识别(WF)可有效从加密流量中识别网站。然而,用户常依赖代理协议规避审查,这些协议的多样性构成重大挑战:在一组协议流量上训练的WF模型,在未见过的协议流量上评估时表现极差。我们将此问题归因于代理诱导的特征漂移,即同一网站的流量模式随代理协议变化,导致WF模型无法捕捉的差异及严重性能下降。为解决该问题,我们提出PA3,一个模型无关的预处理框架,用于分析和缓解代理诱导的漂移。PA3首先对协议特定的漂移进行指纹识别,随后利用这些指纹定制代理流量以实现特征对齐,从而缓解漂移并大幅提升WF模型在未见过协议流量上的泛化能力。大量评估表明,PA3显著提升了在未见过协议上的泛化能力,F1分数平均提升0.12(约27%的相对提升),跨模型的绝对增益最高达0.41,缩小了漂移带来的性能差距。在最优情况下,PA3使WF模型在未见过协议的流量上获得0.96以上的F1分数。
英文摘要
Website fingerprinting (WF) based on deep learning can effectively identify websites from encrypted traffic. However, users often rely on proxy protocols to bypass censorship, and the diversity of these protocols poses a major challenge, as WF models trained on traffic from one set of protocols perform poorly when evaluated on that from unseen protocols. We attribute this issue to proxy-induced feature drift, where traffic patterns of the same website vary with the proxy protocol, leading to discrepancies that WF models fail to capture and severe performance degradation. To tackle this issue, we propose PA3, a model-agnostic preprocessing framework to analyze and mitigate the proxy-induced drift. PA3 first fingerprints the protocol-specific drift. These fingerprints are then used to tailor the proxied traffic for feature alignment, which mitigates the drift and considerably improves the generalization of WF models on traffic from unseen protocols. Extensive evaluations demonstrate that PA3 substantially enhances generalization on unseen protocols with an average improvement of 0.12 in F1-score (roughly 27% relative), achieving up to a 0.41 absolute gain across models, which narrows the performance gap introduced by the drift. In the best case, PA3 enables WF models to obtain F1-scores above 0.96 on traffic from unseen protocols.