SoK:流标签从何而来?加密流量基准测试中的标签来源审计
SoK: Where Do Flow Labels Come From? Auditing Label Provenance in Encrypted Traffic Benchmarks
浏览论文内容
中文总结 AI 辅助
本文审计14个加密流量基准,发现其标签侧存在粗粒度继承与过严过滤问题,推导了侧信道特征下分类器平衡准确率上限,为基准构建者和用户提出建议。
中文摘要 AI 辅助
加密流量分类从传输层观测值推断流记录之外的语义,监督式训练依赖于附加到单个流的标签。近期的系统研究审查了模型输入和数据划分,本文则对互补的标签侧进行系统梳理。在审计的14个基准条目中,我们识别出两种反复出现的标签侧策略:粗粒度继承,存在将证据未覆盖的流打上标签的风险;以及过严过滤,仅保留自证明流,存在丢弃相关流的风险。所有被审计条目均未公开可计数的预选总体,且在引用的23个单元格中,下游论文附加到相同标签的任务对象与恢复的记录存在8处不一致。在严格的侧信道特征下,我们推导了受限于这些特征的任何分类器的平衡准确率的表示相关上限:在采用继承策略的公开基准上,该上限范围为0.56至0.76。在过滤侧,我们完整捕获的语料库中仅有24.95%的连接带有自身的可观测SNI;然而被丢弃的连接通过同运行共现特征将宏准确率从0.44提升至0.65。最后,我们为基准构建者和用户提出了建议。
英文摘要
Encrypted traffic classification infers semantics beyond the flow record from transport-layer observables, and supervised training rests on labels that hold for the individual flow they are attached to. Recent systematizations scrutinize model in- puts and data splits; we systematize the complementary label side. Across 14 audited benchmark entries, we identify two recurring label-side strategies: coarse inheritance, which risks labelling flows the evidence does not cover, and overstrict filtering, which keeps only self-attesting flows and risks dis- carding relevant ones. No audited entry exposes a countable pre-selection population, and the task objects downstream papers attach to the same labels disagree with the recovered record in 8 of 23 referenced cells. Under strict side-channel features we derive a representation-relative ceiling on bal- anced accuracy for any classifier restricted to those features: on the public benchmarks that inherit, it ranges from 0.56 to 0.76. On the filtering side, only 24.95% of connections in our fully captured corpus carry an observable SNI of their own; yet the discarded connections raise macro accuracy from 0.44 to 0.65 through same-run co-occurrence features. We end with recommendations for benchmark builders and users.