arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31141cs.NIcs.LG

未知流量检测、校准与捷径依赖:一年内蒸馏加密流量分类器的研究

Unknown-Traffic Detection, Calibration and Shortcut Reliance in Distilled Encrypted-Traffic Classifiers over One Year

Mahmoud Abbasi

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过一年真实TLS流量上的预注册实验,发现知识蒸馏主要传递教师模型的捷径依赖和过度自信,而非未知流量检测能力,后者可通过标签平滑获得。

中文摘要 AI 辅助

知识蒸馏是将加密流量分类器压缩到边缘设备的标准方法,而几乎所有此类工作仅以准确率来评判学生模型。我们追问学生模型还继承了哪些其他特性:未知流量检测、校准、捷径依赖,以及这些特性在一年数据漂移后是否依然存在。相似性本身并不能证明什么,因为软目标也具有正则化作用。因此,我们从一个由两个准确率相同但结构不同的教师模型(一个五成员集成模型和一个单一更宽的模型)中蒸馏出一个101k参数的学生模型,这样学生模型选择跟随哪一个教师模型就可以归因于该教师模型。该设计在查看任何测试结果之前已预先注册。我们在CESNET-TLS-Year22(一年的真实TLS流量)上测试了十个假设,跨越35周内的18个测试窗口。其中两个假设得到支持:学生模型的每流未知分数向其自身教师模型偏移,但仅在常规温度下,而非准确率最优温度下;以及一个依赖捷径的教师模型将其过度自信传递给了从未见过该特征的学生模型。漂移预测在两种评分下均被逆转,差距缩小而非扩大,并且在三个重复实验中,学生模型在能量评分下两次超越教师模型;同样,关于此类教师模型会损害其学生模型检测能力的预测也被逆转,学生模型的检测能力略有提升。捷径依赖由模型大小决定,而非蒸馏。在基于logit的评分下,没有其他特性转移:蒸馏既不优于温度缩放的直接学生模型,也不优于标签平滑。探索性分析表明,这取决于评分规则:使用特征空间检测器时,教师模型检测未知流量的AUROC比直接学生模型高0.073,而能量评分下差异为0.000,常规温度下的学生模型继承了大部分优势。没有教师模型的标签平滑恢复了更多能力。蒸馏传递了教师模型的习惯;看似继承的能力在没有教师模型的情况下也能获得。

英文摘要

Knowledge distillation is the standard way to compress encrypted-traffic classifiers for the edge, and almost all such work judges students by accuracy alone. We ask what else a student inherits: unknown-traffic detection, calibration, shortcut reliance, and whether any survives a year of drift. Resemblance proves little on its own, since soft targets also regularise. We therefore distil one 101k-parameter student from two teachers of equal accuracy but different construction, a five-member ensemble and a single wider model, so that following one rather than the other is attributable to it. The design was pre-registered before any test result was seen. We tested ten hypotheses on CESNET-TLS-Year22, a year of real TLS traffic, across 18 test windows over 35 weeks. Two are supported: a student's per-flow unknown-scores shift toward its own teacher, but only at a conventional temperature, not the accuracy-optimal one; and a shortcut-reliant teacher passes its over-confidence to a student that never sees the feature. The drift prediction is reversed under both scores, the gap narrowing rather than widening and the student overtaking under the energy score in two of three replicates, as is the prediction that such a teacher harms its student's detection, which improves slightly. Shortcut reliance is set by model size, not distillation. Under the logit-based scores nothing else transfers: distillation beats neither a temperature-scaled direct student nor label smoothing. Exploratory analysis shows this turns on the scoring rule: with a feature-space detector the teacher detects unknown traffic 0.073 AUROC better than the direct student, where the energy score sees 0.000, and the conventional-temperature student inherits most of it. Label smoothing, with no teacher, recovers more. Distillation transfers the teacher's habits; what looks like an inherited ability is available without one.

发表机构

  • University of Salamanca(萨拉曼卡大学)
  • AIR Institute(AIR研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑