AI 中文总结
UniDFKD是一种统一的无数据知识蒸馏框架,用架构无关语义先验替代特定架构统计先验,在CNNs和ViTs上实现最优性能,比现有方法平均绝对提升超20%。
AI 中文摘要
无数据知识蒸馏(DFKD)通过合成语义信息丰富的数据,将知识从预训练的教师模型迁移到紧凑的学生模型,无需访问原始训练数据集。现有DFKD方法严重依赖特定架构的统计先验(如批量归一化统计量)来指导数据合成,但这类依赖架构的先验在现代架构(如视觉Transformer(ViTs))中往往缺失,导致合成数据的语义质量下降,进而引发灾难性的性能退化。本文提出UniDFKD,一种统一的无数据知识蒸馏框架,它用显式的、与架构无关的语义先验替代特定架构的统计量。UniDFKD沿三个维度管控整个合成-蒸馏流程:(1)类别语义条件(CSC)通过持续用语言衍生的嵌入调制生成器以捕捉语义多样性,定义要合成的内容;(2)空间语义锚定(SSA)通过将教师模型的空间属性锚定到高斯先验,规定证据的位置;(3)空间语义蒸馏(SSD)通过在预测过程中显式对齐教师与学生模型的空间证据,控制知识的迁移方式。在卷积神经网络(CNNs)和ViTs上开展的大量实验表明,UniDFKD达到了新的最优性能,在同构和异构设置中均比现有方法平均绝对提升超过20%。
英文摘要
Data-Free Knowledge Distillation (DFKD) transfers knowledge from a pretrained teacher model to a compact student model by synthesizing semantically informative data, eliminating the need for access to the original training dataset. Existing DFKD methods rely heavily on architecture-specific statistical priors (e.g., Batch Normalization statistics) to guide data synthesis, however, such architecture-dependent priors are often absent in modern architectures such as Vision Transformers (ViTs), resulting in degraded semantic quality of the synthesized data and consequently catastrophic performance degradation. In this paper, we propose \emph{UniDFKD}, a unified data-free knowledge distillation framework that replaces architecture-specific statistics with explicit, architecture-agnostic semantic priors. \emph{UniDFKD} governs the entire synthesis-distillation pipeline along three dimensions: (1) Categorical Semantic Conditioning (CSC) defines \emph{what} to synthesize by persistently modulating the generator with language-derived embeddings to capture semantic diversity; (2) Spatial Semantic Anchoring (SSA) dictates \emph{where} evidence belongs by anchoring the teacher's spatial attributions to a Gaussian prior; and (3) Spatial Semantic Distillation (SSD) controls \emph{how} knowledge is transferred by explicitly aligning teacher-student spatial evidence alongside predictions. Extensive experiments across CNNs and ViTs demonstrate that UniDFKD establishes a new state-of-the-art, outperforming existing methods by an average absolute margin of over 20\% in both homogeneous and heterogeneous settings.