arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

关于瓶颈生成架构中边界寻求蒸馏的失败

On the Failure of Boundary-Seeking Distillation in Bottlenecked Generative Architectures

Mohamed Amine Kina

arXiv 2607.15919首次发表:更新:

AI 中文总结

研究无数据知识蒸馏中边界寻求原则能否用于自动编码器蒸馏,通过在MNIST数据集实验发现其在瓶颈生成架构中不适定,CAKE方法会产生梯度冲突,流形感知合成绕过冲突,为无数据生成蒸馏建立有效基线。

AI 中文摘要

无数据知识蒸馏将教师模型中编码的知识转移到学生模型,而无需访问原始训练数据。先前的工作如对比归纳知识提取(CAKE)通过在教师决策边界附近合成样本,实现了分类器的无数据知识蒸馏。本文通过在MNIST数据集上的实验,研究这种边界寻求原则是否适用于自动编码器蒸馏。为了进行直接比较,将连续重建重新表述为密集的、逐特征分类任务,使解码器能够输出分类对数。结果表明,在瓶颈生成架构中,边界寻求目标从根本上是不适定的。CAKE在单个实例级目标上运行,但解码器充当由共享低维瓶颈约束的紧密耦合的特征级分类器阵列。独立采样这些耦合输出的对比目标会违反学习到的潜在流形的几何结构,并产生严重的梯度冲突,而不是信息丰富的边界样本。流形感知合成完全绕过了这些冲突,并为无数据生成蒸馏建立了有效的基线。

英文摘要

Data-free knowledge distillation transfers the knowledge encoded in a teacher model to a student model without access to the original training data. Prior work such as Contrastive Abductive Knowledge Extraction (CAKE) achieves this for classifiers by synthesizing samples near the teacher's decision boundary. In this work, we investigate whether this boundary-seeking principle extends to autoencoder distillation through experiments on the MNIST dataset . To enable a direct comparison, we reformulate continuous reconstruction as a dense, per-feature classification task, allowing the decoder to output categorical logits. We show that boundary-seeking objectives are fundamentally ill-posed in bottlenecked generative architectures. CAKE operates on a single, instance-level objective, but a decoder acts as an array of tightly coupled, feature-level classifiers constrained by a shared low-dimensional bottleneck. Independently sampling contrastive targets for these coupled outputs violates the geometry of the learned latent manifold and produces severe gradient conflicts instead of informative boundary samples. Manifold-aware synthesis bypasses these conflicts entirely and establishes an effective baseline for data-free generative distillation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑