arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SQuaT:基于学生感知量化教师特征的自监督知识蒸馏

SQuaT: Self-Supervised Knowledge Distillation via Student-Aware Quantized Teacher Features

HyeonJun Lee, Hyeonsik Jo, Jinwoo Chung, Jangho Kim

arXiv 2608.10709首次发表:更新:

发表机构

Kookmin University(国民大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对量化感知训练结合知识蒸馏的损失下界问题,提出SQuaT框架,通过学生量化参数量化教师特征消除下界,在低比特设置下性能优于基线,适用多种架构。

AI 中文摘要

量化感知训练(QAT)可实现量化模型的部署,且精度下降极小。但实际场景中,受隐私、版权或成本限制,训练标签往往不可用。知识蒸馏(KD)是解决该问题的常用方法,不过我们发现,现有将QAT与KD结合的研究存在一个根本局限:蒸馏过程中,教师模型与量化学生模型的范围不匹配会产生无法实现的残差,导致蒸馏损失存在不可降低的下界。基于此,我们提出SQuaT(Student-Aware Quantized Teacher Features,学生感知量化教师特征),这是一种结合KD的无标签QAT框架,其核心是在蒸馏过程中,通过应用学生模型的量化参数来量化教师模型的特征,从理论上消除该损失下界。我们在多种设置下开展了全面实验,结果显示SQuaT始终优于强大的基线方法,尤其在极低比特(如1比特和2比特)设置下提升显著。此外,对多种模型设计选择的大量评估表明,该方法不依赖特定架构假设,可广泛适用于各类架构和量化设置。源代码可在指定URL获取。

英文摘要

Quantization-Aware Training (QAT) enables the deployment of quantized models with minimal accuracy degradation. However, in practical scenarios, training labels are often unavailable due to privacy, copyright, or cost constraints. Knowledge Distillation (KD) is a common approach to address this challenge, but we observe that prior work combining QAT with KD suffers from a fundamental limitation: during distillation, the range mismatch between the teacher and the quantized student model induces an unattainable residual, resulting in an irreducible lower bound on the distillation loss. Motivated by this observation, we propose SQuaT (Student-Aware Quantized Teacher Features), a label-free QAT framework with KD that theoretically eliminates this lower bound by applying the student's quantization parameters to quantize the teacher's features during distillation. Through comprehensive experiments across diverse settings, we demonstrate that SQuaT consistently outperforms strong baselines, with particularly pronounced gains in extreme low-bit (e.g., 1- and 2-bit) settings. Furthermore, extensive evaluations across various model design choices show that our approach does not rely on specific architectural assumptions, making it broadly applicable across diverse architectures and quantization settings. The source code is available at https://github.com/lcdbsa522/SQuaT.

CommentsAccepted at AISTATS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑