arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11789cs.CV

基于时序负知识迁移的抗捷径蒸馏

Anti-Shortcut Distillation via Temporal Negative Knowledge Transfer

  • Neubility Inc.(纽比蒂公司)
  • Perception AI(感知人工智能)

机构由 AI 辅助整理,请以论文原文为准。

Syed Muhammad Raza, Omer Tariq, Jeongbae Son

AI总结:

提出抗捷径蒸馏(ASD)框架,利用教师模型优化轨迹识别捷径方向,结合时序对比损失与捷径抑制损失,在多数据集的教师-学生对中提升了模型准确率与腐败鲁棒性。

AI中文摘要:

知识蒸馏(KD)通过让紧凑的学生模型向已收敛的教师模型学习来进行训练,但它未关注教师模型自身学习到的需要抑制的方向:虽存在排斥和感知偏差的目标,但 none 利用教师模型自身的优化轨迹来确定学生应规避的内容。我们发现,缺失的信号已编码在教师模型的优化轨迹中:早期教师模型强调但已收敛教师模型弱化的特征,正是值得推动学生模型远离的捷径方向。我们将这一观察实例化为抗捷径蒸馏(ASD),这是一种推拉式 KD 框架,将已收敛教师模型 $\boldsymbol{T}_{final}$ 作为正语义锚点,将早检查点教师模型 $\boldsymbol{T}_{early}$ 作为时序负参考。ASD 结合了两种损失:时序对比损失($\boldsymbol{\textit{L}}_{tc}$),在 InfoNCE 目标中将早期教师模型特征作为同一样本的负样本,与批次内和记忆库中的最终教师模型特征相对;以及捷径抑制损失($\boldsymbol{\textit{L}}_{ss}$),惩罚学生模型向 $\boldsymbol{E}[\boldsymbol{\nabla}\boldsymbol{h}\boldsymbol{\nabla}\boldsymbol{h}^{\top}]$ 的 top 特征向量投影,$\boldsymbol{E}[\boldsymbol{\nabla}\boldsymbol{h}\boldsymbol{\nabla}\boldsymbol{h}^{\top}]$ 是早期到最终特征位移的未中心化二阶矩矩阵。在 CIFAR-100、ImageNet-100 和 TinyImageNet 上的 13 组教师-学生对实验中,ASD 在超过 10 组对中获得最高的干净 top-1 准确率,在 12 组对中优于标准 KD。在 CIFAR-100-C 对抗腐败鲁棒性上,ASD 在最具挑战性的跨架构对(WRN-40-2→ShuffleNet-V2)上获得最低的平均腐败误差(86.1 mCE)。机制诊断证实了预期的几何结构:ASD 学生模型与捷径方向系统地反对齐,而其在鲁棒子空间上的投影显著更大(0.45 对比 0.12)。

英文摘要:

Knowledge distillation (KD) trains a compact student by attracting it towards a converged teacher. It is silent about which directions the teacher itself learned to suppress: repulsive and bias-aware objectives exist, but none exploits the teacher's own trajectory to identify what the student should avoid. We observe that the missing signal is already encoded in the teacher's optimization trajectory: features that an early-stage teacher emphasizes but that a converged teacher attenuates are precisely the shortcut directions worth pushing the student away from. We instantiate this observation as \textbf{A}nti-\textbf{S}hortcut \textbf{D}istillation (ASD), a push--pull KD framework that treats the converged teacher $\Tfinal$ as a positive semantic anchor and an early-checkpoint teacher $\Tearly$ as a temporal negative reference. ASD couples two losses: a temporal contrastive loss ($\Ltc$) that places the early-teacher feature as a same-sample negative against in-batch and memory-bank final-teacher features in an InfoNCE objective; and a shortcut suppression loss ($\Lss$) that penalizes student projection onto the top eigenvectors of $\E[\Dh\Dh^{\top}]$, the uncentered second-moment matrix of early-to-final feature displacements. Across 13 teacher--student pairs on CIFAR-100, ImageNet-100, and TinyImageNet, ASD attains the highest clean top-1 accuracy on more than 10 pairs and outperforms standard KD on 12. On CIFAR-100-C corruption robustness, ASD obtains the lowest mean Corruption Error ($86.1$\,mCE) on the most challenging cross-architecture pair (WRN-40-2$\to$ShuffleNet-V2). Mechanistic diagnostics confirm the intended geometry: the ASD student is systematically anti-aligned with the shortcut direction, while its projection onto the robust subspace is substantially larger ($0.45$ vs.\ $0.12$).

补充信息

↑