arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10723cs.CV

保留网格的知识蒸馏:数据稀缺下向视觉Transformer迁移卷积归纳偏置

Grid-Preserving Knowledge Distillation: Transferring Convolutional Inductive Bias to Vision Transformers under Data Scarcity

Junyong Choi, Cheolhyeon Park, Jaehoon Cho

中文总结 AI 辅助

本文提出iBKD蒸馏框架,在数据稀缺场景下通过保留网格的核心模块向ViT迁移卷积归纳偏置,在多基准上优于相关方法且优势随数据减少而扩大。

中文摘要 AI 辅助

当训练数据稀缺时,视觉Transformer(ViT)的性能会逊于卷积神经网络(CNN),从CNN教师模型中蒸馏卷积归纳偏置是一种有效且不改变部署模型的解决方案。然而,通用特征蒸馏在该场景下的迁移效果很差:它在CNN到CNN的流水线中继承的池化、展平及对数几率空间投影会丢弃编码了局部性与平移等变性的空间网格,且与卷积学生模型不同,ViT无法自行重建该结构。本文提出iBKD,一种在整个迁移路径中保留网格的蒸馏框架,其核心模块为归纳偏置注意力模块:该模块通过学习到的权重将每个学生层聚合到教师网格,利用通道和可变形空间注意力强化结构线索,并通过在网格间而非令牌集间操作的卷积交叉注意力注入这些线索。该模块仅在训练时使用,因此部署后的模型是未修改的ViT,无推理开销。在7种Transformer骨干网络和6个数据稀缺基准上,iBKD的性能优于局部性引导方法和通用知识蒸馏基线,且随着训练数据减少,其优势差距会进一步扩大。

英文摘要

Vision Transformers demonstrate remarkable global modeling capacity but often underperform in data-scarce regimes. Distilling convolutional inductive biases from a CNN teacher provides an effective remedy while leaving the deployed model unchanged. However, general-purpose feature distillation transfers little in this setting. In CNN-to-CNN distillation, pooling, flattening, and logit-space projections remove the spatial grid that encodes locality and translation equivariance. Unlike a convolutional student, a ViT cannot readily reconstruct this structure on its own. In this paper, we propose iBKD, a distillation framework that preserves the spatial grid throughout the entire transfer process. Its core module, the Inductive Bias Attention Module, aggregates features from all student layers onto the teacher's grid using learned weights. It then enhances structural cues through channel and deformable spatial attention and injects them via convolutional cross-attention operating directly between spatial grids rather than token sets. The module is used only during training, leaving the deployed model as an unmodified ViT with no inference overhead. Across seven Transformer backbones and six data-scarce benchmarks, iBKD consistently outperforms both locality-guidance methods and general knowledge distillation baselines, with its advantage increasing as the amount of training data decreases.

↑