arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

魔鬼在于频谱偏差:面向鲁棒表示蒸馏的频谱平衡特征匹配

The Devil is in the Spectrum Bias: Spectrum-Balanced Feature Matching for Robust Representation Distillation

Kuniaki Saito, Yoshitaka Ushiku

arXiv 2609.34106首次发表:更新:

发表机构

OMRON SINIC X Corporation(欧姆龙SINIC X公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对特征匹配蒸馏偏向主导频谱方向的问题,提出频谱平衡特征匹配(SpecMatch),自适应强调欠优化方向,在视觉和蛋白质任务中显著提升下游性能。

AI 中文摘要

大型视觉基础模型在广泛的下游任务中展现出了显著的迁移能力。为了高效部署此类模型,特征匹配已成为一种流行的知识蒸馏方法,无需标注数据即可将教师表示迁移至较小的学生模型。然而,我们表明,传统的基于L2距离的特征匹配目标固有地偏向于重建教师表示的主导频谱方向,而欠优化那些通常包含任务相关信息的低方差方向。为解决此问题,我们提出了频谱平衡特征匹配(SpecMatch),这是一种简单目标,能够自适应地强调欠优化的频谱方向,同时保留主导方向的相对重要性。SpecMatch易于实现,且引入的计算开销可忽略不计。在图像识别上的大量实验表明,SpecMatch在多样任务中持续提升下游适应性,包括图像分类、异常检测、医学图像分析和领域泛化。特别地,SpecMatch在42种教师-学生及训练设置组合中的40种上优于传统特征匹配,并在所有设置中持续优于原始学生模型。我们进一步证明,所提出的目标可泛化至视觉之外,在六个蛋白质理解任务中提升了下游性能。

英文摘要

Large visual foundation models have demonstrated remarkable transferability across a wide range of downstream tasks. To deploy such models efficiently, feature matching has become a popular knowledge distillation approach that transfers teacher representations to smaller student models without requiring labeled data. However, we show that the conventional feature matching objective with L2-distance is inherently biased toward reconstructing dominant spectral directions of the teacher representation, while under-optimizing low-variance directions that often contain task-relevant information. To address this, we propose Spectrum-Balanced Feature Matching, SpecMatch, a simple objective that adaptively emphasizes under-optimized spectral directions while preserving the relative importance of dominant directions. SpecMatch is easy to implement and introduces negligible computational overhead. Extensive experiments on image recognition demonstrate that SpecMatch consistently improves downstream adaptation across diverse tasks, including image classification, anomaly detection, medical image analysis, and domain generalization. In particular, SpecMatch outperforms conventional feature matching in 40 of 42 teacher--student and training-setting combinations, while consistently improving over the original student model in all settings. We further demonstrate that the proposed objective generalizes beyond vision, improving downstream performance across six protein understanding tasks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑