发表机构
University of Warsaw(华沙大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对视觉基础模型架构庞大的问题,提出可解释性引导的注意力头软剪枝框架SAPER,在ImageNet-1K上实现了优于RAPTOR的精度-效率权衡。
AI 中文摘要
以DINOv2为代表的视觉基础模型能学习到高表达性的表征,但依赖庞大且不透明的架构,需要大量计算能力与内存。为解决该问题,本文提出一种基于注意力图拉普拉斯特征向量的光谱分析方法及新的注意力头可视化技术;结合视觉Transformer块结构的最新研究结论,对注意力头进行语义聚类以识别功能冗余。基于上述发现,本文提出SAPER(Soft Attention PrunER),这是一种基于LapSum Soft Top-K方法的端到端可微剪枝框架。在ImageNet-1K上开展的大量实验表明,SAPER实现了优异的精度-效率权衡,在保留强分类性能的同时,其浮点运算量(FLOPs)的减少效果优于竞争力基准RAPTOR。
英文摘要
Vision foundation models, such as DINOv2, learn highly expressive representations but rely on massive, opaque architectures that demand substantial computational power and memory. To provide an interpretable-guided and efficient solution to this issue, we first propose a spectral analysis and new visualization technique for individual attention heads based on the Laplacian eigenvectors of their attention maps. Building upon recent observations regarding the block structure of Vision Transformers, we perform semantic clustering of attention heads and identify functional redundancies. Leveraging these insights, we introduce SAPER (Soft Attention PrunER), an end-to-end differentiable pruning framework based on the LapSum Soft Top-K approach. Extensive experiments on ImageNet-1K demonstrate that SAPER achieves a highly favorable accuracy-efficiency trade-off, outperforming the competitive RAPTOR baseline in FLOPs reduction while preserving strong classification performance.