发表机构
Technical University of Munich; BMW Group; University of Glasgow; TU Wien(慕尼黑工业大学; 宝马集团; 格拉斯哥大学; 维也纳技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对ViTs压缩中Fisher近似保真度与压缩后精度关联弱的问题,提出结构化Fisher近似FACTS及约束秩搜索CoRS,无需微调即在Swin-B上提升Top-1达5.8个百分点。
AI 中文摘要
模型压缩是缓解日益增长的机器学习模型部署挑战的关键。在该研究领域中,基于奇异值分解(SVD)的压缩在计算效率与模型精度之间提供了引人注目的权衡。特别是Fisher加权SVD提供了有原则的、损失感知的压缩。然而,我们发现提高压缩中使用的Fisher近似的保真度对视觉Transformer(ViTs)压缩后的精度预测性较差。受此观察启发,我们提出了FACTS,一种针对使用Fisher加权SVD压缩ViTs的结构化Fisher近似,它在保留token内激活梯度依赖性的同时,强制进行token局部聚合。此外,我们引入了一种快速的约束秩搜索(CoRS),在遵守固定浮点运算(FLOP)约束的同时优化逐层秩分配。在ViTs和混合架构上的大量实验表明,FACTS无需微调即可持续改善精度-效率权衡。值得注意的是,在Swin-B上,它比最强的SVD基线高出多达+5.8个百分点(p.p.)的Top-1精度,我们的搜索方法进一步带来了增益。代码可在该https URL获取。
英文摘要
Model compression is key to mitigate deployment challenges of ever growing machine learning models. In this area of research, singular value decomposition (SVD)-based compression offers a compelling trade-off between computational efficiency and model accuracy. Fisher-weighted SVD in particular provides principled, loss-aware compression. However, we find that improving the fidelity of Fisher approximation used in the compression is poorly predictive of post-compression accuracy for Vision Transformers (ViTs). Motivated by this observation, we propose FACTS, a structured Fisher Approximation tailored to Compressing ViTs with Fisher-weighted SVD, which enforces token-local aggregation while preserving within-token activation-gradient dependence. Additionally, we introduce a fast Constrained Rank Search (CoRS), that optimizes layer-wise rank allocation while adhering to a fixed floating point operation (FLOP) constraint. Extensive experiments across ViTs and hybrid architectures demonstrate that FACTS consistently improves accuracy-efficiency trade-offs without requiring finetuning. Notably, it outperforms the strongest SVD baseline by up to +5.8 percentage points (p.p.) Top-1 on Swin-B, with further gains driven by our search method. Code is available at https://github.com/MoritzTho/FACTS.