FAIR-Pruner: 一种通过差异容忍性实现自动分层剪枝的灵活框架
FAIR-Pruner: A Flexible Framework for Automatic Layer-Wise Pruning via Tolerance of Difference
- School of Statistics and Mathematics, Zhejiang Gongshang University(浙江工商大学统计与数学学院)
- École de technologie supérieure (ÉTS), Université du Québec(魁北克大学埃克森技术学院)
- Southern University of Science and Technology(南方科技大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出FAIR-Pruner,一种无需搜索的自适应分层结构化剪枝框架,通过引入差异容忍度(ToD)来实现非均匀的分层剪枝深度,从而在多个数据集和模型上实现了良好的准确率-压缩率权衡。
AI中文摘要:
结构化剪枝是压缩深度神经网络的标准工具,但其实际性能取决于稀疏性如何分配到各层。我们提出了FAIR-Pruner,一种无需搜索的自适应分层结构化剪枝框架。FAIR-Pruner使用两种在同一层内的排名:一种是去除导向的信号,提出候选单元;另一种是保护导向的信号,识别任务敏感的单元。其核心组件,差异容忍度(ToD),测量去除前缀与保护尾部之间的重叠,并使用共享容忍级别来诱导各层非均匀的剪枝深度。作为默认视觉实例,FAIR-Pruner结合基于Wasserstein的U-Score用于类条件单元分离性,以及基于Taylor的R-Score用于任务级敏感性;相同的ToD分配规则也可以与替代的去除信号配对。理论上,我们通过群体R-Score分析ToD,推导出高R-Score质量进入剪枝集的排名控制,并识别出相同预算比较与均匀剪枝的加法交换条件。在CIFAR-10、CIFAR-100、SVHN和ImageNet上,跨VGG、ResNet、DenseNet、ConvNeXt和DeiT的实验显示了强的准确率-压缩率权衡。在 routed-expert Qwen1.5-MoE-A2.7B-Chat 上的仅剪枝实验进一步检验了在匹配专家预算下的架构扩展性。FAIR-Pruner作为可 pip-install 的开源包发布。
英文摘要:
Structured pruning is a standard tool for compressing deep neural networks, but its practical performance depends on how sparsity is allocated across layers. We propose FAIR-Pruner, a search-free framework for adaptive layer-wise structured pruning. FAIR-Pruner uses two within-layer rankings: a removal-oriented signal that proposes candidate units and a protection-oriented signal that identifies task-sensitive units. Its core component, Tolerance of Difference (ToD), measures the overlap between the removal prefix and the protected tail, and uses a shared tolerance level to induce non-uniform pruning depths across layers. As a default vision instantiation, FAIR-Pruner combines a Wasserstein-based U-Score for class-conditional unit separability with a Taylor-based R-Score for task-level sensitivity; the same ToD allocation rule can also be paired with alternative removal signals. Theoretically, we analyze ToD through the population R-Score, derive rank-based control of the high-R-Score mass entering the pruning set, and identify an additive exchange condition for same-budget comparison with uniform pruning. Experiments on CIFAR-10, CIFAR-100, SVHN, and ImageNet across VGG, ResNet, DenseNet, ConvNeXt, and DeiT show strong accuracy--compression trade-offs. Prune-only experiments on routed-expert Qwen1.5-MoE-A2.7B-Chat further examine architectural extensibility under matched expert budgets. FAIR-Pruner is released as a pip-installable open-source package.