arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向视觉与语言Transformer的损伤感知Bandit剪枝

Damage-Aware Bandit Pruning for Vision and Language Transformers

Salem Ameen, Sunil Vadera

arXiv 2609.05448首次发表:更新:

发表机构

University of Salford(索尔福德大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出损伤感知Bandit剪枝方法,将Transformer结构化单元选择建模为多臂Bandit问题,在固定预算下通过掩蔽损失配对评估,实验表明优于预算贪心方法。

AI 中文摘要

结构化后训练剪枝需要选择完整的功能单元,抑制这些单元只会造成有限的性能下降。我们将语言和视觉Transformer的结构化单元选择问题,在固定候选评估预算下,建模为损伤感知的多臂Bandit问题。注意力头和MLP通道组在校准批次上被临时掩蔽。配对损伤定义为同一批次上掩蔽损失减去基础损失,从而减少批次间差异。平滑有界奖励驱动UCB风格策略或分数Beta Thompson采样,最终掩码通过逐步添加单元来顺序构建。所选单元在原始稠密检查点中被功能置零;因此,所报告的参数效果代表有效的结构抑制,而非物理压缩或实测加速。在WikiText-2、LAMBADA和Imagenette上的实验覆盖GPT-2、OPT、Pythia、Qwen2.5、SmolLM2、ViT-B/16、DeiT-Tiny和Swin-Tiny,并与随机、幅度、静态显著性和预算贪心选择进行比较。在五个随机种子下,Bandit方法在配对语言模型比较中通常相对于预算贪心减少性能下降。论文中强调的28项比较中,23项bootstrap置信区间排除零,11项配对检验p<0.05;在全部116项数据集级检验的Benjamini-Hochberg校正后,6项q<0.05。ViT-B/16和Swin-Tiny的匹配评估结果表明,其增益并非仅由更大的候选评估预算解释。

英文摘要

Structured post-training pruning of transformers requires selecting complete functional units whose suppression causes limited degradation. We formulate structured-unit selection for language and vision transformers as a damage-aware multi-armed bandit problem under a fixed candidate-evaluation budget. Attention heads and MLP channel groups are temporarily masked on calibration batches. Paired damage is the masked loss minus the base loss on the same batch, reducing batch-to-batch variation. A smooth bounded reward drives either a UCB-style policy or fractional-Beta Thompson Sampling, and the final mask is constructed sequentially by adding one unit at each step. The selected units are functionally zeroed in the original dense checkpoint; therefore, the reported parameter effects represent effective structural suppression rather than physical compression or measured speedup. Experiments on WikiText-2, LAMBADA, and Imagenette cover GPT-2, OPT, Pythia, Qwen2.5, SmolLM2, ViT-B/16, DeiT-Tiny, and Swin-Tiny, with comparisons against random, magnitude, static-saliency, and budgeted-greedy selection. Across five seeds, the bandit methods usually reduce degradation relative to budgeted greedy in the paired language-model comparisons. Of 28 comparisons highlighted in the paper, 23 bootstrap confidence intervals exclude zero and 11 paired tests have p < 0.05; six have q < 0.05 after Benjamini-Hochberg correction across the full family of 116 dataset-wise tests. Matched-evaluation results for ViT-B/16 and Swin-Tiny indicate that their gains are not explained solely by a larger candidate-evaluation budget.

Comments23 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑