arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一次剪枝:面向视觉-语言模型的无需重训练的任务无关剪枝

Prune Once: Retraining-Free Task-Agnostic Pruning for Vision-Language Models

Minseok Kang, Hyunwoo Kim, Chanyoung Kim, Minwoo Kim, Jaekoo Lee, Dahuin Jung

arXiv 2608.06901首次发表:更新:

发表机构

Chung-Ang University; Soongsil University; Kookmin University(中央大学; 崇实大学; 国民大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出名为PORTA的无需重训练的任务无关VLM剪枝框架,通过自适应稀疏分配实现高压缩下的竞争力性能,支持高效VLM压缩。

AI 中文摘要

视觉-语言模型(VLMs)通过大规模预训练在各类多模态任务中展现出卓越的泛化能力,但其不断增长的计算与内存需求对受限环境下的部署构成重大挑战。现有剪枝策略往往依赖任务特定准则或面向大语言模型(LLM)的重要性度量,不适用于任务无关剪枝——该场景下剪枝时无任务特定样本可用,且剪枝后的模型仍需具备广泛适用性。本文提出一种名为PORTA的无需重训练的VLM剪枝框架,其基于通用校准数据估算的激活变化量推导任务与模态无关的重要性公式,可可靠捕捉跨模态的特征级表示效用。PORTA还引入自适应稀疏性分配机制,根据输出特征变异性分配各层剪枝比例,避免均匀稀疏性的局限并降低高压缩水平下的性能退化。在CLIP、BLIP、Qwen2-VL等多个VLM架构上开展的大量实验表明,PORTA在高稀疏度下可实现有竞争力的下游性能,且无需任何重训练,支持高效的VLM压缩。代码可在该https URL获取。

英文摘要

Vision-language models (VLMs) have achieved remarkable generalization across diverse multimodal tasks through large-scale pre-training, yet their rapidly increasing computational and memory requirements pose significant challenges for deployment in constrained environments. Existing pruning strategies often depend on task-specific criteria or LLM-oriented importance measures, making them unsuitable for task-agnostic pruning, where no task-specific samples are available at pruning time and the pruned model remains broadly applicable. We introduce a retraining-free VLM pruning framework called PORTA that derives a task- and modality-agnostic importance formulation based on activation variation, estimated from generic calibration data, which reliably captures feature-level representation utility across modalities. PORTA further incorporates an adaptive sparsity allocation mechanism that assigns layer-wise pruning ratios based on output feature variability, avoiding the limitations of uniform sparsity and reducing performance degradation at high compression levels. Extensive experiments across VLM architectures, such as CLIP, BLIP, and Qwen2-VL, demonstrate that PORTA achieves competitive downstream performance under high sparsity without requiring any retraining, supporting efficient VLM compression. Code is available at https://github.com/cau-hai-lab/PORTA.git.

CommentsAccepted to ECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑