arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15229cs.CVcs.AI

Pre-PEFT探测:用于VLM视觉编码器层选择的权重统计与扰动鲁棒性

Pre-PEFT Probing: Weight Statistics and Perturbation Robustness for Layer Selection in VLM Vision Encoders

  • Harbin Institute of Technology(哈尔滨工业大学)

机构由 AI 辅助整理,请以论文原文为准。

Qingtao Xia, Jiahua Bao, Siyao Cheng, Jie Liu

AI总结:

提出预微调探测方法,通过分析VLM视觉编码器各层权重统计与扰动鲁棒性来指导PEFT层选择,实验表明高范数和高条件数的层更易获得微调增益,从而减少可训练参数并保持性能。

AI中文摘要:

我们提出了一种用于参数高效微调(PEFT)层选择的预微调探测方法,旨在适配大型视觉-语言模型(VLM)时,以更少的可训练参数获得更稳定和更高的收益。与通常一次性对所有层应用LoRA和其他适配器的常见做法不同——其中层选择往往依赖于启发式规则——我们聚焦于视觉编码器,直接评估每个Transformer层的“可适配性”。具体而言,我们从两个角度刻画每一层:(i)其Q/K/V投影权重的统计特性(如范数和条件数);(ii)在受控参数扰动下的鲁棒性。然后,我们将这些指标与将PEFT应用于单层所带来的下游性能增益进行系统比较。在涵盖七个基准和五种PEFT变体的实验中,我们观察到一致的相关性:具有较大权重范数和较高条件数的层(或矩阵)通常对扰动更鲁棒,并且更可能产生更大的微调增益。这些结果表明,在微调之前进行分布统计分析和扰动测试可以为适配层选择提供实用信号,从而在保持或提升性能的同时减少可训练参数。

英文摘要:

We propose a pre-fine-tuning probing method for Parameter-Efficient Fine-Tuning (PEFT) layer selection, aiming to obtain more stable and higher gains with fewer trainable parameters when adapting large vision--language models (VLMs). Unlike the common practice of applying LoRA and other adapters to all layers at once---where layer selection often relies on heuristic rules---we focus on the vision encoder and directly evaluate the "adaptability'' of each Transformer layer. Specifically, we characterize each layer from two perspectives: (i) the statistical properties of its Q/K/V projection weights (e.g., norms and condition numbers); (ii) robustness under controlled parameter perturbations. We then systematically compare these indicators with the downstream performance gains brought by applying PEFT to a single layer. Across experiments covering seven benchmarks and five PEFT variants, we observe a consistent correlation: layers (or matrices) with larger weight norms and higher condition numbers are usually more robust to perturbations and are more likely to yield larger fine-tuning gains. These results show that distribution-statistics analysis and perturbation tests before fine-tuning can provide practical signals for adaptation-layer selection, thereby maintaining or improving performance while reducing trainable parameters.

补充信息

↑