ES-VP:用于高效模型适配的能量型动态视觉提示
ES-VP : Energy-Shaped Dynamic Visual Prompting for Efficient Model Adaptation
- Rutgers University(罗格斯大学)
- Zhejiang University(浙江大学)
- University at Buffalo, SUNY(纽约州立大学布法罗分校)
- Adobe Research(奥多比研究院)
- University of Connecticut(康涅狄格大学)
- Washington University(华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出ES-VP方法,以低秩初始化和能量引导动态适配生成图像特异性提示,参数效率更高、泛化性更好,在多架构多数据集上均优于现有SOTA视觉提示方法。
AI中文摘要:
视觉提示(Visual Prompting, VP)已成为一种参数高效的方法,用于将预训练模型适配到下游任务。然而,现有方法在灵活性和效率之间存在权衡:一些方法对所有图像应用固定提示,忽略单个图像的特征;另一些方法则引入辅助网络生成多样化提示,虽然能提升性能,但会显著增加参数使用量,还存在过拟合特定数据集的风险。此外,辅助网络结合预训练模型的固有偏差,限制了方法的可扩展性和泛化性。本文提出能量型视觉提示(Energy-Shaped Visual Prompting, ES-VP),这是一种采用低秩初始化和能量引导动态适配的新方法,可生成图像特异性提示,与单提示方法相比,在更少参数下实现更优性能。ES-VP直接利用预训练模型进行自适应提示生成,确保参数效率并提升泛化性。在五个架构、十五个数据集上开展的大量实验表明,ES-VP始终优于当前最先进(State-of-the-art, SOTA)的单提示和多样化VP方法。例如,在CLIP架构上的四个数据集实验中,ES-VP的准确率较SOTA方法DAM-VP平均提升2.6%,而VP参数使用量减少590倍,从而为高效且可泛化的模型适配建立了新基准。
英文摘要:
Visual prompting (VP) has emerged as a parameter-efficient method for adapting pre-trained models to downstream tasks. However, existing approaches encounter a trade-off between flexibility and efficiency. Some methods apply a fixed prompt to all images, ignoring individual image characteristics, while others introduce auxiliary networks to generate diverse prompts. Although the latter can improve performance, it also significantly increases parameter usage and the potential for overfitting to specific datasets. Furthermore, the auxiliary networks, combined with inherent biases in pre-trained models, limit scalability and generalization. In this paper, we propose Energy-Shaped Visual Prompting (ES-VP), a novel approach that generates image-specific prompts using low-rank initialization and energy-guided dynamic adaptation, achieving superior performance with fewer parameters compared to single-prompt methods. ES-VP directly utilizes the pre-trained model for adaptive prompt generation, ensuring both parameter efficiency and improved generalization. Extensive experiments conducted on five architectures across fifteen datasets demonstrate that ES-VP consistently outperforms current state-of-the-art (SOTA) single and diverse VP methods. For instance, using the CLIP architecture across four datasets, ES-VP outperforms the SOTA method DAM-VP by an average of 2.6\% in accuracy while utilizing 590$\times$ fewer VP parameters, thereby establishing a new benchmark for efficient and generalizable model adaptation.