arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ZOMP:面向视觉-语言模型的零阶多模态提示调优

ZOMP: Zeroth-Order Multi-Modal Prompt Tuning for Vision-Language Models

Sajjad Ghiasvand, Yifan Yang, Mahnoosh Alizadeh, Ramtin Pedarsani

arXiv 2608.08060首次发表:更新:

发表机构

UC Santa Barbara(加州大学圣巴巴拉分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出ZOMP方法,通过跨模态低秩重参数化等技术,在仅前向传递的条件下高效调优CLIP模型,在13个基准上优于现有无反向传播的提示调优方法,泛化能力更强。

AI 中文摘要

对CLIP等视觉-语言模型进行微调通常需要通过完整模型进行反向传播(BP),但在仅可获取前向传递权限的场景下这是不可行的,内存受限的边缘设备、专有模型部署中常存在此类情况。现有的无反向传播的零阶提示调优方法虽无需反向传播,但通常仅针对单模态调整提示,或在足够大的搜索空间上优化,导致收敛需要数千次前向传递,在实际查询预算下不实用。本文提出ZOMP(零阶多模态提示调优),这是一种查询高效、完全仅需前向传递的方法,通过同时扰动随机近似,在冻结的CLIP模型的视觉分支和文本分支中调整深度提示。该方法结合了三个核心要素:跨模态低秩重参数化,通过共享因子连接两个分支并保持有效搜索维度较小;梯度校正动量项,用于稳定有噪声的零阶估计;按查询预算调整的秩调度,随查询预算的消耗释放模型容量。在13个视觉-语言基准上,以5000次查询的匹配预算下,ZOMP在少样本准确率和查询效率上均显著优于现有无反向传播的提示调优方法,且在基类到新类迁移、跨数据集迁移、分布外场景下的泛化能力更强。研究结果表明,联合利用多模态与低秩结构是实现实用、查询高效的无反向传播提示调优的有效途径。

英文摘要

Fine-tuning vision-language models such as CLIP typically requires backpropagation (BP) through the full model, which is infeasible when only forward-pass access is available, as is common for memory-constrained edge devices and proprietary model deployments. Prior BP-free, zeroth-order prompt-tuning methods avoid this requirement but often tune prompts in a single modality or optimize over a search space large enough that convergence requires thousands of forward passes, which is impractical under realistic query budgets. We propose ZOMP (Zeroth-Order Multimodal Prompt tuning), a query-efficient, fully forward-only method that tunes deep prompts in both the vision and text branches of a frozen CLIP model using simultaneous perturbation stochastic approximation. ZOMP combines three ingredients: a cross-modal low-rank reparameterization that ties the two branches through a shared factor and keeps the effective search dimensionality small, a gradient-correction momentum term that stabilizes the noisy zeroth-order estimate, and a budget-indexed rank schedule that unlocks capacity as the query budget is spent. Across 13 vision-language benchmarks under a matched 5,000-query budget, ZOMP consistently outperforms prior BP-free prompt-tuning methods in both few-shot accuracy and query efficiency, and it generalizes better across base-to-new, cross-dataset transfer, and out-of-distribution settings. Our results show that jointly exploiting multimodality and low-rank structure is an effective route to practical, query-efficient BP-free prompt tuning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑