arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向视觉-语言模型的固定Token基的低秩提示学习

Low-Rank Prompt Learning for Vision-Language Models with Fixed-Token Bases

Tanvir Muntakim Tonoy, Sajjad Ghiasvand, Mahnoosh Alizadeh, Ramtin Pedarsani

arXiv 2609.09462首次发表:更新:

发表机构

UC Santa Barbara(加州大学圣塔芭芭拉分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出低秩分解提示矩阵,固定token基仅训练嵌入侧系数,在少样本基准上以更少参数匹配或超越密集CoOp,揭示提示适应主要发生在嵌入侧。

AI 中文摘要

提示学习通过用学习到的连续上下文向量替换手工编写的模板,使CLIP适应下游识别任务,在上下文优化(CoOp)中,这些向量构成一个密集提示矩阵$\mathbf{P}\in\mathbb{R}^{m\times d}$,仅从每个类别的少量样本中训练。我们研究该矩阵是否过参数化,将其分解为$\mathbf{P}=\mathbf{B}\mathbf{A}$,这将可训练的提示参数从$md$减少到$r(m+d)$,并且一旦token侧因子$\mathbf{B}$固定,则减少到$rd$。在七个少样本基准和两个CLIP主干上,低秩提示在参数远少于密集CoOp的情况下匹配或改进其性能,在低样本基到新类泛化上增益最明显。然后我们发现token侧因子根本无需学习:将$\mathbf{B}$固定为高斯、正交、SVD派生甚至随机基,仅训练嵌入侧因子$\mathbf{A}$,与完全可训练分解保持同等性能,且源训练的$\mathbf{B}$相比随机基无优势。提示因子不对称性和局部更新空间维度差距解释了为何固定$\mathbf{B}$远比固定$\mathbf{A}$限制少,仅光滑性保证即可证明在固定$\mathbf{B}$上优化$\mathbf{A}$收敛。在CLIP提示设置中,嵌入侧系数承载适应,而token基可以简单固定。

英文摘要

Prompt learning adapts CLIP to downstream recognition by replacing hand-written templates with learned continuous context vectors, which in Context Optimization (CoOp) form a dense prompt matrix $\mathbf{P}\in\mathbb{R}^{m\times d}$ trained from only a few examples per class. We study whether this matrix is over-parameterized by factorizing it as $\mathbf{P}=\mathbf{B}\mathbf{A}$, which cuts the trainable prompt parameters from $md$ to $r(m+d)$, and to $rd$ once the token-side factor $\mathbf{B}$ is fixed. Across seven few-shot benchmarks and two CLIP backbones, low-rank prompts match or improve dense CoOp at far fewer parameters, with the clearest gains on low-shot base-to-new generalization. We then find that the token-side factor need not be learned at all: fixing $\mathbf{B}$ to a Gaussian, orthogonal, SVD-derived, or even random basis and training only the embedding-side factor $\mathbf{A}$ stays on par with the fully trainable factorization, and a source-trained $\mathbf{B}$ offers no advantage over a random one. A prompt-factor asymmetry and a local update-space dimension gap show why fixing $\mathbf{B}$ is far less restrictive than fixing $\mathbf{A}$, and a smoothness-only guarantee certifies that optimizing $\mathbf{A}$ over a fixed $\mathbf{B}$ converges. In the CLIP prompt setting, the embedding-side coefficients carry the adaptation while the token basis can simply be fixed.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑