PPOM:用于基于CLIP的通用视觉-语言提示调优的补丁网格相位边缘化
PPOM: Marginalizing Patch-Grid Phase for CLIP-Based Generalizable Vision-Language Prompt Tuning
浏览论文内容
中文总结 AI 辅助
本研究针对基于CLIP的视觉-语言提示调优对补丁网格对齐敏感的问题,提出无训练的PPOM算子,通过边缘化相位偏移提升宿主性能,无需重新训练。
中文摘要 AI 辅助
提示调优以少量可训练参数适配基于CLIP的视觉-语言模型,但其预测仍对冻结视觉Transformer施加的空间采样敏感。具体而言,非重叠补丁分词使预测依赖于图像与补丁格网之间的对齐(相位)。为降低预测对补丁网格对齐的敏感性,我们引入补丁相位轨道边缘化(PPOM),这是一种无训练的推理算子,将相位偏移视为冗余变量。给定补丁步长,PPOM评估恒等视图与反射填充平移,将相反偏移配对为水平、垂直和对角对立族,并为这些族及恒等预测分配相等权重,以避免相位集成期间的视图计数偏差。综上,PPOM在提示适配与补丁网格敏感性之间提供确定性接口。在多个提示学习宿主上,PPOM无需重新训练即可提升宿主性能。
英文摘要
Prompt tuning adapts CLIP-based vision-language models with few trainable parameters, yet its predictions remain sensitive to the spatial sampling imposed by a frozen vision transformer. In particular, non-overlapping patch tokenization makes predictions depend on the alignment (phase) between image and the patch lattice. To reduce prediction sensitivity to patch-grid alignment, we introduce Patch-Phase Orbit Marginalization (PPOM), a training-free inference operator that treats phase shift as a nuisance variable. Given a patch stride, PPOM evaluates the identity view and reflection-padded translations, pairs opposite shifts into horizontal, vertical, and diagonal antithetic families, and assigns equal mass to these families and the identity prediction to avoid view-count bias during phase integration. In summary, PPOM provides a deterministic interface between prompt adaptation and patch-grid sensitivity. Across multiple prompt-learning hosts, PPOM improves host performance without re-training.
发表机构
- Shanghai University(上海大学)
- University of Technology Sydney(悉尼科技大学)
机构由 AI 辅助整理,请以论文原文为准。