Understanding the Multi-modal Prompts of the Pre-trained Vision-Language Model
专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV
Comments We find that the statistical information in Figure 2 neglect the statistics for tSOS, so we make corrections. Additionally, we change the statistical samples to those where CLIP misidentify, but prompt tuning identify correctly. At the same time, we also revise some of the descriptions. The changes to the supplementary materials will be updated shortly. arXiv admin note: text overlap with arXiv:2307.06948 by other authors