用于视觉语言模型的U形多粒度学习
U-shaped Multi-granularity Learning for Vision-Language Models
浏览论文内容
中文总结 AI 辅助
研究视觉语言模型提示学习的粒度困境,提出U形多粒度提示学习框架UPrompt,通过构建并行多粒度表示及粗细粒度交互增强与监督,在多个基准实验中验证了其有效性和优势。
中文摘要 AI 辅助
视觉语言模型的提示学习范式有效但面临粒度困境:全局提示缺乏细粒度语义感知,局部提示忽略上下文关联,限制跨任务泛化,这种困境存在于密集预测任务中。受统一跨粒度多级表示的U-Net启发,提出UPrompt,一个用于视觉语言模型的U形多粒度提示学习框架。它在视觉和文本模态中构建并行多粒度表示,通过粗到细的级联增强传播全局上下文以细化局部细节,同时细到粗的分层监督确保跨尺度语义一致性。在17个基准上的广泛实验验证了其有效性。
英文摘要
The prompt learning paradigm for vision-language models is effective yet faces a granularity dilemma: global prompts lack fine-grained semantic awareness, while local prompts ignore contextual associations, limiting cross-task generalization. This dilemma exists in dense prediction tasks. Inspired by U-Net, which unifies multi-level representations across granularities, we propose UPrompt, a U-shaped multi-granularity prompt learning framework for vision-language models. Similar to how U-Net integrates fine and coarse features through symmetric encoder-decoder pathways with cross-level connections, UPrompt constructs parallel multi-granularity representations in both visual and textual modalities, where coarse-to-fine cascaded enhancement propagates global context to refine local details, while fine-to-coarse hierarchical supervision ensures semantic consistency across scales. Extensive experiments on 17 benchmarks validate our effectiveness. UPrompt outperforms MAMET and VPKE by 4.1 and 7.3 rSum on MSCOCO, surpasses CoCoA-Mix by 5.09% in base-to-novel generalization, while maintaining competitive performance with minimal overhead (coarse-grained) and matching PSRC with 1/3 cost (medium-grained).
发表机构
- University of Electronic Science and Technology of China(电子科技大学)
- Hithink Research(恒生电子股份有限公司)
机构由 AI 辅助整理,请以论文原文为准。