arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14966cs.CV

用于视觉语言模型的U形多粒度学习

U-shaped Multi-granularity Learning for Vision-Language Models

Biao Chen, Yunqian Yu, Xiangxu Zhao, Zhongshu Chen, Mengmeng Jing, Lin Zuo

首次发表
浏览论文内容

中文总结 AI 辅助

研究视觉语言模型提示学习的粒度困境,提出U形多粒度提示学习框架UPrompt,通过构建并行多粒度表示及粗细粒度交互增强与监督,在多个基准实验中验证了其有效性和优势。

中文摘要 AI 辅助

视觉语言模型的提示学习范式有效但面临粒度困境:全局提示缺乏细粒度语义感知,局部提示忽略上下文关联,限制跨任务泛化,这种困境存在于密集预测任务中。受统一跨粒度多级表示的U-Net启发,提出UPrompt,一个用于视觉语言模型的U形多粒度提示学习框架。它在视觉和文本模态中构建并行多粒度表示,通过粗到细的级联增强传播全局上下文以细化局部细节,同时细到粗的分层监督确保跨尺度语义一致性。在17个基准上的广泛实验验证了其有效性。

英文摘要

The prompt learning paradigm for vision-language models is effective yet faces a granularity dilemma: global prompts lack fine-grained semantic awareness, while local prompts ignore contextual associations, limiting cross-task generalization. This dilemma exists in dense prediction tasks. Inspired by U-Net, which unifies multi-level representations across granularities, we propose UPrompt, a U-shaped multi-granularity prompt learning framework for vision-language models. Similar to how U-Net integrates fine and coarse features through symmetric encoder-decoder pathways with cross-level connections, UPrompt constructs parallel multi-granularity representations in both visual and textual modalities, where coarse-to-fine cascaded enhancement propagates global context to refine local details, while fine-to-coarse hierarchical supervision ensures semantic consistency across scales. Extensive experiments on 17 benchmarks validate our effectiveness. UPrompt outperforms MAMET and VPKE by 4.1 and 7.3 rSum on MSCOCO, surpasses CoCoA-Mix by 5.09% in base-to-novel generalization, while maintaining competitive performance with minimal overhead (coarse-grained) and matching PSRC with 1/3 cost (medium-grained).

发表机构

  • University of Electronic Science and Technology of China(电子科技大学)
  • Hithink Research(恒生电子股份有限公司)

机构由 AI 辅助整理,请以论文原文为准。

↑