Unleashing the Power of Vision-Language Models for Long-Tailed Multi-Label Visual Recognition
释放视觉-语言模型在长尾多标签视觉识别中的潜力
机构 * School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院) ; Key Laboratory of Computer Network and Information Integration (Southeast University)(计算机网络与信息集成重点实验室) ; Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(Mohamed bin Zayed人工智能大学) ; Carnegie Mellon University(卡内基梅隆大学)
AI总结 本文提出CAPNET框架,通过显式建模标签相关性,结合图卷积网络和可学习软提示,解决长尾多标签视觉识别中的类别不平衡问题,提升模型性能。