发表机构
State Key Laboratory of Virtual Reality Technology and Systems; School of Computer Science and Engineering; Beihang University(虚拟现实技术与系统国家重点实验室; 计算机科学与工程学院; 北京航空航天大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对部分标签多标签分类中统计先验导致的过拟合问题,提出语言驱动的密集语义适配器LDSA,结合密集对比适配器与类特定提示调整的交互解码器,在公开基准上取得当前最佳性能。
AI 中文摘要
基于不完整标注的多标签图像分类是一项具有挑战性的任务,因其在大规模数据集上能实现高效率与低人力消耗的出色平衡,已被广泛研究。主流方法依赖强先验假设从部分标注中恢复缺失语义,但这些统计先验存在不稳定的语义错误,进而导致灾难性过拟合。为此,我们提出一种语言驱动的密集语义适配器(LDSA),用于从多模态预训练CLIP模型中挖掘先验自适应关系。在我们的方法中,首先提出密集对比适配器以构建密集视觉对比约束,将任务特定知识迁移至视觉领域;随后,我们提出一种语言驱动的交互解码器,借助类特定提示调整,使语言代理适配视觉领域。通过所提模块的协同学习,实验结果表明,我们提出的LDSA在公开多标签分类基准上达到了新的当前最佳性能,可解释分析显示,我们的LDSA通过先验自适应学习方案发现了隐式语义关系。
英文摘要
Learning multi-label image classification with incomplete annotations is a challenging task that has been widely studied for its superior trade-off between high efficiency and less labor consumption on large-scale datasets. Predominant methods rely on strong prior assumptions to recover the missing semantics from partial annotations. However, these statistic priors suffer from unstable semantic mistakes and thus lead to catastrophic overfitting. Toward this end, we propose a Language-driven Dense Semantic Adaptor (LDSA) that excavates prior-adaptive relationships from multimodal pretrained CLIP models. In our approach, the densely contrastive adaptor is first proposed to construct dense visual contrastive constraints, transferring the task-specific knowledge to visual domains. We then propose a language-driven interactive decoder with the help of class-specific prompt tuning, which adapts language proxies with visual domains. With the collaborative learning of proposed modules, experimental results demonstrate our proposed LDSA achieves a new state of the art on public multi-label classification benchmarks, and interpretable analyses reveal that our LDSA discovers implicit semantic relationships with the prior-adaptive learning scheme.