双曲视觉-语言模型的层级提示学习
Hierarchical Prompt Learning for Hyperbolic Vision-Language Models
- Aarhus University(奥胡斯大学)
- University of Amsterdam(阿姆斯特丹大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对双曲视觉-语言模型,提出层级提示学习插件,利用父类层级增强提示学习,提升基类到新类泛化与跨数据集迁移,并在双曲空间中实现层级一致的嵌入组织。
AI中文摘要:
双曲视觉-语言模型(VLMs)在天然适合层级结构的几何空间中表示图像和文本特征,但其对下游任务的适配在很大程度上依赖于固定提示。与此同时,现有的提示学习方法将类别标签视为平面集合,未利用可用的分类学结构。我们通过为冻结的双曲VLM提出一种层级提示学习插件来解决这一差距。给定一个固定的离线父类层级,该方法通过一个独立的父类提示学习器、父类级监督、双曲蕴含正则化以及父类反馈逻辑融合来增强类别提示学习器。我们使用CoOp、CoCoOp和MaPLe实例化了该方法,分别得到HyPLO、CoHyPLO和MaHyPLO。在标准的11数据集基准上,所有变体均提升了基类到新类的泛化能力和跨数据集迁移能力,并在域偏移下与各自的提示学习基线保持相当。六个层级指标和嵌入分析表明,该方法产生了更具分类学一致性的预测,并在双曲空间中诱导出父类、类别和图像嵌入的层级一致组织。当新类别必须放置在固定分类学中时,其收益最大;而对于兄弟类别间的细粒度混淆或仅影响图像分布的偏移,其收益最小。
英文摘要:
Hyperbolic vision-language models (VLMs) represent image and text features in a geometry naturally suited to hierarchy, but their adaptation to downstream tasks has largely relied on fixed prompts. Existing prompt learning methods, meanwhile, treat class labels as a flat set and do not exploit available taxonomic structure. We address this gap with a hierarchical prompt learning plug-in for frozen hyperbolic VLMs. Given a fixed offline parent-class hierarchy, it augments a class prompt learner with a separate parent prompt learner, parent-level supervision, hyperbolic entailment regularization, and parent-feedback logit fusion. We instantiate the method with CoOp, CoCoOp and MaPLe, yielding HyPLO, CoHyPLO and MaHyPLO. Across the standard 11-dataset benchmark, all variants improve base-to-new generalization and cross-dataset transfer, and remain comparable to their prompt learning baselines under domain shift. Six hierarchical metrics and embedding analyses show that the method produces more taxonomically consistent predictions and induces a hierarchy-consistent organization of parent, class, and image embeddings in hyperbolic space. Its gains are largest when novel classes must be placed within a fixed taxonomy, and smallest for fine-grained confusions among sibling classes or shifts affecting only the image distribution.