arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Hyper3-CLIP:层级条件双曲视觉-语言训练

Hyper3-CLIP: Hierarchy-Conditioned Hyperbolic Vision-Language Training

Matin Mahmood, Antonio Rueda-Toicen, Mohamed ElBassat, Seifeldin Elkerdany, Weixing Wang, Gerard de Melo

arXiv 2608.29313首次发表:更新:

发表机构

hyper 3 labs; Hasso Plattner Institute; University of Potsdam; Faculty of Computers and Data Science, Alexandria University; Faculty of Computer Science and Engineering, Alamein International University(超3实验室; 哈索·普拉特纳研究所; 波茨坦大学; 亚历山大大学计算机与数据科学学院; 阿拉曼国际大学计算机科学与工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Hyper3-CLIP结合层级条件与双曲几何,通过文本构建多粒度查询层级训练视觉-语言模型,提升图像-文本检索与多标签分类性能,在层级指标上保持竞争力。

AI 中文摘要

采用对比目标训练的类CLIP视觉-语言模型(VLMs)能学习到强大的全局图像-文本表征,但其欧几里得嵌入与全局池化无法编码部分-整体、父子关系等关系结构。双曲VLMs通过基于蕴含的目标解决这一问题,而文本条件变体则通过句子和短语级查询提升细粒度对齐。然而这两类工作相互独立:双曲VLMs使用静态图像和区域特征,查询条件方法则缺乏层级几何结构。本文提出Hyper3-CLIP,一种层级条件双曲VLM,将全局、局部及全局-局部对比学习与查询条件视觉池化相结合。为训练该模型,我们从文本构建轻量查询层级,包含完整标题、句子片段、局部部件描述及提取的短语;每个查询条件视觉块的池化,生成的表征支持图像-文本、整体-部分及父子蕴含损失。查询条件池化仅在训练阶段激活。Hyper3-CLIP在COCO和Flickr上的R@5、R@10检索任务及VOC、COCO上的多标签分类任务均有提升,同时在层级指标上保持竞争力。我们还在固定提示机制下审查了零样本提示敏感性,并研究了训练中使用的局部GRIT部件预算的影响。代码可在该https URL获取。

英文摘要

CLIP-like vision-language models (VLMs) trained with contrastive objectives learn strong global image-text representations, but their Euclidean embeddings and global pooling fail to encode relational structure such as part-whole and parent-child relations. Hyperbolic VLMs address this gap with entailment-based objectives, and text-conditioned variants improve fine-grained alignment through sentence- and phrase-level queries. However, these two lines of work remain separate: hyperbolic VLMs use static image and region features, while query-conditioned methods lack hierarchical geometric structure. We present Hyper3-CLIP, a hierarchy-conditioned hyperbolic VLM that combines global, local, and global-local contrastive learning with query-conditioned visual pooling. To train the model, we construct lightweight query hierarchies from text, comprising full captions, sentence fragments, localized part descriptions, and extracted phrases. Each query conditions the pooling of visual patches, and the resulting representations support image-text, whole-part, and parent-child entailment losses. Query-conditioned pooling is active only during training. Hyper3-CLIP improves R@5 and R@10 retrieval on COCO and Flickr, as well as multi-label classification on VOC and COCO, while remaining competitive on hierarchy metrics. We also audit zero-shot prompt sensitivity under fixed prompt regimes and study the effect of the localized GRIT part budget used during training. Code is available at https://github.com/Hyper3Labs/hyper3-clip.

Comments16 pages, 2 figures, ECCV 2026 Beyond Euclidean Workshop

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑