HyperCLIP++:在双曲空间中微调CLIP以实现开放词汇语义分割
HyperCLIP++: Fine-tuning CLIP forOpen-vocabulary Semantic Segmentation in Hyperbolic Space
浏览论文内容
中文总结 AI 辅助
针对CLIP微调提升开放词汇分割的现象,提出HyperCLIP++,通过双曲空间半径缩放对齐层级,仅微调5%参数即在三个基准上达到最先进性能。
中文摘要 AI 辅助
CLIP作为一种基础视觉-语言模型,已成为开放词汇语义分割的强大工具。虽然已知冻结CLIP的文本编码器能保留其泛化能力,但近期研究表明,联合微调CLIP的文本和图像编码器能显著提升分割性能,尤其是对开放集中的类别。在本工作中,我们从层级对齐的角度解释这一现象,因为在微调过程中,图像嵌入的层级从图像级转变为像素级。我们通过利用天然编码层级结构的双曲空间来实现这一点。我们的关键观察是,在微调期间,CLIP文本嵌入的双曲半径减小,从而促进与视觉数据像素级粒度的更好对齐。基于此,我们提出了HyperCLIP++,一种新颖且参数高效的适配策略。HyperCLIP++通过缩放变换直接调整CLIP嵌入的双曲半径,以实现与目标任务(即分割)的层级对齐。为确保这种层级对齐在两种模态中一致生效,并在训练期间保持其跨模态对齐,HyperCLIP++集成了一个双交叉关系通信(DCRC)模块,用于同步视觉和文本路径之间的这些调整。我们的实验表明,HyperCLIP++在三个基准上实现了最先进的性能,同时仅微调了CLIP总参数的大约5%。更重要的是,我们观察到调整后,CLIP的文本嵌入在不同数据集上表现出相对固定的双曲半径,这表明该分割任务所需的层级水平可能可以用双曲半径来量化。
英文摘要
CLIP, a foundational vision-language model, has emerged as a powerful tool for open-vocabulary semantic segmentation. While freezing CLIP's text encoder is known to preserve its generalization capability, recent studies show that fine-tuning both CLIP's text and image encoders jointly significantly enhances segmentation performance, especially for classes from open sets. In this work, we explain this phenomenon from the perspective of hierarchy alignment, since during fine-tuning, the hierarchical level of image embeddings shifts from image-level to pixel-level. We achieve this by leveraging hyperbolic space, which naturally encodes hierarchical structures. Our key observation is that, during fine-tuning, the hyperbolic radius of CLIP's text embeddings decreases, facilitating better alignment with the pixel-level granularity of visual data. Building on this, we propose HyperCLIP++, a novel and parameter-efficient adaptation strategy. HyperCLIP++ directly adjusts the hyperbolic radius of CLIP's embeddings via scaling transformations to achieve a hierarchy alignment to the target task, i.e., segmentation. To ensure this hierarchy alignment is effected consistently across both modalities and preserves their cross-modal alignment during training, HyperCLIP++ integrates a Dual Cross-Relation Communication (DCRC) module that synchronizes these adjustments between the vision and text pathways. Our experiments show that HyperCLIP++ achieves state-of-the-art performance across three benchmarks while fine-tuning only approximately 5% of CLIP's total parameters. More importantly, we observe that after adjustment, CLIP's text embeddings exhibit a relatively fixed hyperbolic radius across datasets, suggesting that the hierarchical level required for this segmentation task might be quantified using the hyperbolic radius.
发表机构
- Shanghai Jiao Tong University(上海交通大学)
- Meituan(美团)
机构由 AI 辅助整理,请以论文原文为准。