SPACE:面向域自适应的CLIP嵌入语义投影与对齐
SPACE: Semantic Projection and Alignment of CLIP Embeddings for Domain Adaptation
浏览论文内容
中文总结 AI 辅助
SPACE方法利用CLIP嵌入的语义结构,通过奇异值分解构建正交基作为语义锚点,将视觉特征投影到语义子空间,以解决域偏移问题,实现基于语义的跨域对齐。
中文摘要 AI 辅助
在部署视觉模型时,一个根本性挑战是域偏移,即训练数据和测试数据遵循不同分布时,会导致性能下降。当同一语义概念以不同视觉形式出现时,例如照片和素描,这一挑战会被放大,因为尽管存在语义对应关系,但视觉相似性较弱。现有的无监督域自适应方法旨在对齐跨域的分布,但往往忽略了同一类别内样本之间的语义关系。为解决此问题,本文提出了SPACE,一种利用CLIP视觉-语言空间的语义结构进行域自适应的方法。其关键思想是使用文本描述作为语义锚点,通过对类别描述的CLIP嵌入应用奇异值分解,得到一个捕获类别间语义关系的正交基。来自两个域的视觉特征被投影到该语义子空间中,从而基于语义而非外观来对齐图像。
英文摘要
A fundamental challenge in deploying vision models is domain shift, which arises when training and test data follow different distributions, leading to degraded performance. This challenge is amplified when the same semantic concept appears under distinct visual forms, such as photographs and sketches, where visual similarity is weak despite semantic correspondence. Existing unsupervised domain-adaptation methods aim to align distributions across domains but often ignore semantic relationships among samples of the same class. To address this issue, this paper introduces SPACE, a method that exploits the semantic structure of CLIP's vision-language space for domain adaptation. The key idea is to use text descriptions as semantic anchors by applying Singular Value Decomposition to CLIP embeddings of class descriptions, yielding an orthogonal basis that captures semantic relationships among categories. Visual features from both domains are projected into this semantic subspace, aligning images based on meaning rather than appearance.
发表机构
- São Paulo State University (UNESP) School of Sciences(圣保罗州立大学(UNESP)理学院)
机构由 AI 辅助整理,请以论文原文为准。