通过文本条件变换控制嵌入空间
Controlling Embedding Spaces with Text-Conditioned Transformations
浏览论文内容
中文总结 AI 辅助
研究提出文本条件变换视觉嵌入,通过自然语言描述属性类别,网络生成仿射变换强调指定属性,能同时学习多属性,可在推理时访问,为控制嵌入空间提供统一高效框架,在多任务中展现近零推理成本的最优性能。
中文摘要 AI 辅助
像CLIP这样的模型中的多模态嵌入空间具有语义相似性检索和跨模态零样本分类等强大功能。这些嵌入将高级语义压缩到单个向量中,代价是主要表达主导语义而抑制其他重要属性。我们提出了一种视觉嵌入的文本条件变换,使这些属性能够明确访问。给定属性类别的自然语言描述,网络生成强调指定属性的仿射变换。基于文本进行条件设定使其能够同时学习多个属性,通过直观界面在推理时访问它们。网络经过训练,使变换后的嵌入与冻结的潜在空间对齐,可使用现有大规模嵌入进行检索而无需重新编码。应用于全集时,同样机制可变换潜在空间用于属性解缠任务。我们的方法通过直接在潜在空间操作,为控制嵌入空间提供了统一高效的框架,在基于属性的检索和多属性组织任务中展现了近零推理成本的最优性能。
英文摘要
Multimodal embedding spaces in models like CLIP enable powerful capabilities such as semantic similarity retrieval and cross-modal zero-shot classification. These embeddings compress high-level semantics into a single vector, which comes at the cost of primarily expressing a dominant semantics like main object while suppressing other important attributes such as camera angle or color tone. We propose a text-conditioned transformation of visual embeddings that makes such attributes explicitly accessible. Given a natural language description of an attribute category (e.g., "color" or "art style"), a network generates an affine transformation that emphasizes the specified attribute. Conditioning on text enables it to learn many attributes simultaneously, accessing them at inference time through an intuitive interface. The network is trained to align transformed embeddings with the frozen latent space, enabling retrieval using existing large-scale embeddings without any re-encoding. When applied to a full set, the same mechanism transforms the latent space for attribute disentanglement tasks such as multi-clustering. By operating directly in latent space, our method provides a unified and efficient framework for controlling embedding spaces, demonstrating state-of-the-art performance across both attribute-based retrieval and multi-attribute organization tasks with near-zero inference cost. Project page: https://joefioresi718.github.io/ControlEmbed_webpage/
发表机构
- Institute of Artificial Intelligence, University of Central Florida(中佛罗里达大学人工智能研究所)
- Adobe Research(Adobe研究院)
机构由 AI 辅助整理,请以论文原文为准。