arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DIFFCZSL:基于扩散表示正则化的组合零样本学习

DIFFCZSL: Compositional Zero-Shot Learning Regularized by Diffusion Representations

Hangyu Tian, Zhenqi He, Yanghao Wang, Long Chen

arXiv 2608.19871首次发表:更新:

发表机构

The Hong Kong University of Science and Technology(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DIFFCZSL将预训练扩散模型的生成先验注入基于CLIP的组合零样本学习流程,通过对比对齐提升性能,在两类设置下均优于强基线,凸显了扩散表示与视觉-语言模型的互补优势。

AI 中文摘要

组合零样本学习(CZSL)旨在利用从已见组合中学习到的基础概念知识,识别未见的属性-对象组合。尽管近期工作通过利用大型视觉-语言模型在CZSL中取得了令人瞩目的性能,但它们主要依赖判别式表示,可能无法明确保留基础概念与其组合之间的结构化关系。受近期基于扩散的分类器取得的成功及其相对于判别式模型的竞争力的启发,本文研究中间扩散表示是否能为CZSL提供互补线索。为此,我们提出DIFFCZSL,这是一种扩散增强框架,将预训练扩散模型的生成先验注入基于CLIP的CZSL流程中。我们提取中间扩散表示并将其投影到CLIP嵌入空间,以对图像和文本模态提供辅助监督。训练期间,通过CLIP嵌入与扩散特征之间的对比对齐,我们的方法促使嵌入几何朝向更丰富的组合感知语义,且推理时不引入额外成本。在三个公开CZSL基准上的大量实验表明,在闭世界和开世界设置下,该方法均优于强大的基于CLIP的基线。我们的结果凸显了生成式扩散表示与判别式视觉-语言模型在组合泛化方面的互补优势。

英文摘要

Compositional Zero-Shot Learning (CZSL) aims to recognize unseen attribute-object compositions by leveraging knowledge of primitive concepts learned from seen compositions. Although recent works achieve impressive performance in CZSL by leveraging large vision-language models, they primarily rely on discriminative representations that may not explicitly preserve the structured relationships between primitive concepts and their compositions. Motivated by the recent success of diffusion-based classifiers and their competitive performance relative to discriminative models, we investigate whether intermediate diffusion representations can provide complementary cues for CZSL. To this end, we propose DIFFCZSL, a diffusion-augmented framework that injects generative priors from pre-trained diffusion models into CLIP-based CZSL pipelines. We extract intermediate diffusion representations and project them into the CLIP embedding space to provide auxiliary supervision on both image and text modalities. Through contrastive alignment between CLIP embeddings and diffusion features during training, our method encourages the embedding geometry toward richer composition-aware semantics, while introducing no additional cost at inference time. Extensive experiments on three public CZSL benchmarks demonstrate consistent improvements over strong CLIP-based baselines under both closed-world and open-world settings. Our results highlight the complementary strengths of generative diffusion representations and discriminative vision-language models for compositional generalization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑