发表机构
Beijing University of Posts and Telecommunications; Beijing Key Laboratory of Network System and Network Culture; Key Laboratory of Interactive Technology and Experience System(北京邮电大学; 网络系统与网络文化北京市重点实验室; 交互技术与体验系统重点实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出CLEAR框架,通过从粗到细提取条件变体并进行完形填空式推理,在组合零样本学习中实现抽象语义推断,显著提升C-GQA和MIT-States数据集上的性能。
AI 中文摘要
组合零样本学习旨在通过重新组合已学习的基元来识别未见过的组合。近期方法依赖视觉语言模型,并尝试通过多种表示显式建模基元的上下文变化。然而,这类方法受限于固定的变体容量以及抽象语义与具体语义之间的竞争。在本工作中,我们提出一种新视角,将基元变化视为具体视觉线索的上下文驱动激活,而非独立实体。基于此,我们提出CLEAR,一种受人类感知过程启发的完形填空式推理重排序框架。CLEAR以从粗到细的方式从基元候选集中提取条件变体,执行完形填空式推理以推断高层语义,并重新排序预测以纠正对显著具体基元的偏差。大量实验表明,CLEAR持续改进基础模型,并在具有挑战性的C-GQA和MIT-States数据集上优于最先进方法。代码可在以下网址获取:https://this https URL。
英文摘要
Compositional Zero Shot Learning aims to recognize unseen compositions by recombining learned primitives. Recent methods rely on vision language models and attempt to explicitly model contextual variations of primitives through multiple representations. However, such approaches are limited by fixed variant capacity and competition between abstract and concrete semantics. In this work, we present a new perspective that views primitive variations as the context-driven activation of concrete visual cues rather than independent entities. Based on it, we propose CLEAR, a CLoze-style rEAsoning-based Re-ranking framework inspired by human perceptual processes. CLEAR extracts conditional variants from the primitive candidate set in a coarse-to-fine manner, performs cloze-style reasoning to infer high-level semantics, and re-ranks predictions to correct biases toward salient concrete primitives. Extensive experiments demonstrate that CLEAR consistently improves the Base Model and outperforms state-of-the-art methods on the challenging C-GQA and MIT-States datasets. Code is available at https://github.com/buptLwz/CLEAR.
CommentsAccepted to ICME 2026