arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CS-CLIP:组合场景图引导的CLIP用于稳健的组合推理

CS-CLIP: Compositional Scene Graph-guided CLIP for Robust Compositional Reasoning

SeongJun Jeong, Minjoon Jung, Woo Suk Choi, Youwon Jang, Byoung-Tak Zhang

arXiv 2609.08242首次发表:更新:

发表机构

Seoul National University(首尔国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出CS-CLIP,利用场景图构建结构化负样本,解决组合推理中元素特定偏差,实现稳健的最先进性能并保留通用能力。

AI 中文摘要

视觉-语言模型(VLMs)在组合推理基准测试中表现出强大的性能,这些基准要求对对象、属性、关系及其交互的语义扰动进行推理。然而,我们的受控分析揭示,现有的组合感知VLM表现出元素特定的偏差,往往在某些组合元素上表现不如原始CLIP。为解决这一问题,我们提出了组合场景图引导的CLIP(CS-CLIP),它利用场景图识别组合元素,并通过选择性掩蔽构建结构化负样本。我们进一步保留与原始标题最矛盾的负样本,迫使模型依赖组合结构而非表面线索。CS-CLIP在组合推理上达到了最先进的性能,并在各组合元素上表现稳健。它还保留了跨模态检索和下游视觉推理等通用视觉-语言能力,同时比先前方法需要更少的训练样本。

英文摘要

Vision-language models (VLMs) demonstrate strong performance across compositional reasoning benchmarks, which require reasoning over semantic perturbations of objects, attributes, relations, and their interactions. However, our controlled analysis reveals that existing compositionality-aware VLMs exhibit element-specific biases, often underperforming vanilla CLIP on certain compositional elements. To address this, we propose Compositional Scene Graph-guided CLIP (CS-CLIP), which uses scene graphs to identify compositional elements and construct structured negatives via selective masking. We further retain negatives that are most contradictory to the original caption, forcing the model to rely on compositional structure rather than surface cues. CS-CLIP achieves state-of-the-art compositional reasoning with robust performance across compositional elements. It also preserves general vision-language capabilities such as cross-modal retrieval and downstream visual reasoning, while requiring fewer training samples than prior methods.

CommentsAccepted to Findings of EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑