arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BindCLIP: 一种用于组合视觉语言评分的平衡耦合方法

BindCLIP: One Balanced Coupling For Compositional Vision Language Scoring

Liuyang Song, Yi Zhang, Zhongyi Deng, Daqian Yang, Hongbo Zhang

arXiv 2609.23717首次发表:更新:

发表机构

Peking University; Dongguan University of Technology; Sichuan Agricultural University(北京大学; 东莞理工学院; 四川农业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

BindCLIP提出一种平衡的令牌-补丁-深度最优传输耦合,用于组合视觉语言评分,无需额外标签或检测器,在多个基准上超越冻结CLIP,尤其在关系分割上表现最强。

AI 中文摘要

全局视觉-语言相似度将图像和标题压缩为一个向量,保留了语义,但无法体现哪个词对应哪个区域,或这些区域如何排列;模型可以识别每个词和物体,却可能偏好组合上不正确的标题。我们认为,冻结的编码器保留了这种关联结构,因此问题在于读取它,而非在预训练相似度之外重建它。我们引入了BindCLIP,一种基于单一潜在对象的成对评分器:一个平衡的令牌-补丁-深度最优传输耦合,将候选标题和多个视觉深度置于同一个计划中。语义、实体、顺序和空间证据作为该状态的能量被读取,交换候选标题会置换计划,使评分严格反对称。耦合内部的几何细化收缩了候选标题和视觉深度不支持的运动。不使用任务标签、解析器、关系清单或检测器。一个检查点和一个推理路径在官方What'sUp、ARO和SugarCrepe基准上优于冻结的全局CLIP,在关系分割上迁移最强。控制实验排除了补丁访问和标题长度捷径,推理时的消融将空间排列定位到耦合。

英文摘要

Global vision--language similarities compress an image and a caption into one vector, preserving semantics but not which word corresponds to which region or how those regions are arranged; a model can recognize every word and object yet prefer a compositionally incorrect caption. We argue that a frozen encoder retains this association structure, so the problem is to read it rather than to rebuild it beside the pretrained similarity. We introduce BindCLIP, a pairwise scorer built on one latent object: a balanced token--patch--depth optimal-transport coupling that places both candidate captions and several visual depths in a single plan. Semantic, entity, order, and spatial evidence are read as energies of this state, and exchanging the candidates permutes the plan, making the score exactly antisymmetric. A geometric refinement inside the coupling contracts moves that the candidates and the visual depths do not support. No task label, parser, relation inventory, or detector is used. One checkpoint and one inference path improve the official What'sUp, ARO, and SugarCrepe benchmarks over frozen global CLIP, with the strongest transfer on the relation splits. Controls rule out patch access and caption-length shortcuts, and an inference-time lesion localizes spatial arrangement to the coupling.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑