发表机构
University College London(伦敦大学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文揭示CLIP等双编码器的绑定缺陷源于激励机制与编码结构限制,通过数学分析明确了深度、目标、几何三方面的阻碍,为改进双编码器模型提供了理论依据。
AI 中文摘要
CLIP等双编码器模型通过两个独立计算的单位向量的单个内积对图像-文本对进行评分,且在绑定任务上存在缺陷,当被要求区分“红色汽车和蓝色狗”与“蓝色汽车和红色狗”时,其评分通常接近随机水平。本文对该缺陷何时为必然、何时为偶然给出了数学解释。在Kang等人提出的理想编码器框架内,我们首先证明相关公理是可满足的,因此所有不可能性必然源于额外的、可验证的假设。随后我们证明了三个此类阻碍因素:深度方面,对于递归角色绑定编码,交换间隔遵循精确定律m(D)=2b^(-D),其中D为嵌套深度,有限维版本在一个明确标记的浓度估计范围内成立;可解析深度随维度仅呈对数增长,在CLIP规模下为个位数,对应普通语言的嵌套深度。目标方面,无架构约束的限制定理表明,对比目标对绑定的全部奖励受限于训练中将文本与其自身交换进行对比的速率,该速率在网络规模下会消失,且完全反转绑定的成本仅为该速率乘以平均绑定间隔,两者均在模拟中得到验证。几何方面,存在严格的平滑度-绑定边界:与交换相关的两个文本必须嵌入到共享释义锚点的距离越近,绑定间隔就越小,且存在精确常数。在18个已部署的文本编码器上测量仅文本诊断指标,每个模型均处于其上限的25%-35%左右,且诱导的每项目标上限与SugarCrepe子集难度的相关系数r=0.99。因此,已部署双编码器中的绑定缺陷并非当前的维度或平滑度限制,而是激励机制和编码结构的限制,且在修正这些因素后,已证明的深度上限仍然存在。
英文摘要
Dual-encoder models such as CLIP score an image-caption pair by a single inner product of two independently computed unit vectors, and fail at binding, often scoring near chance when asked to distinguish "a red car and a blue dog" from "a blue car and a red dog". We give a mathematical account of when this failure is necessary and when it is contingent. Working within the ideal-encoder framework proposed by Kang et al., we first show the relevant axioms are satisfiable, so every impossibility must enter through an added, checkable hypothesis. We then prove three such obstructions. Depth: for recursive role-binding codes the swap margin obeys an exact law $m(D) = 2b^{-D}$ in the nesting depth D, with a finite-dimension version holding up to one explicitly flagged concentration estimate; the resolvable depth grows only logarithmically in the dimension and is single-digit at CLIP scale, the nesting depth of ordinary language. Objective: architecture-free throttle theorems showing that the contrastive objective's entire reward for binding is bounded by the rate at which training contrasts a caption against its own swap, a rate that vanishes at web scale, and that exactly reversed binding costs only that rate times the mean binding margin; both are verified in simulation. Geometry: a tight smoothness-binding frontier: the closer the two swap-related captions must embed to a shared paraphrase anchor, the smaller the binding margin can be, with an exact constant. Measuring its text-only diagnostic across 18 deployed text encoders, every model sits at roughly 25-35% of its ceiling, and the induced per-item ceiling tracks SugarCrepe's subset difficulty at r = 0.99. Binding failure in deployed dual encoders is thus not a dimension or smoothness limit today, but an incentive and code-structure limit, with a proved depth ceiling that remains once those are fixed.