AI 中文总结
研究双编码器视觉语言模型组合约束失效问题,提出因式推理,将证据提取与约束执行分离,引入LCSE方法,在FACTOR-Bench上实验,提升了准确率并保持检索性能。
AI 中文摘要
双编码器视觉语言模型具有相似性接口,可实现零样本检索,但不满足组合约束。例如,即使概念检测可靠,“雨伞且无人”的查询也会检索到同时包含两者的图像。我们将此归因于接口级的概念袋效应,即相似性分数近似于概念证据的平均池化,而与运算符无关。虽然文本嵌入中存在依赖运算符的信号,但它们太弱或未对齐,无法影响排名。微调无法可靠解决此问题,因为主要瓶颈在于相似性如何聚合证据,而非编码器的表示。我们提出因式推理,将证据提取与约束执行分离,并引入LCSE(逻辑约束分数编辑),一种无需训练的方法,使用冻结编码器的概念分数在外部执行约束。我们还引入了FACTOR-Bench,在其中LCSE的准确率达到85.5%,而最佳微调基线为73.2%;应用于SigLIP 2时为90.7%;同时将NegBench COCO MCQ的准确率从27.2%提高到65.2%,且保持检索性能。
英文摘要
Dual-encoder vision-language models (VLMs) expose a similarity interface that enables zero-shot retrieval but fails compositional constraints: queries like "umbrella and no person" retrieve images containing both, even when concept detection is reliable. We trace this to an interface-level Bag-of-Concepts effect, where similarity scores approximate mean pooling of concept evidence regardless of operators. Although operator-dependent signals exist in text embeddings, they are too weak or misaligned to affect rankings. Fine-tuning does not reliably resolve this failure because the dominant bottleneck is how similarity aggregates evidence rather than what encoders represent. We propose factored inference, which separates evidence extraction from constraint execution, and introduce LCSE (Logic-Constrained Score Editing), a training-free method that executes constraints externally using concept scores from frozen encoders. We also introduce FACTOR-Bench, where LCSE achieves 85.5% accuracy versus 73.2% for the best fine-tuned baseline, 90.7% when applied to SigLIP 2, and improves NegBench COCO MCQ accuracy from 27.2% to 65.2% while preserving retrieval performance.
CommentsAccepted at ICML 2026. 18 pages, 8 figures. Project page: https://sultanmo.github.io/factored-vlm Code and benchmark: https://github.com/SultanMo/factored-vlm