发表机构
The Ohio State University; Netflix Research(俄亥俄州立大学; Netflix 研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出COMPASS框架,通过分解组合描述,量化视觉-语言模型中组合失败的各因素,发现技能退化主要受自身负载影响,跨负载多为正向,表明组合退化不能仅归因于联合推理。
AI 中文摘要
视觉-语言模型(VLMs)通常在组合推理任务上表现不佳,但造成这种性能不足的原因仍不清楚。一个常见的假设是,模型难以整合多个组件,从而引发旨在改善组合绑定的训练干预措施。然而,这一假设从未被直接量化。现有基准仅以组合形式评估描述,这使得在负载增加的情况下,无法将联合推理的成本与识别单个组件的成本分开。我们引入了COMPASS(技能组合分析),一个受控评估框架,旨在隔离和衡量组合失败背后的不同因素。通过将组合描述的性能与其分解对应物进行比较,我们直接量化了87K图像-描述对中组合整合的成本。在多个视觉-语言模型中,这一差距是真实但部分的,仅解释了观察到的性能下降的一部分。这激发了对控制模型行为的其他因素的更细粒度调查。我们分析了单个技能层面的性能:物体检测、属性绑定和关系推理,使用针对技能的扰动,覆盖274K图像-描述对。我们发现了一致的技能特定模式:每种技能主要随其自身原始类型的数量(自负载)而退化,而跨负载效应主要是正向的,表明不同类型的原始提供了有用的基础上下文。这一模式在标准对比编码器、显式训练的组合推理模型以及非对比架构中均成立。这些发现表明,组合退化反映了多个可分离的因素,不能仅归结为联合推理。
英文摘要
Vision-language models (VLMs) often struggle with compositional reasoning tasks, but the reasons for this underperformance remain unclear. A common hypothesis is that models struggle to integrate multiple components, leading to training interventions to improve compositional binding. However, this assumption has never been directly quantified. Existing benchmarks evaluate captions only in their composed form, making it impossible to separate the cost of joint reasoning from the cost of recognizing individual components under increasing load. We introduce COMPASS (COMPositional Analysis of SkillS), a controlled evaluation framework designed to isolate and measure the distinct factors underlying compositional failure. By comparing performance on composed captions with their decomposed counterparts , we directly quantify the cost of compositional integration across 87K image-caption pairs. Across multiple VLMs, this gap is real but partial, accounting for only part of the observed degradation. This motivates a finer-grained investigation into what additional factors govern model behavior. We analyze performance at the level of individual skills: object detection, attribute binding, and relation reasoning, using skill-targeted perturbations across 274K image-caption pairs. We find a consistent skill-specific pattern: each skill degrades primarily with the count of its own primitive type (self-load), while cross-load effects are predominantly positive, suggesting that primitives of different types provide useful grounding context. This pattern holds across standard contrastive encoders, explicitly trained compositional reasoning models, and non-contrastive architectures. These findings show that compositional degradation reflects multiple separable factors that cannot be reduced to joint reasoning alone.