ProtoLIP:从句子级到对象级的证据解缠
ProtoLIP: From Sentence-Level to Object-Level Evidence Disentanglement
浏览论文内容
中文总结 AI 辅助
ProtoLIP提出轻量级原型介导证据层,通过语义族路由实现对象级证据解缠,无需空间标注,提升定位精度并保持检索性能。
中文摘要 AI 辅助
查询条件下的视觉-语言模型通过揭示视觉证据如何随文本查询变化,实现了细粒度的解释。然而,基于完整描述的证据并不一定能分解为对象特定的证据,且暴露的证据图也不一定能识别构成模型预测的证据。在多个VLM架构和独立基准测试中,我们发现对象级查询往往保留来自共现对象和共享上下文的证据。在本文中,我们引入了ProtoLIP,一种轻量级的原型介导证据层,它将可复用的视觉原型组织成文本派生的语义族,并使用查询相关的族路由来约束哪些原型可以提供证据。在没有空间标注或骨干网络重训练的情况下,ProtoLIP改善了跨查询粒度的证据定位和分离,其定位增益可迁移到具有良好对齐的补丁-文本表示的独立预训练VLM。尽管仅使用文本派生的弱监督,ProtoLIP在保持强匹配和竞争力的图像-文本检索的同时,仍能与空间监督的接地模型相媲美。关键在于,ProtoLIP直接根据局部化的原型证据构建其匹配分数,使得该分数能够精确分解为语义族和原型的贡献。
英文摘要
Query-conditioned vision-language models enable fine-grained interpretation by revealing which visual content supports a given textual query and how this evidence changes across queries. However, semantically, sentence-level evidence does not necessarily decompose into object-specific contributions, while spatially, object-level evidence can remain entangled with co-occurring objects and surrounding scene context. Across multiple VLM architectures and independent benchmarks, we observe persistent object-level evidence entanglement. Moreover, exposed evidence maps do not necessarily correspond to the evidence that directly constitutes the model's prediction. To disentangle visual evidence at both semantic and spatial levels, we introduce ProtoLIP, a lightweight prototype-mediated evidence layer that organizes reusable visual prototypes into text-derived semantic families and uses coarse-to-fine evidence routing, where semantic families constrain prototype eligibility and the complete query determines fine-grained prototype contributions. Our studies show that ProtoLIP improves evidence localization and separation across query granularities, achieving average relative gains of 29% in Pointing and 43% in Energy across four object- and phrase-level OOD benchmarks. Its localization gains also transfer to independently pretrained VLMs, with larger improvements observed in several transfer settings. On the primary backbone, ProtoLIP also improves image-text matching discrimination while remaining competitive with a spatially supervised grounding model in object-level localization. Crucially, ProtoLIP constructs its image-text matching score directly from localized prototype evidence, enabling exact decomposition across prototypes, semantic families, and spatial evidence without spatial annotations or backbone retraining.
发表机构
- Tulane University(杜兰大学)
机构由 AI 辅助整理,请以论文原文为准。