原子视觉蕴含:通过原子事实分解与学习选择增强零样本视觉-语言推理
Atomic Visual Entailment: Enhancing Zero-Shot Vision-Language Reasoning through Atomic Fact Decomposition and Learned Selection
浏览论文内容
中文总结 AI 辅助
提出原子视觉蕴含(AVE),通过将假设分解为原子事实并学习选择可信预测,在零样本下无需微调视觉-语言模型即在SNLI-VE上达到0.803准确率,缩小了与微调系统的差距。
中文摘要 AI 辅助
视觉蕴含(VE)任务询问图像是支持、反驳还是无法判定一个文本假设。在标注数据上微调大型视觉-语言模型可获得强结果,而零样本和混合方法仍远远落后。一个VE假设往往捆绑多个视觉声明,但现有零样本方法将其作为一个整体进行推理。我们提出原子视觉蕴含(AVE),该方法将假设分解为原子事实,使用冻结的视觉-语言模型从完整假设及其事实中产生候选预测,并通过一个仅基于这些候选行为训练的轻量级分类器来预测最终标签。我们发现,仅当保留假设上下文时,分解才有帮助:孤立地判断事实比完全不分解更差。完整假设和原子预测产生互补性错误,学习信任哪个比多数投票更能恢复这种互补性,在SNLI-VE上无需微调任何视觉-语言模型即可达到0.803的测试准确率。AVE还能在没有区域级监督的情况下定位其预测背后的视觉证据。这些结果表明,学习信任哪个候选预测可以大大缩小与微调系统的差距,为直接微调视觉-语言模型需要更多标注数据或计算资源的情况提供了实用替代方案。
英文摘要
Visual entailment (VE) asks whether an image supports, contradicts, or leaves undecided a textual hypothesis. Strong results come from fine-tuning large vision-language models on labelled data, while zero-shot and hybrid approaches remain far behind. A VE hypothesis often bundles several visual claims, yet existing zero-shot methods reason over it as a single unit. We propose Atomic Visual Entailment (AVE), which decomposes the hypothesis into atomic facts, produces candidate predictions from both the full hypothesis and its facts using frozen vision-language models, and predicts the final label with a lightweight classifier trained only on how those candidates behave. We find that decomposition helps only when the hypothesis context is preserved: judging facts in isolation is worse than not decomposing at all. Full-hypothesis and atomic prediction make complementary errors, and learning which to trust recovers far more of that complementarity than majority voting, reaching 0.803 test accuracy on SNLI-VE without fine-tuning any vision-language model. AVE also localises the visual evidence behind its prediction without region-level supervision. These results suggest that learning which candidate prediction to trust can close much of the gap to fine-tuned systems, offering a practical alternative where fine-tuning a vision-language model directly would need more labelled data or compute than is available.
发表机构
- Leiden University(莱顿大学)
机构由 AI 辅助整理,请以论文原文为准。