arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

保持一致!利用一致性约束增强大视觉语言模型中的鲁棒视觉推理

Be Consistent! Enhancing Robust Visual Reasoning in LVLMs with Consistency Constraints

Liqiang Jing, Xiong Zhou, Siddharth Varia, Neha Anna John, Xinya Du, Vassilis N. Ioannidis

arXiv 2607.21722首次发表:更新:

发表机构

University of Texas at Dallas; Amazon Web Services(德克萨斯大学达拉斯分校; 亚马逊网络服务公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对LVLMs在视觉推理任务中较脆弱的问题,引入ConVBench基准及逻辑一致性等评估指标,提出ConVLM,通过基于GRPO的强化学习及一致性奖励改进LVLM推理,利用自动生成的问答对和双重奖励设计提升模型表现。

AI 中文摘要

虽然大视觉语言模型(LVLMs)展现出强大的感知能力,但在视觉推理任务中仍较为脆弱。现有基准主要关注符号数学或科学问题以及简单的以视觉为中心的任务,对复杂视觉推理和逻辑一致性的评估有限。我们引入了ConVBench,一个以视觉为中心的复杂推理基准,每个图像与六个类别的两个逻辑等价问题配对。我们定义了两个评估指标,逻辑一致性和鲁棒准确性。还提出了ConVLM,通过基于组相对策略优化(GRPO)的强化学习及新颖的一致性奖励来改进LVLM推理,该方法利用自动生成的逻辑等价问答对和双重奖励设计,框架在有无严格答案监督下均有效。

英文摘要

While Large Vision-Language Models (LVLMs) exhibit strong perceptual capabilities, they remain vulnerable in visual reasoning tasks. Existing benchmarks largely focus on symbolic mathematical or scientific problems and simple vision-centric tasks, offering limited assessment of complex visual reasoning and logical consistency, a critical requirement for reliable reasoning systems. We introduce ConVBench, a complex vision-centric reasoning benchmark in which each image is paired with two logically equivalent questions across six categories: action and state, complex counting, spatial reasoning, causal and intent understanding, commonsense reasoning, and temporal perception. To complement this benchmark, we define two evaluation metrics, logical consistency and robust accuracy, that jointly assess both the correctness and consistency of model responses. We further present ConVLM, which improves LVLM reasoning through Group Relative Policy Optimization (GRPO)-based reinforcement learning with a novel consistency reward. This method leverages automatically generated logically equivalent question-answer pairs and a dual-reward design combining accuracy- and consistency-based signals, encouraging agreement between paired responses. The framework functions effectively with or without strict answer supervision.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑