发表机构
University of Leicester; Nanjing University of Posts and Telecommunications; University of North Carolina at Chapel Hill; Amazon; University of Exeter; University of Illinois Urbana-Champaign; Meta AI; Eindhoven University of Technology(莱斯特大学; 南京邮电大学; 北卡罗来纳大学教堂山分校; 亚马逊; 埃克塞特大学; 伊利诺伊大学厄巴纳-香槟分校; Meta AI; 埃因霍温理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视觉-语言模型自验证依赖单一标准导致可靠性不足的问题,提出MOTIVE多视角自验证框架,通过可靠性分数引导接受或重新思考决策,提升推理可靠性与效率。
AI 中文摘要
视觉-语言模型(VLMs)在多模态推理中表现出色,但仍容易生成看似合理实则错误的答案。自验证提供了一种无需外部评判者即可提高答案可靠性的实用方法,但现有方法通常依赖于单一的验证标准或固定提示,导致可靠性估计不完整且不稳定。我们首先系统分析了验证器能力和提示设计对验证性能的影响。研究结果表明,更强的验证器能提供更可靠的判断,而验证性能对提示选择高度敏感,没有单一提示能在所有任务中持续占优。基于这些发现,我们提出了MOTIVE,一个用于可靠多模态推理的多视角自验证与可靠性引导选择性重新思考框架。MOTIVE从互补的验证视角评估每个候选答案,并通过基于正确性的多视角验证学习来学习一个与正确性对齐的可靠性分数。在推理过程中,该分数控制接受或重新思考的决策,使得可靠答案能直接返回,而不确定的答案则触发历史引导的重新思考。跨多个多模态基准和VLM骨干网络的广泛实验表明,MOTIVE持续优于强大的自验证和自校正基线。进一步的结果显示,可靠的验证能改善接受或重新思考的决策并减少不必要的推理轮次,从而在没有外部评判者的情况下实现更可靠和高效的自验证。
英文摘要
Vision-language models (VLMs) have achieved strong performance in multimodal reasoning, yet they remain prone to generating plausible but incorrect answers. Self-verification offers a practical way to improve answer reliability without relying on external judges, but existing methods typically depend on a single verification criterion or fixed prompt, resulting in incomplete and unstable reliability estimates. We first systematically analyze how verifier capability and prompt design affect verification performance. Our findings show that stronger verifiers provide more reliable judgments, while verification performance is highly sensitive to prompt choice, with no single prompt consistently dominating across tasks. Guided by these findings, we propose \texttt{MOTIVE}, a \textbf{M}ulti-View Self-Verificati\textbf{O}n wi\textbf{T}h Rel\textbf{I}ability-Guided Selecti\textbf{VE} Rethinking framework for reliable multimodal reasoning. \texttt{MOTIVE} evaluates each candidate answer from complementary verification perspectives and learns a correctness-aligned reliability score through correctness-grounded multi-view verification learning. During inference, this score governs an accept-or-rethink decision, allowing reliable answers to be returned directly while uncertain ones trigger history-guided rethinking. Extensive experiments across diverse multimodal benchmarks and VLM backbones demonstrate that \texttt{MOTIVE} consistently outperforms strong self-verification and self-correction baselines. Further results show that reliable verification improves accept-or-rethink decisions and reduces unnecessary reasoning turns, enabling more reliable and efficient self-verification without an external judge.