VERDICT:基于分歧感知共识的多模态推理无训练逐步验证方法
VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus
浏览论文内容
中文总结 AI 辅助
该研究针对多模态大语言模型推理链易出错的问题,提出无训练的VERDICT方法,利用跨模态分歧的闭式解计算共识评分,在六个基准上提升最高达5.95%,性能接近需大量监督的特定领域评论家。
中文摘要 AI 辅助
多模态大语言模型常生成包含细微错误的推理链,进而导致错误答案。现有验证方法存在明显局限:要么需要昂贵的带标签监督且跨任务表现不一致,要么通过简单聚合方式汇总多来源分数,却忽略了一个关键洞见——当这些分数存在分歧时,该分歧本身携带着推理步骤是否真正有效的重要信息。我们将此形式化为不同冻结验证器之间的耦合评分问题,可解释为具有唯一闭式解均衡的协调博弈,其中一致信号对应有效步骤,分歧则揭示不稳定性。为此,我们提出一种无训练、领域无关的逐步验证方法,命名为VERDICT:即通过分歧信息耦合阈值进行验证。据我们所知,VERDICT是首个明确且可操作地构建跨模态分歧结构的无训练验证器,它通过闭式解计算共识分数,实现了对推理步骤的分歧感知过滤和稳定性感知排序。在六个基准测试中,\n方法始终优于基础模型,提升幅度最高达5.95%,且与需要大量监督的特定领域评论家表现相当,证明跨模态一致性可提供鲁棒的验证信号,无需特定领域适配且属于无训练验证。
英文摘要
Multimodal large language models often generate reasoning chains containing subtle errors that lead to incorrect answers. Current verification approaches have notable limitations. Existing approaches either require expensive labelled supervision with inconsistent cross-task performance or aggregate scores from multiple sources by simple aggregations, missing a key insight: when these scores disagree, that disagreement itself carries important information about whether a reasoning step is truly valid or not. We formalise this as a coupled scoring problem among disparate, frozen verifiers, interpretable as a coordination game with a unique closed-form equilibrium where agreement signals valid steps while disagreement reveals instability. Towards this end, we propose a training-free domain-agnostic step-wise verification approach we call VERDICT: VERification via Disagreement-Informed Coupled Thresholding. To our knowledge, VERDICT is the first training-free verifier that makes the structure of cross-modal disagreement explicit and actionable. It computes consensus scores through a closed-form solution, enabling both disagreement-aware filtering and stability-conscious ranking of reasoning steps. Evaluated across six benchmarks, \method consistently improves over the base model by up to +5.95%, and performs competitively with domain-specific critics that demand extensive supervision, demonstrating that cross-modal agreement provides robust verification signals without task-specific adaptation and Training-Free Verification
发表机构
- Indian Institute of Technology, Hyderabad(印度理工学院海得拉巴分校)
- Microsoft Research(微软研究院)
机构由 AI 辅助整理,请以论文原文为准。