发表机构
Rutgers University; Google(罗格斯大学; 谷歌)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CLIMB提出无需训练的推理时框架,通过互补证据池和置信度控制的细化,提升多模态RAG在知识密集型视觉问答中的性能。
AI 中文摘要
多模态大语言模型(MLLM)已展现出强大的视觉推理能力,但知识密集型视觉问答通常需要图像和模型参数化知识之外的外部文本证据。现有的多模态RAG系统通常依赖Top-$K$检索或重排序,这可能返回冗余段落,并且对于答案更新是否充分得到检索证据支持的控制有限。我们提出\textit{CLIMB},一种无需训练、在推理时运行的多模态RAG框架。CLIMB首先使用一种MMR风格的目标构建一个紧凑的互补证据池,该目标平衡查询相关性和段落级冗余。然后,它在此固定池内执行置信度控制的细化:一个R/E/C批评者根据相关性、证据特异性和跨模态对齐对段落进行评分,而一个基于证据的置信度估计器仅在估计置信度增加时接受更新后的答案。这种设计提供了一个简单的停止标准,并减少了不必要的细化,而无需修改底层检索器或MLLM。在Encyclopedic-VQA和InfoSeek上的实验表明,CLIMB持续优于检索增强的多模态基线。消融研究进一步表明,互补池化、基于批评者的评分和迭代的置信度控制细化各自对最终性能有所贡献。
英文摘要
Multimodal large language models (MLLMs) have shown strong visual reasoning abilities, but knowledge-intensive visual question answering often requires external textual evidence beyond the image and the model's parametric knowledge. Existing multimodal RAG systems commonly rely on Top-$K$ retrieval or reranking, which may return redundant passages and provide limited control over whether an answer update is sufficiently supported by the retrieved evidence. We propose \textit{CLIMB}, a training-free inference-time framework for multimodal RAG. CLIMB first constructs a compact complementary evidence pool using an MMR-style objective that balances query relevance and passage-level redundancy. It then performs confidence-controlled refinement within this fixed pool: an R/E/C critic scores passages by relevance, evidence specificity, and cross-modal alignment, while an evidence-grounded confidence estimator accepts an updated answer only when the estimated confidence increases. This design provides a simple stopping criterion and reduces unnecessary refinement without modifying the underlying retriever or MLLM. Experiments on Encyclopedic-VQA and InfoSeek show that CLIMB consistently improves over retrieval-augmented multimodal baselines. Ablations further indicate that complementary pooling, critic-based scoring, and iterative confidence-controlled refinement each contribute to the final performance.
CommentsEMNLP 2026 Findings