复杂城市场景中基于证据的可信多模态推理与评估基准
Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes
- University of Chinese Academy of Sciences (UCAS)(中国科学院大学)
- Tencent CDG(腾讯云与智慧产业事业群)
- Institute of Information Engineering, CAS(中国科学院信息工程研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对复杂城市场景中多模态大语言模型的推理可靠性问题,提出AD2-Bench基准与EGVOR模型,提升了不利条件下的多模态推理稳定性。
AI中文摘要:
尽管多模态大语言模型(MLLMs)在良性场景中展现出令人印象深刻的性能,但在不利条件下的复杂场景中,其认知可靠性会显著下降。在这些场景中,模型常常依赖缺乏充足视觉证据的隐式推理,导致感知与推理脱节。与此同时,现有的面向结果的基准仅评估最终预测,无法诊断底层推理过程中的失败。为解决这一缺口,作者提出AD2-Bench,引入分层视觉诊断框架,将推理分解为结构化的证据链(CoE)。这种细粒度诊断揭示,稳健的多模态推理从根本上依赖于准确的证据获取。基于这一视角,作者从概率角度构建推理过程,识别出两类主要的推理失败原因:空间歧义,即模型无法区分目标对象与背景杂波,导致定位错误;语义不确定性,即退化的视觉特征引发错误的语义解读,导致理解错误。为克服这些证据缺陷,作者进一步提出基于证据的视觉推理(EGVOR),用显式生成证据原子(Evidence Atoms)替代隐式推理,证据原子是结构化的空间-语义三元组,可强化定位与语义理解间的紧密对齐。该模型通过分层课程进行训练,课程从反思性监督构建逐步推进至强化学习,其中明确奖励降低推理方差的行为。大量实验表明,EGVOR在不利条件下大幅提升了推理稳定性,为可信多模态认知提供了更稳健的框架。
英文摘要:
While Multimodal Large Language Models (MLLMs) demonstrate impressive performance in benign scenarios, their cognitive reliability deteriorates significantly in complex scenes under adverse conditions. In these settings, models often rely on implicit inference without sufficient visual evidence, leading to a disconnect between perception and reasoning. Meanwhile, existing outcome-oriented benchmarks evaluate only final predictions and fail to diagnose failures in the underlying reasoning process. To address this gap, the authors propose AD2-Bench, which introduces a Hierarchical Visual Diagnosis framework that decomposes reasoning into a structured Chain of Evidence (CoE). This fine-grained diagnosis reveals that robust multimodal reasoning fundamentally depends on accurate evidence acquisition. Building on this perspective, the authors formulate reasoning from a probabilistic viewpoint and identify two primary causes of reasoning failure: Spatial Ambiguity, where models fail to distinguish target objects from background clutter, resulting in localization errors; and Semantic Uncertainty, where degraded visual features lead to incorrect semantic interpretation, resulting in understanding errors. To overcome these evidence deficiencies, they further propose Evidence-grounded Visual Reasoning (EGVOR), which replaces implicit reasoning with the explicit generation of Evidence Atoms - structured spatial-semantic triplets that enforce tight alignment between localization and semantic understanding. The model is trained through a hierarchical curriculum that progresses from reflective supervision construction to reinforcement learning, where reducing reasoning variance is explicitly rewarded. Extensive experiments demonstrate that EGVOR substantially improves reasoning stability under adverse conditions, providing a more robust framework for trustworthy multimodal cognition.