arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CAER:面向多模态大语言模型的冲突感知证据路由,采用双前缀专家机制

CAER: Conflict-Aware Evidence Routing with Dual Prefix Experts for Multimodal Large Language Models

Zixuan Liu, Juntao Cai, Xiaoxu Cai, Haishuai Wang, Jiajun Bu

arXiv 2607.28991首次发表:更新:

AI 中文总结

本研究提出与主干无关的CAER框架,通过双前缀专家路由机制实现视觉-语言冲突检测与冲突感知生成,在MMMC及AgriConflict数据集上验证可提升开源多模态大语言模型的可靠性。

AI 中文摘要

多模态大语言模型(MLLMs)在多模态理解与生成方面展现出卓越能力,但当文本输入与视觉证据冲突时,仍会出现幻觉,生成与视觉内容不一致的响应。现有方法主要依赖解码策略、额外训练、验证方法或提示技术,却常缺乏细粒度冲突定位与冲突感知生成。本研究提出CAER,一种与主干无关的视觉-语言冲突检测及冲突感知生成框架。CAER引入基于跨度的证据路由器,将声明表征转换为软文本查询,从冻结的视觉令牌中检索对应证据,实现细粒度冲突估计。此外,设计双前缀专家路由机制,为视觉支持与冲突输入学习独立专家,通过显式专家选择实现冲突感知生成。在公开MMMC基准及新整理的AgriConflict数据集上的实验表明,CAER可有效检测视觉-语言冲突,且无需更新主干参数即可提升开源MLLMs的可靠性。

英文摘要

Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in multimodal understanding and generation. However, when textual inputs conflict with visual evidence, they still suffer from hallucinations and produce responses inconsistent with visual content. Existing approaches mainly rely on decoding strategies, additional training, verification methods, or prompting techniques, but often lack fine-grained conflict localization and conflict-aware generation. In this work, we propose CAER, a backbone-agnostic framework for visual-language conflict detection and conflict-aware generation. CAER introduces a span-grounded evidence router that transforms claim representations into soft textual queries and retrieves corresponding evidence from frozen visual tokens, enabling fine-grained conflict estimation. Furthermore, we design a dual-prefix expert routing mechanism that learns separate experts for visually supported and contradicted inputs, enabling conflict-aware generation through explicit expert selection. Experiments on the public MMMC benchmark and our newly curated AgriConflict dataset demonstrate that CAER effectively detects visual-language conflicts and improves the reliability of open-source MLLMs without updating their backbone parameters.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑