融合前的信任:用于受污染多模态RAG的QIMG-7和源感知分辨率
Trust Before Fusion: QIMG-7 and Source-Aware Resolution for Polluted Multimodal RAG
浏览论文内容
中文总结 AI 辅助
研究多模态检索增强生成中受污染内容问题,提出QIMG-7基准。针对朴素多模态融合脆弱问题,提出源感知信任分辨率(SATR),Field-Selector变体效果最佳,结果支持选择性信任,显式文本可靠性建模是收益主要驱动因素。
中文摘要 AI 辅助
多模态检索增强生成(RAG)通常使用干净的证据进行评估,但实际检索可能会返回主题相关但不可靠的内容,如来自损坏元数据、实体交换、排版覆盖、语义编辑、对抗性补丁、混合或风格转移的虚假文本和误导性图像。我们引入了QIMG-7,这是一个用于多句事实问答中多模态检索污染的受控基准,涵盖四个数据集、七个图像攻击家族和16对干净/受污染状态,每种方法有1760个评估行。在四个生成器/门堆栈中,朴素的多模态融合很脆弱。我们提出了源感知信任分辨率(SATR),一种无需训练的方法,通过比较参数化、纯文本和全多模态候选答案,并根据源可靠性在候选答案中进行选择或回退。Field-Selector变体取得了最佳平衡分数0.816,比全多模态提高了11.7分,比级联路由器提高了2.7分。消融实验表明,在这种文本优先的设置中,显式的文本可靠性建模是这些收益的主要驱动因素。总体而言,在具有多模态检索冲突的文本优先事实问答中,我们的结果支持选择性信任而非无条件融合。
英文摘要
Multimodal retrieval-augmented generation (RAG) is often evaluated with clean evidence, yet real retrieval can return topically relevant but unreliable content: false text and misleading images from corrupted metadata, entity swaps, typographic overlays, semantic edits, adversarial patches, blends, or style transfer. We introduce QIMG-7, a controlled benchmark for multimodal retrieval pollution in multi-sentence factual QA, spanning four datasets, seven image-attack families, and 16 paired clean/polluted regimes, for 1,760 evaluation rows per method. Across four generator/gate stacks, naive multimodal fusion is brittle: in the main gpt-4o-mini stack, Full-MM support drops from 0.908 with clean text to 0.490 with polluted text, often making Parametric fallback safer than retrieval. We propose source-aware trust resolution (SATR), a training-free approach that compares Parametric, Text-only, and Full-MM candidate answers and selects among candidate answers or falls back based on source reliability. The Field-Selector variant achieves the best balanced score, 0.816, improving over Full-MM by 11.7 points and over the Cascaded Router by 2.7 points. Ablations show that, in this text-first setting, explicit text-reliability modeling is the dominant driver of these gains. Overall, in text-first factual QA with multimodal retrieval conflict, our results support selective trust rather than unconditional fusion. Artifacts are available at https://github.com/SaadElDine/Trust_Before_Fusion.
发表机构
- Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。