AI 中文总结
针对基于扩散语言模型的视觉检索增强生成,研究发现无条件扩大检索证据会因语义冲突降低准确率,提出无需训练的基于熵的候选过滤框架,可平均提升答案准确率2.62个百分点。
AI 中文摘要
视觉检索增强生成(RAG)通常会扩大检索到的证据集以提升答案页覆盖范围,隐含假设是应将所有可用证据传递给生成器。我们证明该假设不适用于扩散语言模型(DLMs):检索更多页面会提升答案页召回率,但无条件将所有检索到的页面传递给生成器往往会降低答案准确率,主要原因是语义冲突。潜在源分析通过并行去噪中的源一致性损失解释了这种不匹配,其中逐位置提案可将不兼容的视觉源组合成无依据的答案。我们进一步发现,此类干扰在第一步答案块分布中已可见,这使得在解码前评估证据成为可能。为在保留检索覆盖范围的同时限制有害视觉暴露,我们提出了基于熵的候选过滤(ECF),这是一种无需训练的证据接纳框架。为减少单个候选中的无关内容,ECF构建了多粒度证据单元;为识别有益的额外证据,它使用空白控制的块置信度和检索排名来确定是否以及哪个候选应进入最终上下文。在3个多模态DLMs和5个视觉QA基准上,ECF相比最强的固定前k输入平均提升答案准确率2.62个百分点,且在LLaDA2.0-Uni上,相比每个数据集的最佳无训练竞争结果平均提升2.37个百分点。这些结果表明,更广泛的检索通过选择性证据接纳而非无条件证据扩展,有益于视觉DLM-RAG。代码可在this https URL公开获取。
英文摘要
Visual retrieval-augmented generation (RAG) commonly expands the retrieved evidence set to improve answer-page coverage, implicitly assuming that all available evidence should be passed to the generator. We show that this assumption does not hold for diffusion language models (DLMs): retrieving more pages increases answer-page recall, whereas unconditionally passing all retrieved pages to the generator often reduces answer accuracy, primarily because of semantic conflict. A latent-source analysis explains this mismatch through source-coherence loss in parallel denoising, where position-wise proposals can combine incompatible visual sources into unsupported answers. We further find that such interference is already visible in the first-step answer-block distribution, making it possible to assess evidence before decoding. To preserve retrieval coverage while limiting harmful visual exposure, we propose the Entropy-Based Candidate Filter (ECF), a training-free evidence-admission framework. To reduce irrelevant content within individual candidates, ECF constructs multi-granularity evidence units; to identify beneficial additional evidence, it uses blank-controlled block confidence and retrieval rank to determine whether and which candidate should enter the final context. Across three multimodal DLMs and five visual QA benchmarks, ECF improves answer accuracy by 2.62 percentage points on average over the strongest fixed top-$k$ input and, with LLaDA2.0-Uni, by 2.37 percentage points on average over the best competing training-free result for each dataset. These results show that broader retrieval benefits visual DLM-RAG through selective evidence admission rather than unconditional evidence expansion. Code is publicly available at https://github.com/wjkuser/ECF.