arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

弥合多模态RAG中的语义-效用差距:基于生成器在环对齐

Bridging the Semantic-Utility Gap in Multimodal RAG via Generator-in-the-Loop Alignment

Zhan-Lun Chang, Dong-Jun Han, Seyyedali Hosseinalipour, Mung Chiang, Christopher G. Brinton

arXiv 2609.08188首次发表:更新:

发表机构

Elmore Family School of Electrical and Computer Engineering, Purdue University; Yonsei University; University at Buffalo–SUNY(普渡大学埃尔莫尔电气与计算机工程学院; 延世大学; 纽约州立大学布法罗分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出两阶段生成器在环对齐框架,利用VLM生成假设文本弥合模态差距,并以答案监督偏好对微调重排序器,有效弥合多模态RAG中的语义-效用差距。

AI 中文摘要

视觉-语言模型(VLM)通过检索增强生成(RAG)受益于对外部证据的访问。然而,标准检索器和重排序器优化的是语义相似性而非答案效用,从而产生偏好差距:看似相关的文档可能无法帮助生成器产生正确答案。受此启发,我们提出了一种两阶段生成器在环对齐框架,无需人工文档级相关性标注即可弥合这一差距。我们的框架由两个阶段组成:在第一阶段,VLM根据图像-查询对生成一个假设性文本段落,该段落被用作密集文本检索的检索查询,从而弥合图像到文本的模态差距。在第二阶段,使用低秩适配(LoRA)适配的交叉编码器重排序器,利用从冻结的VLM中挖掘的答案监督偏好对进行微调:给定数据集答案标签,如果VLM在将该候选文档作为上下文时产生正确答案,则该文档被标记为正样本,否则标记为负样本。这种生成器引导的信号与多种对齐损失函数兼容,包括对比(三元组)损失、成对直接偏好优化(DPO)和监督微调(SFT),并支持周期性重新挖掘,以随着重排序器的改进刷新偏好对。在VQA-X和A-OKVQA上使用Qwen3.5-2B和Qwen3-VL-4B-Instruct进行的实验表明,我们提出的框架在各种对齐损失和池大小设置下始终优于排序、随机和REPLUG风格似然基线,这表明答案级生成器反馈是偏好对齐的有效监督信号。

英文摘要

Vision-language models (VLMs) augmented with retrieval-augmented generation (RAG) benefit from access to external evidence. However, standard retrievers and rerankers optimize for semantic similarity rather than answer utility, creating a preference gap: documents that appear relevant may not help the generator produce a correct answer. Motivated by this, we propose a two-stage generator-in-the-loop alignment framework that closes this gap without human document-level relevance annotations. Our framework consists of two stages: in Stage 1, a VLM generates a hypothetical text passage from the image-query pair, which is used as the retrieval query for dense text search, bridging the image-to-text modality gap. In Stage 2, a cross-encoder reranker adapted with low-rank adaptation (LoRA) is fine-tuned using answer-supervised preference pairs mined from the frozen VLM: given the dataset answer label, a candidate document is labeled positive if the VLM produces the correct answer when given that document as context, and negative otherwise. This generator-guided signal is compatible with multiple alignment loss functions, including contrastive (triplet) loss, pairwise direct preference optimization (DPO), and supervised fine-tuning (SFT), and supports periodic re-mining to refresh preference pairs as the reranker improves. Experiments on VQA-X and A-OKVQA with Qwen3.5-2B and Qwen3-VL-4B-Instruct show that our proposed framework consistently outperforms rank-order, random, and REPLUG-style likelihood baselines under various alignment losses and pool size settings, suggesting that answer-level generator feedback is an effective supervision signal for preference alignment.

CommentsSubmitted to IEEE Transactions on Artificial Intelligence

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑