MMRAG-RFT: 两阶段强化微调用于可解释的多模态检索增强生成
MMRAG-RFT: Two-stage Reinforcement Fine-tuning for Explainable Multi-modal Retrieval-augmented Generation
AI总结:
MMRAG-RFT通过两阶段强化微调提升多模态检索增强生成的可解释性,实现更清晰的推理逻辑和更优的生成效果。
AI中文摘要:
多模态检索增强生成(MMRAG)通过整合外部多模态知识,实现了高可信度的生成,从而在复杂的多模态场景中表现出色。然而,现有MMRAG方法未能明确检索和响应生成背后的推理逻辑,这限制了结果的可解释性。为解决这一差距,我们提出将强化学习引入多模态检索增强生成,通过两阶段强化微调框架增强多模态大语言模型的推理能力,以实现可解释的多模态检索增强生成。具体而言,在第一阶段,基于规则的强化微调用于对多模态文档进行粗粒度的逐点排序,有效过滤出显著不相关的文档。在第二阶段,基于推理的强化微调用于联合优化细粒度的列表排序和答案生成,引导多模态大语言模型在MMRAG过程中输出可解释的推理逻辑。我们的方法在WebQA和MultimodalQA两个多模态检索增强生成基准数据集上实现了最先进的结果,并通过全面的消融实验验证了其有效性。
英文摘要:
Multi-modal Retrieval-Augmented Generation (MMRAG) enables highly credible generation by integrating external multi-modal knowledge, thus demonstrating impressive performance in complex multi-modal scenarios. However, existing MMRAG methods fail to clarify the reasoning logic behind retrieval and response generation, which limits the explainability of the results. To address this gap, we propose to introduce reinforcement learning into multi-modal retrieval-augmented generation, enhancing the reasoning capabilities of multi-modal large language models through a two-stage reinforcement fine-tuning framework to achieve explainable multi-modal retrieval-augmented generation. Specifically, in the first stage, rule-based reinforcement fine-tuning is employed to perform coarse-grained point-wise ranking of multi-modal documents, effectively filtering out those that are significantly irrelevant. In the second stage, reasoning-based reinforcement fine-tuning is utilized to jointly optimize fine-grained list-wise ranking and answer generation, guiding multi-modal large language models to output explainable reasoning logic in the MMRAG process. Our method achieves state-of-the-art results on WebQA and MultimodalQA, two benchmark datasets for multi-modal retrieval-augmented generation, and its effectiveness is validated through comprehensive ablation experiments.