arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24799cs.IRcs.AI

使用开源语言模型进行科学文档理解的多模态混合检索增强生成

Multimodal Hybrid Retrieval-Augmented Generation for Scientific Document Understanding using Open-Source SLMs

  • Faculty of Electronics, Telecommunications and Information Technology(电子、电信与信息技术学院)
  • University Politehnica of Bucharest(布加勒斯特理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Alexandru-Andrei Saucă, Ana-Luiza Rusnac

AI总结:

针对科学文档理解中大型语言模型的问题,提出多模态混合检索增强生成系统。利用开源视觉语言模型生成摘要,结合多种检索方法并优化,采用查询压缩器确保连贯性。经评估,该系统提高了检索质量,验证了开源模型与先进策略结合的竞争力。

AI中文摘要:

大型语言模型在回答科学文档中的特定领域问题时,未经预训练往往会产生幻觉。检索增强生成等方法部分解决了这一问题,但面临有限的上下文知识、稀疏与密集检索的差异以及检索噪声等挑战。本文提出了一种先进的多模态检索增强生成系统,旨在解决这些挑战并提高信息提取的准确性。该系统引入了多模态摄取管道,利用开源视觉语言模型生成表格和图形的文本摘要。检索阶段将基于HNSW的语义搜索与基于GIN的词汇搜索相结合,并通过互反排名融合和交叉编码器重排进行优化,以最小化检索噪声。为确保多轮交互中的对话连贯性,采用了查询压缩器模块。通过使用MMLongBench基准、BeIR格式的合成数据集和DeepEval框架对摄取、检索和生成阶段进行独立评估。结果表明,与朴素检索增强生成基线相比,检索质量提高了157%,延迟仅增加50毫秒,且Qwen2-VL-2B-Instruct在BERTScore方面取得了与基于云的模型相当的结果。这些发现验证了开源优化的语言模型与先进的检索策略相结合,可以在不依赖基于云的模型的情况下为文档理解提供有竞争力的性能。

英文摘要:

Large Language Models tend to hallucinate when answering domain-specific ques tions from scientific documents without prior fine-tuning. Currently, methods such as Retrieval-Augmented Generation partially solve this problem but face different challenges: limited context knowledge, difference between sparse and dense retrieval, and retrieval noise. This paper presents an Advanced Multimodal Retrieval-Augmented Generation system that aims to solve those challenges and im prove the accuracy of information extraction. The proposed architecture introduces a multimodal ingestion pipeline that leverages an open-source Vision-Language Model (Qwen2-VL-2B-Instruct) to generate textual summaries of tables and fig ures. The retrieval phase integrates HNSW-based semantic search with GIN-based lexical search, unified through Reciprocal Rank Fusion and refined using Cross Encoder reranking to minimize retrieval noise. To ensure conversational coherence across multi-turn interactions, a Query Condenser module is employed. Evaluation is conducted by independently assessing the ingestion, retrieval and generation stages using the MMLongBench benchmark, a BeIR-format synthetic dataset and the DeepEval framework. Moreover, results demonstrate a 157% improvement in retrieval quality over a Naive-RAG baseline, with only 50 ms additional la tency, while Qwen2-VL-2B-Instruct achieved results comparable to cloud-based models in BERTScore. These findings validate that open-source optimized SLMs, paired with advanced retrieval strategies, can provide competitive performance for document understanding without relying on cloud-based models.

补充信息

↑