AI 中文总结
研究针对异构PDF集合检索增强生成的挑战,提出多模态CoLRAG-TF四轴融合架构,集成多种技术。构建多模态索引,经贝叶斯优化确定融合权重。实验显示其检索召回率高,多跳答案相似性提升显著,还适用于视觉输入,提供通用框架。
AI 中文摘要
由于多模态内容、特定领域术语以及跨分散证据进行多跳推理的需求,在异构PDF集合上进行检索增强生成(RAG)仍然具有挑战性。我们提出了多模态CoLRAG-TF,一种四轴融合架构,集成了密集文本嵌入、BM25关键词匹配、知识图谱三重过滤和基于图像的相似性,用于在复杂文档上进行稳健检索。我们的系统构建了一个从43份日本灾难教训PDF中提取的2403个块的多模态索引,由混合OCR管道和基于LLM的字幕生成支持。为了增强组合推理,我们提取了11414个OpenIE三元组并用FAISS索引,实现亚秒级三元组查找和相关信号的分层传播。一个受HippoRAG2启发的从粗到细的检索器(卷→章→块)在最终融合评分之前缩小搜索空间。对融合权重的贝叶斯优化表明,三重轴必须占主导(α_triple = 0.44)以抵消词汇偏差并维持多跳检索质量。在457对基准上评估,多模态CoLRAG-TF实现了0.9909的检索召回率,并且在多跳答案相似性方面比单跳查询提高了71.6%。使用视觉LLM的图像到教训管道进一步证明了该方法对视觉输入的适用性。这些结果表明,三重过滤的多模态融合对于在嘈杂、异构的PDF上进行结构化推理至关重要,并提供了一个适用于灾难领域之外的通用框架。
英文摘要
Retrieval-augmented generation (RAG) over heterogeneous PDF collections remains challenging due to multimodal content, domain-specific terminology, and the need for multi-hop reasoning across dispersed evidence. We present Multimodal CoLRAG-TF, a four-axis fusion architecture that integrates dense text embeddings, BM25 keyword matching, knowledge-graph triple filtering, and image-based similarity for robust retrieval over complex documents. Our system constructs a multimodal index of 2,403 blocks extracted from 43 Japanese disaster lesson PDFs, supported by a hybrid OCR pipeline and LLM-based caption generation. To enhance compositional reasoning, we extract 11,414 OpenIE triples and index them with FAISS, enabling sub-second triple lookup and hierarchical propagation of relevance signals. A HippoRAG2-inspired coarse-to-fine retriever (volume $\to$ chapter $\to$ block) narrows the search space before final fusion scoring. Bayesian optimization over fusion weights reveals that the triple axis must dominate ($α_\text{triple} = 0.44$) to counteract lexical bias and sustain multi-hop retrieval quality. Evaluated on a 457-pair benchmark, Multimodal CoLRAG-TF achieves a Retrieval Recall of 0.9909 and a 71.6$\%$ improvement in multi-hop answer similarity over single-hop queries. An image-to-lesson pipeline using a vision LLM further demonstrates the applicability of the approach to visual inputs. These results show that triple-filtered multimodal fusion is essential for structured reasoning over noisy, heterogeneous PDFs and provides a general framework applicable beyond the disaster domain.
Comments18 pages, 8 tables, 9 figures