GraphLoom:面向多模态KG-RAG的可靠性校准图证据路由
GraphLoom: Reliability-Calibrated Graph Evidence Routing for Multimodal KG-RAG
- School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院)
- Aristotle University(亚里士多德大学)
- Dashub(达舒布)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出GraphLoom框架,通过可靠性校准的图证据路由实现多模态KG-RAG,在ScienceQA等数据集上提升了答案质量与证据忠实度。
AI中文摘要:
多模态检索增强生成(RAG)系统常依赖长非结构化上下文或过度扩展的证据图,这会引入噪声证据、削弱多跳推理能力并增加无依据生成。本文提出GraphLoom,一种用于紧凑且忠实证据路由的可靠性校准多模态知识图谱RAG框架。给定问题及其关联的多模态输入,GraphLoom从接地场景描述、提取的关系三元组及外部常识知识构建实例级多模态知识图谱。GraphLoom未将所有检索到的证据注入生成器,而是通过有界扩展的可靠性感知子图检索,经分层图内存槽与冻结语言模型中的联合图-序列注意力选择性路由高效用证据。为提升复杂推理场景下的鲁棒性,GraphLoom进一步将交错检索与预算校正检索结合,在噪声检索条件下实现自适应多跳证据优化。在ScienceQA、MultiModalQA与OK-VQA数据集上开展评估,含近似噪声外部知识检索的大型干扰证据池。实验结果显示,与强大的多模态RAG、图检索及开源视觉-语言基线相比,GraphLoom在答案质量与证据忠实度上实现持续提升,在MultiModalQA上的检索质量有所改善,且在噪声证据池下性能稳定。基于MiniCheck验证、人工评估与延迟分析的额外分析表明,可靠性校准图证据路由是长上下文多模态证据注入的有效替代方案。
英文摘要:
Multimodal retrieval-augmented generation (RAG) systems often rely on long unstructured contexts or aggressively expanded evidence graphs, which can introduce noisy evidence, weaken multi-hop reasoning, and increase unsupported generation. We present GraphLoom, a reliability-calibrated multimodal knowledge-graph RAG framework for compact and faithful evidence routing. Given a question and its associated multimodal input, GraphLoom constructs an instance-level multimodal knowledge graph from grounded scene descriptions, extracted relational triples, and external commonsense knowledge. Instead of injecting all retrieved evidence into the generator, GraphLoom performs reliability-aware subgraph retrieval with bounded expansion and selectively routes high-utility evidence through hierarchical graph memory slots and joint graph-sequence attention in a frozen language model. To improve robustness in complex reasoning settings, GraphLoom further combines interleaved retrieval with budgeted corrective retrieval, enabling adaptive multi-hop evidence refinement under noisy retrieval conditions. We evaluate GraphLoom on ScienceQA, MultiModalQA, and OK-VQA, including large distractor evidence pools that approximate noisy external knowledge retrieval. Experimental results show consistent gains in answer quality and evidence faithfulness over strong multimodal RAG, graph-retrieval, and open-source vision-language baselines, with improved retrieval quality on MultiModalQA and stable performance under noisy evidence pools. Additional analyses using MiniCheck-based verification, human evaluation, and latency profiling show that reliability-calibrated graph evidence routing provides an effective alternative to long-context multimodal evidence injection.