arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ColGraphRAG:用于多模态GraphRAG的后期交互证据检索

ColGraphRAG: Late-Interaction Evidence Retrieval for Multimodal GraphRAG

Seonok Kim

arXiv 2607.16208首次发表:更新:

AI 中文总结

研究多模态问答中基于图的证据检索问题,核心方法是在ColBERT/ColPali系列中用后期交互多向量评分替换视觉候选排序算子,主要贡献是改进了图相连图像候选检索及下游问答效果,揭示视觉证据作用模式,为后续工作指明方向。

AI 中文摘要

基于图的多模态问答在结构化证据图中组织文本、表格和图像,然而端到端的准确性取决于哪些多模态资产被排到足够高的位置以进入下游推理;对于与图相连的图像,单向量双编码器相似度会丢弃细粒度对齐所需的补丁和令牌级结构。我们评估了在ColBERT/ColPali系列中,用后期交互的MaxSim风格多向量评分替换与图相连图像节点上的视觉候选排序算子,同时保持离线图构建、文本和表格侧检索、结构化提取以及下游推理不变。在多模态问答任务中,这一变化与图相连图像候选的检索阶段点估计改进以及下游问答增益相关,在视觉证据最重要的地方有更大提升,在文本主导问题上有混合趋势;我们将这种模式解释为包含与图相连视觉证据的机制级证据,而更广泛的验证和更精细的图级诊断仍是未来重要工作。

英文摘要

Graph-grounded multimodal question answering organizes text, tables, and images in a structured evidence graph, yet end-to-end accuracy depends on which multimodal assets are ranked highly enough to enter downstream reasoning; for graph-linked images, single-vector bi-encoder similarity can discard patch- and token-level structure needed for fine-grained alignment. We evaluate replacing the visual candidate-ranking operator over graph-linked image nodes with late-interaction MaxSim-style multi-vector scoring in the ColBERT/ColPali lineage, while keeping offline graph construction, text- and table-side retrieval, structured extraction, and downstream reasoning unchanged. On MultimodalQA, this change is associated with improved retrieval-stage point estimates for graph-linked image candidates and downstream QA gains, with larger movement where visual evidence matters most and mixed trends on text-dominant questions; we interpret the pattern as mechanism-level evidence for graph-linked visual evidence inclusion, while broader validation and finer graph-level diagnostics remain important future work.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑