arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20843cs.CLcs.LG

VISPATH: 面向多模态知识图谱问答的视觉意图引导路径推理

VISPATH: Visual-Intent-Guided Path Reasoning for Multimodal Knowledge Graph Question Answering

Jinke Wu, Zhengpin Li, Mengzhe Jia, Yang Li, Wentao Zhang

AI总结:

VISPATH提出视觉意图引导的路径推理框架,通过多模态定位、意图引导路径发现和推理链剪枝,解决多模态知识图谱问答中多跳推理问题,在基准上超越强基线。

AI中文摘要:

知识图谱问答(KGQA)使模型能够通过结构化图推理回答自然语言问题,并已在许多基准测试和应用中取得了实质性进展。近年来,多模态知识图谱问答(MM-KGQA)引起了越来越多的关注,因为许多问题需要联合使用多模态输入和知识图谱证据。然而,现有的MM-KGQA方法通常仅将多模态信息用于起始实体定位或证据检索,之后的多跳推理退化为纯文本图搜索。因此,它们无法利用在中间跳变得重要的多模态线索。为了解决这一局限,我们提出了VISPATH,一个面向MM-KGQA的视觉意图引导路径推理框架。VISPATH首先通过结合多模态定位和图结构线索来识别可靠的起始实体。然后,它通过从输入、问题和当前部分路径重新计算每跳特定的多模态意图来执行意图引导的路径发现,使得每次扩展都受当前推理状态引导。发现的路径进一步通过推理链剪枝进行细化,该剪枝根据候选路径与问题、推理草图和跳特定意图的一致性,将其作为完整证据链进行评估。最后,VISPATH检查所选证据是否足以生成答案。我们进一步构建了VISPATH-Bench,一个用于评估知识图谱上多模态多跳推理的基准,涵盖需要在知识图谱路径上进行二至四跳的问题。在VISPATH-Bench和另外三个多模态问答基准上的大量实验表明,VISPATH始终优于强基线。值得注意的是,以GPT-4o为骨干,VISPATH在VISPATH-Bench上超过了GPT-5.4,平均准确率相对提高了10.6%,在2跳推理上提高了13.1%。

英文摘要:

Knowledge graph question answering (KGQA) enables models to answer natural-language questions through structured graph reasoning and has achieved substantial progress across many benchmarks and applications. Recently, multimodal KGQA (MM-KGQA) has attracted increasing attention because many questions require jointly using multimodal inputs and KG evidence. However, existing MM-KGQA methods typically use multimodal information only for starting entity grounding or evidence retrieval, after which multi-hop reasoning degenerates into text-only graph search. As a result, they cannot exploit multimodal cues that become important at intermediate hops. To address this limitation, we propose VISPATH, a visual-intent-guided path reasoning framework for MM-KGQA. VISPATH first identifies a reliable starting entity by combining multimodal grounding with graph-structural cues. It then performs intent-guided path discovery by recomputing hop-specific multimodal intent from the input, question, and current partial paths, so that each expansion is guided by the current reasoning state. The discovered paths are further refined through reasoning-chain pruning, which evaluates candidate paths as complete evidence chains based on their consistency with the question, reasoning sketch, and hop-specific intent. Finally, VISPATH checks whether the selected evidence is sufficient for answer generation. We further construct VISPATH-Bench, a benchmark for evaluating multimodal multi-hop reasoning over KGs, covering questions that require two to four hops over KG paths. Extensive experiments on VISPATH-Bench and three additional multimodal QA benchmarks show that VISPATH consistently outperforms strong baselines. Notably, with GPT-4o as the backbone, VISPATH surpasses GPT-5.4 on VISPATH-Bench, achieving a 10.6% relative improvement in average accuracy and a 13.1% improvement at 2-hop reasoning.

补充信息

↑