AI 中文总结
该研究针对多跳检索中三元组排序与监督缺失问题,提出KAMR知识对齐多跳检索器,通过构建部分对齐数据集优化对比目标,在多基准与LLM主干上提升了多跳检索和下游问答性能。
AI 中文摘要
基于图的检索增强生成越来越依赖多跳检索,回答查询需要组合多个相连的知识图谱三元组。然而,现有检索器常通过全局语义匹配独立对三元组排序,且许多多跳基准仅提供最终答案,限制了查询与三元组对齐的监督,导致结构上必要但对齐较弱的事实被遗漏。为解决这些问题,我们提出知识对齐多跳检索器KAMR,它区分受查询强约束的锚三元组与与锚弱对齐但结构相连的三元组。为缓解查询-三元组对齐监督的缺失,我们通过掩码三元组元素并提示大语言模型(LLM)生成对应查询构建部分对齐数据集,优化对级和元素级匹配的两个对比目标。推理时,KAMR全局检索锚,再局部扩展收集相连证据。在四个基准、三个LLM主干和十四个基线中,KAMR持续提升多跳检索及下游问答性能。
英文摘要
Graph-based retrieval-augmented generation increasingly relies on multi-hop retrieval, where answering a query requires composing multiple connected knowledge-graph triplets. However, existing retrievers often rank triplets independently via global semantic matching. Moreover, many multi-hop benchmarks provide only final answers, which limits supervision for query--triplet alignment and causes structurally necessary but weakly aligned facts to be missed. To address these issues, we propose a knowledge-aligned multi-hop retriever, KAMR, which distinguishes anchor triplets that are strongly constrained by the query from connected triplets that are weakly aligned yet structurally linked to the anchors. To mitigate the lack of query-triplet alignment supervision, we build a partial alignment dataset by masking triplet elements and prompting an LLM to generate corresponding queries, and optimize two contrastive objectives for pair-level and element-level matching. At inference time, KAMR retrieves anchors globally and then expands locally to collect connected evidence. Across four benchmarks, three LLM backbones, and fourteen baselines, KAMR consistently improves multi-hop retrieval and downstream question answering performance.
CommentsAccepted by COLM'26