arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

知识图谱上的大语言模型问答中的答案路径与接地指令

The Answer Path and the Grounding Instruction in LLM Question Answering over Knowledge Graphs

Arquimedes Canedo

arXiv 2609.10237首次发表:更新:

发表机构

Siemens Digital Industries Software(西门子数字工业软件)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过系统实验发现,在知识图谱问答中,答案路径的存在和接地指令显著影响大语言模型性能,而语法、顺序和子图大小无影响,并纠正了关于图上下文有害的虚假发现。

AI 中文摘要

图检索增强生成流水线会选择将哪些三元组放入提示中、以何种语法书写它们、以何种顺序书写它们,以及一句告诉模型如何处理它们的指令。我们在大语言模型和两个知识图谱问答基准上变化了这四个选择。四个选择中有两个会影响答案,另外两个则没有影响。第一个选择是答案路径(即到达答案所需的三元组)是否出现在提示中。在保持三元组数量固定,并将不在链上的每个三元组替换为来自无关实体的材料时,答案准确率变化了+0.003 F1,而移除该链则损失了图的大部分价值。检索预算应放在召回率上,而我们能测试范围内的精确率则毫无收益。这里没有检索器:子图来自金标准SPARQL,因此精确率描述的是我们构建的上下文,而非系统设置。第二个选择是接地指令。当提示中没有事实时,告诉模型仅使用提供的事实来回答会使F1从0.299降至0.035,下降了8.63倍。这个数字描述的是带有空上下文臂的评估,而非工作流水线,并且一项将指令应用于其上下文臂但未应用于其无上下文基线的实验,制造了一个虚假的发现,即图上下文在深度上是有害的。我们在自己的结果中发现了这样一个发现,并予以撤回。语法、三元组顺序和子图大小在多跳深度上未产生我们可测量的影响。将接地指令与正确上下文进行定价的比较,无法用对格式敏感的评分器来衡量,因为该指令决定了响应格式;我们将其报告为开放对比,而非一个数字。

英文摘要

A graph retrieval-augmented generation pipeline chooses which triples to put in the prompt, a syntax to write them in, an order to write them in, and a sentence telling the model what to do with them. We vary all four over six large language models and two knowledge-graph question answering benchmarks. Two of the four choices move the answer and the other two are flat. The first is whether the answer path, the triples needed to reach the answer, is in the prompt at all. Holding the number of triples fixed and replacing every triple that is not on the chain with material from an unrelated entity changes answer accuracy by +0.003 F1, while removing the chain costs most of what the graph was worth. Retrieval budget belongs on recall, and precision in the range we can test buys nothing. There is no retriever here: subgraphs come from gold SPARQL, so precision describes the context we build, not a system setting. The second is the grounding instruction. With no facts in the prompt, telling a model to answer using only the provided facts drops F1 from 0.299 to 0.035, a factor of 8.63. That figure describes an evaluation with an empty context arm rather than a working pipeline, and an experiment that applies the instruction to its context arm but not to its no-context baseline manufactures a spurious finding that graph context hurts at depth. We found one in our own results and retract it. Syntax, triple order and subgraph size produce no effect we can measure at multi-hop depth. The comparison that would price the grounding instruction against correct context is not measurable with a format-sensitive scorer, because the instruction determines the response format; we report it as an open contrast rather than a number.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑