发表机构
Bosch Center for Artificial Intelligence; LMU Munich; University of Oslo; Munich Center for Machine Learning(博世人工智能中心; 慕尼黑大学; 奥斯陆大学; 慕尼黑机器学习中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出将知识图谱编译为 LoRA 适配器库的参数化知识存储方案,发现其存在存储的知识无法通过语义相似性检索的差距,为参数化知识图谱记忆的应用提供了思路并指出核心开放问题。
AI 中文摘要
图检索增强生成在查询时将检索到的子图放入模型的上下文窗口,产生持续的 token 成本,且每次调用都会暴露源数据。我们研究了一种替代方案:离线将知识图谱编译为 LoRA 适配器库,每个实体对应一个适配器,该适配器作为参数化知识层,通过注入权重而非文本进行查询,实现零查询时上下文成本。在 MetaQA 数据集上,我们发现经子图训练的适配器编码了与上下文无关的事实知识,可泛化到未见过的问题:在单值关系上,适配器的精确匹配得分比几乎无闭卷能力的基础模型(0.007)高出 +0.243,且仅正确的适配器能恢复该知识(与基础模型的神谕差距为 +0.283)。然而,存储的知识无法通过相似性恢复:给定无分子图的查询,基于嵌入和权重空间几何的检索均表现为随机水平,因为语义相邻实体的适配器不包含答案——知识是局部存储的,无法迁移。权重几何与子图语义相关(ρ = +0.329),但与功能检索能力无关。我们量化了其与图检索增强生成相比的字节和上下文 token 成本,并讨论了部署意义。我们的结果表明,参数化知识图谱记忆可用于存储知识,并确定除语义相似性外,通过某种机制选择和组合正确适配器是核心开放问题,这推动了基于查询的学习组合机制的研究。
英文摘要
Graph retrieval-augmented generation places retrieved subgraphs into the model's context window at query time, paying a recurring token cost and exposing source data on every call. We study an alternative: compiling a knowledge graph offline into a bank of LoRA adapters, one per entity, that serve as a parametric knowledge layer queried by injecting weights rather than text, at zero query-time context cost. On the MetaQA dataset, we find that subgraph-trained adapters encode context-free factual knowledge that generalizes to unseen questions: on single-valued relations the adapter gains $+0.243$ exact-match score over a base model that is nearly blind closed-book ($0.007$), and only the correct adapter recovers this knowledge (an oracle gap of $+0.283$ over the base model). However, the stored knowledge is not recoverable by similarity: given a query with no subgraph, embedding-based and weight-space geometry retrieval both perform at chance, because a semantically neighbouring entity's adapter does not contain the answer - knowledge is stored locally and does not transfer. Weight geometry correlates with subgraph semantics ($ρ= +0.329$) but not with functional retrievability. We quantify the byte and context-token costs against graph retrieval-augmented generation and discuss deployment implications. Our results establish that parametric knowledge graph memory is feasible for storing knowledge, and identify selecting and composing the right adapters by a mechanism other than semantic similarity as the central open problem - motivating a learned, query-conditioned composition mechanism.
Comments12 pages, 2 figures, 7 tables, accepted at SKGi 2026; v2: editorial corrections only, v3: bibliography fix