arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从存储到访问:通过显式预激活与隐式推理实现大语言模型中参数化知识的可验证激活

From Storage to Access: Verifiable Activation of Parametric Knowledge in LLMs via Explicit Priming and Implicit Reasoning

Zuocheng Ying, Yang Yang, Yumou Wu, Chuanbo Zhu, Jiarui Wang, Ziqi Wu, Jingming Cai, Junqing Yu, Zikai Song

arXiv 2608.18581首次发表:更新:

发表机构

ByteDance; Huazhong University of Science and Technology(字节跳动; 华中科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出VAKE框架,通过两阶段强化学习将LLM的显式知识提取能力迁移至隐式推理,可激活参数化知识,在多基准上优于基线且具备跨数据集迁移性。

AI 中文摘要

尽管大语言模型(LLMs)在其参数中编码了丰富的事实知识,但可靠地回忆和验证此类知识仍然是事实问答中的一个关键瓶颈。现有的端到端方法将知识提取与推理纠缠在一起,使得难以确定正确答案是源于参数化知识还是输入上下文。为解决这一挑战,我们提出VAKE(Verifiable Activation of Parametric KnowledgE,参数化知识可验证激活),这是一个两阶段强化学习框架,通过显式预激活(Priming)外化潜在的参数化知识,并将获得的提取能力迁移至隐式推理(Reasoning)。给定一个查询和一个不足的检索子图,预激活策略会显式插入桥接三元组作为可验证证据,监督来自独立冻结模型在增强子图上生成答案所得到的奖励。基于预激活阶段学习到的策略,推理阶段训练模型从原始输入中回答,测试通过显式知识提取获得的能力是否能迁移至隐式推理。在7个基准及3B至14B规模的模型上开展的实验表明,VAKE始终优于标准基线,包括在从HotpotQA直接迁移至分布外(OOD)数据集时。基于LLM的评估进一步显示,超过80%的插入三元组提供了无法从检索上下文中推导的事实桥接知识,而超过一半的插入三元组可引出无法通过直接提示获取的知识。这些结果表明,VAKE激活了潜在的参数化知识,而非复制输入上下文或记忆数据集特定关联。

英文摘要

Although Large Language Models (LLMs) encode rich factual knowledge in their parameters, reliably recalling and verifying such knowledge remains a key bottleneck in factual question answering. Existing end-to-end methods entangle knowledge elicitation with reasoning, making it difficult to determine whether correct answers arise from parametric knowledge or the input context. To address this challenge, we propose VAKE (Verifiable Activation of Parametric KnowledgE), a two-stage reinforcement-learning framework that externalizes latent parametric knowledge through explicit Priming and transfers the acquired elicitation capability to implicit Reasoning. Given a query and an insufficient retrieved subgraph, the Priming policy explicitly inserts bridging triples as verifiable evidence, with supervision provided by rewards derived from answers generated by a separate frozen model over the augmented subgraph. Building on the policy learned during Priming, the Reasoning stage trains the model to answer from the original input, testing whether the capability acquired through explicit knowledge elicitation transfers to implicit reasoning. Experiments across seven benchmarks and models from 3B to 14B show that VAKE consistently outperforms standard baselines, including when transferring directly from HotpotQA to OOD datasets. LLM-based evaluation further shows that over 80% of the inserted triples provide factual bridging knowledge not derivable from the retrieved context, while more than half elicit knowledge inaccessible through direct prompting. These results suggest that VAKE activates latent parametric knowledge rather than copying the input context or memorizing dataset-specific associations.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑