arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RING:用于持续大规模知识注入的检索内化生成

RING: Retrieval-Internalized Generation for Continual Large-Scale Knowledge Injection

Shicheng Xu, Liang Pang, Liyi Chen, Zihao Wei, Jingcheng Deng, Yan Gao, Yi Wu, Yao Hu, Huawei Shen, Xueqi Cheng

arXiv 2608.01630首次发表:更新:

发表机构

State Key Laboratory of AI Safety, Institute of Computing Technology, CAS; University of Chinese Academy of Sciences; Xiaohongshu Inc.(中国科学院计算技术研究所人工智能安全国家重点实验室; 中国科学院大学; 小红书科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

RING是一种将检索内化的范式,通过三阶段训练学习参数化记忆的检索策略,在自主构建的News-2025基准上,其知识注入的准确性与效率优于或匹配现有方法。

AI 中文摘要

检索增强生成(RAG)提升了事实性,但在推理阶段增加了延迟和工程开销。我们提出RING(Retrieval-Internalized Generation,检索内化生成),这是一种涵盖架构与训练的整体范式,用于将大规模外部知识注入混合记忆专家(Mixture-of-Memory Experts),并通过强化学习学习对该内部记忆的参数化搜索,完全移除了外部检索器。训练分为三个阶段:持续预训练通过我们提出的双因果注意力(Dual Causal Attention)将新语料库注入知识专家;监督微调教授“先检索后回答”模式;带分层奖励的强化学习优化参数化记忆的路由与搜索策略。与此前将内部记忆与固定或基于规则的检索器配对的参数化注入方法不同,RING直接从任务信号中学习其检索策略。我们从理论上将RING表述为经典RAG目标的无搜索近似。为评估无测试时泄露的真正新知识的大规模注入,我们构建了News-2025基准,该基准严格基于基础大语言模型预训练截止时间之后的新闻构建。RING在准确性和效率上与基于搜索的RAG及参数化注入基准相当或超越它们。

英文摘要

Retrieval-augmented generation (RAG) improves factuality but adds latency and engineering overhead at serving time. We propose RING (Retrieval-Internalized Generation), a holistic paradigm spanning both architecture and training that injects large-scale external knowledge into a \textit{Mixture-of-Memory Experts} and learns parametric search over this internal memory via reinforcement learning, removing the external retriever entirely. Training proceeds in three stages: continued pre-training injects new corpora into a Knowledge Expert via our novel \textit{Dual Causal Attention}; supervised fine-tuning teaches a ``search-then-answer'' pattern; and reinforcement learning with hierarchical rewards optimizes the routing-and-search policy over the parametric memory. Unlike prior parametric injection methods that pair internal memory with a fixed or rule-based retriever, RING {learns} its retrieval policy directly from task signals. We further frame RING theoretically as a search-free approximation to the classical RAG objective. To evaluate large-scale injection of genuinely {new} knowledge without test-time leakage, we further construct News-2025, a benchmark built from news strictly post-dating the base LLM's pretraining cutoff. RING matches or surpasses both search-based RAG and parametric injection baselines in accuracy and efficiency.

Comments16 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑