发表机构
The Hong Kong University of Science and Technology; Hong Kong Baptist University(香港科技大学; 香港浸会大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文综述了通过参数内记忆增强大语言模型的方法,按参数放置与获取时间分类,并探讨了干扰、安全等开放方向。
AI 中文摘要
近期,大语言模型(LLMs)及基于LLM的智能体越来越需要整合预训练之后获取的知识,例如领域事实、用户偏好、文档和交互经验。上下文学习(ICL)和基于ICL的智能体工具虽然灵活,但会消耗上下文容量,并产生随上下文长度增长而增加的重复离散化编码成本。\textbf{参数内记忆}提供了一种互补的载体:可复用的记忆信息以模型参数、适配器或其他参数类对象的形式表示,并在推理时组合到前向传播中。本综述聚焦于在部署时通过此类参数化记忆增强LLM的方法:一个承载记忆的参数对象在推理期间被接入前向传播,无论该对象是在部署前还是部署期间获取的。我们通过两个正交轴来组织该领域:\textbf{参数放置},包括嵌入层、注意力层、前馈网络层,或当使用两层及以上时的混合方式;以及\textbf{参数获取时间},区分在部署期间(在线)获取记忆对象的方法与在部署前(离线)获取的方法。我们厘清了边界,进行了比较,并讨论了干扰、安全性、与ICL的协同设计以及递归自我改进等开放方向。
英文摘要
Recently Large Language Models (LLMs) and LLM-based agents increasingly need to incorporate knowledge acquired after pretraining, e.g., domain facts, user preferences, documents, and interaction experience. In-context learning (ICL) and ICL-based agent harness remain flexible, but they consume context capacity and incur repeated discretized encoding cost that grows with context length. \textbf{In-parameter memory} offers a complementary substrate: reusable memory information is represented in model parameters, adapters, or other parameter-like objects that are composed into the forward pass at inference time. This survey focuses on methods that augment LLMs with such parametric memory at deployment: a memory-bearing parameter object is plugged into the forward pass during inference, whether it is acquired before or during deployment. We organize the landscape with two orthogonal axes: \textbf{Parameter Placement}, which includes Embedding, Attention, FFN layers, or Hybrid when two or more layers are used; and \textbf{Parameter Acquisition Time}, which distinguishes methods whose memory object is acquired during deployment (online) from those acquired before it (offline). We clarify boundaries, conduct comparisons, and discuss open directions in interference, safety, co-design with ICL, and recursive self-improvement.