无限参数大语言模型:从实时数据生成与适配权重
Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data
- Boltzbit Limited(Boltzbit有限公司)
- University of Cambridge(剑桥大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
针对静态权重无法从实时交互中学习的问题,提出无限参数大语言模型,用超网络将实时数据生成低秩权重调制,并通过贝叶斯在线更新,以固定存储实现无限权重,优于上下文内学习。
中文摘要 AI 辅助
缩放定律认为,语言模型随着参数和训练数据的增加而变得更加强大,混合专家(Mixture-of-Experts, MoE)架构正是借助这些定律取得了显著成果,每个词元仅激活庞大存储参数库中的一小部分。然而,这一成功建立在静态预训练数据之上。部署后的模型面对的是一个不同的世界,其中许多能让它更有用的数据并不在其训练集中,而是存在于它当前正在处理的实时交互中,例如用户提供的事实或给出的纠正。传统模型无法从这些数据中学习,因为其权重在训练后被冻结。相反,运行时提供的知识和行为被放入提示中,通过检索或指令实现,并在每次请求时重新读取,一旦请求结束便被丢弃。我们探讨一种架构如何通过将实时交互写入其权重来从中学习。受MoE启发,我们提出了无限参数大语言模型(Infinite-Parameter LLM)。一个紧凑的超网络将运行时给定的数据转化为共享基础网络的低秩调制,因此前馈权重是从实时数据生成的,而非存储在固定库中。与先前的权重生成器一次性读取上下文并冻结不同,我们对生成器的潜在编码保持贝叶斯信念,并在线更新,因此有效权重会随着会话的进行从该演化的信念中重新推导,而非在一次读取后固定。存储占用保持固定,但模型能够编译的权重实际上是无限的。对于运行时提供的知识和行为,将其携带在权重中而非提示中,在计算上是分摊的,释放了上下文窗口,跨轮次持续存在,并且能比上下文内使用更好地泛化。我们指定了一个评估协议,专门针对上下文内学习和检索来测试这一点。
英文摘要
Scaling laws hold that language models grow more capable with more parameters and more training data. Mixture-of-Experts (MoE) architectures are a remarkable demonstration of these laws, activating only a fraction of an enormous parameter bank for each token. But this success is built on static pretraining data --- the facts and corrections supplied by users during live interactions are a significant untapped source of potential improvement for a deployed model, but cannot be exploited by conventional architectures whose weights are frozen after training. Instead, this newfound knowledge must be placed in the context (by instruction or retrieval) and re-read on every request, only to be discarded afterwards. We seek instead to learn from live interactions by dynamically updating model weights. Inspired by MoEs, we propose the \textbf{Infinite-Parameter LLM}. A compact hypernetwork turns the online data into low-rank modulations of a shared base network, so feed-forward weights are generated from live data, not read from static memory. Whereas existing weight generators are held fixed after reading the context once, we form a Bayesian belief over the generator's latent state and update it online, such that the effective weights are re-derived as our belief evolves during the session. Although the model's memory footprint is constant, the feasible space of generated weights is thus effectively infinite. Representing live data in the weights rather than the prompt amortises compute, frees the context window, persists updates across turns, and can generalise better than in-context use. Our evaluation protocol applies this methodology to in-context learning and retrieval.