发表机构
Interval(Interval)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Internalizer是可移植上下文到参数映射超网络,可为2840亿参数DeepSeek v4 Flash生成文档专用LoRA适配器,在未见过文档上准确率远超基础模型,单次前向传播即可实现文档到适配器的转换。
AI 中文摘要
将上下文直接映射到LoRA适配器的超网络可让大语言模型将该上下文存储在其权重中,但现有研究仅在参数规模达140亿的基础模型上验证过该方法。本文提出Internalizer,这是一种达到当前最优水平的可移植上下文到参数映射超网络,可为冻结的2840亿参数DeepSeek v4 Flash生成文档专用的LoRA适配器,目标模型参数规模比以往研究大两个数量级。该超网络的大部分参数位于与模型无关的主干中,每个基础模型仅含薄的入口和出口层,因此可在小型模型上廉价训练后再移植到大型模型。在最多4096个token的未见过文档上,生成的适配器达到84.9%的Top-1准确率和97.8%的Top-5教师强制准确率,而基础模型的对应准确率仅为63.4%和83.5%,上下文窗口中仅包含一个三词指令。超网络训练完成后,单次前向传播即可将任意文档转换为该模型的适配器,该适配器可单独部署以提升速度,或与文档一同置于窗口中以进一步提升准确率。
英文摘要
Hypernetworks that map a context directly to a LoRA adapter let a large language model carry that context in its weights, but prior work has demonstrated them only on base models of up to 14 billion parameters. We present the Internalizer, a state-of-the-art, portable Context-to-Parameter Mapping hypernetwork that generates document-specific LoRA adapters for the frozen 284B-parameter DeepSeek v4 Flash, a target two orders of magnitude larger than in any previous work. Most of its parameters live in a model-agnostic trunk with only thin entry and exit layers per base model, so it trains cheaply against small models before being ported to the large one. On unseen documents of up to 4096 tokens, the generated adapters reach 84.9% top-1 and 97.8% top-5 teacher-forced accuracy against 63.4% and 83.5% for the base model, with nothing in the context window but a three-word instruction. Once the hypernetwork is trained, a single forward pass turns any document into an adapter for such a model, which could be served alone for speed or alongside the document in the window to raise accuracy further.
Comments14 pages, 3 figures, 1 table