发表机构
Google; Google Inc.(谷歌; 谷歌公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出训练超网络将用户上下文映射为个性化LoRA,在设备端仅需前向传播即可合成适配,兼顾ICL的计算可行性与PEFT的权重修改优势,在长文本生成任务上验证了有效性。
AI 中文摘要
设备端大语言模型(LLMs),例如在手机上运行的模型,已具备通过个性化进行改进的成熟条件。移动设备有限的计算资源对模型规模乃至模型质量施加了限制,使得任何可实现的性能提升都极具影响力。与此同时,其个人属性(即与特定用户的紧密耦合)意味着,给定的设备端LLM在时间推移中往往以相似、可预测的模式被使用。本文提出了一种新颖的设备端LLM个性化方法。它训练一个超网络,将用户的上下文令牌映射到适合该用户的低秩适配(LoRA)。一旦训练好的通用组件部署到用户设备上,每个用户便使用超网络在设备端完全合成个性化的LoRA。该方法融合了两种现有LLM定制方法(上下文学习(ICL)和参数高效微调(PEFT))的优点,同时避免了它们的缺点。与ICL类似(不同于PEFT),我们方法的设备端阶段在计算上是可行的,仅需通过神经网络进行前向传播。与PEFT类似(不同于ICL),我们的方法通过权重(即LoRA)修改“目标”基础LLM,避免了扩展输入序列带来的负面后果(如延迟增加)。我们的方法特别适合移动设备场景。除了上述设备端计算和延迟优势外,它还需要极少的额外存储,因为其架构内部部分利用了与待个性化目标LLM相同的LLM权重。我们在多个代表性个性化数据集上展示了LoRA生成超网络的优越性,并与ICL和PEFT等基线进行了比较。值得注意的是,我们的个性化实验聚焦于更具挑战性且研究较少的生成长文本生成任务。
英文摘要
On-device large language models (`LLMs'), e.g. running on mobile phones, are ripe for improvement via personalization. The limited compute resources of mobile devices impose limits on model scale and thus model quality, making any realizable quality gains highly impactful. At the same time, their personal nature (i.e., the close coupling to a particular user) means that a given on-device LLM tends to be used in similar, predictable patterns over the course of time. This paper presents a novel method for personalizing on-device LLMs. It trains a hypernetwork to map a user's context tokens to a low-rank adaptation (`LoRA') well-suited to that user. Once the trained common artifacts are deployed to users' devices, each user uses the hypernetwork to synthesize (entirely on device) a personalized LoRA. This approach blends the benefits while avoiding the drawbacks of two existing approaches to LLM customization: in-context learning (`ICL') and parameter-efficient fine-tuning (`PEFT'). Like ICL (and unlike PEFT), the on-device phase of our approach is computationally feasible, requiring only forward passes through neural networks. Like PEFT (and unlike ICL), our approach modifies the `target' base LLM via weights (the LoRA), avoiding negative consequences (e.g. increased latency) associated with extending the input sequence. Our approach is particularly well-suited to the mobile device regime. Apart from the on-device compute and latency benefits mentioned, it also requires minimal additional storage, as internally its architecture partly leverages the same LLM weights as belong to the target LLM to be personalized. We demonstrate the benefits of LoRA-generating hypernetworks on several representative personalization datasets, comparing against baselines like ICL and PEFT. Of note, our personalization experiments focus on more challenging and less studied long-form text generation tasks.
Comments19 pages, 4 figures