arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于超网络的大语言模型知识注入的缩放定律

Scaling Laws for Hypernetwork-Based Knowledge Injection in Large Language Models

Nischay Dhankhar, Dos Baha, Abulhair Saparov

arXiv 2607.19604首次发表:更新:

发表机构

Nace AI; Purdue University(Nace人工智能公司; 普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究大语言模型训练时知识注入问题,核心方法是用超网络生成LoRA适配器,主要贡献是发现超网络注入能力随规模变化规律,能可靠进行分布外泛化,为训练时适应提供新基础及缩放定律。

AI 中文摘要

将事实性知识可靠且大规模地注入大语言模型仍然是一个开放的挑战。超网络为大规模知识注入提供了一个有前途的解决方案。虽然超网络通常用于测试时适应,我们探索它们在训练时知识注入中的应用。给定大量事实语料库,训练一个超网络来生成一个固定的LoRA适配器,插入目标模型后能使其回答关于这些事实的问题。在这项工作中,我们研究超网络是否可用于训练时知识注入以及这种能力如何随规模变化。我们的设计将超网络的注入能力与目标模型的一般能力解耦,首次能够对超网络架构的缩放定律进行严格研究。我们构建了一个名为MegaWikiQA的大规模数据集,包含来自Wikidata5M中39个领域的数千万个多跳问答示例。结果表明:基于超网络的注入在所有架构轴上呈现出大致的预测幂律缩放;超网络在不断增加的规模上能够可靠地进行分布外泛化,表明超网络是其他训练时适应方法(如LoRA微调)的有前途的替代方案,在所有分布外评估中呈现出更陡的缩放指数。这些结果确立了超网络作为训练时适应的有原则且可扩展的基础,并提供了第一个基于经验的缩放定律来指导大语言模型中用于事实推理的超网络。

英文摘要

Injecting factual knowledge into large language models (LLMs) reliably and at scale remains an open challenge. Hypernetworks provide a promising solution to large-scale knowledge injection. Although hypernetworks are typically applied for test-time adaptation, we explore their use in train-time knowledge injection, where, given a large corpus of facts, we train a hypernetwork to generate a fixed LoRA adapter that, when inserted into the target model, enable the model to answer questions about those facts. In this work, we investigate whether hypernetworks can be used to perform train-time knowledge injection and how this ability varies with scale. The scaling behavior of hypernetworks remains largely unstudied. Our design decouples the hypernetwork's injection capacity from the target model's general capability, enabling, for the first time, a rigorous study of scaling laws for hypernetwork architectures. We characterize how loss, reasoning accuracy, and out-of-distribution (OOD) generalization vary with hypernetwork depth, width, and target network size. We construct a large-scale dataset, called MegaWikiQA, containing tens of millions of multi-hop question-answer examples across 39 domains constructed from examples in Wikidata5M. Our results reveal: (i) hypernetwork-based injection exhibits broadly predictive power law scaling along all architecture axes; and (ii) hypernetworks are capable of reliable OOD generalization at increasing scales, suggesting that hypernetwork provides a promising alternative to other train-time adaptation methods such as LoRA finetuning and full fine-tuning, exhibiting steeper scaling exponents in all OOD evaluations. Together, these results establish hypernetworks as a principled and scalable substrate for train-time adaptation, and provide the first empirically grounded scaling laws to guide hypernetworks for factual reasoning in large language models.

CommentsPreprint, 22 pages, 14 figures. Dataset collection available at https://huggingface.co/collections/nace-ai/hypernetwork-datasets

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑