arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00082cs.CLcs.AI

KItCAT:通过输入损坏进行知识注入的自回归训练

KItCAT: Knowledge Injection via Input Corruption for Auto-regressive Training

  • IBM(国际商业机器公司)
  • Amazon Books Science(亚马逊图书科学部门)

机构由 AI 辅助整理,请以论文原文为准。

Meghanadh Pulivarthi, Kushagra Bhushan, Vineet Kumar, Gaurav Pandey, Jaydeep Sen, Dinesh Raghu, Sachindra Joshi, Yatin Nandwani

AI总结:

本研究提出轻量级知识注入策略KItCAT,通过随机损坏输入序列实现低成本数据增强,在多个数据集和模型系列上的表现优于持续预训练。

AI中文摘要:

大型语言模型(LLMs)在预训练过程中获取了大量知识,但往往缺乏回答手册或技术文档等小众来源问题所需的专业知识,这些小众来源在预训练期间未被见过。持续预训练(CPT)被广泛用于将此类知识注入模型参数中。然而,小众文档很少重复事实,这使得CPT难以稳健地获取此类知识。近期研究通过生成新知识的多个改写来解决该问题,但改写计算成本高昂,且通常需要强大的LLMs。本研究中,我们提出KItCAT:通过损坏的自回归训练进行知识注入,这是一种轻量级训练策略,可减少仅解码器LLMs中对改写的需求。KItCAT通过随机损坏输入序列来增强标准的下一个token预测。在训练期间,随机子集的输入token被替换为其他词汇表token,同时保留原始的下一个token标签不变。这种简单的干预从每个样本生成多样化的训练输入,能够以可忽略的成本实现大规模数据增强。我们表明,KItCAT在多个数据集和模型系列上始终优于CPT。代码可在https URL获取。

英文摘要:

LLMs acquire vast amounts of knowledge during pre-training, but often lack the specialized knowledge needed to answer questions from niche sources such as manuals or technical documents unseen during pre-training. Continued pre-training (CPT) is widely used to inject such knowledge into model parameters. However, niche documents seldom repeat facts, making it difficult for CPT to robustly acquire such knowledge. Recent works address this by generating multiple paraphrases of the new knowledge, but paraphrasing is computationally expensive and typically requires powerful LLMs. In this work, we introduce KItCAT: Knowledge Injection via Corrupted Auto-regressive Training, a lightweight training strategy that reduces the need for paraphrasing in decoder-only LLMs. KItCAT augments standard next-token prediction by stochastically corrupting the input sequence. During training, a random subset of input tokens is replaced with other vocabulary tokens while the original next-token labels are kept unchanged. This simple intervention generates diverse training inputs from each sample, enabling large-scale data augmentation at negligible cost. We show that KItCAT consistently improves over CPT across multiple datasets and model families. Code is available at https://github.com/meghanadhpulivarthi/KItCAT.

补充信息

↑