发表机构
IBM; Indian Institute of Technology, Madras(国际商业机器公司; 马德拉斯印度理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出DKL方法,通过对指令调优语言模型对应的基础大语言模型进行扩展预训练并合并权重,在不损害指令跟随能力的前提下提升RAG准确率,且训练数据需求远低于现有方案。
AI 中文摘要
检索增强生成(RAG)已成为将特定语料库的新知识融入指令跟随大语言模型(Instruct LLM)的事实标准方法。尽管基于RAG的提示能提升事实准确性,但在检索不正确或不完整时会失效,导致幻觉问题。RAFT、PA-RAG等微调方法通过向模型参数注入新知识来增强RAG,但需要生成覆盖整个语料库的大量合成问答数据。对文本语料库进行扩展预训练(EPT)可避免生成全面合成数据的需求,但会损害Instruct LLM的指令跟随能力,需在预训练后进行指令微调(IFT)。然而,IFT成本高昂,且因缺乏指令调优语料库而难以实施。本研究提出面向指令调优语言模型的解耦知识学习(DKL):DKL不对Instruct LLM进行EPT,而是对其对应的基础大语言模型进行EPT以注入新知识,再将注入新知识的权重与Instruct LLM合并,从而在不影响指令跟随能力的前提下赋予模型新知识。DKL是一种轻量方法,无需昂贵的指令微调,依靠模型合并将新知识注入Instruct LLM且不破坏其指令跟随能力。实验结果显示,DKL在检索失败案例上将RAG准确率从54.17提升至79.26,且使用的训练数据远少于现有方法,性能优于现有方案。
英文摘要
RAG has become the de facto method for incorporating new, corpus-specific knowledge into an instruction following LLM (Instruct LLM). Although RAG-based prompting improves factual grounding, it fails when retrieval is incorrect or incomplete, leading to hallucinations. Finetuning methods such as RAFT and PA-RAG enhance RAG by injecting new knowledge into the model's parameters, but require generating a massive amount of synthetic QA that covers the entire corpus. Extended Pre-Training (EPT) on the text corpus avoids the need for comprehensive synthetic data generation but compromises an Instruct LLM's instruction-following capabilities, necessitating instruction fine-tuning (IFT) after pre-training. However, IFT is costly and may be infeasible due to the unavailability of an instruction-tuning corpus. In this work, we propose DKL-Decoupled Knowledge Learning for Instruction-Tuned Language Models. Instead of doing EPT on the Instruct LLM, DKL performs EPT on its corresponding base LLM to infuse new knowledge. These knowledge infused weights are then merged with the Instruct LLM, imparting new knowledge without affecting their instruction-following capabilities. DKL is a lightweight method that avoids expensive instruction fine-tuning and relies on model merging to infuse the new knowledge into the Instruct LLM without destroying its instruction following capabilities. Empirical results show that DKL improves RAG accuracy from 54.17 to 79.26 on retrieval failure cases, while outperforming prior approaches with substantially less training data.
Comments20 pages, 4 figures, 15 tables