AI 中文总结
针对大型语言模型无需检索文档时无法回答相关问题的痛点,提出IAR三阶段后训练框架,在多数据集、多模型上显著提升了文档知识内化的领域与通用性能。
AI 中文摘要
大型语言模型在推理时未检索到源文档的情况下,常常无法回答关于有限文档集合的问题。我们将该场景研究为文档知识内化:将固定语料库转换为可用于无需检索的问答任务的参数化知识。我们提出IAR(Inject, Align, and Recover),这是一个三阶段后训练框架,分别处理结构化文档知识注入、问答行为对齐和通用能力恢复。与传统的持续预训练不同,Inject阶段将源文档转换为续写、重写和指令条件重建目标。Align阶段随后仅使用答案的问答监督来适配已注入知识的模型,而Recover阶段将领域适配后的模型与基础指令模型合并,以恢复通用能力。在Common Corpus(CC)和CCI数据集,以及Llama、Phi、Qwen和SmolLM模型家族上,IAR提升了无需检索的文档内化任务的领域-通用前沿水平。在主要对比中,IAR在8个数据集-模型设置中的7个里,在所有4个报告指标上均优于Vanilla SFT;在领域问答准确率上平均提升3.6个百分点,在IFEval、MMLU和MSBench上的平均通用性能提升12.1个百分点。扩展的CC基线显示,LoRA和FAPM可在单个通用指标上胜出,但在同时达到领先或接近领先的领域内化水平的方法中,IAR仍保持最强的通用性能之一。
英文摘要
Large language models often fail to answer questions about a bounded document collection when the source documents are not retrieved at inference time. We study this setting as document knowledge internalization: converting a fixed corpus into usable parametric knowledge for retrieval-free question answering. We propose IAR (Inject, Align, and Recover), a three-stage post-training framework that separates structured document knowledge injection, QA behavior alignment, and general ability recovery. Unlike conventional continued pretraining, Inject converts source documents into continuation, rewrite, and instruction-conditioned reconstruction objectives. Align then adapts the injected model with answer-only QA supervision, while Recover merges the domain-adapted model with the base instruction model to recover general capabilities. Across Common Corpus (CC) and CCI, and across Llama, Phi, Qwen, and SmolLM model families, IAR improves the domain-primary domain-general frontier for retrieval-free document internalization. In the main comparison, IAR improves over Vanilla SFT on all four reported metrics in 7 of 8 dataset-model settings, with average gains of 3.6 percentage points in domain QA accuracy and 12.1 percentage points in mean general performance across IFEval, MMLU, and MSBench. Extended CC baselines show that LoRA and FAPM can win individual general metrics, but among methods that also reach leading or near-leading domain internalization, IAR retains one of the strongest general profiles.
Comments21 pages, 4 figures. Includes Supplementary Material Sections A--G. Qian Kou and Xiaofeng Shi contributed equally and are co-corresponding authors. Hua Zhou is the project leader