学习从上下文中学习:基于扰动公开文档的合成训练
Learning to Learn from Context: Synthetic Training from Perturbed Public Documents
浏览论文内容
中文总结 AI 辅助
通过扰动公开文档构建合成训练流水线,生成依赖上下文的推理样本,提升学生模型在CL-bench上的表现,达到与万亿参数前沿模型相当的水平。
中文摘要 AI 辅助
现实世界的任务通常要求大型语言模型(LLMs)从复杂的特定任务上下文中学习,而非依赖预训练的参数化知识。这一能力目前仍是LLMs的薄弱环节,而针对此类任务上下文的人工标注成本高昂且难以规模化。高质量的公开文档是一种丰富的替代资源,但公共网络中的大部分内容在预训练阶段已被消耗:直接基于这些文档进行训练会鼓励记忆而非上下文学习。在本工作中,我们尝试利用带有轻微扰动的高质量公开文档,并实证发现LLMs能够成功生成依赖上下文的推理轨迹和答案,这些内容随后被用于训练学生模型。具体而言,我们构建了一个合成流水线,该流水线(i)重写源文档以降低记忆风险,(ii)生成需要基于文档进行推理的问题和评分标准,(iii)以文档为上下文回答问题,以及(iv)仅保留真正依赖文档的样本。在无需任何人工标注的情况下,我们的流水线从3.5k篇文档中生成了约10k个样本,所得学生模型在CL-bench上的性能显著提升。监督微调(SFT)将Qwen3.6-35B-A3B学生模型的得分从13.7%提升至22.8%,随后基于评分标准的强化学习(RL)阶段达到24.6%,与参数量超过万亿的前沿模型Qwen3.8-2.4T(23.9%)在CL-bench上相当。我们还观察到改进广泛迁移至长上下文理解、指令遵循和推理任务,而代码生成和知识保持基本持平。我们希望这项工作提供一种可复现且可扩展的方式,以提升LLMs从上下文中学习的能力,并促进对基于上下文的推理的进一步研究。
英文摘要
Real-world tasks often require large language models (LLMs) to learn from complex task-specific context rather than pretrained parametric knowledge. This capability remains a weakness of LLMs, while human annotation for such task contexts is expensive and difficult to scale. Public high-quality documents are an abundant alternative, but much of the public web has already been consumed during pretraining: training on such documents naively would reward memorization rather than context learning. In this work, we attempt to make use of high-quality public documents with small perturbations and empirically find that LLMs can successfully generate context-dependent reasoning traces and answers, which are then used to train a student model. Specifically, we construct a synthesis pipeline that (i) rewrites source documents to reduce memorization risk, (ii) generates questions and rubrics that require reasoning over the document, (iii) answers the questions with the document as context, and (iv) admits only samples that genuinely depend on the document. Without any human annotators, our pipeline generates about 10k samples from 3.5k documents, and the resulting student model substantially improves the performance on CL-bench. SFT raises a Qwen3.6-35B-A3B student from 13.7% to 22.8%, and a subsequent rubric-reward RL stage reaches 24.6%, on CL-bench comparable with a frontier model of over a trillion parameters, Qwen3.8-2.4T (23.9%). We also observe a broad transfer of improvements to long-context understanding, instruction following, and reasoning, while code generation and knowledge remain mostly flat. We hope this work provides a reproducible and scalable way to improve the ability of LLMs to learn from context, and to facilitate further research on context-grounded reasoning.