arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

保留推理能力的后RL大语言模型微调:基于零基LoRA的方法

Reasoning-Preserving Fine-Tuning of Post-RL LLMs with Null-Basis LoRA

Wenzhi Fang, Nicholas Tzou, Lazar Valkov, Srinivas Chappidi

arXiv 2609.25618首次发表:更新:

发表机构

Samsung Research America; Purdue University(三星美国研究院; 普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对后RL模型微调导致推理能力遗忘的问题,提出NB-LoRA方法,利用推理激活的零空间约束LoRA更新,在保持适应性能的同时保留推理能力。

AI 中文摘要

基于强化学习(RL)的后训练已成为激发大型语言模型(LLMs)推理能力的有效方法。然而,通过后续的监督微调(SFT)将后RL模型适应于新的知识领域或行为,可能会严重覆盖这些能力。现有方法通过经验回放、专门初始化或使用梯度投影的约束优化来缓解此类遗忘,但要么提供有限的保留效果,要么带来大量的训练开销。我们的分析表明,推理激活集中在低维子空间中,为适应留下了大量的零空间容量,并且相应的近似零空间可以从少量示例中可靠地估计。受这些观察的启发,我们提出了零基低秩适应(NB-LoRA),一种参数高效的方法,用于适应后RL LLMs,同时保留其习得的推理能力。我们将推理保留表述为逐层隐藏状态保留约束,并从推理激活中构建一个固定的近似零基。然后通过该基对LoRA更新进行重新参数化,在微调过程中强制执行保留约束。在多个经过RL训练的LLMs和多样化的下游任务上的大量实验表明,NB-LoRA在适应性能上与标准LoRA相当,将推理准确性保持在接近微调前的水平,并将这种保留推广到未见的推理基准上。

英文摘要

Reinforcement learning (RL)-based post-training has become an effective approach for eliciting reasoning capabilities in large language models (LLMs). However, adapting post-RL models to new knowledge domains or behaviors through subsequent supervised fine-tuning (SFT) can severely overwrite these capabilities. Existing approaches mitigate such forgetting through experience replay, specialized initialization, or constrained optimization using gradient projection, but either provide limited preservation or incur substantial training overhead. Our analysis shows that reasoning activations concentrate in low-dimensional subspaces, leaving substantial null-space capacity for adaptation, and that the corresponding approximate null spaces can be reliably estimated from a modest number of examples. Motivated by these observations, we propose Null-Basis Low-Rank Adaptation (NB-LoRA), a parameter-efficient method for adapting post-RL LLMs while preserving their acquired reasoning ability. We formulate reasoning retention as a layer-wise hidden-state preservation constraint and construct a fixed approximate null basis from reasoning activations. LoRA updates are then reparameterized through this basis, enforcing the preservation constraint throughout fine-tuning. Extensive experiments across multiple RL-trained LLMs and diverse downstream tasks show that NB-LoRA matches standard LoRA in adaptation performance, maintains reasoning accuracy near pre-fine-tuning levels, and generalizes this preservation to held-out reasoning benchmarks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑