通过贝叶斯视角统一ICL、SFT与KL正则化的强化学习
Unifying ICL, SFT, KL-Regularized RL Through a Bayesian Lens
浏览论文内容
中文总结 AI 辅助
本研究通过贝叶斯视角将ICL、SFT、KL正则化的RLHF/RLVR等大型语言模型训练评估范式统一,厘清等价性差异,并为现代推理流水线提供启示。
中文摘要 AI 辅助
大型语言模型目前在多种范式下进行训练和评估:监督微调(SFT)、少样本上下文学习(ICL)、KL正则化的RLHF/RLVR、在线策略蒸馏(OPD)以及结合搜索和思维链的测试时推理。这些方法常被视为根本不同,而近期的实证结果——如少样本提示对RL调优的推理模型的混合影响——可能显得令人困惑。本研究提出一种贝叶斯视角,将这些方法置于同一框架下。核心是一个两步模板:(i)基于上下文,利用先验/参考模型和效用信号(对数似然、奖励或优势),在输出或动作上构建(广义)贝叶斯或吉布斯后验q*;(ii)通过前向KL投影将q*近似到参数族中,该投影可在权重内(SFT/RL)或上下文内(ICL)进行。第一部分将少样本ICL和SFT形式化为对贝叶斯后验预测的摊销和权重内投影。第二至第四部分表明,KL正则化的RLHF/RLVR、奖励加权SFT、奖励加权ICL(RW-ICL)以及优势加权SFT(AWSFT)都是对由奖励或优势诱导的后验进行前向KL投影的实例。我们厘清了这些等价性成立的地方(目标和一阶更新)与不成立的地方(学习信号的来源和粒度)。第五部分概述了对现代推理流水线的启示:RLHF/RLVR方案作为“后验设计+投影”,为什么重要性加权KL投影在实践中不可避免地需要冷启动或监督预热,以及DeepSeek-R1和o1风格的推理模型是将测试时贝叶斯搜索与训练时KL摊销相结合。
英文摘要
Supervised fine-tuning (SFT), few-shot in-context learning (ICL), KL-regularized RLHF/RLVR, and on-policy distillation are usually treated as distinct post-training paradigms. We develop a unified Bayesian perspective in which each is an instance of a two-step template: construct a (generalized) Bayes or Gibbs posterior from a reference model and a utility signal (log-likelihood, reward, or advantage), then approximate it by a forward-KL projection onto a parametric family, either in-weights (SFT/RL) or in-context (ICL). This yields a single chain of equivalences: few-shot ICL is an amortized projection onto the Bayes posterior predictive, and reward-weighted SFT, reward-weighted ICL, and advantage-weighted SFT are forward-KL projections of reward-induced Gibbs posteriors. The framework explains why supervised warm-up is practically unavoidable for importance-weighted projections, and interprets R1/o1-style reasoning models as combining test-time Bayesian search with training-time amortization. Matched-budget experiments on Qwen3 models corroborate the picture: operators that share their learning-signal granularity produce nearly identical updates when support is good and diverge when it degrades, and reward-weighted projection performs on par with standard baselines.
发表机构
- Fudan University(复旦大学)
机构由 AI 辅助整理,请以论文原文为准。