arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.27787cs.LGcs.AI

LoRA 支架策略优化(LSPO):零奖励悬崖提示下恢复强化学习梯度的采样时低秩支架

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts

  • NVIDIA(英伟达)

机构由 AI 辅助整理,请以论文原文为准。

Ken Ding

AI总结:

针对数学推理 RLVR 的悬崖提示梯度丢失问题,提出 LSPO 采样时机制,通过 LoRA 适配器恢复梯度,在 DeepMath-103K 数据集上 16 项基准指标均优于 DAPO,平均提升 3.8 分。

AI中文摘要:

数学推理的可验证奖励强化学习(RLVR)存在结构性盲区:对于“悬崖”提示——即一组采样的 rollout 全部失败的提示,组归一化优势恒为零,因此 GRPO 恰好无法在模型能力边界的提示上产生梯度。我们提出 LoRA 支架策略优化(LSPO),这是一种采样时机制,用于恢复该丢失的梯度。每一步 RL 中,LSPO 检测悬崖提示,通过对其真实解执行简短的监督步骤拟合小型低秩(LoRA)适配器,使用基础模型加适配器的模型重新采样悬崖提示,通过重要性采样修正将当前成功的完成结果拼接回 RL 批次,仅对基础模型执行 GRPO 步骤;适配器仅接收监督梯度并在检查点被丢弃,最终得到仅含基础模型的结果。在 DeepMath-103K 数据集上使用 DeepSeek-R1-Distill-Qwen-1.5B 模型,以匹配的 1000 步报告 horizon,在每组 n=5 个配对种子的评估中,LSPO 的 5 种子均值在全部 16 个(基准、pass@k)单元上均匹配或优于 DAPO 基线(15 个严格胜出,1 个平局),在 AIME24/pass@4 上提升达 +10.7 分,在 AIME24 和 AIME26 的 pass@16 上提升 +6.7 分,在 MATH500/pass@1 上提升 +2.4 分;16 个单元的平均提升为 +3.8 分。

英文摘要:

Reinforcement learning from verifiable rewards (RLVR) for mathematical reasoning suffers from a structural blind spot: on "cliff" prompts-those on which every sampled rollout in a group fails-the group-normalized advantage is identically zero, so GRPO produces no gradient on precisely the prompts at the frontier of the model's capability. We introduce LoRA Scaffolded Policy Optimization (LSPO), a sampling-time mechanism that recovers this lost gradient. Each RL step, LSPO detects cliff prompts, fits a small low-rank (LoRA) adapter by a brief supervised step on their ground-truth solutions, re-rolls the cliffs with the base-plus-adapter model, splices the now-successful completions back into the RL batch with an importance-sampling correction, and takes a GRPO step on the base alone; the adapter receives only the supervised gradient and is discarded at checkpoint, yielding a base-only model. On DeepMath-103K with DeepSeek-R1-Distill-Qwen-1.5B, evaluated over n=5 paired seeds per arm at a matched 1000-step reporting horizon, LSPO's 5-seed mean matches or beats a DAPO baseline on all 16 (benchmark, pass@k) cells (15 strict wins and one exact tie), with gains of up to +10.7 points on AIME24/pass@4, +6.7 points on AIME24 and AIME26 at pass@16, and +2.4 points on MATH500/pass@1; averaged over the 16 cells the improvement is +3.8 points.

↑