arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39321cs.LGcs.AI

小型语言模型的GRPO训练动态

GRPO Training Dynamics for Small Language Models

  • Nutanix

机构由 AI 辅助整理,请以论文原文为准。

Rajat Ghosh, Vaishnavi Bhargava, Henry Wong, Aryan Singhal, Debojyoti Dutta

AI总结:

本研究系统分析了1.5B至7B参数的小型语言模型在GRPO微调中的训练动态,涵盖数学、编码和科学MCQ任务,发现其在数学基准上表现优异但在其他领域有限,并提出了LoRA和奖励塑形的改进建议。

AI中文摘要:

组相对策略优化(GRPO)已成为一种针对推理密集型任务的内存高效的强化微调(RFT)技术。然而,GRPO在小型语言模型(SLMs)上的训练动态仍鲜为人知,这限制了其在开放和资源受限环境中的可靠采用和可复现性。在本工作中,我们对参数量从1.5B到7B的SLMs进行了GRPO微调的系统性研究,并在实际单节点8xA100计算预算下进行。我们的研究涵盖多个模型家族和推理领域,包括数学、编码和科学中的多项选择题回答(MCQ)。在这些设置中,我们分析了组大小如何影响策略收敛、训练稳定性和下游基准性能。我们进一步刻画了GRPO训练期间张量级别的更新动态,并研究了LoRA目标模块和层的选择是否能够提升GRPO调优模型的性能。虽然我们初步的GRPO调优模型在约80%的数学基准评估中优于其基础对应模型,但它们在MCQ和代码推理任务上表现出有限的能力。在机制性评估的指导下,我们优化了LoRA和奖励塑形配置,以提升后两个领域的性能。这些发现为SLMs的GRPO训练提供了实用指导。

英文摘要:

Group Relative Policy Optimization (GRPO) has emerged as a memory-efficient reinforcement fine-tuning (RFT) technique for reasoning-intensive tasks. How- ever, GRPO training dynamics on small language models (SLMs) remain poorly understood, limiting its reliable adoption and reproducibility in open and resource- constrained environments. In this work, we present a systematic study of GRPO fine-tuning for SLMs ranging from 1.5B to 7B parameters under a practical single- node 8xA100 compute budget. Our study spans multiple model families and reasoning domains, including mathematics, coding, and multiple-choice question answering (MCQ) in science. Across these settings, we analyze how group size affects policy convergence, training stability, and downstream benchmark per- formance. We further characterize tensor-level update dynamics during GRPO training and investigate whether the choice of LoRA target modules and layers can improve the performance of GRPO-tuned models. While our initial GRPO-tuned models outperform their base counterparts on approximately 80% of mathematical benchmark evaluations, they demonstrate limited capability on MCQ and code reasoning tasks. Guided by our mechanistic evaluations, we refined our LoRA and reward-shaping configurations to improve performance in latter domains. These findings provide practical guidance for GRPO training for SLMs.

↑