发表机构
University of Waterloo; Khulna University of Engineering & Technology; Mohamed bin Zayed University of Artificial Intelligence(滑铁卢大学; 库尔纳工程技术大学; 穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究小规模语言模型强化学习不稳定问题,识别出三种失败模式,提出容量余量假设,采用合并并重新初始化适配器技术等方法,所提系统稳定收敛,提升偏好胜率,优于指令调整基线且减少训练数据。
AI 中文摘要
使用强化学习对参数范围在70 - 500M的小型语言模型(SLMs)进行对齐通常被认为是不稳定的,但其潜在失败机制尚未得到系统研究。在当前最优(SOTA)研究中,使用近端策略优化(PPO)训练了十五种(模型,语料库)配置。实验涵盖了Pythia - 70M、160M、410M以及SmolLM2 - 135M、360M在TinyStories、CNN/DailyMail和Wikitext - 103语料库上的情况。识别出小规模语言模型的三种可重现失败模式:标准PEFT/TRL管道中静默的LoRA参数冻结、使用bfloat16时重要性比率的数值溢出以及由于奖励模型错误导致的灾难性策略崩溃。通过合并并重新初始化适配器技术、PPO更新期间的float32精度以及包含奖励白化、重要性比率保护和权重回滚的三层安全机制解决了这些问题。本文提出了容量余量假设,即SLM规模下PPO性能取决于流畅的监督模型($\text{PPL}<20$)和有区分性的奖励信号,而非模型参数数量。所提出的系统在所有实验中稳定收敛,在具有流畅先验和信息丰富奖励信号的配置中,相对于SFT基线提高了偏好胜率。此外,它在需要显著更少训练数据的情况下优于指令调整基线。所有检查点、偏好数据集和训练脚本均已公开发布。
英文摘要
The alignment of Small Language Models (SLMs) in the 70--500M parameter range using reinforcement learning is often considered unstable, though the underlying failure mechanisms have not been systematically investigated. In the State-of-the-Art (SOTA) research, fifteen (model, corpus) configurations were trained using Proximal Policy Optimization (PPO). The experiments included Pythia-70M, 160M, 410M and SmolLM2-135M, 360M on the TinyStories, CNN/DailyMail, and Wikitext-103 corpora. Three reproducible failure modes were identified in small-scale language models: silent LoRA parameter freezing in standard PEFT/TRL pipelines, numerical overflow in importance ratios when using bfloat16, and catastrophic policy collapse due to reward-model error. These issues were addressed using a merge-and-reinitialize adapter technique, float32 precision during PPO updates, and a three-layer safety mechanism comprising reward whitening, importance-ratio guarding, and weight rollback. In this paper, a capacity-headroom hypothesis is proposed, which states that PPO performance at the SLM scale depends on both a fluent supervised model ($\text{PPL}<20$) and a discriminative reward signal, rather than on the number of model parameters. The proposed system converged stably in all experiments and improved preference win rate over the SFT baseline in configurations with a fluent prior and an informative reward signal. Furthermore, it outperformed instruction-tuned baselines while requiring significantly less training data. All checkpoints, preference datasets, and training scripts are publicly released$^§$.
CommentsProceedings of the 2026 IEEE International Conference on Systems, Man, and Cybernetics (SMC), Bellevue, WA, USA