发表机构
Huawei(华为)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对大语言模型强化学习训练因训练与推理差异导致的不稳定问题,提出自适应控制强化学习(ACRL)方法,可维持训练-推理差异在合理范围,稳定训练,增加策略熵,提升探索与准确性,实验表现优于BF16基线和重要性采样修复。
AI 中文摘要
大语言模型的强化学习训练常因训练与推理之间的差异而不稳定。这种训练-推理差异源于两个主要因素:训练和推理引擎之间的架构分离,以及推理中使用低精度量化而训练中使用高精度计算。为解决高训练-推理差异导致的训练不稳定问题,我们提出了其自适应控制的原理和方法。我们提出了自适应控制强化学习(ACRL),它能将训练-推理差异自适应地维持在合理范围内以确保稳定的强化学习训练。除了稳定之外,ACRL还能固有地增加策略熵,从而增强探索并提高准确性。实验结果表明,当推理引擎使用FP8量化时,ACRL始终将训练-推理差异维持在合理范围内并稳定强化学习训练。此外,ACRL不仅与BF16基线的准确性相匹配,还优于重要性采样(IS)修复。
英文摘要
Reinforcement Learning (RL) training for Large Language Models (LLMs) often suffers from instability due to the discrepancy between training and inference. This training-inference discrepancy stems from two primary factors: an architectural separation between training and inference engines, and the use of low-precision quantization in inference versus higher-precision computation in training. To address training instability issues caused by high training-inference discrepancy, we present the principles and methods for its adaptive control. We propose Adaptive Control Reinforcement Learning (ACRL), which adaptively maintains the training-inference discrepancy within a reasonable range to ensure stable RL training. Beyond stabilization, ACRL inherently increases policy entropy, thereby enhancing exploration and improving accuracy. The experimental results show that when the inference engine utilizes FP8 quantization, ACRL consistently maintains the training-inference discrepancy within a reasonable range and stabilizes RL training. Furthermore, ACRL not only matches the accuracy of the BF16 baseline but also outperforms importance sampling (IS) fixes.