发表机构
TU Darmstadt; Hessian Center for Artificial Intelligence; National Research Center for Applied Cybersecurity ATHENE(达姆施塔特工业大学; 黑森人工智能中心; 国家应用网络安全研究中心 ATHENE)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探讨代码LLM的离线RL后训练,利用现有数据集避免在线采样,数小时内显著提升零样本代码生成性能,并在0.5B至7B参数模型上普遍有效。
AI 中文摘要
使用强化学习(RL)进行后训练是代码生成大语言模型(LLMs)开发中的关键阶段,因为它能确保模型遵循指令并生成功能正确的代码。该过程通常需要从基于Transformer的LLMs中生成计算密集型的代码样本,并进行大量的GPU-CPU通信以进行序列验证。为应对这些计算挑战,本研究探讨了是否可以通过利用现有数据集而非生成新样本,完全离线地执行基于RL的后训练。研究结果表明,仅需数小时的训练,无需在线采样即可显著提升LLMs的零样本代码生成性能。此外,离线RL在0.5B至7B参数规模的多种模型上均带来了性能提升,尽管提升幅度因模型家族而异。
英文摘要
Post-training with reinforcement learning (RL) is a critical phase in the development of code-generating large language models (LLMs), as it ensures adherence to instructions and the production of functionally correct code. This process typically requires computationally intensive code sample generation from Transformer-based LLMs and substantial GPU-CPU communication for sequence verification. To address these computational challenges, this work examines whether RL-based post-training can be performed entirely offline by leveraging existing datasets rather than generating new samples. The findings indicate that, with only a few hours of training, zero-shot code generation performance of LLMs can be substantially improved without online sampling. Additionally, offline RL produces performance gains across models ranging from 0.5B to 7B parameters, although the extent of improvement varies among model families.