arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2502.18449cs.SEcs.AIcs.CL

SWE-RL:通过强化学习提升大语言模型推理能力

SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution

  • Meta AI
  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

Yuxiang Wei, Olivier Duchenne, Jade Copet, Quentin Carbonneaux, Lingming Zhang, Daniel Fried, Gabriel Synnaeve, Rishabh Singh, Sida I. Wang

更新

AI总结:

SWE-RL通过强化学习提升大语言模型在软件工程领域的推理能力,实现了41.0%的解决率,超越了现有中等规模LLM的性能。

AI中文摘要:

最近发布的DeepSeek-R1展示了强化学习(RL)在增强大语言模型(LLM)通用推理能力方面的巨大潜力。虽然DeepSeek-R1和其他后续工作主要集中在将RL应用于竞争编程和数学问题,但本文介绍了SWE-RL,这是首个将基于强化学习的LLM推理扩展到现实软件工程领域的做法。通过轻量级的基于规则的奖励(例如,真实解与LLM生成解之间的相似度分数),SWE-RL使LLM能够通过学习大量的开源软件演变数据来自主恢复开发者的推理过程和解决方案——这记录了软件的整个生命周期,包括代码快照、代码变更以及诸如问题和拉取请求等事件。在Llama 3的基础上训练,我们得到的推理模型Llama3-SWE-RL-70B在SWE-bench Verified上实现了41.0%的解决率——一个经过人工验证的现实GitHub问题集合。据我们所知,这是目前中等规模(<100B)LLM的最佳性能报告,甚至可以与领先的专有LLM如GPT-4o相媲美。令人惊讶的是,尽管仅在软件演变数据上进行强化学习,Llama3-SWE-RL甚至展现出了泛化推理能力。例如,它在五个跨领域任务上表现有所提升,即函数编码、库使用、代码推理、数学和一般语言理解,而监督微调基线在平均上甚至导致性能下降。总体而言,SWE-RL为通过大规模软件工程数据的强化学习来提升LLM推理能力开辟了新方向。

英文摘要:

The recent DeepSeek-R1 release has demonstrated the immense potential of reinforcement learning (RL) in enhancing the general reasoning capabilities of large language models (LLMs). While DeepSeek-R1 and other follow-up work primarily focus on applying RL to competitive coding and math problems, this paper introduces SWE-RL, the first approach to scale RL-based LLM reasoning for real-world software engineering. Leveraging a lightweight rule-based reward (e.g., the similarity score between ground-truth and LLM-generated solutions), SWE-RL enables LLMs to autonomously recover a developer's reasoning processes and solutions by learning from extensive open-source software evolution data -- the record of a software's entire lifecycle, including its code snapshots, code changes, and events such as issues and pull requests. Trained on top of Llama 3, our resulting reasoning model, Llama3-SWE-RL-70B, achieves a 41.0% solve rate on SWE-bench Verified -- a human-verified collection of real-world GitHub issues. To our knowledge, this is the best performance reported for medium-sized (<100B) LLMs to date, even comparable to leading proprietary LLMs like GPT-4o. Surprisingly, despite performing RL solely on software evolution data, Llama3-SWE-RL has even emerged with generalized reasoning skills. For example, it shows improved results on five out-of-domain tasks, namely, function coding, library use, code reasoning, mathematics, and general language understanding, whereas a supervised-finetuning baseline even leads to performance degradation on average. Overall, SWE-RL opens up a new direction to improve the reasoning capabilities of LLMs through reinforcement learning on massive software engineering data.

补充信息

↑