arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习在LLM的强化学习中永不放弃以解决难题

Learning to Solve Hard Problems in RL for LLMs by Never Giving Up

Michael Noukhovitch, Hamish Ivison, Nathan Lambert, Aaron Courville

arXiv 2609.13443首次发表:更新:

发表机构

Mila, Université de Montréal; Allen Institute for AI; University of Washington; Trillium Labs(米拉,蒙特利尔大学; 艾伦人工智能研究所; 华盛顿大学; 特里利厄姆实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM强化学习中简单问题提升大而难题提升小的马太效应,提出自适应采样方法NGU,动态分配计算资源,持续采样直至正确,在数学和编码基准上提升难题解决性能。

AI 中文摘要

我们证明,使用强化学习(RL)训练大型语言模型(LLM)并不会在数据集上均匀地提升性能。对于LLM已经擅长解决的简单问题,RL带来了显著的改进,但对于难题,改进则很小。我们将此称为LLM强化学习中的马太效应,这一命名源自经济学和网络科学中的累积优势现象,概括为“富者愈富”。天真的解释是难题需要更多的计算资源来找到解决方案。我们认为,现代RL方法通过在简单问题上浪费过多计算而加剧了这一问题,应动态重新分配计算资源的使用。我们引入了“永不放弃”(NGU),一种简单的自适应采样方法,该方法持续为一个问题生成样本,直到生成正确的样本。通过利用异步RL,这种方法自然地使用较少的样本来筛选出简单问题,并将更多计算分配给解决更难的问题。我们研究了影响NGU的设计选择,如离策略鲁棒性,并制定了一套最佳实践。在数学基准Deepscaler上,NGU提升了每单位计算资源的性能,尤其是在更困难的问题上。在最近的编码任务Manufactoria中,使用逐测试奖励的标准GRPO未能完全解决那些包含一系列简单和困难测试的问题。NGU则迭代改进,解决越来越难的测试,直到学会完全解决编码问题。

英文摘要

We demonstrate that training LLMs with RL does not improve performance equally across a dataset. RL shows large improvements on easy problems that an LLM is already good at solving, but small improvements on hard problems. We call this the Matthew Effect in RL for LLMs, after the phenomenon of cumulative advantage from economics and network science summarized as "the rich get richer". The naive explanation is that hard problems require more compute to find a solution. We argue that modern RL methods are exacerbating the issue by wasting too much compute on easy problems and instead should dynamically reallocate how they use compute. We introduce Never Give Up (NGU), a simple adaptive sampling method that keeps generating samples for a problem until one is correct. By leveraging asynchronous RL, this naturally uses fewer samples to filter out easy problems and allocates more compute to solving harder problems. We investigate the design choices that affect NGU, such as off-policy robustness, and develop a set of best practices. On the math benchmark Deepscaler, NGU improves performance per compute, especially on harder problems. On a recent coding task, Manufactoria, standard GRPO with a per-test reward fails to fully solve problems that have a range of easy and difficult tests. NGU iteratively improves, solving harder and harder tests, until it learns to fully solve coding problems.

CommentsBlog post mnoukhov.github.io/posts/ngu and code available at github.com/mnoukhov/never-give-up

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑