arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.13394cs.CLcs.LG

GFlowRL:将分布匹配强化学习扩展到大型语言模型

GFlowRL: Scaling Distribution-Matching RL to Large Language Models

Xiaodong Liu, Michael Xu, Jack W. Stokes, Paul Smolensky, Doug Burger, Jianfeng Gao

首次发表
浏览论文内容

中文总结 AI 辅助

研究旨在将GFlowNet风格的RL扩展到大型语言模型,提出GFlowRL算法,去除辅助分区网络,用批内蒙特卡罗估计替代学习的分区函数,并通过两个稳定器实现奖励分布匹配,在多个基准测试中表现出色,能稳定扩展到不同架构。

中文摘要 AI 辅助

生成流网络(GFlowNets)为大型推理模型提供了一种有前景的替代奖励最大化强化学习(RL)的方法,通过匹配奖励分布鼓励多样化推理路径。近期工作在数学和代码方面有进展,但将GFlowNet风格的RL扩展到现代训练后管道仍困难。经系统分析发现,可由训练所需的展开组计算的批内蒙特卡罗估计替代学习的分区函数。我们提出GFlowRL,一种简化的GFlowNet风格RL算法,去除了辅助分区网络,通过两个稳定器实现奖励分布匹配目标。GFlowRL在数学、代码和对抗性红队基准测试中超越所有对手,在14B规模达到Codeforces评级2048,在AdvBench和HarmBench上获得最高平均ASR@1,优于先前SOTA多轮攻击者。该方法可扩展到高达235B参数的所有评估的混合专家(MoE)配置。据我们所知,GFlowRL是首个能在密集和稀疏架构上稳定扩展的GFlowNet风格RL算法。代码将在:此https URL

英文摘要

Generative Flow Networks (GFlowNets) offer a promising alternative to reward-maximizing reinforcement learning (RL) for large reasoning models, encouraging diverse reasoning paths by matching reward distributions rather than collapsing to dominant modes. Recent work shows promise on math and code, but scaling GFlowNet-style RL to modern post-training pipelines remains difficult: as model size, rollout horizon, reward noise, and distributed-systems complexity grow together, a learned prompt-conditional partition function becomes a source of gradient instability and engineering overhead rather than a useful normalizer. Through systematic analysis, we find that the learned partition function, previously treated as essential, can be replaced by an in-batch Monte Carlo estimate computed from the rollout group already required for training. We propose GFlowRL, a streamlined GFlowNet-style RL algorithm that removes the auxiliary partition network entirely while preserving the reward-distribution-matching objective, completed by two stabilizers: importance-sampling correction for rollout/trainer drift and asymmetric flow-gap clipping for outlier residuals. GFlowRL exceeds all counterparts on math, code, and adversarial red-teaming benchmarks, reaching a Codeforces rating of 2048 at the 14B scale (within 25 Elo of o3-mini) and attaining the highest average ASR@1 on AdvBench and HarmBench, outperforming the previous SOTA multi-turn attacker in a regime where FlowRL, a prior GFlowNet-style method, diverges. The same recipe transfers to all evaluated MoE configurations up to 235B parameters, where FlowRL again fails to converge. To our knowledge, GFlowRL is the first GFlowNet-style RL algorithm to scale stably across both dense and sparse architectures. Code will be at: https://github.com/microsoft/gflowrl

发表机构

  • Microsoft Research(微软研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑