TGRL:面向大语言模型高效探索的温度分组强化学习
TGRL: Temperature-Grouped Reinforcement Learning for Efficient Exploration in LLMs
浏览论文内容
中文总结 AI 辅助
提出温度分组强化学习(TGRL),利用温度诱导的多样性作为训练信号,通过奖励对比和JS散度分配词元级信用,在不扩大采样预算下显著加速探索并提升多领域基准性能。
中文摘要 AI 辅助
在具有可验证奖励的强化学习(RLVR)中,高效探索常常仍是核心瓶颈。尽管温度控制与测试时缩放策略能够提升大语言模型(LLMs)的采样多样性,但它们要么在采样时扩大样本预算,要么未量化探索的收益。为此,我们提出温度分组强化学习(TGRL),将温度诱导的多样性转化为明确的训练信号。对于每个提示,TGRL将其采样组划分为低温和高温子集,通过它们的奖励对比估计探索增益,并利用同一logits下相应温度缩放的下一词元分布之间的Jensen-Shannon(JS)散度,将该组级信号分配为词元级信用。值得注意的是,在不扩大采样预算的情况下,TGRL达到与强RLVR基线相当的准确率,速度最高提升36%。在来自不同领域的11个基准测试中,TGRL广泛优于强RLVR基线:在32B规模下,六个数学基准的平均分提高1.6%;CodeForces评分提高196.7分,LiveCodeBench Pass@16提高4.4%;ALFWorld/WebShop成功率分别提高6.3%/4.9%。全面的消融实验和墙钟时间分析证实了所有提出组件的有效性。代码可在该https URL获取。
英文摘要
Efficient exploration often remains a central bottleneck in reinforcement learning with verifiable rewards (RLVR). Although temperature control and test-time scaling strategies can increase rollout diversity of large language models (LLMs), they either expand the sample budget at rollout time or leave the benefit of exploration unquantified. To this end, we propose Temperature-Grouped Reinforcement Learning (TGRL), which turns temperature-induced diversity into an explicit training signal. For each prompt, TGRL partitions its rollout group into low- and high-temperature subsets, estimates exploration gain through their reward contrast, and allocates this group-level signal as token-level credit using Jensen--Shannon (JS) divergence between the corresponding temperature-scaled next-token distributions induced by the same logits. Notably, TGRL reaches equivalent accuracy up to 36% faster than strong RLVR baselines without expanding the rollout budget. Across 11 benchmarks from diverse domains, TGRL broadly improves over strong RLVR baselines: it improves the six-benchmark math average by 1.6% at 32B, raises CodeForces rating by 196.7 points and LiveCodeBench Pass@16 by 4.4%, and improves ALFWorld/WebShop success rates by 6.3%/4.9%. Comprehensive ablations and wall-clock analysis confirm the efficacy of all proposed components. Code is available at https://github.com/1229095296/TGRL/tree/main.
发表机构
- University of Chinese Academy of Sciences(中国科学院大学)
- Meituan(美团)
- MAIS&NLPR, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所模式识别国家重点实验室与智能信息处理重点实验室)
机构由 AI 辅助整理,请以论文原文为准。