arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GMTS:基于梯度幅度的令牌选择改进用于大语言模型推理的RLVR训练

GMTS: Gradient Magnitude-based Token Selection Improves RLVR Training for LLM Reasoning

Outongyi Lv, Yuanwei Zhang, Xiaoqun Zhang

arXiv 2608.30632首次发表:更新:

发表机构

School of Mathematical Sciences, Shanghai Jiao Tong University; Institute of Natural Sciences, Shanghai Jiao Tong University(上海交通大学数学科学学院; 上海交通大学自然科学研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对RLVR训练中高熵令牌重要性机制不明的问题,提出GMTS方法量化令牌重要性,实验显示其在多领域、多模型规模下性能优于熵基选择。

AI 中文摘要

强化学习(RL),尤其是带有可验证奖励的强化学习(RLVR),近来已成为提升大语言模型(LLMs)推理能力的核心范式,在各类推理任务中展现出显著效果。近期研究表明,高熵令牌在模型训练中发挥着极为重要的作用,仅使用熵值最高的20%令牌进行训练即可带来显著的性能提升。然而,这类高熵令牌为何有益仍未得到充分理解。本研究发现,尽管单个答案内的高熵令牌往往与较大的梯度幅度相关,但考虑到答案级奖励信号的变化,仅熵值无法一致地反映不同答案间令牌的重要性。基于这一观察,我们引入了基于梯度幅度的令牌选择(GMTS)方法来量化令牌重要性,该方法利用熵-梯度关联来近似梯度幅度排名以进行令牌选择。我们发现,在三个推理领域及多种模型规模下,使用GMTS排名前20%的令牌进行训练,其性能始终优于基于熵的令牌选择,这表明GMTS能为RLVR训练提供更细粒度的令牌贡献估计。

英文摘要

Reinforcement learning (RL), particularly RL with Verifiable Rewards (RLVR), has recently emerged as a central paradigm for enhancing large language models' (LLMs) reasoning abilities, demonstrating remarkable effectiveness across reasoning tasks. Recent studies suggest that high-entropy tokens play an exceptionally important role in model training, since training with only the highest 20% entropy tokens yields significant performance gains. However, why such high-entropy tokens are beneficial remains insufficiently understood. In this work, we find that although high-entropy tokens within one answer tend to correlate with large gradient magnitude, entropy alone fails to consistently reflect token importance across different answers, considering the variations in the answer-level reward signals. Based on this observation, we introduce the Gradient Magnitude-based Token Selection (GMTS) method to quantify token importance, which leverages the entropy-gradient connection to approximate gradient-magnitude rankings for token selection. We find that training on the top 20% tokens ranked by GMTS consistently outperforms entropy-based token selection across three reasoning domains and various model sizes, suggesting that GMTS provides a more fine-grained estimate of token contribution for RLVR training.

CommentsFindings of the 2026 Conference on Empirical Methods in Natural Language Processing

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑