Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
机构 * Qwen Team, Alibaba Inc.(通义团队,阿里巴巴公司) ; LeapLab, Tsinghua University(清华大学跃实验室)
专题命中 其他推理 :reasoning(title,abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted to NeurIPS 2025. 25 pages, 17 figures, 2 tables