arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

理解用于大语言模型推理的进化策略:比GRPO更广泛的推理覆盖范围

Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO

Yunpeng Ba, Zhi Zheng, Yue Xie, Jiaqing Li, Xialiang Tong, Tao Zhong, Mingxuan Yuan, Zhichao Lu, Xuyang Wu, Zhenkun Wang

arXiv 2608.27351首次发表:更新:

发表机构

Southern University of Science and Technology; National University of Singapore; Harbin Institute of Technology, Weihai; Huawei Noah’s Ark Lab; City University of Hong Kong(南方科技大学; 新加坡国立大学; 哈尔滨工业大学(威海); 华为诺亚方舟实验室; 香港城市大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究用于LLM推理的进化策略(ES),发现ES比GRPO有更广泛的推理覆盖范围,开发了结合两者优势的GRPO-ES策略,还揭示了ES的功能稀疏性等特性,确立其为独特的推理后训练范式。

AI 中文摘要

进化策略(Evolution Strategies, ES)近期成为一种内存高效的大语言模型(Large Language Model, LLM)推理后训练范式。然而,ES的优化行为仍未得到充分研究,因此难以明确其与主流后训练范式(如组相对策略优化(Group Relative Policy Optimization, GRPO))相比的优势范围。通过系统研究ES的动态特性与机制,本文首先从理论和实证层面明确了ES相较于GRPO的性能优势:ES可实现更广泛的推理覆盖范围,从而更好地挖掘预训练LLM的推理能力。理论上,我们证明了ES群体中验证器投影的Jensen-Shannon多样性有助于提升Pass@K性能;实证上,与表现出熵崩溃的GRPO不同,ES在提升Pass@1的同时,获得了比GRPO更高的Pass@K。我们进一步开发了一种顺序GRPO-ES训练策略,结合GRPO在Pass@1上的优势与ES在Pass@K上的增益。其次,我们发现尽管整体模型参数发生了显著漂移,但ES的任务性能增益仅由一小部分幅度较大的更新贡献。这种功能稀疏性表明,大规模参数移动并不一定意味着功能的广泛变化,而保留的评估进一步显示其不一定会导致灾难性遗忘。最后,我们研究了超参数设计对ES有效性的影响,证明在更大的LLM中,ES需要更小的群体规模。这些发现确立了ES是一种独特的推理后训练范式,而非GRPO的低效、内存高效替代方案。

英文摘要

Evolution Strategies (ES) have recently emerged as a memory-efficient post-training paradigm for LLM reasoning. However, the optimization behavior of ES remains understudied, making it hard to define its advantage scope compared to mainstream post-training paradigms (e.g., Group Relative Policy Optimization (GRPO)). By systematically investigating ES dynamics and mechanisms, this paper first identifies a performance advantage of ES over GRPO, theoretically and empirically showing that ES can lead to broader reasoning coverage, thereby better exploiting the reasoning capabilities of pretrained LLMs. Theoretically, we show that verifier-projected Jensen-Shannon diversity across the ES population is helpful to higher Pass@K performances. Empirically, unlike GRPO, which exhibits entropy collapse, ES improves Pass@1 while attaining higher Pass@K than GRPO. We further develop a sequential GRPO-ES training strategy that combines GRPO's strength in Pass@1 with ES's gains in Pass@K. Second, we find that despite substantial whole-model parameter drift, the task-performance gains of ES are only contributed to a sparse subset of larger-magnitude updates. This functional sparsity suggests that large parameter movement need not imply widespread functional change, and held-out evaluations further show that it does not necessarily lead to catastrophic forgetting. Finally, we study how hyperparameter design affects the effectiveness of ES, demonstrating that ES requires a smaller population size in a larger LLM. These findings position ES as a distinct reasoning post-training paradigm rather than a less effective, memory-efficient alternative to GRPO.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑