发表机构
Southern University of Science and Technology; National University of Singapore; Nanjing University(南方科技大学; 新加坡国立大学; 南京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出Hyper-ES框架,通过梯度微调获取下降方向张成子空间,再用CMA-ES优化DARE-TIES合并系数,在数学推理任务上性能优于GRPO-LoRA,梯度更新量减少10%。
AI 中文摘要
进化策略(Evolution Strategy, ES)是资源受限场景下大语言模型(Large Language Model, LLM)推理的一种有前景的替代梯度微调的方案,但直接将ES应用于数十亿参数规模的LLM时效果极差。在这类高维参数空间中,大多数随机扰动几乎与有用的更新方向正交,导致优化不稳定。本文提出Hyper-ES,这是一种基于子空间的ES框架,既避免了ES在全参数搜索中的缺陷,又发挥了其在低维优化中的优势。Hyper-ES不要求ES从LLM参数空间的随机扰动中发现有用方向,而是先执行少量低成本的基于梯度的微调,以获取下降方向;尽管每个方向单独只能提供有限的改进,但它们张成的空间形成了一个紧凑的适配子空间,可捕获有用的推理更新。随后,Hyper-ES应用CMA-ES在该子空间内优化分层DARE-TIES合并系数,使ES能够在有意义的下降方向的组合上搜索,而非在任意全模型扰动上搜索。我们在三个Qwen2.5-Instruct和DeepSeek-R1-Distill主干模型上,针对六个数学推理数据集对Hyper-ES进行评估,结果显示Hyper-ES的表现始终优于GRPO-LoRA,性能提升1%,同时所需的空间密集型梯度更新量减少10%。代码见此https://URL。
英文摘要
Evolution Strategy (ES) is a promising alternative to gradient-based fine-tuning for resource-constrained Large Language Model (LLM) reasoning. However, directly applying ES to billion-parameter LLMs is highly ineffective. In such high-dimensional parameter spaces, most random perturbations are nearly orthogonal to useful update directions, leading to unstable optimization. We propose Hyper-ES, a subspace-based ES framework that avoids the weakness of ES in full-parameter search while exploiting its strength in low-dimensional optimization. Instead of asking ES to discover useful directions from random perturbations in the LLM parameter space, Hyper-ES first performs a small number of inexpensive gradient-based fine-tuning runs to obtain descent directions. Although each direction may provide only a limited improvement on its own, their span forms a compact adaptation subspace that captures useful reasoning updates. Hyper-ES then applies CMA-ES to optimize layer-wise DARE-TIES merging coefficients within this subspace, allowing ES to search over combinations of meaningful descent directions rather than over arbitrary full-model perturbations. We evaluate Hyper-ES on three Qwen2.5-Instruct and DeepSeek-R1-Distill backbones across six mathematical reasoning datasets. Results show that Hyper-ES consistently outperforms GRPO-LoRA by 1% while requiring 10% fewer space-consuming gradient updates. Code at https://github.com/kuangrepi/Hyper-ES.
Comments19 pages, 4 figures, 14 tables. Code: https://github.com/kuangrepi/Hyper-ES