发表机构
SGIT AI Lab(SGIT AI实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DART-ES通过难度感知重加权与定向重放改进进化策略微调,无需额外模型或反向传播,在多个基准上提升准确率并显著降低计算开销。
AI 中文摘要
进化策略(ES)仅使用前向计算即可实现大语言模型(LLMs)的全参数高效微调。然而,标准ES在所有问题上均匀地平均奖励,并将问题级别的种群反馈压缩为单个标量,这使得难以捕捉每个问题的学习价值如何随模型能力而变化。为解决这一局限性,我们提出了用于进化策略的难度感知重加权与定向重放(DART-ES)。DART-ES根据每个问题在扰动种群中的通过率来估计其局部可解性,并聚合历史观测以构建动态难度状态。这一共享状态共同指导连续难度重加权和稀有可解样本的重放,从而在不引入额外难度模型或反向传播的情况下改进扰动方向评估和训练数据分配。大量实验表明,DART-ES取得了良好的微调性能。DART-ES在所有五个基础模型上均优于ES,并将平均准确率从72.07%提升至73.53%,超过了GRPO在GSM8K上达到的73.26%。在五个具有挑战性的数学推理基准上,DART-ES的平均准确率达到49.20%,而ES为48.34%,并且与使用RL训练的强7B模型保持竞争力。进一步实验表明,在指令微调、代码生成和14B模型的Countdown任务上均取得一致改进,展示了跨任务的强泛化能力和对更大模型的可扩展性。除性能提升外,DART-ES在系统效率方面也显示出明显优势。与GRPO相比,它每步运行时间减少了15.2%–50.2%,每GPU峰值内存使用量减少了21.1%–51.1%。尽管执行全参数微调,DART-ES所需的运行时间和GPU内存也少于GRPO+LoRA。
英文摘要
Evolution Strategies (ES) enable memory efficient full parameter fine-tuning of large language models (LLMs) using only forward computation. However, standard ES uniformly averages rewards across problems and compresses problem level population feedback into a single scalar, making it difficult to capture how the learning value of each problem changes with model capability. To address this limitation, we propose Difficulty-Aware Reweighting and Targeted Replay for Evolution Strategies (DART-ES). DART-ES estimates the local solvability of each problem from its pass rate across the perturbation population and aggregates historical observations to construct a dynamic difficulty state. This shared state jointly guides continuous difficulty reweighting and rare solvable sample replay, thereby improving perturbation direction evaluation and training data allocation without introducing an additional difficulty model or backpropagation. Extensive experiments show that DART-ES achieves good fine-tuning performance. DART-ES outperforms ES on all five base models and improves the average accuracy from 72.07\% to 73.53\%, exceeding the 73.26\% achieved by GRPO on GSM8K. Across five challenging mathematical reasoning benchmarks, DART-ES achieves an average accuracy of 49.20\%, compared with 48.34\% for ES and remains competitive with strong 7B models trained with RL. Further experiments show consistent gains in instruction tuning, code generation and the Countdown task with a 14B model, demonstrating strong generalization across tasks and scalability to larger models. Beyond performance gains, DART-ES also shows clear advantages in system efficiency. It reduces runtime per step by 15.2\%--50.2\% and peak memory usage per GPU by 21.1\%--51.1\% compared with GRPO. Despite performing full parameter fine-tuning, DART-ES also requires less runtime and GPU memory than GRPO+LoRA.