AI 中文总结
针对资源受限智能体LLM后训练的高内存、长耗时问题,提出CoPES方法,在GPU小时预算下,其验证准确率恢复率、内存效率均优于标准ES和LoRA-GRPO,在多任务基准上表现更优。
AI 中文摘要
使用工具的大语言模型(LLM)智能体会产生长且多轮的轨迹,这使得基于梯度的后训练内存密集。进化策略(ES)支持无需反向传播的内存高效全参数后训练,最终可达到基于梯度的强化学习(RL)的性能。然而,资源受限场景通常仅提供少量GPU,因此ES的高GPU小时需求会导致训练时间过长,令人难以承受。为解决该问题,我们提出协同参数子空间进化策略(CoPES),这是一种协同协同进化方法,它将全参数空间分解为低维子空间,并对其进行协同搜索以提高优化效率。我们对Qwen3.5-4B工具使用智能体进行数学任务后训练,并在5个不同难度的基准上进行评估。在全参数GRPO最佳验证检查点的GPU小时预算下,CoPES恢复了GRPO 92%的验证准确率提升,而标准ES仅恢复67%;同时,其理论GPU内存需求不到全参数GRPO的八分之一。在5个基准的所有评估pass@k指标上,它始终优于标准ES和基于LoRA的GRPO。额外实验进一步显示CoPES在问答任务上的优势。这些结果表明,在资源受限条件下,智能体LLM后训练的内存需求与训练时间之间的权衡得到了改善。代码开源在该httpsURL。
英文摘要
Tool-using large language model (LLM) agents produce long, multi-turn trajectories, making gradient-based post-training memory-intensive. Evolution strategies (ES) enable memory-efficient full-parameter post-training without backpropagation and can eventually match the performance of gradient-based reinforcement learning (RL). However, resource-constrained settings typically offer only a few GPUs, so the high GPU-hour requirements of ES translate into prohibitively long training times. To address this, we introduce Cooperative Parameter-subspace Evolution Strategy (CoPES), a cooperative coevolutionary method that decomposes the full parameter space into lower-dimensional subspaces and searches over them cooperatively to improve optimization efficiency. We post-train a Qwen3.5-4B tool-using agent for the math task and evaluate it on five benchmarks of varying difficulty. Under the GPU-hour budget of full-parameter GRPO's best validation checkpoint, CoPES recovers 92% of GRPO's validation-accuracy gain, versus 67% for standard ES, while its theoretical GPU memory requirement is less than one-eighth that of full-parameter GRPO. It consistently outperforms standard ES and LoRA-based GRPO on all evaluated pass@k metrics across the five benchmarks. Additional experiments further show the advantage of CoPES on the question-answering task. These results demonstrate an improved trade-off between memory requirements and training time for agentic LLM post-training under resource constraints. The code is open-sourced in https://github.com/MetaronWang/CoPES
Comments14 pages,9 figures, submit to AAAI 2027