arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17310cs.LG

Agentic ESOpt:以最小GPU资源微调长视野大语言模型智能体

Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements

Zhi Zheng, Rongsheng Chen, Yunpeng Ba, Zhenkun Wang, Yee Whye Teh, Wee Sun Lee

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出Agentic ESOpt框架,用进化策略微调长视野LLM智能体,仅需最小GPU资源,在WebArena-Lite上使Qwen-3.5-27B性能提升6.69%,提示-参数协同演化在多数设置中优于基准。

中文摘要 AI 辅助

强化学习(RL)在单轮大语言模型(LLM)微调中已展现出应用前景,但长视野智能体推理会产生日益增多的分支交互与稀疏奖励,暴露出RL的若干局限:其基于反向传播的训练栈开销过大,难以用于微调更大规模的LLM,且更长视野的轨迹会使RL中的信用分配难度大幅提升。本文认为进化策略(ES)是微调长视野LLM智能体的更优选择。与智能体RL相比,ES具备三项核心优势:1)模型可扩展性:ES仅需最小的推理级GPU内存即可实现全参数优化,使微调大型LLM成为可能;2)灵活性:其轻量型黑盒反馈接口使ES微调易于与提示空间演化(如技能优化、测试时计算)结合;3)长视野可扩展性:ES执行轨迹级参数归因,无需在各视野间分解奖励,随视野长度增加,其可扩展性优于智能体RL。基于这一见解,本文提出Agentic ESOpt,这是一款专为灵活的参数-上下文协同演化设计的全参数智能体微调框架。每一步中,Agentic ESOpt会在当前LLM参数附近采样扰动,通过奖励评估生成的智能体,并应用在线奖励加权更新。为优化探索-适应的权衡,Agentic ESOpt进一步引入了扰动尺度σ的余弦衰减调度。在WebArena-Lite上,Qwen-3.5-27B的全参数优化使无技能基准提升了6.69%;在测试时自动启发式设计任务中,Agentic ESOpt执行在线提示-参数协同演化,在36个设置中的28个里优于其匹配基准。

英文摘要

Reinforcement Learning (RL) has been promising in single-turn LLM fine-tuning. However, long-horizon agentic reasoning introduces increasingly branching interactions and sparse rewards, exposing several limitations of RL: its heavyweight backpropagation-based training stack makes it impractical to fine-tune larger LLMs, and longer-horizon trajectories make credit assignment in RL substantially harder. This paper argues that evolution strategies (ES) can be a better choice for fine-tuning long-horizon LLM agents. Compared with agentic RL, ES offers three key advantages: 1) Model Scalability: ES enables full-parameter optimization with only minimal, inference-level GPU memory, making it possible to fine-tune large LLMs. 2) Flexibility: its lightweight, black-box feedback interface makes ES fine-tuning easy to compose with prompt-space evolution (e.g., skill optimization & test-time compute); and 3) Long-Horizon Scalability: ES performs trajectory-level parameter attribution without decomposing rewards across horizons, yielding better scalability than Agentic RL as the horizon length grows. Based on this insight, we propose Agentic ESOpt, a full-parameter agentic fine-tuning framework tailored to flexible parameter--context co-evolution. At each step, Agentic ESOpt samples perturbations around the current LLM parameters, evaluates the resulting agents with rewards, and applies an online reward-weighted update. To improve the exploration--adaptation trade-off, Agentic ESOpt further introduces a cosine decay schedule of the perturbation scale $σ$. On WebArena-Lite, full-parameter optimization of Qwen-3.5-27B improves the No Skill baseline by 6.69%. In test-time automatic heuristic design, Agentic ESOpt performs online prompt--parameter co-evolution, improving its matched baseline in 28 of 36 settings.

发表机构

  • National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

↑