arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越能力边界:用于自进化大语言模型智能体的零阶优化

Beyond the Capability Boundary: Zeroth-Order Optimization for Self-Evolving LLM Agents

Bingzhen Liu, Xiaomeng Fan, Yuwei Wu, Zhi Gao, Mingyang Gao, Chuanhao Li, Yunde Jia

arXiv 2608.09292首次发表:更新:

发表机构

Beijing Institute of Technology; Shenzhen MSU-BIT University; Alaya Lab(北京理工大学; 深圳北理莫斯科大学; 阿莱亚实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出零阶自进化框架,通过扰动LLM的LoRA参数形成闭环自进化循环,结合并行扰动推理等机制,在多基准实验中显著提升了LLM智能体在困难样例上的表现。

AI 中文摘要

自进化方法通过从底层大语言模型(LLM)中采样轨迹并学习这些轨迹来提升LLM智能体的能力。然而,这些方法难以学习超越智能体固有能力边界的内容,因为智能体无法在困难样例上采样到正确的轨迹以实现进一步改进。在本文中,我们提出了一种零阶自进化框架,该框架通过扰动LLM参数使智能体能够在无任何轨迹标注的情况下适配困难样例,从而超越其能力边界。具体而言,我们扰动LLM的LoRA参数,运行智能体,计算扰动后与原始参数下的损失,并利用损失差估计梯度以进一步更新LoRA参数。我们使用更新后的LLM采样轨迹进行监督微调,以打破智能体的能力边界,形成闭环自进化循环。我们引入并行扰动推理机制和自适应查找机制来降低零阶优化的时间消耗,同时采用困惑度(perplexity)损失,其能提供平滑且稳定的零阶损失值。在多个深度研究基准上的实验表明,我们的方法获得了显著更多的成功轨迹,且始终优于强基线,尤其在困难样例上表现突出。代码及已发布的制品可在该https URL获取。

英文摘要

Self-evolving methods improve the capabilities of LLM agents by sampling trajectories from the underlying LLMs and learning from these trajectories. However, these methods struggle to learn beyond the inherent capability boundary of the agents, since the agents cannot sample correct trajectories on difficult examples for further improvements. In this paper, we propose a zeroth-order self-evolution framework that enables agents to learn beyond their capability boundary by perturbing LLM parameters to adapt to difficult examples without any trajectory annotations. Specifically, we perturb LoRA parameters of LLMs, run the agent, compute the losses under the perturbed and original parameters, and use the loss difference to estimate gradients and further update the LoRA parameters. We sample trajectories using the updated LLMs for supervised fine-tuning to break through the capability boundary of the agents, forming a closed self-evolution loop. We introduce a parallel perturbation inference mechanism and an adaptive lookup mechanism to reduce time consumption in zeroth-order optimization, with an answer perplexity loss that provides smooth and stable zeroth-order loss values. Experiments on multiple deep research benchmarks show that our method obtains substantially more successful trajectories and consistently outperforms strong baselines, especially on difficult examples. The code and released artifacts are available at https://github.com/hidk1911/ZOForLLMAgents.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑