发表机构
Renmin University of China; Tencent(中国人民大学; 腾讯)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对智能体技能自我进化中的不稳定和低效问题,提出Adam启发的SkillAdam框架,通过优化记忆和波动驱动的编辑预算实现稳定高效更新,在七个基准上达到最优性能。
AI 中文摘要
智能体技能提供了一种轻量级的方式,为冻结的语言模型智能体配备领域知识和程序性指导,然而获取高质量技能仍然成本高昂且难以扩展。专家编写的技能需要大量人力。最近的技能自我进化方法自动化了一个迭代循环,利用执行反馈来修订技能,但其启发式更新策略往往导致优化不稳定和迭代效率低下。我们识别出实现稳定高效的技能自我进化面临的两个挑战。方向稳定性要求有效修正能够累积,而不是被迭代局部的反馈覆盖。更新适应性要求每次修订的范围反映近期案例级改进的一致性。我们引入了SkillAdam,一个受Adam启发的框架,用于优化离散且不可微的技能文档。作为Adam一阶矩的功能类比,一个优化记忆记录了已识别的问题和先前解决方案尝试的结果,以稳定更新方向。作为Adam二阶矩的功能类比,一个波动驱动的编辑预算跟踪近期案例级改进的历史加权变化,并自适应地控制更新幅度。在跨越短时和长时任务的七个基准上,SkillAdam实现了最先进的性能,并具有更稳定的优化动态。它还以显著更少的优化迭代和更低的成本获得了更强的技能,优于先前方法。代码仓库:此https URL
英文摘要
Agent skills provide a lightweight way to equip frozen language-model agents with domain knowledge and procedural guidance, yet obtaining high-quality skills remains costly and difficult to scale. Expert-written skills require substantial human effort. Recent skill self-evolution methods automate an iterative loop that uses execution feedback to revise skills, but their heuristic update strategies often yield unstable optimization and low iteration efficiency. We identify two challenges in realizing stable and efficient skill self-evolution. Direction Stability requires effective corrections to accumulate rather than be overwritten by iteration-local feedback. Update Adaptivity requires the scope of each revision to reflect the consistency of recent case-level improvements. We introduce SkillAdam, an Adam-inspired framework for optimizing discrete and non-differentiable skill documents. As a functional analogue of Adam's first moment, an optimization memory records identified problems and the outcomes of prior solution attempts to stabilize the update direction. As a functional analogue of Adam's second moment, a volatility-driven edit budget tracks the history-weighted variation of recent case-level improvements and adaptively controls the update magnitude. Across seven benchmarks that span short- and long-horizon tasks, SkillAdam achieves state-of-the-art performance with more stable optimization dynamics. It also obtains stronger skills with substantially fewer optimization iterations and lower cost than prior methods. Code repository: https://github.com/ruc-datalab/SkillAdam
Comments17 pages, 4 figures, 6 tables