arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SkillProx:基于近端文本梯度下降的自进化智能体技能

SkillProx: Self-Evolving Agent Skills via Proximal Textual Gradient Descent

Mingxuan Zheng, Yujin Zhou, Chuxue Cao, Boqin Yin, Yuyao Zhang, Jiapeng Sun, Shuaishuai Gong, Sirui Han, Yike Guo

arXiv 2608.07449首次发表:更新:

AI 中文总结

SkillProx是结合闭环诊断进化与近端优化的自进化智能体技能框架,可提升LLM智能体在分布内外基准的平均准确率3.0个百分点,且两种核心组件具有互补效果。

AI 中文摘要

大型语言模型(LLM)智能体正日益通过积累技能中的程序性知识来适应重复性任务,这些技能是轻量、可复用的文本制品,会被加载到智能体的上下文窗口中且无需更新模型权重。现有方法通过迭代任务执行、失败诊断和轨迹引导的文本空间更新来优化技能,但这些框架缺乏明确的诊断-结果反馈,且将删除操作视为通用编辑,而非用于巩固积累知识的专用机制。本文提出SkillProx,这是一种受近端梯度启发的前向-后向框架,将闭环诊断进化与感知效用的近端优化相结合。受平衡任务损失与技能复杂度的复合目标驱动,前向阶段在同一任务批次上重新执行诊断驱动的编辑,回退性能下降的情况,并将测得的结果反馈至后续诊断;后向阶段将生成的技能分解为可审计的知识单元,使用冻结的留一法效用审计估计各单元的贡献,并应用经验证门控的巩固、降级或移除操作。在多个骨干LLM的分布内与分布外基准测试中,实验结果显示SkillProx相较于最强的基于梯度的基线方法,平均准确率提升了3.0个百分点;组件消融实验则证明了闭环诊断与近端优化的互补作用。

英文摘要

LLM agents increasingly adapt to recurring tasks by accumulating procedural knowledge in skills. These skills are lightweight, reusable textual artifacts that are loaded into the agent's context without weight updates. Recent methods refine skills through iterative task execution, failure diagnosis, and trajectory-guided text-space updates. However, existing frameworks lack explicit diagnosis--outcome feedback and treat deletion as a generic edit operation rather than a dedicated mechanism for consolidating accumulated knowledge. We introduce SkillProx, a proximal-gradient-inspired forward--backward framework that couples closed-loop diagnostic evolution with utility-aware proximal refinement. Motivated by a composite objective balancing task loss and skill complexity, the forward stage re-executes diagnosis-driven edits on the same task batch, rolls back regressions, and feeds measured outcomes into subsequent diagnoses. The backward stage decomposes the resulting skill into auditable knowledge units, estimates their contributions using a frozen leave-one-out utility audit, and applies validation-gated consolidation, demotion, or removal. Experiments on in-distribution and out-of-distribution benchmarks across multiple backbone LLMs show that SkillProx improves average accuracy by 3.0 percentage points over the strongest gradient-based baseline. Component ablations demonstrate the complementary effects of closed-loop diagnosis and proximal refinement.

Comments23 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑