发表机构
Peking University; Tencent; University of Edinburgh; Northwestern University; Tsinghua University(北京大学; 腾讯; 爱丁堡大学; 西北大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出无需训练的SE-GoS框架,通过拓扑、边权重和描述三种演化从执行轨迹改进技能图,在SkillsBench上提升奖励并减少输入令牌,优于静态基线。
AI 中文摘要
现代大语言模型智能体越来越依赖可复用的技能,然而随着技能库扩展到数千个条目,有效的检索成为瓶颈。技能图(Graph-of-Skills, GoS)通过利用依赖感知的图结构实现可扩展的技能检索来解决这一挑战,而SkillDAG进一步证明了技能图可以在线积累基于执行的结构。然而,这些方法未解决是否可以将历史执行轨迹系统地提炼为更好的检索图,从而泛化到未见任务的问题。我们提出了自演化技能图(Self-Evolving Graph-of-Skills, SE-GoS),这是一个无需训练的框架,它从执行轨迹中演化现有的GoS图,同时保留原始检索流程。SE-GoS执行三种互补的更新:拓扑演化,从执行证据中发现和修剪技能关系;边权重演化,基于历史有效性强化检索相关的关系;以及描述演化,利用执行反馈优化面向检索的技能描述。在SkillsBench上使用三种大语言模型,SE-GoS相对于完整技能加载一致地提高了任务奖励,同时减少了输入令牌,增益因模型家族而异。在一个代表性设置中,一轮演化将奖励从52.4%提高到59.4%,同时相对于完整技能加载将输入令牌减少约三分之一,并且生成的图迁移到不相交的保留分割上,比静态GoS基线提高了5.4个百分点。这些结果表明,技能图可以从执行经验中改进,而无需模型训练、检索算法更改或技能内容修改,将静态检索图转变为演进的检索基础设施。
英文摘要
LLM agents use large libraries of reusable skills. At thousands of skill entries, retrieval becomes the bottleneck. Graph-of-Skills (GoS) retrieves dependency-aware bundles from a typed skill graph, and SkillDAG shows that such a graph can accumulate execution-backed structure online. Neither asks whether execution traces can be distilled into a better retrieval graph that generalizes to unseen tasks. We present \textbf{Self-Evolving Graph-of-Skills (SE-GoS)}, which treats the retrieval graph as an index rather than a learned representation: the graph is maintained from execution traces while the retrieval pipeline, the skill library, and the model stay fixed. SE-GoS applies three updates: (1) \textbf{topology}, which induces relations from execution evidence and retracts an avoid edge only after repeated successful co-use; (2) \textbf{edge-weight}, which softly attenuates unsupported semantic edges and reinforces incoming edges to used skills; and (3) \textbf{node-description}, which updates retrieval-facing descriptions stored on graph nodes ranked too low. On SkillsBench, one evolution round lifts average reward from 52.4\% to 59.4\%, above full-library loading, vector retrieval, static GoS, and SkillDAG, and this ordering repeats on all three backbones. Retrieval over the evolved graph spends about two-thirds of the input tokens that loading the full library costs. Repeating the round does not help. The same graph improves a held-out split it never saw from 52.9\% to 58.3\%, so what it accumulates transfers rather than memorizes traces. Skill graphs can therefore be improved from execution experience without model training, retrieval-algorithm changes, skill-content modifications, or a model judging which skills are related.
Comments19 pages, 1 figure, 7 tables