arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

动态智能体技能:不断演进的技能库的生命周期调查与分类法

Dynamic Agent Skills: A Lifecycle Survey and Taxonomy of Evolving Skill Libraries

Yubo Li

arXiv 2607.10113首次发表:更新:

发表机构

Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究大型语言模型智能体中技能库随时间的变化,通过分类法、生命周期架构和技能记录模式等工具组织文献,合成证据分级模式,指出相关问题并提出报告标准及开放问题。

AI 中文摘要

大型语言模型智能体越来越多地在模型之外存储可重复使用的程序,这些程序通常被称为“技能”。本分类法驱动的调查探讨了此类技能库如何随时间变化。在2023年至2026年的124篇论文审核集中,我们将动态技能系统综合为“生命周期管理、经过验证、不断演进的工件存储”。智能体从交互中收集证据,提出技能更新,验证并接纳候选技能,对其进行组织以便检索和组合,修复或修剪陈旧条目,并通过出处和回滚来管理共享。我们围绕三种调查工具组织文献:六种意义的分类法区分当前论文中结构不同的“技能”工件;八阶段生命周期架构识别证据获取、提议、验证/接纳、存储、检索/组合、维护、提炼/可移植性和治理背后反复出现的设计决策;轻量级技能记录模式和十种操作符词汇提供比较库更新的通用术语。我们合成了带有明确警告的证据分级模式:接纳和修复反复重要,验证器质量对基于技能的强化学习有重大影响,随着库的增长,平面检索可能会退化,当前基准仍未充分报告库的轨迹、使用-效用差距和安全表面。最后,我们提出了具体的报告标准和开放问题,用于将动态技能评估为不断变化的库,而不是静态提示或工具集合。

英文摘要

Large language model agents increasingly store reusable procedures outside the model. These reusable procedures are often called \emph{skills}: they may be code functions, natural-language instructions, SKILL.md packages, workflow graphs, or learned adapters that a future agent can retrieve and invoke. This taxonomy-driven survey asks how such skill libraries change over time. Across a $124$-paper $2023$--$2026$ audit set, we synthesize dynamic skill systems as \emph{lifecycle-managed, verified, evolving artifact stores}: agents collect evidence from interaction, propose skill updates, verify and admit candidates, organize them for retrieval and composition, repair or prune stale entries, and govern sharing through provenance and rollback. We organize the literature around three survey tools. First, a $\text{six}$-sense taxonomy distinguishes the structurally different artifacts called ``skills'' in current papers. Second, an $\text{eight}$-stage lifecycle architecture identifies the recurring design decisions behind evidence acquisition, proposal, verification/admission, storage, retrieval/composition, maintenance, distillation/portability, and governance. Third, a lightweight skill-record schema and $\text{ten}$-operator vocabulary provide common terms for comparing library updates without elevating them into a separate method contribution. Using this structure, we synthesize evidence-graded patterns with explicit caveats: admission and repair are repeatedly important, verifier quality materially affects skill-aware RL, flat retrieval can degrade as libraries grow, and current benchmarks still under-report library trajectories, usage--utility gaps, and safety surfaces. We close with concrete reporting standards and open problems for evaluating dynamic skills as changing libraries rather than static prompt or tool collections.

CommentsAccepted by TMLR (2026.07), OpenReview Link: https://openreview.net/forum?id=cjU3YbcRr8

Journal refTransactions on Machine Learning Research, 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑