谁在维护智能体技能?一项关于人类主导、AI辅助技能维护的纵向研究
Who Maintains Agent Skills? A Longitudinal Study of Human-Governed, AI-Assisted Skill Maintenance
浏览论文内容
中文总结 AI 辅助
本研究通过挖掘五个公共AI技能仓库的873次提交,揭示技能维护是人类主导、AI辅助的循环,而非自主流水线,并发布了语料库和编码手册。
中文摘要 AI 辅助
终身大语言模型智能体越来越依赖外部技能工件,作为随时间保存和复用能力的一个要素。这些技能(通常是可移植的Markdown文件,如 http this URL 所示)描述了何时以及如何应用某项能力,并且随着部署过程中工具和使用模式的变化,必须对其进行修正、扩展和整合。近期工作试图自动化技能策展,但主要针对自动化基线进行评估,并将人类维护视为一个未测量的瓶颈。我们直接研究这一缺失的过程。我们挖掘了五个公共AI技能仓库的完整提交历史,这是一个有目的的AI工具组织样本,涵盖2025年10月至2026年6月间的873次提交、143个技能文件和254次实质性的创建后编辑。我们使用预注册的治理、操作和触发证据编码手册对每次编辑进行编码。发现了三个结果。首先,每次实质性编辑都由一个具名的人类账户撰写或合并,而62%带有AI共同作者尾注,且仓库间差异很大。其次,这些编辑是真正的策展:一个经审计的样本显示,大多数编辑改变了技能内容,且编码操作以添加和修正为主。第三,一个预注册的规则相似性轴未能通过其可靠性门槛;从提交工件中可靠地编码规则相似性仍然是一个开放的测量问题。我们发布了语料库、编码手册、挖掘脚本以及一个供自动化技能策展者使用的重放协议。对于自我进化的智能体而言,公共技能维护目前看起来不像一个自主流水线,而更像一个人类主导、AI辅助的循环,未来的策展者必须以此为基准并在其中运作。
英文摘要
Lifelong LLM agents increasingly rely on external skill artifacts as one element for preserving and reusing capabilities over time. These skills (usually portable Markdown files such as SKILL.md) describe when and how to apply a capability and must be corrected, expanded, and consolidated as tools and usage patterns shift over deployment. Recent work seeks to automate skill curation, but it largely evaluates against automated baselines and treats human maintenance as an unmeasured bottleneck. We study that missing process directly. We mine the full commit histories of five public AI-skill repositories, a purposive sample of AI-tooling organizations, covering 873 commits, 143 skill files, and 254 substantive post-creation edits from October 2025 to June 2026. We code each edit with pre-registered governance, operation, and trigger-evidence codebooks. Three findings emerge. First, every substantive edit is authored or merged through a named human account, while 62% carry an AI co-author trailer, with large repository-level variation. Second, these edits are genuine curation: an audited sample shows that most change skill content, and the coded operations are dominated by additions and corrections. Third, a pre-registered rule-likeness axis fails its reliability gate; reliably coding rule-likeness from commit artifacts remains an open measurement problem. We release the corpus, codebooks, mining scripts, and a replay protocol for automated skill curators. For self-evolving agents, public skill maintenance currently looks less like an autonomous pipeline than a human-governed, AI-assisted loop that future curators must measure against and operate within.
发表机构
- Megagon Labs
机构由 AI 辅助整理,请以论文原文为准。