Repo2Skill-Evo:仓库技能在静默中过时
Repo2Skill-Evo: Repository Skills Go Stale in Silence
浏览论文内容
中文总结 AI 辅助
Repo2Skill-Evo研究发现,在57个真实仓库的105次版本转换中,前沿智能体难以可靠维护仓库技能,其平均@3 macro F1仅29.9%-69.7%,仓库技能会在无明确信号的情况下静默过时。
中文摘要 AI 辅助
大型语言模型(LLM)智能体越来越多地在不断演进的软件仓库上运行,其成功取决于仓库特定的程序知识:调用哪些API、运行哪些脚本、当前版本期望哪些约定。智能体技能将这种知识外部化为可重用单元,现有研究表明它们可以提升智能体性能,但这种提升是否持久仍不清楚。技能的版本特异性使其有用的同时也使其脆弱:发布后,它可能在没有任何明确信号的情况下过时,同时继续提供过时的指导,将知识外部化为技能会使其衰减变得不可见。我们研究智能体能否保持这种外部化知识的时效性。Repo2Skill-Evo将每个版本转换视为一项技能维护任务:给定V1技能集和官方V1到V2的补丁,智能体必须更新过时的技能内容,同时保留仍然有效的指导。在57个真实仓库和105个选定的版本转换中,每个评估的转换都会使部分V1技能集失效。然而,六个前沿智能体在基于补丁的删除指标下,平均@3 macro F1仅达到29.9%-69.7%,该指标平衡了过时内容的召回率与过度编辑的精确率。在运行中,两个相反的错误占主导:技能集中受影响文件的覆盖不完整,导致过时内容未被触及;而过度编辑与更高的召回率相关,但精确率更低。仓库技能在静默中过时,即使是前沿智能体也无法可靠地维护它们。
英文摘要
Large language model (LLM) agents increasingly operate over evolving software repositories, where success depends on repository-specific procedural knowledge: which APIs to call, which scripts to run, and which conventions the current release expects. Agent skills externalize this knowledge into reusable units, and prior work shows that they can improve agent performance. What remains unclear is whether that improvement is durable. The same version specificity that makes a skill useful also makes it fragile: after a release, it may become stale without raising any explicit signal, while continuing to provide obsolete guidance. Externalizing knowledge into a skill can therefore make its decay invisible. We study whether agents can keep this externalized knowledge current. Repo2Skill-Evo casts each release transition as a skill-maintenance task: given a V1 skill set and the official V1-to-V2 patch, an agent must update obsolete skill content while preserving guidance that remains valid. Across 57 real-world repositories and 105 selected release transitions, every evaluated transition invalidates part of the V1 skill set. Yet six frontier agents reach only 29.9%-69.7% avg@3 macro F1 under a patch-grounded removal metric that balances stale-content recall against over-editing precision. Across runs, two opposing errors dominate: incomplete coverage of affected files in the skill set leaves stale content untouched, while overbroad editing is associated with higher recall but lower precision. Repository skills go stale in silence, and even frontier agents cannot reliably maintain them.
发表机构
- ByteDance(字节跳动)
- Peking University(北京大学)
- Beijing Jiaotong University(北京交通大学)
机构由 AI 辅助整理,请以论文原文为准。