arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11647cs.SEcs.AI

多一个技能:编码智能体中共安装技能的冲突问题

One Skill Too Many: How Co-Installed Skills Conflict in Coding Agents

Chaoliang Yan, Zihao Xu, Yuekang Li, Shangzhi Xu, Yi Liu, Gelei Deng, Siqi Ma

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对编码智能体中共安装相似技能的冲突问题,通过实证研究揭示冲突规律,提出基准应评估独有核心功能、平台应管控技能首次读取的解决方案。

中文摘要 AI 辅助

编码智能体配备了智能体技能,这些技能是目录,其对应的URL告知模型何时以及如何执行任务。由于技能来自独立来源(团队、开发者、插件、复制的集合),已安装的技能可能与执行相同任务的相似技能共存,模型仅通过名称和描述在它们之间进行选择。发生冲突时,已安装的技能会失去核心功能(例如禁止操作git),因为相似技能会替代它运行或改变自身行为。任务仍会完成,因此仅检查任务完成情况的基准会忽略此类情况。我们对这类冲突开展了首次实证研究:从20947个代码库的快照中挖掘出822109个候选相似技能对,让大语言模型(LLM)对3754个分层样本进行判断,并在三个模型上运行312个已确认的技能对(共6368次运行、169294次工具调用、542个智能体小时)。我们报告五项发现:(1)易发生冲突的技能很常见:近四分之一的已安装技能与执行相同任务的技能共存,37%的被判断技能属于复制集合;(2)这类技能对大多涉及规范性技能,其次是能力型技能;(3)相似技能在不降低任务完成率的情况下,会在五分之一的运行中替代已安装技能,且先打开相似技能的运行会失去已安装技能独有的超过三分之一的核心功能;(4)安装位置决定哪个技能运行,列表顺序几乎无影响,且在被替代的运行中,最终回复仅在0.9%的情况下会提及所使用的技能;(5)冲突在首次读取技能时就已确定,几乎总是在任何文件被修改之前,在该读取处设置预工具钩子可将独有的核心功能的保真度恢复到先打开已安装技能的运行水平。因此,基准应评估独有的核心功能,平台应管控首次读取环节并显示运行的是哪个技能。

英文摘要

Coding agents are extended with agent skills, directories whose SKILL.md tells the model when and how to perform a task. Because skills come from independent sources (teams, developers, plugins, copied collections), an installed skill can be co-installed with a similar skill doing the same job, and the model picks between them by name and description alone. In a conflict, the installed skill loses core functions (e.g., a ban on touching git) because the similar skill runs instead or changes what it does. The task still passes, so benchmarks that check only task completion miss such cases. We present the first empirical study of such conflicts. From snapshots of 20,947 repositories, we mine 822,109 candidate similar-skill pairs, have an LLM judge a stratified sample of 3,754, and run 312 confirmed pairs on three models (6,368 runs, 169,294 tool calls, 542 agent-hours). We report five findings. (1) Conflict-prone skills are common: nearly one in four installed skills is co-installed with one that does the same job, and 37% of judged skills sit inside copied collections. (2) Most such pairs involve normative skills, then capability skills. (3) Without lowering task completion, a similar skill takes one in five runs from the installed skill, and runs that open the similar skill first lose over a third of the exclusive core functions that only the installed skill fulfills. (4) Install location decides which skill runs, listing order barely matters, and the final reply names the skill used in only 0.9% of substituted runs. (5) Conflicts are decided at the first skill read, almost always before any file is changed, and a pre-tool hook at that read restores fidelity on exclusive core functions to the level of runs that open the installed skill first. Benchmarks should thus score exclusive core functions, and platforms should guard the first read and show which skill ran.

发表机构

  • University of New South Wales(新南威尔士大学)
  • Griffith University(格里菲斯大学)
  • Nanyang Technological University(南洋理工大学)
  • The University of Wollongong(伍伦贡大学)

机构由 AI 辅助整理,请以论文原文为准。

↑