发表机构
The University of Arizona; California State University, Long Beach; Nanyang Technological University; Institute of Information Engineering, Chinese Academy of Science; Wuhan University; East Carolina University(亚利桑那大学; 加州州立大学长滩分校; 南洋理工大学; 中国科学院信息工程研究所; 武汉大学; 东卡罗来纳大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究分析编码智能体技能修订的影响,发现多数修订改变规则,添加规则可提升模型合规性和智能体行动率及最终正确性,但真实加载场景下收益部分保留,且成本有限。
AI 中文摘要
智能体技能(Agent Skills),即告知LLM编码智能体项目如何工作的this http URL文件,像代码一样被修订,但修订对智能体的影响尚不明确。基于3159个技能中的2608对首次/最后修订,我们刻画了技能如何演化以及它们如何与智能体框架配置共同变化。然后我们聚焦于规则变更,即添加或移除可自动检查的规则(如“运行allium check”)的修订。我们测量了这些修订对21个模型在单次回答中的影响,以及对4个智能体在沙盒环境中的影响,并测量了其中20个模型和4个智能体的成本。大多数修订(55%)改变了规则或流程,且修订技能的提交比同规模的其他提交更频繁地更改harness文件(如this http URL)。在16个开放权重模型中,添加规则使单次回答的合规性平均提高+0.41。在4个智能体中,智能体采取所需行动的比例平均提高+0.23(范围+0.16至+0.36),对于由盲审评估的3个智能体,最终正确性平均提高+0.10(范围+0.06至+0.14)。收益主要来自命名旧技能未提及的命令或路径的规则。真实工具仅在智能体决定需要时才加载技能正文。在该设置下,4个智能体平均保留了约一半的行动收益(51%),3个开放模型保留了约38%。一次修订为单次回答增加18-19%的输入令牌,且对智能体回合无可检测成本,而加载技能正文使回合令牌平均增加50%。
英文摘要
Agent Skills, the SKILL.md files that tell an LLM coding agent how a project works, are revised like code, yet what a revision does to the agent is unknown. From 2,608 first/last revision pairs of 3,159 Skills, we characterize how Skills evolve and how they change together with the configuration of the agent's harness. We then focus on rule changes, revisions that add or remove a rule we can check automatically, such as "run allium check". We measure their effect on 21 models in single answers and on four agents in a sandbox, and their cost on 20 of these models and the four agents. Most revisions (55%) change a rule or procedure, and commits that revise a Skill change harness files such as CLAUDE.md more often than other commits of the same size. Across 16 open-weight models, an added rule raises compliance in a single answer by +0.41 on average. Across the four agents, the rate at which the agent takes the required action rises by +0.23 on average (+0.16 to +0.36), and for the three agents that blind judges assessed, final correctness rises by +0.10 on average (+0.06 to +0.14). The gain comes mainly from rules that name a command or path the old Skill did not mention. Real tools load a Skill's body only when the agent decides it needs it. In that setting the four agents keep about half of the action gain on average (51%), and the three open models about 38%. A revision adds 18-19% input tokens to a single answer and no detectable cost to an agent episode, while loading a Skill's body raises the tokens of an episode by 50% on average.