智能体技能可能有害:LLM智能体中技能诱导故障的实证研究
Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents
- Huazhong University of Science and Technology(华中科技大学)
- Microsoft Research(微软研究院)
- Microsoft(微软)
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文通过提出差异分析框架,在SkillsBench等数据集上发现LLM智能体的技能会引发功能故障与效率回归,并构建SkillTriage工具,为安全复用技能提供研究方向。
AI中文摘要:
智能体技能是为大语言模型(LLM)智能体扩展可复用指导的事实机制,一项技能可影响智能体的任务执行,包括规划、工具使用、问题解决与验证。过往研究对智能体技能的结果不一:部分技能可提升任务成功率,而另一些技能无效果、增加token用量与执行时间,甚至降低成功率。本文通过将任务故障与成本回归归因于特定加载技能,对技能诱导的智能体故障展开全面分析。我们提出一种差异分析框架,通过对比目标技能引导的运行与无技能或语义匹配的技能参考运行(该参考运行可解决相同任务或以更低成本解决),将故障或回归归因于某一技能。我们在SkillsBench与SWE-Skills-Bench上实例化该框架,共得到307个技能诱导故障,其中包括125个功能故障与182个效率回归。我们还构建了SkillTriage,这是一种分类法引导的归因工具,可标准化配对案例、提取差异证据并生成分类报告。我们的主要发现包括:(1)技能诱导的功能故障极少由明显不相关的技能导致;相反,看似相关的技能常使智能体错误实现或遗漏任务要求的实现元素。(2)技能诱导的效率回归无法仅用提示长度解释。(3)“过度流程”类别的最大来源为过度验证与繁重实现流程,分别贡献67与30个案例,这表明技能常将验证检查表与构建流程变为强制工作。基于这些发现,我们提出更安全、更具成本意识的技能复用的研究主题与工具改进方向。
英文摘要:
Agent skills are the de facto mechanism for extending LLM agents with reusable guidance. A skill can shape the agent's task execution, including planning, tool use, problem-solving, and validation. Prior work reported mixed results of agent skills: some skills improve task success rates, while others have no effect, increase token use and execution time, and even reduce success rates. This paper presents a comprehensive analysis of skill-induced agent failures by attributing task failures and cost regressions to specific loaded skills. We introduce a differential analysis framework that attributes a failure or regression to a skill by comparing a target skill-guided run against a no-skill or semantically matched skill reference run that solves the same task, or solves it more cheaply. We instantiate this framework on SkillsBench and SWE-Skills-Bench, yielding 307 skill-induced failures, including 125 functional failures and 182 efficiency regressions. We also build SkillTriage, a taxonomy-guided attribution tool that normalizes paired cases, extracts differential evidence, and produces triage reports. Our major findings include: (1) Skill induced functional failures are rarely caused by obviously irrelevant skills; instead, seemingly relevant skills often make the agent incorrectly implement or omit task-required implementation elements. (2) Skill-induced efficiency regressions are not explained by prompt length alone. (3) The largest sources within Excessive Procedure are excessive verification and heavy implementation pipelines, contributing 67 and 30 cases, respectively. This shows that skills often turn validation checklists and construction recipes into mandatory work. Based on our findings, we propose research topics and tooling improvements for safer and more cost-aware skill reuse.