arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

信号还是噪声?智能体在Web开发中的技能基准研究

Signal or Noise? A Benchmark Study of Agent Skills in Web Development

Ziyue Yang, Fan Ding

arXiv 2608.23067首次发表:更新:

发表机构

Baidu(百度)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究构建WebDev-Skills-Bench基准,通过实证分析Web开发智能体技能的效果,发现技能注入会降低任务性能、增加成本,提出需将技能视为特定三元组假设并建立对应评估标准。

AI 中文摘要

智能体技能是可复用的过程模块,越来越多地被注入编码智能体会话中,以编码框架约定、反模式和可复用工具。然而,由于每个注入的技能都会扩展每个查询的提示,有效的技能基准不仅必须确定智能体是否能解决任务,还必须确定是否应该注入该技能。我们引入WebDev-Skills-Bench,并将其用于对31个公开WebDev技能在50个Web-Bench项目和1000个有序任务上的受控实证研究。该基准比较了四种匹配条件,包括长度匹配的无关对照组和留一法组件消融。为了将技能效应与提示长度伪影隔离开,我们仅将此http URL放入提示中,同时将辅助文件挂载到智能体工作区。在四个模型中,目标技能注入使平均Pass@2降低了1.3%至4.2%,降低了任务完成深度,并使令牌成本增加了72%至394%,仅在17%至36%的技能-项目对中产生增益。长度匹配对照组揭示了两种失败模式:一些模型受长度干扰,即长度相等的无关技能会重现大部分损失;而另一些模型受内容误导,即提示长度是中性的,但技能内容仍使Pass@2降低了1.1%至1.4%。进一步分析表明,损失集中在简单的早期任务上,技能排名在模型间的可迁移性较弱,且在有用技能中,反模式的表现优于以示例为主的内容。这些发现将匹配的技能重新定义为关于特定技能-项目-模型三元组的假设,而非可移植资产,将注入重新定义为每次部署的路由决策,并使长度匹配对照组和按模型审计成为智能体技能评估的最低标准。

英文摘要

Agent Skills are reusable procedural modules that are increasingly injected into coding-agent sessions to encode framework conventions, anti-patterns, and reusable tools. However, because each injected Skill expands the prompt of every query, an effective Skill benchmark must determine not only whether an agent can solve a task, but whether the Skill should have been injected at all. We introduce WebDev-Skills-Bench and use it for a controlled empirical study of 31 public WebDev Skills on 50 Web-Bench projects and 1,000 ordered tasks. The benchmark compares four matched conditions, including a length-matched irrelevant control and leave-one-out component ablations. To isolate Skill effects from prompt-length artifacts, we place only SKILL.md in the prompt while mounting auxiliary files into the agent workspace. Across four models, target Skill injection reduces mean Pass@2 by 1.3% to 4.2%, lowers task completion depth, and increases token cost by 72% to 394%, with gains in only 17% to 36% of Skill-project pairs. Length-matched controls reveal two failure modes: some models are length-distracted, where an equally long irrelevant Skill reproduces most of the loss, while others are content-misled, where prompt length is neutral but Skill content still lowers Pass@2 by 1.1% to 1.4%. Further analysis shows that losses concentrate on easy early tasks, Skill rankings transfer weakly across models, and anti-pattern rules outperform example-heavy content within helpful Skills. These findings recast a matched Skill as a hypothesis about a particular Skill-project-model triple rather than a portable asset, reframing injection as a per-deployment routing decision and making length-matched controls and per-model audits a minimum standard for Agent-Skill evaluation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑