arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

智能体技能下游效用的实证研究

An Empirical Study of Agent Skills' Downstream Utility

Yu Cheng, Dehai Zhao, Zhongxin Liu, Qing Huang, Zhenchang Xing, Xiaoxue Ren

arXiv 2610.08875首次发表:更新:

发表机构

Zhejiang University; Jiangxi Normal University(浙江大学; 江西师范大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过87个任务实证发现,智能体技能效用受配置影响,按操作支持重排序可提升首选通过率4.35-5.80个百分点,并提炼出17项技能编写实践以指导开发者。

AI 中文摘要

智能体技能将程序性指导与资源打包以供复用,但相关技能并不必然提升任务性能。现有研究刻画了技能内容并评估其下游性能,却对效用如何依赖于内容、执行配置及多技能组织方式提供的解释有限。我们在87个SkillsBench任务上开展了一项实证研究,将下游效用定义为在相同模型-框架配置下,同一任务上使用技能与不使用技能(No-Skill)的通过率之差。我们在九种配置下比较相同的技能,随后在三种选定配置下考察备选已发布技能及固定技能集的组织方式。我们从包含37,596个技能的精选语料库中检索市场候选技能。基于LLM辅助的内容、执行轨迹及最终工件分析,并经作者复核,将所提供的支持与实际使用及任务结果相关联。相同的技能在36.78%的任务上对某些配置有帮助,却对另一些配置有害,轨迹显示推荐程序可能成为执行负担。相关性排名忽略了更有用的候选。在评估的候选集内,按所需操作的支持度进行重排序,在三种配置下将首选通过率提升了4.35至5.80个百分点。我们提炼出17项编写实践,将可执行程序与恢复、任务要求的保持及最终工件的检查联系起来。阶段计划(Stage Plan)与依赖有向无环图(Dependency DAG)优于仅使用顺序,且DAG的额外收益集中在提供五或六个技能的任务上。这些发现指导开发者评估可用的操作支持,在保持任务要求的同时允许程序调整,并在组织技能时明确工件依赖关系。

英文摘要

Agent Skills package procedural guidance and resources for reuse, but a relevant Skill does not necessarily improve task performance. Existing studies characterize Skill content and evaluate downstream performance, yet provide limited explanations of how utility depends on content, execution configuration, and multi-Skill organization. We conduct an empirical study on 87 SkillsBench tasks, defining downstream utility as the pass-rate difference from No-Skill on the same tasks under the same model--harness configuration. We compare the same Skills across nine configurations, then examine alternative published Skills and organizations of fixed Skill sets under three selected configurations. We retrieve marketplace candidates from a curated corpus of 37,596 Skills. LLM-assisted analysis of content, execution traces, and final artifacts, followed by author review, relates provided support to actual use and task outcomes. The same Skills help some configurations and hurt others on 36.78\% of tasks, with trajectories showing that recommended procedures can become an execution burden. Relevance rankings overlook more useful candidates. Within the evaluated candidate sets, reranking by support for required operations raises first-choice pass rates by 4.35--5.80 percentage points across the three configurations. We derive 17 authoring practices linking executable procedures to recovery, preservation of task requirements, and checks on final artifacts. Stage Plan and Dependency DAG outperform use order alone, with DAG's additional benefits concentrated in tasks supplied with five or six Skills. These findings guide developers to assess usable operation support, allow procedure adaptation while preserving task requirements, and make artifact dependencies explicit when organizing Skills.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑