arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.08453cs.AIcs.CR

是什么阻碍了智能体技能的可复用性?来自13.8万个SKILL.md文件的证据

What Keeps Agent Skills from Being Reusable? Evidence from 138K SKILL.md Files

发表机构纽约城市大学研究生中心 · 俄亥俄州立大学 · 纽约城市大学亨特学院
查看机构详情
  • The Graduate Center, City University of New York(纽约城市大学研究生中心)
  • The Ohio State University(俄亥俄州立大学)
  • Hunter College, City University of New York(纽约城市大学亨特学院)

机构由 AI 辅助整理,请以论文原文为准。

Chi Zhang, Yimin Liu, Xinze Chen, Ping Ji

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过分析13.8万个SKILL.md文件,发现91.8%的Agent Skills存在可复用性相关缺陷,主要为打包问题,验证了路由元数据对检索可靠性的影响,并提出了结合规范提示等的质量保证生成工作流。

中文摘要 AI 辅助

根据当前标准,Agent Skills是结合了指令与支撑资源的http URL文件,使大语言模型(LLM)智能体能够在单次对话之外复用程序。然而,许多公开技能似乎源自单一任务、代码仓库或对话,即便它们被作为可复用组件共享。我们基于官方规范和最佳实践指南,采用两层缺陷分类法,分析了来自20556个代码仓库的138133个公开http URL文件中的这一差距。我们发现,91.8%的技能至少包含一个检测到的缺陷,在宽松和严格阈值下的估计值保持稳定(88.8%-94.6%)。主要缺陷是普通打包问题而非特殊攻击:路由元数据薄弱、主体臃肿或不可操作、资源组织差。对2万个技能进行的确定性路由压力测试显示了其功能影响:具有有效路由元数据的技能,相比存在路由缺陷的技能,能从启动描述中更可靠地被检索到。缺陷率因平台和来源而异:符合规范的技能缺陷更少,而AI标记的技能则表现出更多安全和可移植性问题。轻量执行与修复实验支持了一种质量保证生成工作流,该工作流结合了符合规范的提示、轻量代码检查(linting)、自动修复以及安全门控。

英文摘要

Under the current standard, Agent Skills are SKILL.md files that combine instructions with supporting resources, enabling Large Language Model (LLM) agents to reuse procedures beyond a single conversation. Yet many public skills appear to originate from a single task, repository, or conversation, even when they are shared as reusable components. We analyze this gap across 138,133 public SKILL.md files from 20,556 repositories using a two-tier defect taxonomy grounded in the official specification and best-practice guidance. We find that 91.8% of skills contain at least one detected defect, with stable estimates across lenient and strict thresholds (88.8-94.6%). The dominant failures are ordinary packaging problems rather than exotic attacks: weak routing metadata, bloated or non-actionable bodies, and poor resource organization. A deterministic routing stress test over 20,000 skills shows the functional impact: skills with valid routing metadata are retrieved more reliably from startup descriptions than skills with routing defects. Defect rates vary by platform and provenance: specification-aware skills contain fewer defects, while AI-marked skills show more safety and portability problems. Lightweight enforcement and repair experiments support a quality-assured generation workflow combining spec-aware prompting, lightweight linting, automated repair, and safety gating.

补充信息

↑