聚会之后:病毒式传播的智能体技能生态系统留下了什么
After the Party: Growth, Governance, and Security Scanning in the OpenClaw Agent Skill Ecosystem
浏览论文内容
中文总结 AI 辅助
本文研究OpenClaw智能体技能注册表繁荣后的遗留问题,发现技能下载高度集中、人工审查缺失且安全扫描器一致性差,提出治理需依赖稳健测量与独立验证。
中文摘要 AI 辅助
AI智能体日益通过智能体技能(即自然语言指令)来行动,这些指令引导宿主智能体执行shell、网络、凭据、文件和进程操作,而公共注册表则大规模分发这些技能。在2026年上半年,OpenClaw AI智能体病毒式传播,其公共技能注册表迅速膨胀:可观察的技能存量在91天内几乎翻倍,且6月份可见的列表中大多数是在短短两个月内创建的。到我们研究窗口结束时,这股浪潮已过顶峰,月度列表创建量和核心仓库活动均从春季峰值回落。本文基于OpenClaw的Git历史、其GitHub issue和pull request以及三个ClawHub注册表快照,衡量了这一繁荣期留下了什么。注意力高度集中:排名前10%的技能获得了全部下载量的46.93%。在控制创建队列和技能年龄后,没有简单的技能特征(如大小或下载量)能继续稳定预测列表的存续。人工审查并未持续:77.86%的技能拥有零星标和零评论,而85.06%的可读技能带有特权证据。自动化清理尚未就绪:三个安全扫描器在其共同覆盖的61,990个技能中,对23,702个技能存在分歧。经过人工裁决后,加权扫描器相对于参考标准的敏感性介于21.67%至61.06%之间。治理快速增长的智能体技能注册表不能依赖简单的元数据或单一扫描器评分;这需要稳健、透明的测量和独立验证。
英文摘要
AI agents increasingly act through agent skills, i.e., natural-language instructions, that direct a host agent toward shell, network, credential, file, and process actions, and public registries distribute them at scale. In the first half of 2026, the OpenClaw AI agent went viral, and its public skill registry boomed: the observable stock nearly doubled in 91 days, and a majority of the listings visible in June were created in just two months. By the end of our study window, the wave had crested, and monthly listing creation and core-repository activity were falling from their spring peaks. This paper measures what the boom left behind, drawing on the OpenClaw Git history, its GitHub issues and pull requests, and three ClawHub registry snapshots. Attention is concentrated: the top 10% of skills received 46.93% of all downloads. No simple skill features (like size or download counts) remained a stable predictor of continued listing once creation cohort and skill age were controlled. Human scrutiny did not stay: 77.86% have zero stars and zero comments, while 85.06% of the readable skills carry privilege evidence. And automated cleanup is not ready: the three security scanners disagreed on 23,702 of the 61,990 skills they all cover. After human adjudication, weighted scanner sensitivity against the reference standard ranged from 21.67% to 61.06%. Governing fast-growing agent-skill registries cannot rely on simple metadata or single scanner scores; it requires robust, transparent measurement and independent validation.
发表机构
- Monash University(莫纳什大学)
机构由 AI 辅助整理,请以论文原文为准。