arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12001cs.CRcs.SE

扫描技能,治理行动:用运行时后果控制组合注册表裁决

Scan the Skill, Govern the Action: Composing Registry Verdicts with Runtime Consequence Control

Rohit Taneja, Travis Weber

首次发表
浏览论文内容

中文总结 AI 辅助

针对智能体技能注册表扫描器无法回答操作许可问题,提出确定性解析器与信任账本设计,测量显示大量技能违反安全控制,门控有效阻止违规操作。

中文摘要 AI 辅助

智能体技能注册表会对其发布的内容进行筛查。OpenClaw 的安全团队报告称,其扫描器在合并后的阳性结果上最多有 10.4% 的重叠,且 81.9% 被标记的技能仅由单个扫描器捕获。我们以此为前提,认为开放的问题不是哪个扫描器正确,而是哪个扫描器被询问了什么问题。每个扫描器回答的都是“此技能是否恶意?”这一其构建时针对的问题。而“此操作是否被允许在此处、由该操作员、在当前时刻执行?”并非其设计用来表达的问题。我们报告了在 66,192 个公开的 ClawHub 技能版本上进行的三项测量。第一,有 705 个技能(来自 135 个不同的发布者)被所有扫描器及注册表的评判器评为干净,但其指示的操作却违反了 CIS Control 2.7 和 NIST SP 800-53 CM-11 的规定。对其中 100 个进行的人工审计将我们的检测器精确率定为 92%,且未发现恶意意图的标记。其中一位发布者贡献了 705 个中的 506 个,因此我们报告了包含该计数的分布。这可通过公开工件复现。第二,在一个沙箱中,一个实时智能体遵循真实技能文档执行了 144 条命令,其中 34.7% 携带了该文档中未出现的后果类别。第三,在 53 个已清除的技能中,这些技能记录了无干净记录可获得的行动,智能体在 23 个案例中尝试了该行动,而门控阻止了全部 23 个。测试平台和每条记录的命令均已发布。我们针对这一差距提出了一种设计:一个决策路径中无模型的确定性解析器,其输出馈送给一个按(资源,类别)划分的信任账本,该账本的提升阈值源自操作员声明的风险容忍度。十次干净批准无法排除在 95% 置信度下 25.9% 的真实失败率。我们在一系列操作员策略下对门控的中断进行了定价,而不是引用单一的误报率,因为摩擦是策略的属性,而非门控的属性。我们发布了一个包含 64 个案例的混淆基准;我们的方案解决了其中的 52%。

英文摘要

Agent skill registries screen what they publish. OpenClaw's security team reported that its scanners overlap on at most 10.4% of combined positives, and 81.9% of flagged skills are caught by one scanner alone. We take that as given, and suggest the open question is not which scanner is right but which is being asked. Each answers a form of "is this skill malicious?", which is what it was built for. "Is this action permitted here, by this operator, right now?" is not one it is designed to express. We report three measurements over 66,192 public ClawHub skill versions. First, 705 skills across 135 distinct publishers that every scanner and the registry's judge rate clean nonetheless instruct an action prohibited by CIS Control 2.7 and NIST SP 800-53 CM-11. A hand audit of 100 puts our detector at 92% precision and found no marker of malicious intent. One publisher contributes 506 of the 705, so we report the distribution with the count. This is reproducible from public artifacts. Second, of 144 commands a live agent executed in a sandbox while following real skill documentation, 34.7% carried a consequence class absent from that document. Third, over 53 cleared skills documenting an action no clean record earns, the agent reached for one in 23 and the gate stopped all 23. The harness and every recorded command are released. We offer one design for that gap: a deterministic resolver with no model in the decision path, feeding a per-(resource, class) trust ledger whose promotion thresholds derive from the operator's stated risk tolerance. Ten clean approvals cannot exclude a true failure rate of 25.9% at 95% confidence. We price the gate's interruptions across a spectrum of operator policies rather than quote one false-positive rate, since friction is a property of the policy, not the gate. We release a 64-case obfuscation benchmark; ours resolves 52%.

发表机构

  • Pheo Inc(Pheo公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑