arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自动策略,而非自动技能:面向物理世界的编译智能体技能

Auto-Policy, not Auto-Skill: Compiled Agent Skills for the Physical World

Zhonghao Zhan, Hamed Haddadi

arXiv 2608.25091首次发表:更新:

发表机构

Imperial College London(帝国理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对恶意智能体技能造成物理伤害的问题,提出位于技能制品内部的类型化权限层Edge Skillguard,可高效拒绝借用权限请求,验证了其在多场景下的防护效果。

AI 中文摘要

自演化技能利用(AutoSkills、Hermes智能体)自动生成更多 advisory orchestration(咨询式编排),其报道的收益是效率而非安全性。这忽略了实际差距:技能描述智能体应如何行为,而策略决定哪种行为可成为动作。当前格式用markdown和脚本覆盖前者,后者则留给模型。生成更多技能会扩大差距而非安全性,尤其当错误调用可解锁门或转移资金时。本文记录了两类相邻攻击:恶意技能破坏云软件,以及越狱大语言模型(LLM)控制的机器人造成物理伤害。它们的交集——恶意智能体技能造成物理伤害——直接存在但未被报道。我们将此类攻击命名为借用权限:技能格式未给接收智能体提供类型化方式拒绝智能体间权限声明,因此恶意或误用的技能可通过附加一个权限来驱动执行器。我们提出Edge Skillguard,这是一个类型化权限层,位于技能制品内部而非工作流引擎等工具之间,具备对世界状态和传感器证据的防护。在实时边缘控制平面测试台上,该防护对5种攻击变体的60/60个借用权限请求均予以拒绝,且未阻止良性请求;该结果在5倍规模及Tailscale网格上的多主机环境中依然成立。这些结果表明,高风险技能应将类型化调用策略与过程知识共同打包,使物理动作依赖机器可核查的证据而非对等智能体的声明。

英文摘要

Self-evolving Skill harnesses (AutoSkills, Hermes Agent) generate more advisory orchestration automatically; their reported gains are efficiency, not safety. This misses the actual gap: a Skill describes how an agent should behave; a Policy decides which behavior is allowed to become an action. Today's format covers the first with markdown and scripts; the second is left to the model. Generating more Skills scales the gap, not the safety, especially when a wrong invocation can unlock a door or move money. Two adjacent attacks are documented: malicious skills compromising cloud software, and jailbroken LLM-controlled robots causing physical harm. Their intersection, malicious agent skills causing physical harm, follows directly but has not been reported. We name this class Borrowed Authority: Skills format gives the receiving agent no typed way to reject an inter-agent permission claim, so a malicious or misused Skill can drive actuation by attaching one. We propose Edge Skillguard, a typed authority layer that lives inside the Skill artifact rather than between tools as workflow engines do, with guards over world state and sensor evidence. On a live edge control-plane testbed, the guards reject 60/60 borrowed-authority requests across five attack variants without blocking benign requests, and the result holds at 5x scale and across hosts over a Tailscale mesh. These results suggest that high-risk Skills should co-package typed invocation policy with procedural knowledge, so that physical actions depend on machine-checkable evidence rather than peer-agent claims.

CommentsPresented at the 1st Workshop on Agent Skills (Agent Skills '26), ACM CAIS 2026, San Jose, May 26, 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑