arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29543cs.AI

物理智能体的安全技能退役

Safe Skill Retirement for Physical Agents

Zhonghao Zhan, Xiao Ma, Hamed Haddadi

首次发表
浏览论文内容

中文总结 AI 辅助

针对物理智能体技能退役,提出匹配权限反事实与双门退役证书,实验显示任务基准可移除94%条款但产生未授权效应,需审计权限契约。

中文摘要 AI 辅助

智能体技能将程序性指导与执行条件捆绑在一起,这些条件管辖着权限、用户同意和实时环境状态。当模型能力提升时,维护者会修剪在授权基准任务上看似冗余的指令。然而,授权的维护测试可能使休眠的安全条件得不到测试。这种不匹配在物理和隐私敏感效应上造成了未测量的支持缺口。我们引入了匹配的权限反事实,这些反事实保持所请求的动作、工具参数和预期效果不变,同时系统地改变单个治理谓词。我们通过一个双门退役证书来形式化这一评估,该证书要求候选缩减在声明裕度内保持授权效用,同时产生零未授权的受保护效应。在跨越四个前沿和本地模型配置、十二个技能包(2592个评估单元)的受控实验中,任务认证的缩减移除了超过94%的技能条款并保持了授权完成,但在每个技能包中都产生了未授权的受保护效应。边界执行在声明的审计上消除了受保护效应,但对一个配置未能通过效用门。一个有界的组合协议在所有四个配置中通过了两道门,且效用裕度为零。对一个只读的Home Assistant摄像头链路的端到端检查验证了在真实设备上的提议、决策和效应测量。这些结果表明,虽然任务基准可以证明退役程序性指导的合理性,但退役决策需要明确审计治理物理动作的权限契约。

英文摘要

Agent skills bundle procedural guidance with execution conditions governing authority, user consent, and live environment state. When model capabilities advance, maintainers prune instructions that appear redundant on authorized benchmark tasks. However, authorized maintenance tests can leave dormant safety conditions untested. This mismatch creates an unmeasured support gap over physical and privacy-sensitive effects. We introduce matched authority counterfactuals that hold the requested action, tool parameters, and intended effect fixed while systematically varying a single governing predicate. We formalize this evaluation via a two-gate retirement certificate requiring a candidate reduction to preserve authorized utility within a declared margin while producing zero unauthorized protected effects. In controlled experiments spanning four frontier and local model configurations across twelve skill bundles (2,592 evaluation cells), task-certified reductions remove over 94% of skill clauses and preserve authorized completion, yet produce unauthorized protected effects in every bundle. Boundary enforcement eliminates protected effects on the declared audit but fails the utility gate for one configuration. One bounded combined protocol passes both gates across all four configurations, with zero utility headroom. An end-to-end check on one read-only Home Assistant camera chain verifies proposal, decision, and effect measurements on a real device. These results demonstrate that while task benchmarks can justify retiring procedural guidance, retirement decisions require explicitly auditing the authority contracts governing physical actions.

发表机构

  • Imperial College London(伦敦帝国理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑