发表机构
North South University; HelicanHQ(南北大学; HelicanHQ)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对智能体运行时中智能体可编辑自身配置的问题,提出“可改能力、不可改限制”的遏制下限规则,并在部署环境中验证其能有效阻止受保护字段被写入。
AI 中文摘要
许多智能体运行时为智能体提供了编辑自身配置的工具。其中部分配置赋予智能体能力,例如启用某个工具。其他部分则设定智能体的限制:它可写入哪些目录、谁可以向它发送消息、它监听哪个网络地址、调用者如何认证,以及阻止危险写入的闸门。如果智能体能够编辑这些限制,一个普通的请求就可能扩大它们。我们在一个已部署的、与模型无关的运行时中对此进行了研究。我们提出一条规则:智能体可以更改赋予能力的字段,但绝不能更改设定其限制的字段。我们将该规则作为配置工具内部的遏制下限加以实施,并测量了有和没有该规则时的不同结果。在没有该下限的情况下,一个前沿模型在72个允许其更改设置的普通请求中的25个请求中写入了受保护的值,且通常是在请求未提及该字段的情况下。写在系统提示中的禁令以可预测的方式失效。一个列出受保护字段名的提示阻止了所有使用这些字段名的请求(36个中保存了0个,而无提示时为36个中保存了17个),但未能阻止那些仅描述目标的请求(36个中保存了10个,而无提示时为36个中保存了8个)。一个描述被禁止效果的提示则产生了相反的效果。有了该下限,167次受保护写入中0次被保存,尽管模型在其中65个案例中尝试了受保护的写入。对工具中其他路径的搜索只发现了一条,即一个固定的shell,而下限的范围声明已将其排除。该研究涵盖了两个模型和一个智能体。我们说明了这能支持什么,不能支持什么。
英文摘要
Many agent runtimes give the agent a tool for editing its own configuration. Some of that configuration grants abilities, such as enabling a tool. Other parts set the agent's limits: which directories it may write to, who may send it messages, which network address it listens on, how callers authenticate, and the gate that blocks risky writes. If the agent can edit those limits, a single ordinary request can widen them. We study this in a deployed, model-agnostic runtime. We propose a rule: the agent may change fields that grant abilities, and may never change fields that set its limits. We enforce the rule as a containment floor inside the configuration tool and measure what happens with and without it. Without the floor, a frontier model wrote a protected value on 25 of 72 ordinary requests that gave it permission to change settings, often when the request never named the field. Prohibitions written in the system prompt failed in a predictable way. A prompt that listed the protected field names stopped every request that used those names (0 of 36 saved, against 17 of 36 with no prompt) and did not stop the requests that only described the goal (10 of 36 saved, against 8 of 36). A prompt that described the forbidden effects did the reverse. With the floor, 0 of 167 protected writes were saved, although the models attempted a protected write in 65 of those cases. A search for other routes through the tool found only one, a pinned shell, which the floor's scope statement already excludes. The study covers two models and a single agent. We state what that does and does not support.
Comments14 pages, 1 figure, 5 tables