arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

特权LLM智能体的职责分离:一种受治理的执行架构及其安全-效用权衡度量

Separation of Duties for Privileged LLM Agents: A Governed Execution Architecture with Measured Security-Utility Trade-offs

Qishuai Jing

arXiv 2609.38224首次发表:更新:

AI 中文总结

针对特权LLM智能体,提出一种在模型外部治理行动路径的架构,通过结构化意图和一次性凭证实现职责分离,将攻击成功率从98.3%降至7.7%,并量化了安全与效用的权衡。

AI 中文摘要

大型语言模型智能体日益被授予真实特权(执行命令、修改文件、调用API),因此,一个出错的智能体已经造成了实际影响。现有防御措施集中于智能体的输入,而从候选行动到特权副作用之间的路径仍缺乏直接研究。我们认为,这条路径必须在模型外部加以治理,并研究了一种在智能体与操作系统之间插入四个角色(规划器、策略门、执行器、审计器)的架构。两个核心选择是:行动以结构化意图形式到达,因此裁决从不解析shell语法;批准是一次性凭证,绑定到将要运行的确切字节。我们在包含八个变体的313个案例基准上进行了评估,使用三个托管LLM的仅提示基线,在150个分层案例上进行五次重复(共2,250次尝试调用;2,249次完成)。有效攻击成功率从直接执行下的98.3%降至部署后的7.7%。针对真实实现的重新执行产生了类似的总体成功率(在66个可沙箱评估的载荷上为7.6%),但案例级别存在显著分歧,修正后的错误拒绝率为11.1%。最终阶段的大幅降低(从30.8%降至7.7%)归因于操作系统沙箱,基准测试发现了四个实现缺陷,均非设计审查所发现。

英文摘要

Large language model agents are increasingly granted real privileges (executing commands, modifying files, calling APIs), so an agent that errs has already acted. Existing defences concentrate on the agent's inputs, while the path from a candidate action to privileged side effects remains less directly studied. We argue that this path must be governed outside the model, and study an architecture interposing four roles (planner, policy gate, executor, auditor) between agent and operating system. Two choices are central: actions arrive as structured intents, so adjudication never parses shell syntax; and approval is a one-shot credential bound to the exact bytes that will run. We evaluate on a 313-case benchmark across eight variants, with prompt-only baselines from three hosted LLMs on 150 stratified cases, five repetitions (2,250 attempted calls; 2,249 completed). Effective attack success falls from 98.3% under direct execution to 7.7% deployed. Re-execution against the real implementation yields a similar aggregate rate (7.6% over 66 sandbox-evaluable payloads) but substantial case-level disagreement, at a corrected false-denial rate of 11.1%. A substantial final-stage reduction (from 30.8% to 7.7%) is attributable to the operating-system sandbox, and the benchmark found four implementation defects, none by design review.

Comments12 pages, 2 figures. Supplementary material and a reproducibility artifact are provided as ancillary files

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑