arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27146cs.AIcs.SE

当工具输出成为命令:在工具增强型大语言模型智能体中分离动作诱导与运行时授权

When Tool Outputs Become Commands: Separating Action Induction from Runtime Authorization in Tool-Augmented LLM Agents

Xiaokun Guo, Zhen Xu, Dongdong Huo, Yanqiu Zhang, Wei Wang, Qinfu Yang, Dongjin Yu, Yu Wang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对工具增强型LLM智能体中工具输出成为命令的风险,提出SARA框架分离动作诱导与执行授权,在AgentDojo等数据集上有效限制攻击成功率并保持任务效用。

中文摘要 AI 辅助

工具增强型大语言模型智能体必须依赖不可信的运行时观测结果来完成开放式任务;然而,当工具输出不再仅仅提供数据,而是开始指定具体动作时,它们实际上变成了能够产生超出用户意图的现实世界副作用的“命令”。我们认为这种风险源于将动作诱导与执行授权混为一谈。为解决这一区分问题,我们提出了SARA,它将动作诱导与执行授权视为不同的运行时角色,并将动作来源与执行权限分离开来。在观测侧,一个上下文隔离的动作探针(Action Probe)暴露动作诱导语义,并在各步骤中持续记录动作来源作为审查信号;在执行侧,仅针对用户目标和来自授权成功执行的审计证据,且满足目标、执行链和参数级别的支持时,才授权实际的工具调用。为在多步执行中保持这种分离,SARA应用无历史提升(No-History-Promotion),以防止历史重演将动作来源洗白为执行权限。在AgentDojo和AgentDyn数据集上,SARA在四个主要评估设置中将攻击成功率(ASR)限制在不超过0.63%,同时保持有竞争力的任务效用,并在额外的智能体主干模型上持续降低ASR。

英文摘要

Tool-augmented LLM agents must rely on untrusted runtime Observations to complete open-ended tasks; however, when tool outputs no longer merely provide data but begin to specify concrete actions, they effectively become ``commands'' that can drive real-world side effects beyond user intent. We argue that this risk arises from conflating action induction with execution authorization. To address this distinction, we propose SARA, which treats action induction and execution authorization as distinct runtime roles and separates action provenance from execution authority. On the Observation side, a context-isolated Action Probe exposes action-inducing semantics and persistently records action-origin provenance across steps as a review signal; on the execution side, actual tool calls are authorized only against the user objective and audited evidence from authorized successful executions, while satisfying goal, execution-chain, and argument-level support. To preserve this separation across multi-step execution, SARA applies No-History-Promotion to prevent historical recurrence from laundering action origins into execution authority. Across AgentDojo and AgentDyn, SARA limits ASR to no more than \(0.63\%\) across four primary evaluation settings while maintaining competitive task utility, and consistently reduces ASR across additional Agent backbones.

发表机构

  • Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
  • School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络空间安全学院)

机构由 AI 辅助整理,请以论文原文为准。

↑