arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34518cs.CRcs.AI

SEAD:基于状态视角的工具使用智能体的攻击与防御

SEAD: A State-Based Perspective on Attack and Defense in Tool-Using Agents

Xinjie Shen, Junran Wang, Rongzhe Wei, Pan Li

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出SEAD框架,将工具使用智能体的攻击与防御视为部分可观测状态控制,并设计DART攻击方法与SAGE防御方法,显著提升攻击成功率并有效拦截有害操作。

中文摘要 AI 辅助

语言模型智能体越来越多地使用工具来作用于外部系统。早期的操作可能改变文件、权限、数据库记录或其他状态,从而使后续看似常规的操作变得有害。然而,可见的交互可能无法揭示评估该操作所需的基础状态。我们将攻击与防御形式化为SEAD中的部分可观测状态控制,并从这个共享的执行过程中推导出它们的设计需求。由于攻击者提供指令而目标选择具体操作,DART将有害目标分解为局部可行的步骤,并利用实际工具执行的反馈来引导轨迹搜索。防御者必须在执行前基于不完整的状态证据做出决策。因此,SAGE可以在允许或阻止每个操作(包括在阻止后提出的操作)之前,通过只读查询调查相关状态。我们构建了一个环境可验证的数据集,整合了受控的初始状态、可重放的工具环境和特定于任务的可执行检查。在四个目标模型上,DART相对于竞争基线将语义攻击成功率提高了18.8至35.9个百分点,并在可执行验证下保持一致的增长。在记录的轨迹上,SAGE保留了95.79%的良性轨迹,同时在危害使能边界拦截了92.73%的有害路径。在在线攻防评估中,它将DART的可执行攻击成功率从48.0%降低到4.0%。SAGE在四种攻击方法中保持有效,并泛化到域外环境。我们的代码和数据可在该https URL获取。

英文摘要

Language-model agents increasingly use tools to act on external systems. Earlier actions can alter files, permissions, database records, or other state, making a later routine-looking action harmful. Yet the visible interaction may not reveal the underlying state needed to assess that action. We formulate attack and defense as partially observed state control in SEAD, deriving their design requirements from this shared execution process. Because attackers supply instructions while the target chooses concrete actions, DART decomposes harmful goals into locally plausible steps and uses feedback from actual tool execution to guide trajectory search. The defender must decide before execution with incomplete state evidence. SAGE can therefore investigate relevant state through read-only queries before allowing or blocking each action, including those proposed after a block. We construct an environment-verifiable dataset integrating controlled initial states, replayable tool environments, and task-specific executable checks. Across four target models, DART improves semantic attack success by 18.8--35.9 percentage points over the competing baseline, with consistent gains under executable verification. On recorded trajectories, SAGE preserves 95.79% of benign trajectories while intercepting 92.73% of harmful paths by the harm-enabling boundary. In online attack-defense evaluation, it reduces DART's executable attack success from 48.0% to 4.0%. SAGE remains effective across four attack methods and generalizes to out-of-domain environments. Our code and data is available at https://github.com/EverywhereSafety/SEAD.

补充信息

↑