arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

$S^3$:通过多阶段防御提升智能体安全性

$S^3$: Improving Agent Safety through Multi-Stage Defense

Zibo Xiao, Haoyu Wang, Jun Sun

arXiv 2608.02683首次发表:更新:

发表机构

Singapore Management University(新加坡管理大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对LLM智能体工作流的跨阶段安全风险问题,提出多阶段防御框架$S^3$,引入阶段特定安全技能,构建MSRB基准,实验显示其安全有效性和效用保留均优于现有基线。

AI 中文摘要

大语言模型(LLM)智能体依赖包含记忆、规划、工具执行等阶段的多阶段智能体工作流来完成复杂任务,但风险可能在不同阶段出现、跨步骤传播,难以检测和缓解。现有安全方法仅保护孤立阶段且难以集成,导致智能体在整个工作流中缺乏全面保护。为解决这些局限,我们引入阶段特定安全技能(Stage-Specific Safety Skills),这一统一抽象将异构安全设计表示为具有显式阶段语义的可复用、可组合组件;还开发了自动化转换流水线,可将现有安全设计转换为可复用安全技能,并建立社区驱动的安全技能库。基于该抽象,我们提出$S^3$,这是一个多阶段防御框架,其中一个守卫智能体协调阶段特定安全技能,在整个智能体工作流中进行风险检测与缓解。我们还构建了多阶段风险基准(Multi-Stage Risk Benchmark, MSRB),以评估工作流各阶段的代表性风险。实验结果显示,$S^3$在安全有效性和效用保留两方面均持续优于代表性的最先进基线。这些结果表明,阶段特定安全技能有望成为构建弹性、可信智能体系统的可扩展且可组合的基础。

英文摘要

Large Language Model (LLM) agents rely on multi-stage agentic workflows, with stages such as memory, planning, and tool execution, to accomplish complex tasks. However, risks may emerge at different stages, propagate across steps, and become difficult to detect and mitigate. Existing safety methods protect only isolated stages and are difficult to integrate, leaving agents without comprehensive protection throughout the workflow. To address these limitations, we introduce Stage-Specific Safety Skills, a unified abstraction that represents heterogeneous safety designs as reusable and composable components with explicit stage semantics. We further develop an automated transformation pipeline that converts existing safety designs into reusable safety skills and establish a community-driven safety skill library. Building on this abstraction, we propose $S^3$, a multi-stage defense framework in which a guard agent orchestrates stage-specific safety skills for risk detection and mitigation throughout the agentic workflow. We also construct the Multi-Stage Risk Benchmark (MSRB) to evaluate representative risks across workflow stages. Experimental results show that $S^3$ consistently outperforms representative state-of-the-art baselines in both safety effectiveness and utility preservation. These results demonstrate the potential of stage-specific safety skills as a scalable and composable foundation for building resilient and trustworthy agent systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑