StepGuard:利用可扩展监督与安全-效用平衡学习步骤级护栏
StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing
浏览论文内容
中文总结 AI 辅助
该研究针对LLM智能体交互的安全风险,提出步骤级护栏模型StepGuard,通过自动数据引擎StepGen与动态平衡学习方法Balance-GRPO训练,在降低攻击成功率的同时仅小幅影响效用。
中文摘要 AI 辅助
基于大语言模型(LLM)的智能体可通过工具调用与外部环境交互,但该能力也带来文件修改、信息泄露、未授权操作等安全风险。现有护栏常评估完整轨迹,对步骤级动作的执行前监测探索不足。我们提出StepGuard,一种步骤级护栏模型,可审核已完成的智能体轨迹并在工具动作执行前进行检查。为训练StepGuard,我们引入StepGen,一种自动数据引擎,生成具有相同上下文但风险步骤动作不同的安全与不安全轨迹。为进一步减少过度防御与防御不足,我们提出Balance-GRPO,其基于安全与不安全动作的观测准确率动态平衡两者间的学习。实验表明,StepGuard在开放权重护栏模型中达到最高平均准确率,性能可与GPT-5.4媲美。当用于防护AgentDojo和AgentDyn上的智能体时,StepGuard相比无护栏设置将平均攻击成功率降低77.3%,同时平均效用仅下降2.8个百分点。
英文摘要
LLM-based agents can interact with external environments through tool invocation, but this capability also introduces security risks such as file modification, information leakage, and unauthorized actions. Existing guardrails often evaluate completed trajectories, leaving pre-execution monitoring of step-level actions underexplored. We propose StepGuard, a step-level guard model that can audit completed agent trajectories and check tool actions before they are executed. To train StepGuard, we introduce StepGen, an automatic data engine that generates safe and unsafe trajectories with the same context but different actions at the risky step. To further reduce over-defense and under-defense, we propose Balance-GRPO, which dynamically balances learning between safe and unsafe actions based on their observed accuracy. Experiments show that StepGuard achieves the highest average accuracy among open-weight guard models, with performance comparable to GPT-5.4. When used to guard agents on AgentDojo and AgentDyn, StepGuard reduces mean attack success rate by 77.3% relative to the no-guard setting, while mean utility drops by only 2.8 percentage points.
发表机构
- Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
- Beihang University(北京航空航天大学)
- Fudan University(复旦大学)
- Renmin University of China(中国人民大学)
- KAUST(阿卜杜拉国王科技大学)
机构由 AI 辅助整理,请以论文原文为准。