RePolicy:用于智能体安全保障中安全策略调用的强化学习方法
RePolicy: Reinforcement Learning for Safety-Policy Invocation in Agent Safeguards
- Zhejiang University(浙江大学)
- Zhongguancun Academy(中关村学院)
- University of Science and Technology of China(中国科学技术大学)
- National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对现有智能体安全保障方法难以适应未见轨迹和变化策略上下文的问题,提出RePolicy方法,结合PolicyTraj-20K与带可验证奖励和策略上下文扰动的GRPO,在六个智能体安全基准上取得良好安全检测与策略调用效果。
AI中文摘要:
保障语言模型智能体需要在依赖上下文的安全策略下评估完整的执行轨迹。现有的策略感知安全保障主要依赖提示或监督微调,限制了其对未见轨迹和变化策略上下文的适应能力。我们提出RePolicy,一种通过强化学习学习安全策略调用的智能体安全保障方法。给定智能体轨迹和动态策略库,RePolicy调用适用的策略并利用其内容生成基于策略的理由和安全判断。我们构建PolicyTraj-20K以支持监督初始化,随后采用带有可验证奖励和策略上下文扰动的GRPO。在六个智能体安全基准上的实验表明,RePolicy在不同策略上下文下实现了强劲的整体安全检测性能和稳健的策略调用。
英文摘要:
Safeguarding language model agents requires assessing complete execution trajectories under context-dependent safety policies. Existing policy-aware safeguards mainly rely on prompting or supervised fine-tuning, limiting their ability to adapt to unseen trajectories and changing policy contexts. We propose RePolicy, an agent safeguard that learns safety-policy invocation through reinforcement learning. Given an agent trajectory and a dynamic policy library, RePolicy invokes the applicable policy and uses its content to produce a policy-grounded rationale and safety judgment. We construct PolicyTraj-20K to support supervised initialization, followed by GRPO with verifiable rewards and policy-context perturbation. Experiments across six agent safety benchmarks show that RePolicy achieves strong overall safety-detection performance and robust policy invocation under varying policy contexts.