arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AdaGuard:具有用户自定义策略的自适应防护模型

AdaGuard: An Adaptive Guard Model with User-defined Policies

Yunhao Feng, Yifan Ding, Yuxiang Xie, Zheng Li, Mingrui Lao, Zeyuan Wang, Yanming Guo

arXiv 2609.34241首次发表:更新:

发表机构

National University of Defense Technology; Fudan University(国防科技大学; 复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

AdaGuard提出自适应防护模型,结合AdaptiveSafety数据集和SafePO强化学习,在用户自定义策略下准确识别智能体违规行为,4B模型在基准上达到高准确率。

AI 中文摘要

防护模型支持语言模型智能体的安全部署,但固定的风险分类限制了其适应不同应用和任务需求的能力。在用户自定义策略下,检测违规需要同时解释适用规则和智能体行为,因为相同的行为在不同策略下可能得到不同的判定。为了支持学习这一能力,我们引入了AdaptiveSafety,一个包含10,939个训练样本和1,000个测试样本的数据集,覆盖1至100条规则的策略。该数据集结合了来自多个来源的轨迹与策略和行为反事实,每个样本都配有解释和完整的违规规则集。这些反事实揭示了改变合规性的变化,而结构增强则提供了在规则重排序和标识符重映射下的一致性监督。基于这种监督,我们提出了SafePO,一种强化学习算法,用于在平衡解释性推理和最终判定之间细化违规识别。SafePO使用结构化奖励来评估预测正确性,在响应级别保留组相对优势,并采用单独训练的价值模型来调节解释和判定区域内的词元权重。独立的归一化控制了它们在训练中的相对贡献,尽管长度不同。通过监督初始化后接SafePO,我们开发了AdaGuard,一个包含0.6B、4B和8B参数的防护模型家族,在推理时根据提供的策略评估智能体轨迹。我们的4B模型在AdaptiveSafety上实现了89.30%的二元准确率,在DynaBench上实现了71.82%。项目仓库可在该https URL获取。

英文摘要

Guard models support the safe deployment of language model agents, but fixed risk taxonomies limit their ability to accommodate requirements that vary across applications and tasks. Under user-defined policies, detecting violations requires interpreting both the applicable rules and the agent's behavior, since identical actions can receive different judgments under different policies. To support learning this capability, we introduce AdaptiveSafety, a dataset of 10,939 training examples and 1,000 test examples covering policies with 1--100 rules. The dataset combines trajectories from multiple sources with policy and behavioral counterfactuals, pairing each example with an explanation and the complete set of violated rules. These counterfactuals expose changes that alter compliance, while structural augmentations provide supervision for consistency under rule reordering and identifier remapping. Building on this supervision, we propose SafePO, a reinforcement learning algorithm for refining violation identification while balancing explanatory reasoning and final verdicts. SafePO uses structured rewards to assess prediction correctness, retains group-relative advantages at the response level, and employs a separately trained value model to modulate token weights within explanation and verdict regions. Separate normalization controls their relative contribution to training despite differences in length. Through supervised initialization followed by SafePO, we develop AdaGuard, a family of 0.6B, 4B, and 8B guard models that assess agent trajectories under policies supplied at inference time. Our 4B model achieves binary accuracies of 89.30\% on AdaptiveSafety and 71.82\% on DynaBench. The project repository is available at https://github.com/Yunhao-Feng/AdaGuard

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑