AdaGuard:通过支持推理的LLM作为评判护栏增强安全性与策略合规性
AdaGuard: Enhancing Safety and Policy Compliance with Reasoning-Enabled LLM-As-A-Judge Guardrails
浏览论文内容
中文总结 AI 辅助
AdaGuard提出一种自适应LLM作为评判框架,通过动态策略执行和推理预算分配,兼顾安全合规与低延迟,媲美更大模型。
中文摘要 AI 辅助
企业级生成式人工智能应用需要能够适应多样化风险姿态、不断演变的策略以及不同延迟约束的稳健安全机制。当前的护栏解决方案往往存在僵化问题,依赖固定策略集,且提供的透明度或推理灵活性有限。我们提出AdaGuard,一种自适应的LLM作为评判框架,旨在通过动态策略执行和自适应推理预算分配来应对这些挑战。AdaGuard基于监督微调(SFT)和强化学习(GRPO)构建,能够在运行时泛化到用户定义的安全与合规策略,而无需频繁更新模型。我们方法的一个核心创新是能够动态推断输入-策略对的复杂性,使模型能够在高速黑盒推理与可解释的、支持推理的审核之间切换。这种灵活性使开发者能够在严格的延迟要求与可操作的透明度需求之间取得平衡。这种自适应能力使AdaGuard能够与数倍于其规模的其它护栏模型和前沿模型相抗衡,而其自动推理模式以极低的延迟恢复了始终开启推理的准确性。
英文摘要
Enterprise generative AI applications require robust safety mechanisms that can accommodate diverse risk postures, evolving policies, and varying latency constraints. Current guardrail solutions often suffer from rigidity, relying on fixed policy sets and offering limited transparency or reasoning flexibility. We present Adaguard, an adaptive LLM-as-a-Judge framework designed to address these challenges through dynamic policy enforcement and adaptive reasoning-budget allocation. Built using supervised fine-tuning (SFT) and reinforcement learning (GRPO), AdaGuard generalizes to user-defined safety and compliance policies at runtime without requiring frequent model updates. A core innovation of our approach is the ability to dynamically infer the complexity of input-policy pairs, allowing the model to switch between high-speed black-box inference and explainable, reasoning-enabled moderation. This flexibility enables developers to balance stringent latency requirements with the need for actionable transparency. This adaptive capability allows AdaGuard to rival other guardrail and frontier models several times its size, while its auto-reasoning mode recovers the accuracy of always-on reasoning at a fraction of the latency
发表机构
- Capital One(第一资本)
机构由 AI 辅助整理,请以论文原文为准。