Persona Guardrail:面向智能体系统的生产级防御框架
Persona Guardrail: A Production-Grade Defense Framework for Agentic Systems
- Uber(优步)
- University of Texas at Dallas(德克萨斯大学达拉斯分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出Persona Guardrail生产级防御框架,通过语义允许/阻止列表同步验证输入输出,并引入PAGE基准,将整体准确率提升至95.9%,域外检测率提升至93.5%,误批准率降至4.7%。
AI中文摘要:
基于大语言模型的智能体正越来越多地被部署到企业环境中,通过与企业知识库、工具和外部服务交互来执行特定领域的任务。现有的运行时防护措施主要针对提示注入和其他特定攻击行为,且基于黑盒威胁模型,但在确保智能体在其预期功能范围内运行方面提供的保障有限。因此,生产环境中的智能体仍然容易受到恶意请求和域外查询的攻击,而现有防御措施往往难以区分这些请求。我们提出了Persona Guardrail,一个生产级运行时防御框架,通过同步的输入和输出验证,由语义允许列表和阻止列表规范驱动,为面向客户的智能体AI系统强制执行明确的功能边界。我们还引入了PAGE(Persona-Aware Guardrail Evaluation,角色感知防护栏评估),一个基准测试,用于在用户和智能体轮次中评估良性、对抗性和域外交互下的功能特定防护栏。与基于通用LLM的防护栏相比,Persona Guardrail将整体准确率从85.7%提高到95.9%,将域外检测率从57.3%提高到93.5%,并将误批准率从25.0%降低到4.7%。目前已在生产中部署,Persona Guardrail在满足延迟预算的同时,在现实生产工作负载下维持了极低的误拦截和误放行率。这些结果表明,Persona Guardrail为保护智能体AI系统提供了一个实用、可扩展且生产就绪的基础。
英文摘要:
Large language model-based agents are increasingly deployed to perform domain-specific tasks by interacting with enterprise knowledge, tools, and external services. Existing runtime guardrails primarily target prompt injection and other attack-specific behaviors under a black-box threat model, but provide limited guarantees that agents operate within their intended functionality. As a result, production agents remain vulnerable to malicious requests and out-of-domain queries that existing defenses often fail to distinguish. We present Persona Guardrail, a production-grade runtime defense framework that enforces explicit functional boundaries for customer-facing agentic AI systems through synchronous input and output validation driven by semantic allowlist and blocklist specifications. We also introduce PAGE (Persona-Aware Guardrail Evaluation), a benchmark for evaluating function-specific guardrails across benign, adversarial, and out-of-domain interactions on both user and agent turns. Compared with a generic LLM-based guardrail, Persona Guardrail improves overall accuracy from 85.7% to 95.9%, increases out-of-domain detection from 57.3% to 93.5%, and reduces the false-approved rate from 25.0% to 4.7%. Currently deployed in production, Persona Guardrail meets its latency budget while sustaining a very low false-block and false-allow rate under realistic production workloads. These results demonstrate that Persona Guardrail provides a practical, scalable, and production-ready foundation for securing agentic AI systems.