arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SingGuard-NSFA:通过生成式推理和实时分类实现的可扩展智能体人工智能护栏

SingGuard-NSFA: Extensible Guardrails for Agentic AI via Generative Reasoning and Real-Time Classification

SingGuard Team

arXiv 2607.13081首次发表:更新:

发表机构

Ant Group(蚂蚁集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对智能体人工智能系统操作威胁,提出NSFA分类法构建基准套件,开发双模式方法,发布四个不同参数模型,在基准测试和跨源评估中表现出色,证明方法可扩展性与通用性。

AI 中文摘要

我们提出了nsfaguard,这是一种用于保护智能体人工智能系统免受操作威胁的护栏框架,如提示注入、敏感信息提取、恶意代码请求、危险工具滥用和资源耗尽。我们首先引入NSFA分类法,将185种风险变体组织成一个基于CIA三元组的层次结构,并根据三个成熟的OWASP指南进行交叉验证。基于此分类法,我们构建了一个涵盖133种语言的基准套件,包括针对用户查询和智能体响应的超过93K个专门构建的样本,以及从五个公共智能体安全数据集改编的3435个跨源样本。为了在实践中检测这些操作威胁,我们开发了一种双模式方法,将基于SFT的生成式推理用于可解释的离线审计,与冻结主干上的判别式分类头相结合,实现约50毫秒的实时检测。我们发布了四个参数分别为0.8B、2B、4B和9B的模型,在专门构建的基准上均达到≥94%的F1,比最强的竞争护栏高出6到12个绝对百分点。在跨源评估中,9B模型以更平衡的精确率-召回率权衡达到91.29%的F1。此外,消融实验表明分类头可以为护栏配备超出其原始范围的风险检测能力并实现最优性能。这些结果证明了该方法的可扩展性及其作为插件增强的通用性。

英文摘要

We present nsfaguard, a guardrail framework for securing agentic AI systems against operational threats, such as prompt injection, sensitive information extraction, malicious code requests, dangerous tool misuse, and resource exhaustion. We first introduce the NSFA taxonomy, which organizes 185 risk variants into a CIA-triad-grounded hierarchy and is cross-validated against three well-established OWASP guidelines. Based on this taxonomy, we construct a benchmark suite spanning 133 languages, comprising over 93K purpose-built samples targeting both user queries and agent responses, along with 3,435 cross-source samples adapted from five public agent-security datasets. To detect these operational threats in practice, we develop a dual-mode approach combining SFT-based generative reasoning for interpretable offline auditing with discriminative classification heads on the frozen backbone, enabling real-time detection at approximately 50,ms. We release four models with 0.8B, 2B, 4B, and 9B parameters, all achieving $\geq$94% F1 on purpose-built benchmarks and surpassing the strongest competing guardrails by 6 to 12 absolute points. On cross-source evaluation, the 9B model attains 91.29% F1 with a more balanced precision--recall trade-off. Moreover, ablation experiments show that classification heads can equip a guardrail with risk detection capabilities beyond its original scope and achieve state-of-the-art performance. These results demonstrate the extensibility of the approach and its generality as a plug-in enhancement.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑