AI 中文总结
Niyam-AI是一种基于零知识证明的意图绑定AI智能体框架,通过SHA-256承诺意图合约、Judge模型验证和zk-SNARK确保证明,在Agent-SafetyBench评估中较多款基线护栏实现显著安全性能提升。
AI 中文摘要
赋予AI智能体发送邮件、查询数据库或执行命令的能力十分有用,但前提是智能体不会被诱骗执行不当操作。提示注入、推理幻觉和不安全工具调用是自主LLM智能体的主要攻击面。现有防御措施依赖于在攻击者目标所在的同一机器上运行的软件检查,如系统提示或策略过滤器,无法提供可验证的执行证明。我们引入Niyam-AI框架,该框架可实现可证明的安全执行。会话开始时,允许的工具和约束会被锁定到通过SHA-256承诺的意图合约中。每次工具调用都会被隔离的Judge模型拦截并验证;通过验证后,会通过EZKL生成zk-SNARK证明。工具仅在证明验证后执行,允许第三方在不访问Judge模型权重的情况下确认执行情况。我们使用Agent-SafetyBench中的2000个真实场景,对NeMo Guardrails、Meta的Llama Prompt Guard 2和OpenAI的GPT-OSS-Safeguard进行5折分层交叉验证评估Niyam-AI,得到F1分数为88.5%,假阳性率为1.1%(自助法95%置信区间:[85.19%, 91.88%],N=1000)。McNemar精确配对检验证实了显著改进:Niyam-AI在390个不一致场景中优于NeMo(vs 20次失败),在115个场景中优于Prompt Guard 2(vs 13次失败),在384个场景中优于GPT-OSS-Safeguard(vs 19次失败),所有情况下p值均小于0.0001。每次批准的操作的证明生成为2260.6±218.4毫秒,验证耗时53.1±11.8毫秒。Niyam-AI提供了一种既高度准确又数学可验证的护栏——尽管这反映了针对Agent-SafetyBench调整的分类器与针对零样本基线的评估,这一区别在第IV.C节中讨论。
英文摘要
Autonomous LLM agents with tool execution capabilities introduce severe security risks through prompt injection, goal hijacking, and unauthorized action invocation. Existing guardrails rely on unverified, host local software filters system prompts, semantic classifiers, policy engines that share the execution environment of the untrusted agent, offering no guarantee to an external observer that a safety policy was correctly evaluated. A compromised host produces no evidence of its own failure. This paper presents NiyamAI, an intent bound runtime guardrail architecture providing cryptographically verifiable execution integrity for autonomous agents. At session initialization, permitted tools and operational constraints are sealed into an immutable Intent Contract under a SHA256 commitment. Every tool invocation is intercepted by a deterministic authority gate and classified by a dedicated neural Judge (11->8->2 feedforward network). For each authorized action, NiyamAI generates a succinct zkSNARK proof certifying correct policy evaluation under the committed contract; execution proceeds only after that proof verifies. Across 2,000 AgentSafetyBench scenarios under 5fold stratified crossvalidation with out of fold scoring, NiyamAI achieves 88.8% F1 at a 1.0% false positive rate (bootstrap 95% CI [85.5%, 92.1%]), against 66.8% for Llama Prompt Guard 2, 46.2% for GPTOSSSafeguard, and 40.4% for NeMo Guardrails; McNemar's exact test confirms each margin at p < 0.0001. Proof generation adds 1.7 s per approved action, verification 51 ms, with an 18.6 KB proof verifiable by any third party without access to model parameters. We further subject NiyamAI's own enforcement mechanism to 18 adversarial vectors across six classes, disclosing two implementation vulnerabilities identified and remediated during development.