arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22926cs.AIcs.CR

SAGE:用于高影响力生成式人工智能验证生命周期控制的安全至上深度防御护栏

SAGE: Safety-First Defense-in-Depth Guardrails for Verified Lifecycle Control of High-Impact Generative AI

  • Quandary Peak Research(Quandary Peak研究公司)

机构由 AI 辅助整理,请以论文原文为准。

Mahdi Eslamimehr

AI总结:

研究高影响力生成式人工智能的生命周期控制问题,提出SAGE安全至上、授权分离架构,结合多种技术手段。通过形式化结果和实验验证,确立安全优先级等,给出有害合规估计,还介绍了预先注册的扩展测试方法。

AI中文摘要:

高影响力生成式人工智能使灾难性滥用成为一个生命周期控制问题,而非仅仅是提示过滤问题。SAGE是一种安全至上、授权分离的架构,其中可信的灾难性启用风险在考虑效用、延迟或商业目标之前限制可接受性。它结合了签名发布清单、多样的检测器、稳健的风险范围、最低风险默认值、输出检查、三值监测、受保护的审计链、遏制和回滚。形式化结果确立了安全优先级、保守的检测器界限、单调的发布门控、防篡改记录和授权削减;两个PRISM抽象验证了在明确假设下的授权分离和生命周期不变性。一项冻结的、供应商对称的研究向四个GPT、四个Claude和两个Gemini快照各发送了84个案例:840次调用产生了794个目标响应、46个提供者错误和449个涵盖375个响应的成功判断。八个快照具有完整的判断域覆盖。有害合规估计较低;差异主要源于良性效用和安全重定向。涉及Claude、Gemini或GPT-5快照与GPT-5 mini和GPT-5 nano快照的七个多重性调整对比得到支持,而Claude或Gemini快照与GPT-5或GPT-5.5之间的任何测试对比在校正后均未通过。观察到的有害合规范围是在没有工具、检索、历史或人工裁决的情况下,每个提示一代的保守的、协议约束的视图;它不是操作辅助的上限。一个预先注册的扩展指定了如何使用锁定分割、重复采样、多轮和沙盒工具条件以及领域专家评分来测试更宽的最佳-最差差距。

英文摘要:

High-impact generative AI makes catastrophic misuse a lifecycle-control problem, not merely a prompt-filtering problem. SAGE is a safety-first, authorization-separated architecture in which credible catastrophic-enablement risk constrains admissibility before utility, latency, or commercial objectives are considered. It combines signed release manifests, diverse detectors, robust risk envelopes, least-risk defaults, output checking, three-valued monitoring, protected audit chains, containment, and rollback. Formal results establish safety priority, conservative detector bounds, monotone release gating, tamper-evident records, and an authorization cut; two PRISM abstractions verify authorization separation and lifecycle invariants under explicit assumptions. A frozen, vendor-symmetric study sent 84 cases to each of four GPT, four Claude, and two Gemini snapshots: 840 calls yielded 794 target responses, 46 provider errors, and 449 successful judgments covering 375 responses. Eight snapshots had complete judged domain coverage. Harmful-compliance estimates were low; variation arose mainly from benign utility and safe redirection. Seven multiplicity-adjusted contrasts involving Claude, Gemini, or GPT-5 snapshots and the GPT-5 mini and GPT-5 nano snapshots were supported, while no tested contrast between the Claude or Gemini snapshots and GPT-5 or GPT-5.5 survived correction. The observed harmful-compliance range is a conservative, protocol-bound view from one generation per prompt with no tools, retrieval, history, or human adjudication; it is not an upper bound on operational assistance. A preregistered extension specifies how to test a wider best-worst gap using a locked split, repeated sampling, multi-turn and sandboxed-tool conditions, and domain-expert scoring.

补充信息

↑