智能体AI的运行时治理:基于可信溯源与故障闭锁执行的行动边界控制
Runtime Governance for Agentic AI: Action-Boundary Control with Trusted Provenance and Fail-Closed Execution
浏览论文内容
中文总结 AI 辅助
该研究提出Aegis运行时治理系统,通过可信决策层调解智能体AI的工具行动提案,在沙堡语料库评估中成功阻止风险提案转化为治理副作用。
中文摘要 AI 辅助
智能体AI系统会请求可修改文件、发送消息、启动作业或更改工作流状态的工具行动,这将安全问题从有害文本生成转向有害操作副作用。提示级治理可塑造模型行为,但无法构建执行边界。我们提出Aegis运行时治理系统,将模型输出视为行动提案,在工具执行前通过可信决策层进行调解:模型提案,可信运行时决策。Aegis会根据活动策略状态评估提案、在服务器端解析溯源、在不确定性下采用故障闭锁机制,并通过基于法定人数的非单边授权路径Senate-style Settlement路由选定案例。我们在包含5个运行族、42个任务、3种条件且每个族重复10次的重复沙堡语料库上评估Aegis:在6300行数据中,提示-策略条件产生79条风险比较路径泄漏行;在2100行Aegis治理行中,系统记录到零治理的模拟工具应用和零治理的风险副作用完成;1832条Aegis尝试治理行均保留了Aegis解析的可信溯源,1019条Senate-settled行均具备法定人数及最终签名计数证据。这些结果未证明通用自主智能体安全性,而是支持更狭义的系统主张:在本次评估的沙堡语料库中,运行时行动边界治理可防止观测到的风险提案成为治理副作用。
英文摘要
Agentic AI systems request tool actions that can modify files, send messages, launch jobs, or change workflow state. This shifts the safety problem from harmful text generation to harmful operational side effects. Prompt-level governance can shape model behavior, but it does not create an execution boundary. We introduce Aegis, a runtime governance system that treats model outputs as action proposals and mediates them through a trusted decision layer before tool execution. The model proposes; the trusted runtime decides. Aegis evaluates proposals against active policy state, resolves provenance server-side, fails closed under uncertainty, and routes selected cases through Senate-style settlement, a quorum- based non-unilateral authorization path. We evaluate Aegis on a repeated sandbox corpus spanning five run families, 42 tasks, three conditions, and ten repeats per family. Across 6,300 rows, prompt-policy conditioning produced 79 risky comparator-path leakage rows. Across 2,100 Aegis-governed rows, the system recorded zero governed mock-tool applications and zero governed risky side-effect completions. All 1,832 Aegis-attempted governed rows preserved trusted Aegis-resolved provenance, and all 1,019 Senate-settled rows had quorum and final signed tally evidence. These results do not prove general autonomous-agent safety. They support the narrower systems claim that, in this evaluated sandbox corpus, runtime action-boundary governance prevented observed risky proposals from becoming governed side effects.
发表机构
- SPQR Technologies Inc.(SPQR科技公司)
机构由 AI 辅助整理,请以论文原文为准。